Method, apparatus and storage medium for face rigid body model and fixation point detection
By generating a rigid face model, the problem of decreasing gaze point calculation accuracy when faces move or rotate is solved, and high-precision gaze point detection is achieved.
Patent Information
- Application Number
- CN202011499370.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-18
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2040-12-18
AI Technical Summary
In the prior art, when faces move or rotate, the polynomial model obtained by historical 9-point calibration is no longer applicable, resulting in a significant decrease in the calculation accuracy of the gaze point.
By obtaining the face image taken by the image sensor, determining the characteristic points of the face, establishing the left and right eye vectors, and mapping them into the three-dimensional face model, and generating a face rigid body model to adapt to the changes in the face.
It realizes accurate calculation of the gaze point when the face moves or rotates, improving the accuracy and stability of gaze point detection.
Smart Images

Figure CN114724200B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of machine vision, and in particular, to a method, apparatus, and storage medium for detecting a rigid face model and a fixation point. Background Art
[0002] At present, most spatial gaze tracking systems adopt the method of polynomial model mapping. Generally, a 9-point calibration method is used, that is, the human eye sequentially fixates on 9 calibration points on the screen (equivalent to a two-dimensional plane), so as to obtain the coordinate mapping model between the pupil and the screen, and use this model to calculate the fixation point coordinates on the screen.
[0003] Although the 9-point calibration method has high accuracy when the face remains stationary, when the face moves or rotates, after the face pose changes, the polynomial model obtained by the historical 9-point calibration is no longer applicable to the gaze calculation under the current face pose, resulting in a significant decrease in the fixation point calculation accuracy. Summary of the Invention
[0004] The present invention provides a method, apparatus, and storage medium for detecting a rigid face model and a fixation point, so as to solve the above technical problems existing in the prior art.
[0005] In a first aspect, to solve the above technical problems, the technical solution of the method for generating a rigid face model provided by an embodiment of the present invention is as follows:
[0006] Obtain a first visible light face image of the current user taken by an image sensor facing forward, and obtain a preset number of feature points of the face of the current user from the first visible light face image;
[0007] Based on the preset number of feature points, determine the three-dimensional coordinates of the left eye center, the three-dimensional coordinates of the right eye center, and the three-dimensional coordinates of the face center of the current user;
[0008] Respectively establish a left eye vector and a right eye vector from the three-dimensional coordinates of the left eye center and the three-dimensional coordinates of the right eye center to the three-dimensional coordinates of the face center, and map them into a three-dimensional face model to obtain the rigid face model of the current user.
[0009] A possible implementation manner, obtaining a first visible light face image of the current user taken by an image sensor facing forward, includes:
[0010] Display a first fixation area at the position closest to the imaging device in the display screen;
[0011] Continuously capture a second visible light face image of the current user when the current user is fixating on the first fixation area until it is determined that the face image of the current user is captured facing forward;
[0012] For each capture, determine the second intersection point between the face orientation of the current user and the display screen based on the second visible-light face image, as well as the second face pose; wherein, the second face pose is used to represent the pose angle of the current user's face.
[0013] Determine whether the second intersection point is within the first fixation region and whether the second face pose is within a set error range.
[0014] If the second intersection point is within the first fixation region and the second face pose is within the set error range, determine that the face image of the current user has been captured frontally, and use the corresponding second visible-light face image as the first visible-light face image.
[0015] A possible implementation manner, determining whether the second face pose is within the set error range includes:
[0016] Abstract the error value of the second face pose as a line segment starting from the center of the first fixation region, and use the boundary of the first fixation region as the limit value of the error range.
[0017] Determine whether the line segment exceeds the first fixation region.
[0018] If the line segment exceeds the first fixation region, determine that the second face pose is not within the error range.
[0019] If the line segment does not exceed the first fixation region, determine that the second face pose is within the error range.
[0020] A possible implementation manner, determining the three-dimensional coordinates of the left eye center, the three-dimensional coordinates of the right eye center, and the three-dimensional coordinates of the face center of the current user based on the preset number of feature points includes:
[0021] Convert the pixel coordinates corresponding to the preset number of feature points into three-dimensional coordinates in the world coordinate system.
[0022] Determine the left-eye region feature points and the right-eye region feature points from the preset number of feature points.
[0023] Determine the center points of the left-eye region feature points and the right-eye region feature points, and use the center point of the left-eye region feature points as the three-dimensional coordinates of the left eye center, and use the center point of the right-eye region feature points as the three-dimensional coordinates of the right eye center.
[0024] According to the preset eyeball radius and the three-dimensional coordinates of the left eye center and the three-dimensional coordinates of the right eye center, determine the three-dimensional coordinates of the left eye center of the eyeball and the three-dimensional coordinates of the right eye center of the eyeball.
[0025] Determine the three-dimensional coordinates of the center of the human face according to the three-dimensional coordinates corresponding to the preset number of feature points.
[0026] A possible implementation manner, after obtaining the rigid body model of the human face of the current user, further includes:
[0027] Display a second fixation area at the center of the display screen, and obtain a third visible light human face image and a third infrared human face image captured when the current user fixates on the second fixation area;
[0028] Determine the third human face pose of the current user based on the third visible light human face image, and detect the center positions of the left eye pupil and the right eye pupil in the third infrared image to obtain the three-dimensional coordinates of the left eye pupil center and the coordinates of the right eye pupil center; wherein, the third human face pose is used to represent the pose angle of the current user's human face.
[0029] Update the pose of the rigid body model of the human face through the third human face pose, and combine the three-dimensional coordinates of the left eye pupil center and the coordinates of the right eye pupil center to obtain the current fixation point of the current user.
[0030] Determine whether the current fixation point is within the second fixation area;
[0031] If it is not within the second fixation area, move the current fixation point into the second fixation area, and fine-tune the left eye vector and the right eye vector in the rigid body model of the human face according to the moving distance to obtain a corrected rigid body model of the human face.
[0032] A possible implementation manner, detecting the center positions of the left eye pupil and the right eye pupil in the third infrared image to obtain the three-dimensional coordinates of the left eye pupil center and the coordinates of the right eye pupil center, includes:
[0033] Convert the third infrared image into a fourth infrared image in the same coordinate system as the third visible light image;
[0034] Obtain the left eye area and the right eye area from the fourth infrared image, and use a pupil detection algorithm to obtain the pixel coordinates of the left eye pupil center and the pixel coordinates of the right eye pupil center from the left eye area and the right eye area respectively;
[0035] Convert the pixel coordinates of the left eye pupil center and the pixel coordinates of the right eye pupil center into the three-dimensional coordinates of the left eye pupil center and the three-dimensional coordinates of the right eye pupil center in the world coordinate system respectively.
[0036] A possible implementation manner, updating the pose of the rigid body model of the human face through the third human face pose, and combining the three-dimensional coordinates of the left eye pupil center and the coordinates of the right eye pupil center to obtain the current fixation point of the current user, includes:
[0037] Adjust the rotation angles of the left eye vector and the right eye vector in the rigid face model according to the third face pose to obtain an adjusted left eye vector and an adjusted right eye vector;
[0038] Obtain the three-dimensional coordinates of the current left eye center of the eyeball and the three-dimensional coordinates of the current right eye center of the eyeball from the adjusted left eye vector and the adjusted right eye vector;
[0039] Use the extension line of the connection between the three-dimensional coordinates of the current left eye center of the eyeball and the three-dimensional coordinates of the left eye pupil as the left eye line of sight, and use the extension line of the connection between the three-dimensional coordinates of the current right eye center of the eyeball and the three-dimensional coordinates of the right eye pupil as the right eye line of sight;
[0040] Determine the two intersection points of the left eye line of sight and the right eye line of sight with the display screen respectively, and use the center point of the connection between the two intersection points as the current fixation point.
[0041] In a second aspect, an embodiment of the present invention provides a method for detecting a fixation point, including:
[0042] Obtain a fourth visible light face image and a fourth infrared face image of the current user, obtain a preset number of feature points of the current user's face from the fourth visible light face image, and obtain the current left eye pupil center pixel coordinates and the current right eye pupil center pixel coordinates from the fourth infrared face image;
[0043] Calculate the preset number of feature points by using a pose measurement method based on PnP to obtain the current face pose of the current user;
[0044] Update the rigid face model as described in the first aspect based on the current face pose to obtain the three-dimensional coordinates of the current left eye center of the eyeball and the three-dimensional coordinates of the right eye center of the eyeball; convert the current left eye pupil center pixel coordinates and the current right eye pupil center pixel coordinates into three-dimensional coordinates in the world coordinate system to obtain the three-dimensional coordinates of the current left eye pupil center and the three-dimensional coordinates of the current right eye pupil center;
[0045] Determine the current fixation point of the current user on the display screen according to the three-dimensional coordinates of the current left eye center of the eyeball, the three-dimensional coordinates of the right eye center of the eyeball, the three-dimensional coordinates of the current left eye pupil center and the three-dimensional coordinates of the current right eye pupil center.
[0046] A possible implementation manner of obtaining the current left eye pupil center pixel coordinates and the current right eye pupil center pixel coordinates from the fourth infrared face image includes:
[0047] Convert the fourth infrared image into a fifth infrared image in the same coordinate system as the fourth visible light image;
[0048] Obtain the left eye region and the right eye region from the fifth infrared image, and use a pupil detection algorithm to obtain the left eye pupil center pixel coordinates and the right eye pupil center pixel coordinates from the left eye region and the right eye region respectively.
[0049] A possible implementation manner is to update the face rigid body model as described in the first aspect based on the current face pose, and obtain the current three-dimensional coordinates of the left eye ball center and the right eye ball center, including:
[0050] Adjust the rotation angles of the left eye vector and the right eye vector in the face rigid body model according to the current face pose to obtain an adjusted left eye vector and an adjusted right eye vector;
[0051] Obtain the current three-dimensional coordinates of the left eye ball center and the current three-dimensional coordinates of the right eye ball center from the adjusted left eye vector and the adjusted right eye vector.
[0052] A possible implementation manner is to determine the current fixation point of the current user on the display screen according to the current three-dimensional coordinates of the left eye ball center, the three-dimensional coordinates of the right eye ball center, the current three-dimensional coordinates of the left eye pupil center, and the current three-dimensional coordinates of the right eye pupil center, including:
[0053] Use the extension line of the connection between the current three-dimensional coordinates of the left eye ball center and the three-dimensional coordinates of the left eye pupil as the left eye line of sight;
[0054] Use the extension line of the connection between the current three-dimensional coordinates of the right eye ball center and the three-dimensional coordinates of the right eye pupil as the right eye line of sight;
[0055] Determine the two intersection points of the left eye line of sight and the right eye line of sight with the display screen respectively, and use the center point of the connection of the two intersection points as the current fixation point.
[0056] In a third aspect, an embodiment of the present invention provides a device for generating a face rigid body model, including:
[0057] An acquisition module, configured to acquire a first visible light face image of the current user captured by the image sensor facing forward, and acquire a preset number of feature points of the face of the current user from the first visible light face image;
[0058] A determination module, configured to determine the three-dimensional coordinates of the left eye ball center, the three-dimensional coordinates of the right eye ball center, and the three-dimensional coordinates of the face center of the current user based on the preset number of feature points;
[0059] An acquisition module, configured to respectively establish a left eye vector and a right eye vector from the three-dimensional coordinates of the left eye center and the three-dimensional coordinates of the right eye center to the three-dimensional coordinates of the face center, and map them into the three-dimensional face model to obtain the rigid face model of the current user.
[0060] A possible implementation, the acquisition module is further configured to:
[0061] Display a first fixation area at the position closest to the imaging device on the display screen;
[0062] Continuously capture a second visible light face image of the current user when the current user is fixating on the first fixation area until it is determined that the face image of the current user is captured frontally;
[0063] Each time a capture is performed, determine a second intersection point between the face orientation of the current user and the display screen, and a second face pose according to the second visible light face image; wherein, the second face pose is used to represent the pose angle of the current user's face;
[0064] Determine whether the second intersection point is within the first fixation area and whether the second face pose is within a set error range;
[0065] If the second intersection point is within the first fixation area and the second face pose is within the set error range, determine that the face image of the current user is captured frontally, and use the corresponding second visible light face image as the first visible light face image.
[0066] A possible implementation, the acquisition module is further configured to:
[0067] Abstract the error value of the second face pose as a line segment starting from the center of the first fixation area, and use the boundary of the first fixation area as the limit value of the error range;
[0068] Determine whether the line segment exceeds the first fixation area;
[0069] If the line segment exceeds the first fixation area, determine that the second face pose is not within the error range;
[0070] If the line segment does not exceed the first fixation area, determine that the second face pose is within the error range.
[0071] A possible implementation, the determination module is further configured to:
[0072] Convert the pixel coordinates corresponding to the preset number of feature points into three-dimensional coordinates in the world coordinate system;
[0073] Determine the left-eye region feature points and the right-eye region feature points from the preset number of feature points;
[0074] Determine the center points of the left-eye region feature points and the right-eye region feature points, and use the center point of the left-eye region feature points as the three-dimensional coordinates of the left-eye center, and use the center point of the right-eye region feature points as the three-dimensional coordinates of the right-eye center;
[0075] According to the preset eyeball radius and the three-dimensional coordinates of the left-eye center and the three-dimensional coordinates of the right-eye center, determine the three-dimensional coordinates of the left-eye eyeball center and the three-dimensional coordinates of the right-eye eyeball center;
[0076] According to the three-dimensional coordinates corresponding to the preset number of feature points, determine the three-dimensional coordinates of the face center.
[0077] A possible implementation manner, after obtaining the rigid face model of the current user, the obtaining module is further configured to:
[0078] Display a second fixation area at the center of the display screen, and obtain a third visible-light face image and a third infrared face image captured when the current user fixates on the second fixation area;
[0079] Based on the third visible-light face image, determine the third face pose of the current user, and detect the center positions of the left-eye pupil and the right-eye pupil in the third infrared image to obtain the three-dimensional coordinates of the left-eye pupil center and the coordinates of the right-eye pupil center; wherein, the third face pose is used to represent the pose angle of the current user's face;
[0080] Update the pose of the rigid face model through the third face pose, and combine the three-dimensional coordinates of the left-eye pupil center and the coordinates of the right-eye pupil center to obtain the current fixation point of the current user;
[0081] Judge whether the current fixation point is within the second fixation area;
[0082] If it is not within the second fixation area, move the current fixation point into the second fixation area, and finely adjust the left-eye vector and the right-eye vector in the rigid face model according to the moving distance to obtain a corrected rigid face model.
[0083] A possible implementation manner, the obtaining module is further configured to:
[0084] Convert the third infrared image into a fourth infrared image in the same coordinate system as the third visible-light image;
[0085] Obtain the left eye region and the right eye region from the fourth infrared image, and use a pupil detection algorithm to obtain the left eye pupil center pixel coordinates and the right eye pupil center pixel coordinates from the left eye region and the right eye region respectively;
[0086] Convert the left eye pupil center pixel coordinates and the right eye pupil center pixel coordinates into the left eye pupil center three-dimensional coordinates and the right eye pupil center three-dimensional coordinates in the world coordinate system respectively.
[0087] A possible implementation, the obtaining module is further configured to:
[0088] Adjust the rotation angles of the left eye vector and the right eye vector in the rigid face model according to the third face pose to obtain the adjusted left eye vector and the adjusted right eye vector;
[0089] Obtain the current left eye center of the eye three-dimensional coordinates and the current right eye center of the eye three-dimensional coordinates from the adjusted left eye vector and the adjusted right eye vector;
[0090] Use the extension line of the connection between the current left eye center of the eye three-dimensional coordinates and the left eye pupil three-dimensional coordinates as the left eye line of sight, and use the extension line of the connection between the current right eye center of the eye three-dimensional coordinates and the right eye pupil three-dimensional coordinates as the right eye line of sight;
[0091] Determine the two intersection points of the left eye line of sight and the right eye line of sight with the display screen respectively, and use the center point of the connection of the two intersection points as the current fixation point.
[0092] In a fourth aspect, an embodiment of the present invention provides a device for detecting a fixation point, including:
[0093] An acquisition module, configured to acquire a fourth visible light face image and a fourth infrared face image of a current user, acquire a preset number of feature points of the current user's face from the fourth visible light face image, and acquire the current left eye pupil center pixel coordinates and the current right eye pupil center pixel coordinates from the fourth infrared face image;
[0094] A calculation module, configured to calculate the preset number of feature points by using a pose measurement method based on PnP to obtain the current face pose of the current user;
[0095] An obtaining module, configured to update the rigid face model as described in the first aspect based on the current face pose to obtain the current left eye center of the eye three-dimensional coordinates and the right eye center of the eye three-dimensional coordinates; convert the current left eye pupil center pixel coordinates and the current right eye pupil center pixel coordinates into three-dimensional coordinates in the world coordinate system to obtain the current left eye pupil center three-dimensional coordinates and the current right eye pupil center three-dimensional coordinates;
[0096] A determination module, configured to determine a current fixation point of the current user on a display screen according to the current three-dimensional coordinates of the left-eye eyeball center, the three-dimensional coordinates of the right-eye eyeball center, the current three-dimensional coordinates of the left-eye pupil center, and the current three-dimensional coordinates of the right-eye pupil center.
[0097] A possible implementation, the obtaining module is further configured to:
[0098] Convert the fourth infrared image into a fifth infrared image in the same coordinate system as the fourth visible light image;
[0099] Obtain a left-eye region and a right-eye region from the fifth infrared image, and respectively obtain the left-eye pupil center pixel coordinates and the right-eye pupil center pixel coordinates from the left-eye region and the right-eye region by using a pupil detection algorithm.
[0100] A possible implementation, the obtaining module is further configured to:
[0101] Adjust the rotation angles of the left-eye vector and the right-eye vector in the rigid face model according to the current face pose, and obtain an adjusted left-eye vector and an adjusted right-eye vector;
[0102] Obtain the current three-dimensional coordinates of the left-eye eyeball center and the current three-dimensional coordinates of the right-eye eyeball center from the adjusted left-eye vector and the adjusted right-eye vector.
[0103] A possible implementation, the determination module is specifically configured to:
[0104] Use the extension line of the connection between the current three-dimensional coordinates of the left-eye eyeball center and the three-dimensional coordinates of the left-eye pupil as the left-eye line of sight;
[0105] Use the extension line of the connection between the current three-dimensional coordinates of the right-eye eyeball center and the three-dimensional coordinates of the right-eye pupil as the right-eye line of sight;
[0106] Determine two intersection points of the left-eye line of sight and the right-eye line of sight with the display screen respectively, and use the center point of the connection of the two intersection points as the current fixation point.
[0107] In a fifth aspect, an embodiment of the present invention further provides a device for detecting a fixation point, including:
[0108] At least one processor, and
[0109] A memory connected to the at least one processor;
[0110] Wherein, the memory stores instructions executable by the at least one processor, and the at least one processor executes the instructions stored in the memory to execute the method described in the first aspect or the second aspect above.
[0111] In a sixth aspect, an embodiment of the present invention further provides a readable storage medium, including:
[0112] a memory,
[0113] wherein the memory is used to store instructions, and when the instructions are executed by a processor, the device including the readable storage medium completes the method described in the first aspect or the second aspect above. Description of the Drawings
[0114] Figure 1 is a flowchart of a method for generating a human face rigid body model provided by an embodiment of the present invention;
[0115] Figure 2 is a schematic diagram of a current user gazing at a first gaze area provided by an embodiment of the present invention;
[0116] Figure 3 is a partial schematic diagram of a display screen located in the first gaze area provided by an embodiment of the present invention;
[0117] Figure 4 is a schematic diagram that the intersection point of the human face orientation and the display screen is not within the first gaze area provided by an embodiment of the present invention;
[0118] Figure 5 is a schematic diagram that the intersection point of the human face orientation and the display screen is within the first gaze area provided by an embodiment of the present invention;
[0119] Figure 6 is a schematic diagram of the positional relationship between the current user and the display screen provided by an embodiment of the present invention;
[0120] Figure 7 is a schematic diagram of establishing a world coordinate system provided by an embodiment of the present invention;
[0121] Figure 8 is a schematic diagram of calculating the intersection point of the human face orientation and the display screen provided by an embodiment of the present invention;
[0122] Figure 9 is a schematic diagram of calibrating a human face rigid body model provided by an embodiment of the present invention;
[0123] Figure 10 is a flowchart of a method for detecting a gaze point provided by an embodiment of the present invention;
[0124] Figure 11 is a schematic structural diagram of a device for generating a human face rigid body model provided by an embodiment of the present invention;
[0125] Figure 12 is a schematic structural diagram of a device for detecting a gaze point provided by an embodiment of the present invention. Detailed Embodiments
[0126] The embodiments of the present invention provide a method, device, and storage medium for generating a rigid face model and detecting a fixation point, so as to solve the above-mentioned technical problems existing in the prior art.
[0127] In order to better understand the above technical solutions, the technical solutions of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific features in the embodiments of the present invention are detailed descriptions of the technical solutions of the present invention, rather than limitations on the technical solutions of the present invention. Without conflict, the technical features in the embodiments of the present invention and the embodiments can be combined with each other.
[0128] Please refer to Figure 1 , the embodiments of the present invention provide a method for generating a rigid face model, and the processing process of this method is as follows.
[0129] Step 101: Obtain a first visible light face image of the current user captured by the image sensor in the forward direction, and obtain a preset number of feature points of the face of the current user from the first visible light face image.
[0130] In the embodiments provided by the present invention, the image sensor includes a visible light image sensor and an infrared image sensor. The visible light image sensor is mainly used to collect RGB face images, and the infrared image sensor is mainly used to collect pupil images. The visible light sensor and the infrared image sensor are arranged on one side of the display screen. For example, they can be arranged on the upper side or the lower side, the left side or the right side of the display screen. Preferably, they can be arranged at the central position on one side, such as the central position on the lower side.
[0131] Before photographing the current user, it is also necessary to calibrate the image sensor, that is, calibrate the visible light image sensor and the infrared image sensor to obtain the internal parameter matrix of each image sensor. The calibration method is as follows: First, according to Zhang's calibration method, place the calibration board within the range of 0.3m to 1m in front of the image sensor and the infrared image sensor to ensure that each image sensor can capture a clear calibration board image; then, rotate or twist the calibration board. After each rotation or twist, both image sensors collect an image. After taking 20 images at different angles, each image sensor obtains 20 calibration images; finally, each image sensor calculates its own internal parameter matrix based on its own 20 calibration images; the above calibration usually only needs to be performed once. After obtaining the internal parameter matrix of the camera, there is no need to repeat the calibration.
[0132] After calibrating the image sensor, it is also necessary to set the parameters of the image sensor. The parameters of the image sensor include resolution, field of view (FOV), frame rate, etc. For example, a visible light image sensor can set its resolution to 640×480, FOV to 60°, and frame rate to 60fps. An infrared image sensor can set its resolution to 2560*1920, FOV to 50°, and frame rate to 30fps.
[0133] After completing the above calibration and parameter setting of the image sensor, image capture can be performed.
[0134] To obtain the first visible light face image of the current user captured by the image sensor facing forward, the following method can be used:
[0135] Display a first fixation area at the position closest to the capture device on the display screen; continuously capture the second visible light face image of the current user when the current user is fixating on the first fixation area until it is determined that the face image of the current user is captured facing forward; each time a capture is performed, determine the second intersection point of the face orientation of the current user and the display screen based on the second visible light face image, and the second face pose; where the second face pose is used to represent the pose angle of the current user's face; determine whether the second intersection point is within the first fixation area and whether the second face pose is within the set error range; if the second intersection point is within the first fixation area and the second face pose is within the set error range, determine that the face image of the current user is captured facing forward, and use the corresponding second visible light face image as the first visible light face image.
[0136] Among them, determining whether the second face pose is within the set error range can be achieved by the following method:
[0137] Abstract the error value of the second face pose as a line segment starting from the center of the first fixation area, and use the boundary of the first fixation area as the boundary value of the error range; determine whether the line segment exceeds the first fixation area; if the line segment exceeds the first fixation area, determine that the second face pose is not within the error range; if the line segment does not exceed the first fixation area, determine that the second face pose is within the error range.
[0138] Please refer to Figure 2 and Figure 3 , Figure 2 which is a schematic diagram of the current user fixating on the first fixation area provided by an embodiment of the present invention, Figure 3 which is a partial schematic diagram of the display screen located in the first fixation area provided by an embodiment of the present invention.
[0139] In Figure 2The direction of the user's gaze is indicated by a dotted line. The area within the elliptical dotted line is the first fixation area. Three line segments can extend outward from the center of this first fixation area. These three line segments are respectively used to represent the measurement errors of the pitch, yaw, and roll corresponding to the face pose. The radius of the first fixation area is used to represent the maximum value of the set error range corresponding to these measurement errors. For example, for each unit increase in the error of the pitch angle, the corresponding line segment increases by one pixel in length, and the other two line segments are similar. As Figure 3 shown, the line segment extending longitudinally from the center of the first fixation area is used to represent the pitch angle corresponding to the face pose, the line segment extending horizontally from the center of the first fixation area is used to represent the yaw angle corresponding to the face pose, and the line segment extending at a 45° angle to the horizontal from the center of the first fixation area is used to represent the roll angle corresponding to the face pose. Among them, the first fixation area can be circular, and the value range of the radius of this circle is 3 to 6 pixels.
[0140] After the first fixation area is displayed at the position closest to the imaging device on the display screen, when the current user gazes at this first fixation area, the image sensor continuously captures the face image of the current user to obtain a second visible light face image until it is determined that the face image of the current user is captured directly.
[0141] Suppose a second visible light face image 1 is currently captured. This second visible light face image 1 can be analyzed and processed to determine the second intersection point 1 of the face orientation of the current user and the display screen, and the second face pose 1; then it is judged whether the second intersection point 1 is within the first fixation area, and whether the second face pose 1 is within the set error range (that is, by judging whether the three line segments extending from the center of the first fixation area are within the first fixation area), so as to facilitate prompting the current user to adjust their face pose through the first fixation area and its corresponding three line segments, enabling the image sensor to capture the face image of the current user directly.
[0142] Please refer to Figure 4 for the schematic diagram provided by the embodiment of the present invention where the intersection point of the face orientation and the display screen is not within the first fixation area. Suppose the judgment result is that the second intersection point 1 ( Figure 4 indicated by a black dot) is not within the first fixation area, then continue to capture the face image of the current user to obtain a second face image 2.
[0143] At this time, analyze and process this second visible light face image 2 again to determine the second intersection point 2 of the face orientation of the current user and the display screen, and the second face pose 2; then, judge whether the second intersection point 2 is within the first fixation area, and whether the second face pose 2 is within the set error range.
[0144] Please refer to Figure 5 This is a schematic diagram of the intersection point between the face orientation and the display screen within the first fixation area provided by an embodiment of the present invention. Assume that at this time, the second intersection point 2 (indicated by a black dot in the same way) is within the first fixation area, and the second face pose 2 is within the set error range, then it is determined that the face image of the current user is captured forward. At this time, the second visible light face image 2 is used as the first visible light face image, and correspondingly, the second face pose 2 is used as the first face pose corresponding to the first visible light face image.
[0145] In the above processing process, to determine the intersection point between the face orientation and the display screen, the following method can be used:
[0146] Obtain a preset number of feature points of the current user's face (such as 68 feature points of the face) from the second visible light face image, calculate the face pose data (composed of a rotation matrix and a translation matrix) through a pose measurement method based on PnP, and then obtain the rotation angle of the face by combining the Rodriguez transformation; according to the face rotation angle, obtain the vector of the face orientation, and combine the distance from the face to the display screen to calculate the intersection point between the face orientation and the display screen.
[0147] Please refer to Figure 6 This is a schematic diagram of the positional relationship between the current user and the display screen provided by an embodiment of the present invention. In Figure 6 , taking a 15-inch display screen as an example (the length and width dimensions of the display screen are 332×187mm), to ensure that within the viewing range of 40 - 70 cm (assuming that the current user is facing the center of the display screen, the viewing range is the distance between the current user and the center of the display screen), in Figure 6 , assuming that the distance between the current user and the center of the display screen is 550mm, the image sensor can capture the face image better. It is necessary to set the optical axis direction and position of the image sensor. Taking the image sensor being set at the middle position on the lower side of the display screen as an example, the optical axis can be determined by calculating the angle α between the optical axis and the horizontal plane:
[0148]
[0149] Please refer to Figure 7 This is a schematic diagram of establishing a world coordinate system provided by an embodiment of the present invention. Taking the center of the image sensor as the origin, the positive direction of the X-axis is horizontally to the right, the positive direction of the Z-axis is the optical axis direction, and the direction perpendicular to the XOZ plane is the positive direction of the Y-axis to establish a world coordinate system, and construct the plane equation of the display screen in this world coordinate system.
[0150] To construct the plane equation of the display screen, first obtain the coordinates of the four vertices of the display screen: A(-166, 31, 184), B(166, 31, 184), C(-166, 0, 0), D(166, 0, 0), and the center coordinate of the display screen E(0, 15.5, 92). The normal vector of the plane where the display screen is located is Substituting into the point-normal form plane equation, the plane equation of the display screen can be obtained: 15.5×(y - sinα) + 92×(z - cosα) = 0.
[0151] After that, perform face detection on the second visible light image to obtain the corresponding face region. Then, perform face feature (LandMark) point detection on this face region to obtain 68 feature points of the face, and use the pose measurement method based on PnP to obtain the rotation matrix (R) and translation matrix (T) representing the face pose. For the rotation vector Perform the Rodriguez transformation to obtain the rotation matrix R. The transformation formula is as follows:
[0152]
[0153] where r x 、r y 、r z are the corresponding values of the unit vector r of R in x, y, and z respectively.
[0154] It should be noted that the rotation matrix can be understood as obtained by rotating around a rotation vector by a certain angle. Therefore, the rotation matrix can be described by the rotation variable, and the two can be converted through the Rodriguez transformation.
[0155]
[0156] The unit vector r of
[0157]
[0158] The Rodriguez formula is:
[0159]
[0160] where R is the rotation matrix and I is the identity matrix.
[0161] After obtaining the above rotation matrix, the face pose can be calculated, that is, the angles of face rotation (pitch angle, yaw angle, roll angle). The calculation formulas are as follows:
[0162]
[0163] Among them, q0 is one of the intermediate variables for calculating pitch, yaw, and roll, and R(0,0), R(1,1), and R(2,2) are the values of the three elements on the diagonal of the rotation matrix R, respectively.
[0164]
[0165] Among them, q1 is one of the intermediate variables for calculating pitch, yaw, and roll, and R(2,1) and R(1,2) are the values of the element in the 2nd row and 1st column, and the element in the 1st row and 2nd column of the rotation matrix R, respectively.
[0166]
[0167] Among them, q2 is one of the intermediate variables for calculating pitch, yaw, and roll, and R(0,2) and R(2,0) are the values of the element in the 0th row and 2nd column, and the element in the 2nd row and 0th column of the rotation matrix R, respectively.
[0168]
[0169] Among them, q3 is one of the intermediate variables for calculating pitch, yaw, and roll, and R(1,0) and R(0,1) are the values of the element in the 1st row and 0th column, and the element in the 0th row and 1st column of the rotation matrix R, respectively.
[0170]
[0171] yaw = arcsin(2×(q0q2 + q1q3)) (11);
[0172]
[0173] Assume that the three-dimensional coordinates of the face center are F1(x1, y1, z1), and the face orientation vector is F2(x2, y2, z2) (here the coordinates of F2 are indicated with F1 as the coordinate origin). Through pitch, yaw, and roll determined by the above formulas (10) - (11), we can obtain:
[0174] x2 = cos(pitch)×sin(yaw) (13);
[0175] y2 = sin(pitch) (14);
[0176] z2 = cos(pitch)×cos(yaw) (15);
[0177] Combined with the coordinates of F1, the coordinates of F2 in the world coordinate system can be obtained as (x1 + x2, y1 + y2, z1 + z2).
[0178] It should be noted that the three-dimensional coordinates of the face center can be obtained by acquiring 68 feature points, determining the coordinate origin corresponding to the three-dimensional face model, and using the coordinates of this coordinate origin in the Figure 7 world coordinate system shown as the face center coordinates.
[0179] Please refer to Figure 8 the schematic diagram for calculating the intersection point of the face orientation and the display screen provided by the embodiment of the present invention. In Figure 8 , F is the intersection point of the face orientation and the display screen, F1 is the face center point, F2 is a point on the face orientation (which can also be understood as a vector in the coordinate system with F1 as the coordinate origin), G is a point perpendicular to the display screen from F1, and G2 is the point of F2 perpendicular to the line F1G.
[0180] Through the previous calculations, the screen equation of the display screen in the Figure 7 world coordinate system shown can be obtained as: 15.5×(y - sinα) + 92×(z - cosα) = 0. Let a = 0, b = 15.5, c = 92, and d = -93.3.
[0181] According to the similarity of triangles, it can be obtained that:
[0182]
[0183] Among them,
[0184] The distance from F1 to F2 is:
[0185]
[0186] Furthermore, it can be obtained that:
[0187] Regarding as the normal vector of the plane where the display screen is located, it can be obtained that:
[0188]
[0189] Then:
[0190]
[0191] Among them, O is the Figure 7 origin of the world coordinate system shown.
[0192] In the above manner, the coordinates of the intersection point of the face orientation and the display screen, the three-dimensional coordinates of the above-mentioned second intersection point 1 and second intersection point 2, and the three-dimensional coordinates of the intersection point of the face orientation and the display plane that need to be determined subsequently can all be determined according to the above manner.
[0193] When obtaining the first visible light face image of the current user captured by the image sensor, the three-dimensional coordinates of the second intersection point between the face orientation determined by the above method and the display screen can be used to determine whether the second intersection point is within the first fixation area of the display screen. Combining whether the determined second face pose is within the set error range can determine whether the captured second visible light face image is a face image captured in the forward direction. Furthermore, the second visible light face image captured in the forward direction is used as the first visible light face image.
[0194] After obtaining the first visible light face image, step 102 can be executed.
[0195] Step 102: Determine the three-dimensional coordinates of the left eye center, the three-dimensional coordinates of the right eye center, and the three-dimensional coordinates of the face center of the current user based on a preset number of feature points.
[0196] Determining the three-dimensional coordinates of the left eye center, the three-dimensional coordinates of the right eye center, and the three-dimensional coordinates of the face center of the current user based on a preset number of feature points can be achieved by the following method:
[0197] Convert the pixel coordinates of the preset number of feature points into three-dimensional coordinates in the world coordinate system; determine the feature points in the left eye area and the feature points in the right eye area from the preset number of feature points; determine the center points of the feature points in the left eye area and the feature points in the right eye area, and use the center point of the feature points in the left eye area as the three-dimensional coordinates of the left eye center, and use the center point of the feature points in the right eye area as the three-dimensional coordinates of the right eye center; according to the preset eyeball radius and the three-dimensional coordinates of the left eye center and the right eye center, determine the three-dimensional coordinates of the left eye center of the eyeball and the three-dimensional coordinates of the right eye center of the eyeball; according to the three-dimensional coordinates corresponding to the preset number of feature points, determine the three-dimensional coordinates of the face center.
[0198] First, the feature points belonging to the left eye area and the feature points belonging to the right eye area can be determined from the preset number of feature points (assumed to be 68 feature points). The pixel coordinates of the center point of the left eye area are determined according to the pixel coordinates of the feature points in the left eye area, and the pixel coordinates of the center point of the right eye area are determined according to the pixel coordinates of the feature points in the right eye area. Then, the pixel coordinates of the center point of the left eye area and the pixel coordinates of the center point of the right eye area are respectively converted into three-dimensional coordinates in the world coordinate system. Since this conversion method is a prior art, it will not be elaborated here.
[0199] For example, the feature points in the left eye area include 6 feature points L1 to L6, and the feature points in the right eye area include 6 feature points R1 to R6. Then the center point of the left eye area (denoted as L center ), the center point of the right eye area (denoted as R center ) can be expressed as:
[0200] L center=(L1 + L2 + L3 + L4 + L5 + L6) / 6 (22);
[0201] R center =(R1 + R2 + R3 + R4 + R5 + R6) / 6 (23);
[0202] It should be noted that formulas (22) and (23) represent the calculation of the pixel coordinates of the corresponding feature points.
[0203] By converting the pixel coordinates of the center point of the left eye region and the pixel coordinates of the center point of the right eye region into three-dimensional coordinates in the world coordinate system respectively, the three-dimensional coordinates of the left eye center and the three-dimensional coordinates of the right eye center can be obtained, which are denoted as L w (zl, yl, zl), R w (xr, yr, zr).
[0204] Then, according to the preset eyeball radius (usually 12 mm) and the three-dimensional coordinates of the left eye center and the three-dimensional coordinates of the right eye center, calculate the three-dimensional coordinates of the left eye eyeball center as L w (zl, yl, zl - 12), and the three-dimensional coordinates of the right eye eyeball center R w (xr, yr, zr - 12).
[0205] The three-dimensional coordinates of the face center can be calculated through a preset number of corresponding three-dimensional coordinates, which is the prior art and will not be elaborated here.
[0206] After determining the three-dimensional coordinates of the left eye eyeball center, the three-dimensional coordinates of the right eye eyeball center, and the three-dimensional coordinates of the face center, step 103 can be executed.
[0207] Step 103: Establish the left eye vector and the right eye vector from the three-dimensional coordinates of the left eye eyeball center and the three-dimensional coordinates of the right eye eyeball center to the three-dimensional coordinates of the face center respectively, and map them into the three-dimensional face model to obtain the rigid face model of the current user.
[0208] Using the three-dimensional coordinates of the left eye eyeball center, the three-dimensional coordinates of the right eye eyeball center, and the three-dimensional coordinates of the face center calculated in step 102, the left eye vector between the left eye eyeball center and the face center, and the right eye vector between the right eye eyeball center and the face center can be calculated:
[0209] Left eye vector: L w -F1 = (xl - x1, yl - y1, zl - z1) (24);
[0210] Right eye vector: R w -F1 = (xr - x1, yr - y1, zr - z1) (25);
[0211] Mapping the above left-eye vector and right-eye vector into the three-dimensional face model can obtain the rigid face model of the current user.
[0212] In the embodiment provided by the present invention, in order to improve the accuracy of the eyeballs in the rigid face model, after obtaining the rigid face model, the rigid face model can also be corrected. Specifically, the following method can be adopted:
[0213] Display a second fixation area at the center of the display screen, and obtain a third visible-light face image and a third infrared face image taken when the current user fixates on the second fixation area; determine the third face pose of the current user based on the third visible-light face image, and detect the center positions of the left and right eye pupils in the third infrared image to obtain the three-dimensional coordinates of the left-eye pupil center and the right-eye pupil center coordinates; wherein, the third face pose is used to represent the pose angle of the current user's face; update the pose of the rigid face model through the third face pose, and combine the three-dimensional coordinates of the left-eye pupil center and the right-eye pupil center coordinates to obtain the current fixation point of the current user; determine whether the current fixation point is within the second fixation area; if not within the second fixation area, move the current fixation point into the second fixation area, and finely adjust the left-eye vector and right-eye vector in the rigid face model according to the moving distance to obtain the corrected rigid face model.
[0214] Please refer to Figure 9 for the schematic diagram of correcting the rigid face model provided by the embodiment of the present invention.
[0215] Display a second fixation area at the center of the display screen ( Figure 9 the elliptical dotted area shown in Figure 9 is actually circular, but appears elliptical from the
[0216] angle in
[0217] The third infrared image is converted into a fourth infrared image in the same coordinate system as the third visible light image; the left eye area and the right eye area are obtained from the fourth infrared image, and the left eye pupil center pixel coordinates and the right eye pupil center pixel coordinates are obtained from the left eye area and the right eye area respectively using a pupil detection algorithm; the left eye pupil center pixel coordinates and the right eye pupil center pixel coordinates are converted into the left eye pupil center three-dimensional coordinates and the right eye pupil center three-dimensional coordinates in the world coordinate system respectively.
[0218] Since the resolution and FOV of the visible light image sensor and the infrared image sensor are different, and their optical axes are also different, the third infrared image needs to be converted into a fourth infrared image in the same coordinate system as the third visible light image. The formula for mapping the third visible light image to the world coordinate system and the formula for mapping the third infrared image to the world coordinate system can be solved simultaneously to obtain the formula for mutual conversion of coordinates between the third visible light image and the third infrared image:
[0219]
[0220] Wherein, (u1, ν1) is the pixel coordinate of a pixel in the third visible light image, (u2, ν2) is the pixel coordinate of a pixel in the third infrared image, (u 01 , ν 01 ) is the pixel coordinate of the center of the third visible light image, (u 02 , ν 02 ) is the pixel coordinate of the center of the third infrared image, (f x1 , f y1 ) is the focal coordinate of the visible light image sensor, (f x2 , f y2 ) is the focal coordinate of the infrared image sensor, (u1, ν1) and (u2, ν2) correspond to the same point in the world coordinate system, and the z-axis coordinate of this point is z w , k is the multiple of the lateral resolution of the infrared image sensor and the lateral resolution of the visible light infrared image sensor (taking the resolution set in step 101 as an example, k=4); d is the distance between the centers of the visible light image sensor and the infrared image sensor.
[0221] In order to reduce the amount of calculation, the infrared left eye area and the infrared right eye area can also be obtained from the third infrared image first, and then the above-mentioned coordinate transformation is performed on them to convert them into the left eye area and the right eye area in the same coordinate system as the third visible light image, and then the pupil detection algorithm (such as an algorithm based on a gradient vector field) is used to detect the pupils in the left eye area and the right eye area to obtain the pixel coordinates of the left eye pupil center and the pixel coordinates of the right eye pupil center, and they are converted into three-dimensional coordinates in the world coordinate system to obtain the three-dimensional coordinates of the left eye pupil center and the three-dimensional coordinates of the right eye pupil center.
[0222] The attitude of the human face rigid body model is updated through the third human face pose, and the current fixation point of the current user is obtained by combining the three-dimensional coordinates of the left eye pupil center and the coordinates of the right eye pupil center, which can be achieved in the following ways:
[0223] Adjust the rotation angles of the left eye vector and the right eye vector in the human face rigid body model according to the third human face pose to obtain the adjusted left eye vector and the adjusted right eye vector; obtain the current three-dimensional coordinates of the left eye ball center and the current three-dimensional coordinates of the right eye ball center from the adjusted left eye vector and the adjusted right eye vector; extend the line connecting the current three-dimensional coordinates of the left eye ball center and the three-dimensional coordinates of the left eye pupil as the left eye line of sight, and extend the line connecting the current three-dimensional coordinates of the right eye ball center and the three-dimensional coordinates of the right eye pupil as the right eye line of sight; determine the two intersection points of the left eye line of sight and the right eye line of sight with the display screen respectively, and take the center point of the line connecting the two intersection points as the current fixation point.
[0224] Assume that the human face pose in the human face rigid body model obtained in step 103 is (α0, β0, γ0), and α0, β0, γ0 are the pitch angle (pitch), yaw angle (yaw), and roll angle (roll) in sequence. In this human face rigid body model, the three-dimensional coordinates of the human face center are F(x, y, z), the three-dimensional coordinates of the left eye ball center are L(xl, yl, zl), and the three-dimensional coordinates of the right eye ball center are R(xr, yr, zr).
[0225] From the human face pose (α1, β1, γ1) obtained from the above third visible light human face image, and the human face pose (α0, β0, γ0), the three-dimensional coordinates of the human face center are F(x, y, z), the three-dimensional coordinates of the left eye ball center are L(xl, yl, zl), and the three-dimensional coordinates of the right eye ball center are R(xr, yr, zr) obtained in step 103, the current coordinates of the left eye ball center and the right eye ball center of the current user (corresponding to the third visible light image) can be calculated.
[0226] Taking the calculation of the left eye ball center coordinates as an example for illustration:
[0227] When the pitch angle changes, the angle by which the vector PL from the left eye ball center to the human face center rotates around the X-axis is α0 - α1, and the unit vector of the X-axis is XA(1, 0, 0). Therefore, the rotated vector PL_x is:
[0228] PL_x = PL × cos(α0 - α1) + (XA × PL)sin(α0 - α1) + XA(XA · PL)(1 - cos(α0 - α1)) (27);
[0230] Similarly, when the yaw angle changes, the angle by which the vector PL_x rotates around the Y-axis is β0 - β1, the unit vector of the Y-axis is YA(0, 1, 0), and the vector PL_xy obtained after rotation is:
[0231] PL_xy = PL_x × cos(β0 - β1) + (YA × PL_x)sin(β0 - β1) + YA(YA·PL_x)(1 - cos(β0 - β1)) (28);
[0233] Similarly, when the roll angle changes, the vector PL_xy rotates around the Z-axis to obtain the vector PL_xyz:
[0234] PL_xyz = PL_xy × cos(γ0 - γ1) + (ZA × PL_xy)sin(γ0 - γ1) + ZA(ZA·PL_xy)(1 - cos(γ0 - γ1 ))(29);
[0236] Using the same method as calculating the left eye eyeball center coordinates, the right eye eyeball center coordinates can be calculated, denoted as PR_xyz. Since both PL_xyz and PR_xyz are vectors starting from the face center, by adding the three-dimensional coordinates F(x, y, z) of the face center to PL_xyz and PR_xyz respectively, the three-dimensional coordinates of the left eye eyeball center and the right eye eyeball center in the world coordinate system corresponding to the current (corresponding to the third visible light face image) can be obtained.
[0237] After that, the extension line of the connection between the current left eye eyeball center three-dimensional coordinates and the left eye pupil three-dimensional coordinates is used as the left eye line of sight, and the extension line of the connection between the current right eye eyeball center three-dimensional coordinates and the right eye pupil three-dimensional coordinates is used as the right eye line of sight; the two intersection points of the left eye line of sight and the right eye line of sight with the display screen are determined, and the center point of the connection between the two intersection points is used as the current fixation point.
[0238] To ensure that the used face rigid body model is within the error range, before calculating the current left eye eyeball center three-dimensional coordinates and the current right eye eyeball center three-dimensional coordinates, it is also necessary to update the pose of the face rigid body model through the third face pose, and determine whether the distances from the left eye eyeball center and the right eye eyeball center to the face center remain unchanged.
[0239] The distance from the left eye eyeball center to the face center (denoted as D L ) is:
[0240]
[0241] The distance from the right eye eyeball center to the face center (denoted as D R ) is:
[0242]
[0243] D before and after updating the pose of the human face rigid body model R and D L If the values remain unchanged, it indicates that the human face rigid body model is within the error range, and the updated pose of the human face rigid body model can be used to calculate the three-dimensional coordinates of the current left eye center of the eyeball and the three-dimensional coordinates of the current right eye center of the eyeball. Otherwise, a new human face rigid body model needs to be regenerated to calculate the three-dimensional coordinates of the current left eye center of the eyeball and the three-dimensional coordinates of the current right eye center of the eyeball.
[0244] It should be noted that the image sensors (infrared image sensor and visible light image sensor) can be independent of the display screen or integrated into the display screen, and no specific limitation is made.
[0245] Please refer to Figure 10 In an embodiment of the present invention, a method for gaze point detection is provided. The specific implementation of the gaze point detection method can be referred to the description of the relevant embodiment part in the human face rigid body model generation method, and the repeated parts will not be elaborated. The method includes:
[0246] Step 1001: Obtain the fourth visible light human face image and the fourth infrared human face image of the current user, and obtain a preset number of feature points of the current user's human face from the fourth visible light human face image, and obtain the current left eye pupil center pixel coordinates and the current right eye pupil center pixel coordinates from the fourth infrared human face image.
[0247] Step 1002: Calculate the preset number of feature points by using the pose measurement method based on PnP to obtain the current human face pose of the current user.
[0248] Step 1003: Update the human face rigid body model as described in steps 101 to 103 based on the current human face pose to obtain the three-dimensional coordinates of the current left eye center of the eyeball and the three-dimensional coordinates of the right eye center; convert the current left eye pupil center pixel coordinates and the current right eye pupil center pixel coordinates into three-dimensional coordinates in the world coordinate system to obtain the three-dimensional coordinates of the current left eye pupil center and the three-dimensional coordinates of the current right eye pupil center.
[0249] Step 1004: Determine the current gaze point of the current user on the display screen according to the three-dimensional coordinates of the current left eye center of the eyeball, the three-dimensional coordinates of the right eye center, the three-dimensional coordinates of the current left eye pupil center, and the three-dimensional coordinates of the current right eye pupil center.
[0250] Since the above method for determining the current gaze point is the same as the method for determining the current gaze point recorded in step 103, it will not be elaborated here.
[0251] A possible implementation manner, obtaining the current left eye pupil center pixel coordinates and the current right eye pupil center pixel coordinates from the fourth infrared face image, includes:
[0252] Converting the fourth infrared image into a fifth infrared image in the same coordinate system as the fourth visible light image;
[0253] Obtaining the left eye region and the right eye region from the fifth infrared image, and respectively obtaining the left eye pupil center pixel coordinates and the right eye pupil center pixel coordinates from the left eye region and the right eye region by using a pupil detection algorithm.
[0254] A possible implementation manner, updating the face rigid body model as described in steps 101 to 103 based on the current face pose, and obtaining the current three-dimensional coordinates of the left eye center of the eyeball and the right eye center of the eyeball, includes:
[0255] Adjusting the rotation angles of the left eye vector and the right eye vector in the face rigid body model according to the current face pose to obtain an adjusted left eye vector and an adjusted right eye vector;
[0256] Obtaining the current three-dimensional coordinates of the left eye center of the eyeball and the current three-dimensional coordinates of the right eye center of the eyeball from the adjusted left eye vector and the adjusted right eye vector.
[0257] A possible implementation manner, determining the current fixation point of the current user on the display screen according to the current three-dimensional coordinates of the left eye center of the eyeball, the three-dimensional coordinates of the right eye center of the eyeball, the current three-dimensional coordinates of the left eye pupil center, and the current three-dimensional coordinates of the right eye pupil center, includes:
[0258] Taking the extension line of the connection between the current three-dimensional coordinates of the left eye center of the eyeball and the three-dimensional coordinates of the left eye pupil as the left eye line of sight;
[0259] Taking the extension line of the connection between the current three-dimensional coordinates of the right eye center of the eyeball and the three-dimensional coordinates of the right eye pupil as the right eye line of sight;
[0260] Determining two intersection points of the left eye line of sight and the right eye line of sight with the display screen respectively, and taking the center point of the connection of the two intersection points as the current fixation point.
[0261] Based on the same inventive concept, an embodiment of the present invention provides a device for generating a face rigid body model. For the specific implementation manner of the face rigid body model generation method of this device, reference can be made to the description in the embodiment part of the face rigid body model generation method. Repeated parts will not be elaborated again. Please refer to Figure 11 , and this device includes:
[0262] An acquisition module 1101, configured to acquire a first visible light face image of the current user captured by the image sensor in the forward direction, and acquire a preset number of feature points of the face of the current user from the first visible light face image;
[0263] A determination module 1102, configured to determine a three-dimensional coordinate of the center of the left eyeball, a three-dimensional coordinate of the center of the right eyeball, and a three-dimensional coordinate of the center of the face of the current user based on the preset number of feature points;
[0264] An obtaining module 1103, configured to respectively establish a left eye vector and a right eye vector from the three-dimensional coordinate of the center of the left eyeball and the three-dimensional coordinate of the center of the right eyeball to the three-dimensional coordinate of the center of the face, and map them into a three-dimensional face model to obtain a face rigid body model of the current user.
[0265] A possible implementation manner, the acquisition module 1101 is further configured to:
[0266] Display a first fixation area at the position closest to the imaging device on the display screen;
[0267] Continuously capture a second visible light face image of the current user when the current user is fixating on the first fixation area until it is determined that the face image of the current user is captured in the forward direction;
[0268] Each time a capture is performed, determine a second intersection point between the face orientation of the current user and the display screen, and a second face pose from the second visible light face image; wherein, the second face pose is used to represent the pose angle of the face of the current user;
[0269] Determine whether the second intersection point is within the first fixation area and whether the second face pose is within a set error range;
[0270] If the second intersection point is within the first fixation area and the second face pose is within the set error range, determine that the face image of the current user is captured in the forward direction, and use the corresponding second visible light face image as the first visible light face image.
[0271] A possible implementation manner, the acquisition module 1101 is further configured to:
[0272] Abstract the error value of the second face pose as a line segment starting from the center of the first fixation area, and use the boundary of the first fixation area as the boundary value of the error range;
[0273] Determine whether the line segment exceeds the first fixation area;
[0274] If the line segment exceeds the first fixation area, determine that the second face pose is not within the error range;
[0275] If the line segment does not exceed the first fixation area, it is determined that the second face pose is within the error range.
[0276] In a possible implementation, the determining module 1102 is further configured to:
[0277] Convert the pixel coordinates corresponding to the preset number of feature points into three-dimensional coordinates in the world coordinate system;
[0278] Determine the left-eye region feature points and the right-eye region feature points from the preset number of feature points;
[0279] Determine the center points of the left-eye region feature points and the right-eye region feature points, and use the center point of the left-eye region feature points as the three-dimensional coordinates of the left-eye center, and use the center point of the right-eye region feature points as the three-dimensional coordinates of the right-eye center;
[0280] According to the preset eyeball radius and the three-dimensional coordinates of the left-eye center and the three-dimensional coordinates of the right-eye center, determine the three-dimensional coordinates of the left-eye eyeball center and the three-dimensional coordinates of the right-eye eyeball center;
[0281] According to the three-dimensional coordinates corresponding to the preset number of feature points, determine the three-dimensional coordinates of the face center.
[0282] In a possible implementation, after obtaining the face rigid body model of the current user, the obtaining module 1103 is further configured to:
[0283] Display a second fixation area at the center of the display screen, and obtain a third visible-light face image and a third infrared face image captured when the current user fixates on the second fixation area;
[0284] Based on the third visible-light face image, determine the third face pose of the current user, and detect the center positions of the left-eye pupil and the right-eye pupil in the third infrared image to obtain the three-dimensional coordinates of the left-eye pupil center and the coordinates of the right-eye pupil center; wherein, the third face pose is used to represent the pose angle of the current user's face;
[0285] Update the pose of the face rigid body model through the third face pose, and combine the three-dimensional coordinates of the left-eye pupil center and the coordinates of the right-eye pupil center to obtain the current fixation point of the current user;
[0286] Determine whether the current fixation point is within the second fixation area;
[0287] If it is not within the second fixation area, move the current fixation point into the second fixation area, and finely adjust the left-eye vector and the right-eye vector in the face rigid body model according to the moving distance to obtain a corrected face rigid body model.
[0288] A possible implementation, the obtaining module 1103 is further configured to:
[0289] Convert the third infrared image into a fourth infrared image in the same coordinate system as the third visible light image;
[0290] Obtain the left eye region and the right eye region from the fourth infrared image, and use a pupil detection algorithm to obtain the left eye pupil center pixel coordinates and the right eye pupil center pixel coordinates from the left eye region and the right eye region respectively;
[0291] Convert the left eye pupil center pixel coordinates and the right eye pupil center pixel coordinates into the left eye pupil center three-dimensional coordinates and the right eye pupil center three-dimensional coordinates in the world coordinate system respectively.
[0292] A possible implementation, the obtaining module 1103 is further configured to:
[0293] Adjust the rotation angles of the left eye vector and the right eye vector in the human face rigid body model according to the third human face pose to obtain an adjusted left eye vector and an adjusted right eye vector;
[0294] Obtain the current left eye eyeball center three-dimensional coordinates and the current right eye eyeball center three-dimensional coordinates from the adjusted left eye vector and the adjusted right eye vector;
[0295] Use the extension line of the connection between the current left eye eyeball center three-dimensional coordinates and the left eye pupil three-dimensional coordinates as the left eye line of sight, and use the extension line of the connection between the current right eye eyeball center three-dimensional coordinates and the right eye pupil three-dimensional coordinates as the right eye line of sight;
[0296] Determine the two intersection points of the left eye line of sight and the right eye line of sight with the display screen respectively, and use the center point of the connection of the two intersection points as the current fixation point.
[0297] Based on the same inventive concept, an embodiment of the present invention provides a fixation point detection device. For the specific implementation of the fixation point detection method of this device, reference can be made to the description in the embodiment part of the fixation point detection method. Repeated parts will not be elaborated. Please refer to Figure 12 , this device includes:
[0298] An acquisition module 1201, configured to acquire a fourth visible light human face image and a fourth infrared human face image of the current user, acquire a preset number of feature points of the current user's human face from the fourth visible light human face image, and acquire the current left eye pupil center pixel coordinates and the current right eye pupil center pixel coordinates from the fourth infrared human face image;
[0299] A calculation module 1202, configured to calculate the preset number of feature points by using a PnP-based pose measurement method to obtain the current face pose of the current user;
[0300] An obtaining module 1203, configured to update the face rigid body model according to any one of claims 1-7 based on the current face pose, to obtain the three-dimensional coordinates of the current left eye center of the eyeball and the three-dimensional coordinates of the right eye center of the eyeball; to convert the current left eye pupil center pixel coordinates and the current right eye pupil center pixel coordinates into three-dimensional coordinates in the world coordinate system, to obtain the three-dimensional coordinates of the current left eye pupil center and the three-dimensional coordinates of the current right eye pupil center;
[0301] A determination module 1204, configured to determine the current fixation point of the current user on the display screen according to the three-dimensional coordinates of the current left eye center of the eyeball, the three-dimensional coordinates of the right eye center of the eyeball, the three-dimensional coordinates of the current left eye pupil center, and the three-dimensional coordinates of the current right eye pupil center.
[0302] A possible implementation manner, the obtaining module 1201 is further configured to:
[0303] Convert the fourth infrared image into a fifth infrared image in the same coordinate system as the fourth visible light image;
[0304] Obtain the left eye area and the right eye area from the fifth infrared image, and respectively obtain the left eye pupil center pixel coordinates and the right eye pupil center pixel coordinates from the left eye area and the right eye area by using a pupil detection algorithm.
[0305] A possible implementation manner, the obtaining module 1203 is further configured to:
[0306] Adjust the rotation angles of the left eye vector and the right eye vector in the face rigid body model according to the current face pose to obtain an adjusted left eye vector and an adjusted right eye vector;
[0307] Obtain the three-dimensional coordinates of the current left eye center of the eyeball and the three-dimensional coordinates of the current right eye center of the eyeball from the adjusted left eye vector and the adjusted right eye vector.
[0308] A possible implementation manner, the determination module 1204 is specifically configured to:
[0309] Use the extension line of the connection line between the three-dimensional coordinates of the current left eye center of the eyeball and the three-dimensional coordinates of the left eye pupil as the left eye line of sight;
[0310] Use the extension line of the connection line between the three-dimensional coordinates of the current right eye center of the eyeball and the three-dimensional coordinates of the right eye pupil as the right eye line of sight;
[0311] Determine two intersection points of the left-eye line of sight and the right-eye line of sight with the display screen respectively, and take the center point of the line connecting the two intersection points as the current fixation point.
[0312] Based on the same inventive concept, an apparatus for fixation point detection is provided in an embodiment of the present invention, including: at least one processor, and
[0313] a memory connected to the at least one processor;
[0314] wherein, the memory stores instructions executable by the at least one processor, and the at least one processor executes the methods for generating a human face rigid body model or fixation point detection as described above by executing the instructions stored in the memory.
[0315] Based on the same inventive concept, an embodiment of the present invention also provides a readable storage medium, including:
[0316] a memory,
[0317] the memory is used to store instructions, and when the instructions are executed by a processor, it enables the device including the readable storage medium to complete the methods for generating a human face rigid body model or fixation point detection as described above.
[0318] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the embodiments of the present invention can take the form of an all-hardware embodiment, an all-software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0319] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0320] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instruction means embodying the functionality specified in the flowchart(s) Figure 1 a flowchart or flowcharts and / or block(s) Figure 1 a block or blocks.
[0321] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functionality specified in the flowchart(s) Figure 1 a flowchart or flowcharts and / or block(s) Figure 1 a block or blocks.
[0322] It will be apparent to those skilled in the art that various modifications and variations can be made to the present invention without departing from the spirit and scope of the invention. Thus, if these modifications and variations of the present invention come within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. A method for generating a human face rigid body model, characterized in that, Including: Display a first fixation area at the position on the display screen closest to the shooting device; Continuously capture a second visible-light face image of the current user when the current user is fixating on the first fixation area until it is determined that the face image of the current user is captured head-on; Each time a capture is made, determine a second intersection point between the face orientation of the current user and the display screen, and a second face pose based on the second visible-light face image; wherein, the second face pose is used to represent the pose angle of the current user's face; Judge whether the second intersection point is within the first fixation area and whether the second face pose is within a set error range; If the second intersection point is within the first fixation area and the second face pose is within the set error range, determine that the face image of the current user is captured head-on, and use the corresponding second visible-light face image as the first visible-light face image; Obtain a preset number of feature points of the current user's face from the first visible-light face image; Based on the preset number of feature points, determine the three-dimensional coordinates of the center of the left eyeball, the three-dimensional coordinates of the center of the right eyeball, and the three-dimensional coordinates of the center of the face of the current user; Respectively establish a left eye vector and a right eye vector from the three-dimensional coordinates of the center of the left eyeball and the three-dimensional coordinates of the center of the right eyeball to the three-dimensional coordinates of the center of the face, and map them into the three-dimensional face model to obtain the rigid face model of the current user.
2. The method according to claim 1, wherein Judging whether the second face pose is within the set error range includes: Abstract the error value of the second face pose as a line segment starting from the center of the first fixation area, and use the boundary of the first fixation area as the boundary value of the error range; Judge whether the line segment exceeds the first fixation area; If the line segment exceeds the first fixation area, determine that the second face pose is not within the error range; If the line segment does not exceed the first fixation area, determine that the second face pose is within the error range.
3. The method according to claim 1, characterized in that Based on the preset number of feature points, determining the three-dimensional coordinates of the center of the left eyeball, the three-dimensional coordinates of the center of the right eyeball, and the three-dimensional coordinates of the center of the face of the current user includes: Convert the pixel coordinates corresponding to the preset number of feature points into three-dimensional coordinates in the world coordinate system; Determine the feature points in the left eye area and the feature points in the right eye area from the preset number of feature points; Determine the center points of the feature points in the left eye area and the feature points in the right eye area, and use the center point of the feature points in the left eye area as the three-dimensional coordinates of the center of the left eye, and use the center point of the feature points in the right eye area as the three-dimensional coordinates of the center of the right eye; According to the preset eyeball radius and the three-dimensional coordinates of the center of the left eye and the three-dimensional coordinates of the center of the right eye, determine the three-dimensional coordinates of the center of the left eyeball and the three-dimensional coordinates of the center of the right eyeball; According to the three-dimensional coordinates corresponding to the preset number of feature points, determine the three-dimensional coordinates of the center of the face.
4. The method according to claim 1, characterized in that, After obtaining the rigid face model of the current user, it further includes: Display a second fixation area at the center of the display screen, and obtain a third visible-light face image and a third infrared face image captured when the current user is fixating on the second fixation area; Determine the third face pose of the current user based on the third visible light face image, and detect the central positions of the left eye pupil and the right eye pupil in the third infrared image to obtain the three-dimensional coordinates of the left eye pupil center and the coordinates of the right eye pupil center; wherein, the third face pose is used to represent the pose angle of the current user's face. Update the pose of the face rigid body model through the third face pose, and combine the three-dimensional coordinates of the left eye pupil center and the coordinates of the right eye pupil center to obtain the current fixation point of the current user. Determine whether the current fixation point is within the second fixation area. If it is not within the second fixation area, move the current fixation point into the second fixation area, and fine-tune the left eye vector and the right eye vector in the face rigid body model according to the moving distance to obtain a corrected face rigid body model.
5. The method according to claim 4, wherein Detect the central positions of the left eye pupil and the right eye pupil in the third infrared image to obtain the three-dimensional coordinates of the left eye pupil center and the coordinates of the right eye pupil center, including: Convert the third infrared image into a fourth infrared image in the same coordinate system as the third visible light image. Obtain the left eye area and the right eye area from the fourth infrared image, and use a pupil detection algorithm to obtain the pixel coordinates of the left eye pupil center and the pixel coordinates of the right eye pupil center from the left eye area and the right eye area respectively. Convert the pixel coordinates of the left eye pupil center and the pixel coordinates of the right eye pupil center into the three-dimensional coordinates of the left eye pupil center and the three-dimensional coordinates of the right eye pupil center in the world coordinate system respectively.
6. The method according to claim 5, characterized in that Update the pose of the face rigid body model through the third face pose, and combine the three-dimensional coordinates of the left eye pupil center and the coordinates of the right eye pupil center to obtain the current fixation point of the current user, including: Adjust the rotation angles of the left eye vector and the right eye vector in the face rigid body model according to the third face pose to obtain an adjusted left eye vector and an adjusted right eye vector. Obtain the three-dimensional coordinates of the current left eye ball center and the three-dimensional coordinates of the current right eye ball center from the adjusted left eye vector and the adjusted right eye vector. Use the extension line of the connection between the three-dimensional coordinates of the current left eye ball center and the three-dimensional coordinates of the left eye pupil as the left eye line of sight, and use the extension line of the connection between the three-dimensional coordinates of the current right eye ball center and the three-dimensional coordinates of the right eye pupil as the right eye line of sight. Determine the two intersection points of the left eye line of sight and the right eye line of sight with the display screen respectively, and use the center point of the connection between the two intersection points as the current fixation point.
7. A method for fixation point detection, characterized in that, including: Obtain the fourth visible light face image and the fourth infrared face image of the current user, obtain a preset number of feature points of the current user's face from the fourth visible light face image, and obtain the pixel coordinates of the current left eye pupil center and the pixel coordinates of the current right eye pupil center from the fourth infrared face image. Calculate the preset number of feature points using a pose measurement method based on PnP to obtain the current face pose of the current user. Update the human face rigid body model as described in any one of claims 1-6 based on the current human face pose to obtain the three-dimensional coordinates of the current left eye center of the eyeball and the three-dimensional coordinates of the right eye center of the eyeball; Convert the current left eye pupil center pixel coordinates and the current right eye pupil center pixel coordinates into three-dimensional coordinates in the world coordinate system to obtain the three-dimensional coordinates of the current left eye pupil center and the three-dimensional coordinates of the current right eye pupil center; Determine the current fixation point of the current user on the display screen according to the three-dimensional coordinates of the current left eye center of the eyeball, the three-dimensional coordinates of the right eye center of the eyeball, the three-dimensional coordinates of the current left eye pupil center and the three-dimensional coordinates of the current right eye pupil center.
8. The method according to claim 7, wherein Obtain the current left eye pupil center pixel coordinates and the current right eye pupil center pixel coordinates from the fourth infrared human face image, including: Convert the fourth infrared image into a fifth infrared image in the same coordinate system as the fourth visible light image; Obtain the left eye region and the right eye region from the fifth infrared image, and use a pupil detection algorithm to obtain the left eye pupil center pixel coordinates and the right eye pupil center pixel coordinates from the left eye region and the right eye region respectively.
9. The method according to claim 7, wherein Update the human face rigid body model as described in any one of claims 1-6 based on the current human face pose to obtain the three-dimensional coordinates of the current left eye center of the eyeball and the three-dimensional coordinates of the right eye center of the eyeball, including: Adjust the rotation angles of the left eye vector and the right eye vector in the human face rigid body model according to the current human face pose to obtain an adjusted left eye vector and an adjusted right eye vector; Obtain the three-dimensional coordinates of the current left eye center of the eyeball and the three-dimensional coordinates of the current right eye center of the eyeball from the adjusted left eye vector and the adjusted right eye vector.
10. The method according to claim 7, wherein Determine the current fixation point of the current user on the display screen according to the three-dimensional coordinates of the current left eye center of the eyeball, the three-dimensional coordinates of the right eye center of the eyeball, the three-dimensional coordinates of the current left eye pupil center and the three-dimensional coordinates of the current right eye pupil center, including: Use the extension line of the connection between the three-dimensional coordinates of the current left eye center of the eyeball and the three-dimensional coordinates of the left eye pupil as the left eye line of sight; Use the extension line of the connection between the three-dimensional coordinates of the current right eye center of the eyeball and the three-dimensional coordinates of the right eye pupil as the right eye line of sight; Determine the two intersection points of the left eye line of sight and the right eye line of sight with the display screen respectively, and use the center point of the connection of the two intersection points as the current fixation point.
11. An apparatus for generating a human face rigid body model, characterized in that, Include: An acquisition module, configured to display a first fixation area at a position closest to the imaging device in the display screen; continuously capture a second visible-light face image of the current user when the current user is fixating on the first fixation area with an image sensor until it is determined that a face image of the current user is captured head-on; each time an image is captured, determine a second intersection point between the face orientation of the current user and the display screen, and a second face pose according to the second visible-light face image; wherein the second face pose is used to represent the pose angle of the current user's face; determine whether the second intersection point is within the first fixation area and whether the second face pose is within a set error range; if the second intersection point is within the first fixation area and the second face pose is within the set error range, determine that a face image of the current user is captured head-on, and use the corresponding second visible-light face image as a first visible-light face image; and obtain a preset number of feature points of the current user's face from the first visible-light face image. A determination module, configured to determine three-dimensional coordinates of the center of the left eyeball, three-dimensional coordinates of the center of the right eyeball, and three-dimensional coordinates of the center of the face of the current user based on the preset number of feature points. An obtaining module, configured to respectively establish a left-eye vector and a right-eye vector from the three-dimensional coordinates of the center of the left eyeball and the three-dimensional coordinates of the center of the right eyeball to the three-dimensional coordinates of the center of the face, and map them into a three-dimensional face model to obtain a rigid body model of the face of the current user.
12. A device for fixation point detection, characterized in that, including: An acquisition module, configured to acquire a fourth visible-light face image and a fourth infrared face image of the current user, obtain a preset number of feature points of the current user's face from the fourth visible-light face image, and obtain current pixel coordinates of the center of the left pupil and current pixel coordinates of the center of the right pupil from the fourth infrared face image. A calculation module, configured to calculate the preset number of feature points by using a pose measurement method based on PnPd to obtain the current face pose of the current user. An obtaining module, configured to update the rigid body model of the face according to any one of claims 1-6 based on the current face pose to obtain three-dimensional coordinates of the center of the current left eyeball and three-dimensional coordinates of the center of the right eyeball. Convert the current pixel coordinates of the center of the left pupil and the current pixel coordinates of the center of the right pupil into three-dimensional coordinates in the world coordinate system to obtain three-dimensional coordinates of the center of the current left pupil and three-dimensional coordinates of the center of the current right pupil. A determination module, configured to determine the current fixation point of the current user on the display screen according to the three-dimensional coordinates of the center of the current left eyeball, the three-dimensional coordinates of the center of the right eyeball, the three-dimensional coordinates of the center of the current left pupil, and the three-dimensional coordinates of the center of the current right pupil.
13. A device for detecting a fixation point, characterized in that, including: at least one processor, and a memory connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the at least one processor executes the instructions stored in the memory to execute the method according to any one of claims 1-10.
14. A readable storage medium, characterized in that, including a memory, The memory is used to store instructions, which, when executed by a processor, cause the device including the readable storage medium to perform the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Sight line calibration method, playing method of display device and sight line calibration system
CN110458122A
Electronic book reader control method and system based on human eye fixation point detection
CN110531853A
Eye gaze tracking method and apparatus and computer-readable recording medium
US20150293588A1