Man-machine interaction sight tracking method

The image data of the interactive object is obtained through the light source device and the camera device, and the fitting is combined with the pupil spot position and the head posture rotation matrix to solve the problem of insufficient line of sight tracing accuracy in the prior art, achieving a higher precision line of sight tracing effect.

CN119992634APending Publication Date: 2025-05-13SUZHOU CHANGFENG AVIATION ELECTRONICS
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411824549.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-12
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Existing gaze tracing techniques have shortcomings in accuracy, especially when dealing with head pose changes of interactive objects, it is difficult to accurately capture the gaze position at the edge of the display area.

Method used

By configuring the light source device and the camera device, the image data of the interactive object is obtained, the reference coordinate system is constructed, the eye area is intercepted and enlarged, the pupil and spot positions are determined, the head posture is represented, and the three-dimensional rotation matrix is ​​obtained, and the pupil spot position difference value is combined with the head posture rotation matrix to accurately locate the gaze position of the interactive object.

Benefits of technology

The accuracy of gaze tracing is significantly improved, especially the positioning accuracy of the edge position of the display area, and can more accurately capture the gaze position of the interactive object.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005183913350000052
    Figure BDA0005183913350000052
  • Figure BDA0005183913350000061
    Figure BDA0005183913350000061
  • Figure BDA0005183913350000073
    Figure BDA0005183913350000073
Patent Text Reader

Abstract

The invention provides a man-machine interaction sight tracking method, which comprises the following steps of: intercepting and amplifying eye positions in a whole image, training a Yolov5 algorithm by adopting a data set to mark pupil positions and light spot positions from the image, and determining the position of a pupil according to a central position coordinate and a height value and a width value of a detection frame. Marking the pupil position and the light spot position; then, according to a better embodiment of the invention, a 3 * 3 rotation matrix is adopted, the head posture of the interaction object is marked, and the rotation angle of the head posture is obtained; finally, in the mapping process of the coordinates of the display area, the rotation angle of the head posture and the difference value of the pupil light spot position are fitted, so that the tracking precision is remarkably improved, and especially the positioning accuracy of the edge position of the display area is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of human-computer interaction, and specifically to a method for human line of sight tracking in the field of human-computer interaction. Background Art

[0002] In the field of human-computer interaction, compared with the more common and mature voice recognition and gesture recognition solutions, gaze tracking is still a special technology that needs to be explored. In terms of application, in addition to infrared pupil identification and human-computer interaction for special groups, there is an increasing demand for gaze tracking technology in some application scenarios involving safety, sterility or high professionalism.

[0003] The purpose of eye tracking is to capture the trajectory of the interactive object's eye movement on the display device and the position where the eye stays and looks. To this end, the mainstream approach in the prior art is to configure at least one infrared light source and at least one camera device on the display device, and the infrared light emitted by the light source is irradiated on the pupil, cornea and other positions of the interactive object to immediately generate light spots, and then the camera device is selected to capture the position of these light spots, and then the known pupil-corneal reflection method is used for mapping and fitting to obtain the interactive object's focus position on the display device. Although, in fact, the position where the interactive object looks is not a single point, but often a specific area, or a specific object in a certain area, in the above-mentioned existing methods, the mapping and fitting basis of the pupil-corneal reflection, or the basis for describing the change of the interactive object's gaze position, all rely on the coordinateization of a point on the screen, that is, taking the center of the display area of ​​the display device as the origin, constructing a coordinate system according to the display pixel points, or other applicable methods, so that the position where the interactive object looks can be corresponded to a point in the coordinate system by the center of the position, and represented by specific coordinates. When the gaze position changes, the spot position of the infrared light source on the pupil and cornea also shifts accordingly. Therefore, based on the displacement of the spot position, the point coordinates corresponding to the new gaze position of the interactive object in the display area are fitted and calculated, ultimately achieving target line of sight tracking.

[0004] However, the accuracy of existing methods for eye tracking is still very limited. One of the manifestations is that, for example, in related products such as eye trackers, the tracking results of the interactive object's eye line always show dynamic changes within a certain range of the display area. The main reasons include the following two aspects:

[0005] 1) The surface of the cornea at the front of the human eyeball is spherical, so the light spot formed on the pupil and iris after the infrared light source passes through the cornea does not converge into a point, but appears as a small light band. It should be understood that the position and shape of the light band will further deform with the rotation of the eyeball. For example, when the line of sight is at the center of the display area, since the pupil of the interactive object is facing the infrared light source at this time, the generated light spot will overlap with the pupil; and when the line of sight falls on the edge of the display area, the light spot will be stretched and deformed, and appear blurred locally or as a whole. In this way, for extreme line of sight points such as the aforementioned, or when dealing with the deformation of the light spot during the change of line of sight points, the result of line of sight capture by existing algorithms can often only be a dynamically changing line of sight range;

[0006] 2) When the line of sight shifts, in addition to the rotation of the human eyeball, the head posture of the interactive object often changes, especially when observing large-sized display devices at close range. Without the help of auxiliary equipment, the focus on the edge area of ​​the display device will inevitably rely on the rotation of the head. The known tracking algorithms are limited by themselves. In the process of extracting features from the pupil and the light spot on the cornea, the factors of head posture are not and cannot be taken into account.

[0007] Based on this, an improved method should be provided to correct the influence of head posture swing on the tracking results during line of sight tracking, so as to solve the technical problems existing in the above-mentioned existing algorithms. Summary of the invention

[0008] In view of this, the present invention provides a human-computer interaction sight tracking method to solve at least one of the above problems.

[0009] In order to solve the above technical problems, the steps of the human-computer interaction eye tracking method provided by the present invention are: the steps of configuring at least one light source device and one camera device in the display area, the light source device generates illumination when working, and the camera device photographs the interactive object under the illumination to obtain image data of the interactive object; constructing a first reference coordinate system with two mutually perpendicular directions in the image data, and according to the strategy of zooming in and then intercepting or intercepting and then zooming in, selecting the eye area of ​​the interactive object from the image data, and determining the position coordinates of the pupil and the light spot under the eye area relative to the first reference coordinate system; The head posture of the object is represented, and a three-dimensional rotation matrix representing the head posture angle is obtained; a plurality of sampling points covering the display area are configured, each sampling point samples the head posture of the interactive object at least once, and then a position difference between a pupil position and a light spot position is obtained, and fitting parameters between the pupil position, the light spot position, the three-dimensional rotation matrix, and the coordinates of the gaze position of the interactive object in the display area are obtained according to the position difference and the three-dimensional rotation matrix; according to the fitting parameters, the pupil position and the light spot position of the current interactive object are fitted with the three-dimensional rotation matrix to obtain the coordinates of the gaze position of the interactive object in the display area.

[0010] Compared with the existing methods and algorithm models, in the preferred embodiment of the present invention, the eye position in the overall image is firstly intercepted and enlarged, and the Yolov5 algorithm is trained with a data set so that it can mark the pupil position and the light spot position from the image, and mark the pupil position and the light spot position by the center position coordinates and the detection frame height value and width value; then, the preferred embodiment of the present invention uses a 3×3 rotation matrix to mark the head posture of the interactive object and obtain the rotation angle of the head posture; finally, in the process of mapping the display area coordinates, the rotation angle of the head posture is fitted with the difference between the pupil light spot position to significantly improve the accuracy of tracking, especially to improve the accuracy of positioning the edge position of the display area. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 is a schematic diagram showing the state of annotating the eye position in the overall image of the interactive object;

[0012] Figure 2 This is a partial enlarged view showing Figure 1 The middle frame selected part is enlarged and processed to mark the status of the pupil and the light spot;

[0013] Figure 3 It is a partial enlarged diagram, showing the state of the pupil position and the light spot position in the image being marked and detected by the updated algorithm model according to the preferred embodiment of the present invention;

[0014] Figure 4is a state diagram showing the result of posture calculation after collecting two sets of pictures in each orientation;

[0015] Figure 5 Schematic diagram showing the coordinate positions output by the algorithm model in a specific embodiment. DETAILED DESCRIPTION

[0016] Based on the above factors, in order to improve the accuracy of the gaze tracking algorithm, adding head posture correction to the existing algorithm is an easy-to-think-of improvement idea. However, this improvement idea is still subject to many limitations.

[0017] The capture and calculation of coordinates by the existing algorithm is based on the analysis of the information of the change of the light spot position vector before and after the sight line position is transferred. Among them, after the camera device obtains the image of the head area of ​​the interactive object, the pupil and the light spot position are analyzed by the algorithm respectively. However, for the overall image including the head of the interactive object, the pupil and the light spot near the pupil position account for a very small proportion of the overall image. If the head posture factor is further superimposed on this basis, many experimental and practical results have proved that the pupil and light spot positions cannot even be captured after superimposing the head posture. Therefore, although the existing algorithm can be optimized by introducing the step of training after marking the head posture so that it can take into account the change of head posture, based on the above reasons, this will cause a large detection error in the tracking result. In expectation, the preferred embodiment of the present invention finally attempts to establish a set of algorithms to fit the pupil position, the light spot position and the head posture position, and seek the configuration parameters between each position, but in view of the aforementioned technical bottleneck, the pupil and light spot capture method should be corrected first to further solve the problem of head posture affecting sight tracking.

[0018] Embodiments of the present invention will be described below with reference to the accompanying drawings. It will be appreciated by those skilled in the art that the described embodiments may be modified in various ways without departing from the spirit and scope of the present invention. Therefore, the drawings and description are illustrative in nature and are not intended to limit the scope of protection of the claims. In addition, in this specification, the drawings are not drawn to scale and the same reference numerals represent the same parts.

[0019] It should be noted that the expressions “first” and “second” used in the embodiments of the present invention are intended to distinguish two non-identical entities with the same name or non-identical parameters. It can be seen that “first” and “second” are only for the convenience of expression and should not be understood as limitations on the embodiments of the invention. The subsequent embodiments will not explain this one by one.

[0020] The preferred embodiment of the present invention first realizes that since it is difficult to extract the pupil and the light spot from the overall facial image, the image near the eyes of the interactive object can be extracted, or in other words, the local position (corresponding to the eye features) in the overall image can be enlarged. The preferred embodiment of the present invention also realizes that when the display area of ​​the display device is set to be certain and the positions of the infrared optical fiber and the camera device are fixed, even if there are individual differences in the height and facial size of the interactive object, in the overall image, the area where the face of the interactive object is located and the area where the eye position is located always remain relatively fixed, wherein the face of the interactive object will be in the approximate center of the overall image, and in the facial image, the eye position falls in the approximate upper middle position. According to the technical guidance of the first two aspects, the preferred embodiment of the present invention selects the existing Yolov5 algorithm and makes adaptive corrections to it. One of the directions of this correction is how to capture specific elements in the image. This can be to enable the algorithm model to capture the eye area of ​​the interactive object according to a certain image size, and then enlarge the captured image according to a certain ratio. It can also be to enable Yolov5 to enlarge the original facial image according to a certain ratio, and then capture the eye area of ​​the interactive object from the enlarged overall image according to a certain preset range size.

[0021] Specifically, the Yolov5 algorithm's training of labeling images and targets is to determine whether the detection box of the algorithm model matches the range of the hit target correctly (incorrect matching results will be discarded). For example, let the Yolov5 algorithm capture the eye area of ​​the interactive object according to a certain image size, and then enlarge the captured image according to a certain ratio (first capture and then enlarge). Figure 1 , Figure 1 is a schematic diagram showing the state of marking the eye position in the overall image of the interactive object. Figure 1 In the two images shown, the detection box of the Yolov5 algorithm (the line box in the figure) defines the two eyes of the interactive object in the image. The annotation content of the hit range (detection box) includes target category information c, center coordinate x, center coordinate y, width w and height h. In expectation, when the match is correct, the center position of the detection box is the center of the eye area. With this position as the midpoint, the preset pixel range is selected to enlarge the eye area to obtain the following Figure 2The enlarged image displayed. Those skilled in the art should know that the selection of the aforementioned pixel range and preset range can be adjusted according to the clarity of the camera device and the normal viewing distance of the interactive object. For example, in a preferred embodiment of the present invention, when the eye range is initially captured, a square area with a pixel size set to 160 pixels is configured as a detection frame, and then the captured image is enlarged to a size of 416×416 pixels. As the clarity of the camera device is improved, the captured image can be within a more precise range and can be enlarged to a larger pixel size. The strategy of first enlarging and then capturing and the strategy of first capturing and then enlarging may have slight differences in specific enlargement parameters and captured data, but in the final effect, there is no difference between the two that will affect the results. Compared with the strategy of first capturing and then enlarging, when enlarging and then capturing, the eye position can be directly captured according to the range of 416×416 pixel values ​​after selecting it, but it should be understood that no matter which strategy is adopted, the pixel value of the detection frame should be smaller than the pixel size after enlargement under normal circumstances. In this way, the enlarged image of the eye position of the interactive object is cut out from the overall image. Next, the pupil and the light spot in the enlarged image need to be identified and marked.

[0022] Generally, the light spot will appear as a bright spot near the pupil. In the annotation information, different target category information is used to distinguish the pupil position and the light spot position. For example, in some preferred embodiments of the present invention, the target category information of the pupil position is set to 0, and the target category information of the light spot position is set to 1. Of course, this configuration method of the label category information is only schematic. The purpose is to use different target category information labels to distinguish different objects and different object positions in the image. Therefore, the specific configuration method of the target category information should not be regarded as a limiting description of the scope of the preferred embodiment of the present invention. Those skilled in the art can still choose the configuration of the category information as needed. For example, replacing the category information in the above preferred embodiment, or increasing the number of bits of the target category information, etc., should be regarded as not departing from the scope of the preferred embodiment of this scheme.

[0023] The target category information alone is not enough, because the algorithm still cannot recognize pupils and light spots. Therefore, after determining the target category information and labeling method, the target detection algorithm is optimized and trained. Figure 2. The algorithm needs to select the pupil position and the spot position generated by the optical fiber of the infrared light source from the enlarged image. Therefore, the optimization direction of the algorithm is mainly to adjust the loss function of the known Yolov5 algorithm, so that the algorithm, which was originally unable to recognize the pupil and the spot, can recognize the pupil and the spot in the image after adjustment. In other words, the confidence value of the correct target selected by the prediction box of the Yolov5 algorithm is improved. The training process and goal is to determine whether the position selected by the algorithm box matches the actual position of the pupil and the spot manually marked. In actual performance, that is, the predicted box and the real box appear to be roughly or completely overlapped. For example, it is defined that when the predicted box output by the neural network matches the real marked box, the algorithm indicator factor is set to 1, otherwise it is set to 0. It is also defined that for the enlarged and / or cropped image, B is used to represent the number of prediction boxes. In some preferred embodiments of the present scheme, regardless of whether the strategy of first zooming in and then cropping or the strategy of first cropping and then zooming in is selected, the value of B is always 2. This is because, although the objects cropped in the two strategies are different, the number of croppings is the same. Regardless of the strategy, it should include two prediction boxes corresponding to the two eye positions of the interactive object, and two prediction boxes corresponding to the pupil position and the light spot position in one eye position respectively.

[0024] It should be noted that since it is expected to identify the eye position, pupil position and spot position of the interactive object in sequence, the algorithm model of the preferred embodiment of the present invention will include three outputs. Considering that the size of the image after the cut and enlargement is set to 416×416 pixel values, when configuring the algorithm output channel, the prediction box size of the eye position in the portrait is set to 52×52 pixel values, the prediction box size of the pupil position is set to 26×26 pixel values, and the prediction box size of the spot position is set to 13×13 pixel values. It is not difficult to see that, in principle, the selection size of the cut and enlarged image is better based on the common size supported by the Yolov5 algorithm itself, and the size of the specific prediction box is also set to a divisor of the size supported by the algorithm. Of course, in addition to other sizes supported by the Yolov5 algorithm itself (for example, 608×608 pixel values), the size of the cut and enlarged image and the prediction box size output according to the common divisor of the image pixel value can also be adaptively adjusted, but on the premise of ensuring the accuracy and precision of the algorithm, as well as the prediction detection and training speed. Similarly, the ratio between the pupil position and the light spot position prediction frames can also be appropriately adjusted. For example, in this illustrative embodiment, the values ​​of 416 / 8, 416 / 16 and 416 / 32 are selected as the size of the channel output, corresponding to the size of the eye position, pupil and light spot prediction frames. If the camera device is replaced, or the image clarity changes due to the change of the fixed interaction position of the interactive object, the specifications of the prediction frame can also be adjusted as needed. The specific selection will not be repeated here.

[0025] Then, looking back at the target annotation content in the previous article, it is a polynomial including target category information c, center coordinates x, center coordinates y, width w and height h. Therefore, in addition to the target category information, the loss function also needs to calculate the coordinates of the selection range and the width and height losses of the selection range.

[0026] First, for the coordinate center loss, that is, whether the range defined by the detection box is close enough (overlaps) to the area where the target is located. In terms of judgment, if the predicted box and the real box on the image are judged to be matched, it is expected that at least two conditions should be met between the two. First, in the coordinate system constructed by the image plane, the distance from the center point of the predicted box to the center point of the real box should be within the set range; second, in the same coordinate system, the size specifications of the above two should also be within the set range. The size specifications here specifically refer to the width and height values ​​of the predicted box and the real box. It should be noted that for the error of the center point of the prediction box, the algorithm model will output three prediction boxes corresponding to different targets through different channels, and the sizes of the three prediction boxes are different. Obviously, the actual errors generated by prediction boxes of different sizes at the same error ratio are not the same. Therefore, it is not possible to rely on a single standard to judge the loss error between prediction boxes of different specifications. The actual error difference between prediction boxes of different sizes should be considered. In a preferred embodiment of the present invention, the difference between the empirical parameters of the existing Yolov5 algorithm and the area of ​​the real selection box is used as the weight to adjust the proportion of the specification error of the different-sized boxes to the total coordinate center loss, so as to balance and adapt to different judgment scenarios between prediction boxes of different sizes. Then, for the width and height loss of the selection range, the difference in the height and width of the box can also be calculated separately according to the same optimization idea of ​​the center point.

[0027] After calculating the loss of the coordinate center point and the width and height of the selection box, it is necessary to continue to calculate the algorithm and the output results, that is, the confidence loss. Due to the initial training stage, most of the predicted boxes of the algorithm model cannot match the real boxes, so a weight value is set to adjust the proportion of the confidence loss of the predicted boxes that do not match the real boxes in the total confidence loss. In some preferred embodiments of the present invention, the weight value is 0.5. Finally, the loss of the target category is calculated, that is, it is determined whether the target category finally output by the prediction box is correct. Therefore, for any hit target on the image captured by the camera device, the revised algorithm model, if defined using As the center coordinates of its real box, and define (x i ,y i ) is the center coordinate of the predicted box output by the algorithm, then the loss function of the algorithm can be specifically expressed as:

[0028]

[0029] in, It is an indicator factor, indicating whether the predicted box output by the neural network matches the true labeled box. If the predicted box matches the true box, it is 1, otherwise it is 0. Its meaning is the indicator factor The value of is opposite. B is the number of true annotation boxes. In human eye detection, it represents two eyes; in pupil spot detection, it represents a pupil and a spot. The first term in the formula is the coordinate center loss mentioned above, where is the coordination parameter of the coordinate center error weight, which is used to adjust the proportion of the loss of prediction boxes of different sizes in the total coordinate center loss. The second item is the coordinate width and height loss, which is calculated for the width and height coordinate loss of all prediction boxes that match the real box. The third and fourth items are the confidence loss, which calculates the loss of whether all prediction boxes contain the real labeled target. Set λ noobj It is used to adjust the proportion of the confidence loss of the predicted box that does not match the real box in the total confidence loss. In this embodiment, the value is 0.5. If a predicted box i matches the real box j, then otherwise is the confidence of the algorithm output. The fifth item is the category loss, which calculates the loss of whether all detection boxes can correctly judge the object category. is the category to which the real box belongs, is the category of the algorithm prediction box. So far, by optimizing the loss function of the existing Yolov5 algorithm, the update of its neural network parameters can be completed. The updated algorithm model can identify the pupil position and the spot position in the image, which also completes the first aspect of the present invention to solve its technical problem. The identification and detection results can be found in Figure 3 ,The values ​​in the figure show the confidence values ​​of the pupil position and the light spot position respectively.

[0030] The pupil position and light spot position obtained according to the above content are a part of the area in the coordinate system composed of the portrait pattern. The range of this area is also the area defined by the prediction box. The identification of this area is obtained by combining the coordinates of the center point of the range with the width and height of the prediction box. Therefore, looking back at the overall idea of ​​introducing head posture changes in line of sight tracking proposed in this scheme, that is, the coordinates obtained in the first aspect of the technical guidance of the present invention to represent the pupil position and the light spot position are fitted with the coordinates representing the head posture, so that the fitted line of sight tracking result is closer to the actual gaze position of the interactive object in the display area. Therefore, the second aspect of the present invention includes two steps: one, capturing and representing the head posture; two, fitting the head posture data with the line of sight landing point coordinate data.

[0031] The head posture algorithm usually needs to capture and record the head posture position and head posture changes of the interactive object through a 3D imaging device. However, from the perspective of system architecture complexity and system equipment cost, adding 3D imaging equipment is not a better idea. The preferred embodiment of the present invention is still based on the aforementioned camera device to obtain the three-dimensional posture angle through a plane image. In theory, the existing methods of representing the rotation of an object in three-dimensional space can be used here to capture and represent the head posture and changes of the interactive object.

[0032] For example, the existing Euler angle method describes the rotation of a rigid body around three axes in a reference coordinate system through roll angle (Roll), pitch angle (Pitch) and yaw angle (Yaw). The Euler angle is applied to this scheme, that is, the angle values ​​of the three angles are used to estimate and represent the posture of the head in the plane. When the head posture changes, the three angle values ​​corresponding to the two positions before and after can be compared to calculate the posture position after the head posture changes. Then, the posture position is fitted with the pupil position and the spot position obtained above, and the result of line of sight tracking is finally obtained. However, when a certain rotation angle is a positive or negative 90° angle, the directions of the two rotation axes of the Euler angle will be aligned with each other (universal joint lock). In this case, the rotation state obtained by aligning the rotation axes is not unique, and the final expected result of the three axes that can be rotated is not a one-to-one mapping, and often a many-to-one mapping situation occurs, which will directly affect the learning and judgment results of the posture of the neural network selected to identify the head posture.

[0033] In order to solve the above-mentioned defects of applying Euler angles to this solution, the conventional further attempt is to select the quaternion representation method to replace the Euler angle solution. Quaternion uses four elements q = (q0, q1, q2, q3) to represent the rotation of the target, where (q1, q2, q3) defines the direction of the target rotation axis, and q0 represents the amplitude of the target rotation. Compared with the Euler angle solution, the advantage of the quaternion solution is that the rotation in the complex plane is represented by the multiplication of complex numbers, and the rotation amplitude is superimposed with the rotation angle, avoiding the universal joint lock problem caused by the alignment of the rotation axis. However, since the rotation of the rigid body itself has the possibility of bidirectional rotation (polar symmetry problem), that is, in this method, q and -q represent the same rotation, therefore, the quaternion solution also has the situation that the results are not unique in some rotation states, which will cause different rotations to be regressed to the same result in the regression training of the algorithm model, which is obviously not expected by the algorithm model.

[0034] It is not difficult to see that the defects of the above two existing methods are essentially because neural network learning requires the input results to be differentiable, or the input results to be continuous. In specific scenarios, it requires that the change in head posture does not have a non-unique solution between the rotation position and angle. Therefore, the preferred embodiment of the present application uses a rotation matrix method to represent the rotation change of the head posture, in order to obtain a method that can provide a continuous and unique smooth representation of the head posture change.

[0035] In a preferred embodiment of the present invention, a 3×3 rotation matrix R is used to represent the rotation of the head posture. The rotation matrix contains 9 parameters, which is relatively large, and it also needs to satisfy the orthogonality constraint (R T R=I,where R T is the transposed matrix, I is the identity matrix), and then the matrix is ​​simplified by Gram-Schmidt orthogonalization or SVD decomposition. After simplification, the representation of the original 3×3 matrix is ​​simplified to a method containing two three-dimensional vectors and a total of six parameters. The simplification process is as follows:

[0036]

[0037] gGS represents the simplification process of the three-dimensional rotation matrix. a1, a2 and a3 are the three-dimensional vectors that identify the current head posture of the interactive object, which are combined to represent the 3x3 rotation matrix before simplification. and It is also a three-dimensional vector, representing the six parameters simplified by orthogonalization or SVD decomposition. The simplified parameters are then reverse mapped to obtain the final rotation information representing the head posture:

[0038]

[0039] fGS is the reverse mapping process, and The six parameters jointly represent the output of the neural network. b1, b2 and b3 are the three-dimensional rotation matrices of the algorithm that finally represent the head posture: The final rotation information of the restored head is:

[0040]

[0041] b3=b1×b2

[0042] It can be seen in the expression of b3 that the third-dimensional vector can be expressed as the product of the first two vectors. In this way, the final output of the neural network is six parameters. In the selection of a specific neural network, the preferred embodiment of the present invention selects CSPDarknet53 as the head posture feature extraction network. The deep structure of this network can learn the fine features of the head posture, and the CSP module helps to better capture diverse angle information. In actual use, the head output of the network is adapted to the algorithm to output a 6-dimensional rotation vector. The loss function selects the L2 loss function, and the training data set selects the network data set 300W-LP, which contains 66,225 facial data. The actual head posture estimation effect after training is as follows. Figure 4 As shown, Figure 4 The state diagram shows the result of posture calculation after collecting two sets of pictures in each direction. The axial line in the figure shows the head posture of the interactive object in each direction. So far, three sets of data of pupil position, light spot position and head posture position have been obtained. Next, these three sets of data are fitted.

[0043] The fitting process is essentially to integrate the three position data and associate them with a point coordinate in the device display area. The point coordinate is the actual gaze position of the interactive object in the display area. In the fitting calculation process, the fitting parameters of the three positions are finally obtained. The combination of the position data adjusted according to the three fitting parameters is the final coordinate value.

[0044] The overall idea of ​​the preferred embodiment of the present invention for this part is that as the interactive object's gaze position changes in the display area, its pupil position and the position of the light spot transmitted on its pupil by the light source device at a fixed position on the display device will also change accordingly. According to the aforementioned annotation method, the changes in the pupil and light spot positions can be quantified by the position change of the center point of the detection frame. For example, let the center point of the pupil be (x1, y1) and the center point of the light spot be (x2, y2), and calculate the position difference between the two △x=x1-x2, △y=y1-y2, and use △x and △y as independent variables. According to the technical idea of ​​the present invention, the quantified coordinate change needs to be superimposed with the rotation information to finally form a vector, and the rotation information finally obtained in the head posture part is a rotation matrix including six parameters. Therefore, the rotation matrix is ​​converted into Euler angle expression, and the conversion process is:

[0045] p=atan2(-R[2][0],sqrt(R[2][1]*R[2][1]+R[2][2]*R[2][2]))

[0046] y=atan2(R[2][1],R[2][2])

[0047] r = atan2(R[1][0],R[0][0])

[0048] Among them, R is a three-dimensional vector obtained according to the reverse mapping, thus obtaining a vector containing three angle values ​​"roll angle (Roll), pitch angle (Pitch) and yaw angle (Yaw)". Finally, the angle value is combined with the aforementioned position change to first form a vector representing the position change and rotation change.

[0049] However, the vector is still only a parameter representing the relative position and rotational posture in the reference coordinate system. To complete the line of sight tracking, the vector still needs to be finally mapped to a specific point value coordinate in the display area. Then, for the correspondence between the vector and the coordinates, or for the determination of the aforementioned fitting parameters, in the preferred embodiment of the present invention, an attempt is made to select several points in the display area (screen range) as references, calculate the coordinate positions mapped by the sampling points selected by the above-mentioned vector reference, and finally obtain the least squares method.

[0050] Then, first configure the sampling points in the display area. The selection of sampling points needs to consider the following aspects:

[0051] 1) The number of sampling points should not be too small, because a small number of sampling points cannot cover all positions in the display area. Even if multiple samplings are used, the error between the mapping result and the true value is still difficult to estimate;

[0052] 2) The sampling points should be evenly distributed in the display area to achieve accurate mapping results in all directions. However, considering the data processing pressure that may be caused by the sampling times of each sampling point, the number of sampling points should not be too large;

[0053] 3) In the uniform distribution of sampling points, more sampling points should be distributed near the edge of the display area to improve the accuracy of line of sight tracking at the edge of the display area.

[0054] Schematically, in some preferred embodiments of the present invention, 12 sampling points are configured in a display area with a resolution of 2560×1440, and the sampling points are evenly distributed along the horizontal and vertical directions of the display area and cover the entire display area. In a reference coordinate system constructed with the center point of the display as the origin, the coordinates of the 12 sampling points are: (80,80), (880,80), (1680,80), (2480,80), (80,720), (880,720), (1680,720), (2480,720), (80,1360), (880,1360), (1680,1360) and (2480,1360), and each sampling point samples the head posture three times. It is worth mentioning that since the relative positions of the sampling points in the display area are different, for the same head posture, there will be different degrees of differences between the collected data of different sampling points. Therefore, if the change of pupil position and spot position, combined with the rotation information converted into Euler angle, is defined as (△x,△y,r,p,y) here, the least square method is used to solve in two directions of the reference coordinate system of the display area so that the results mapped by the acquisition results of each sampling point roughly meet the same linear conditions. The processing process can be expressed as follows:

[0055]

[0056] in, is the model prediction value, φ 1 (x),φ 2 (x)…,φ n (x) is the basis function, (β 1 ,β 2 ,…,β n ) is the model parameter to be solved, which can also be regarded as a function such as Aβ=y. is the dimension, and this embodiment includes 36 sets of sampled data, so m is 36, and n is the number of fitting parameters to be determined. Of course, because the way of configuring the fitting equation is different, the specific value of n here will also be adjusted according to the way the equation is constructed. Since minimizing the sum of squares of the errors can make the estimated model closest to the actual situation, the sum of squares of the errors between the coordinate values ​​of the calibration points and the predicted values ​​of the model will be minimized, which can be expressed in the horizontal and vertical directions as follows:

[0057]

[0058] in, The fitting coefficients can then be obtained by using the matrix singular value decomposition (SVD) method. Specifically, the singular value decomposition of matrix A is performed:

[0059] A=U∑VT

[0060] in, U is the left singular vector matrix, ∑ is the singular value matrix, V is the right singular vector matrix. Using the SVD method, the solution of the normal equation can be converted to the following form:

[0061] β=V∑ + U T y

[0062] Among them, ∑ + is the pseudo-inverse of ∑. After obtaining the fitting parameter β, the calibration phase of the gaze tracking algorithm is completed. When deployed, the (△x,△y,r,p,y) obtained by the front target detection and head posture estimation module is used to construct the matrix A, and the coordinates (X,Y) of the gaze area of ​​the current user can be calculated.

[0063] Taking a more specific embodiment as an example, for the user image data collected by the camera device, according to the strategy of first intercepting and then enlarging, the human eye position area is first marked to obtain the eye position marking information (x, y, w, h), where x is the horizontal center coordinate, y is the vertical center coordinate, w is the width of the marking box and h is the height of the marking box. With this point as the center point, an area with a size of 160×160 pixels is selected and enlarged to 416×416 pixels to meet the needs of the Yolov5 algorithm model. In this specific embodiment, the eye area (xe, ye, we, he) is obtained according to the decoding formula in the previous order. Then the eye area is cropped, and the cropped image is sent to the pupil spot detection algorithm to obtain the pupil spot area. Let the pupil area be (x1, y1, w1, h1) and the spot area be (x2, y2, w2, h2).

[0064] Next, according to the head posture algorithm weight file obtained by the aforementioned optimization training, the image is sent to the head posture estimation neural network again to obtain the network output (a1, a2), and the matrix R = (b1, b2, b3) is obtained according to the fGS formula, which represents the rotation matrix of the head posture, and then converted into Euler angles (r, p, y) again.

[0065] Then, the sampling points on the screen are sampled multiple times (12 sampling points, each sampling point is sampled 3 times). Each time the pupil position and the spot position are sampled, the difference between the center positions of the two is calculated (△x=x1-x2 and △y=y1-y2), and it is combined with the Euler angle (r, p, y) obtained in the previous step to represent the head posture to form a vector (△x, △y, r, p, y), thereby obtaining 36 sets of data. Construct an equation fitting (△x, △y, r, p, y) and the screen point coordinates (X, Y), and solve it according to the least squares method to obtain the fitting parameter β.

[0066] After getting the fitting parameter β, the user’s pupil spot position and head posture are obtained in real time (△x,△y,r,p,y), and the screen coordinate position (X,Y) is calculated to complete the entire gaze tracking algorithm. The final effect is as follows: Figure 5 As shown in the figure, the smaller circles indicate the positions of the sampling points, which are distributed at the edge of the display area as much as possible, while the larger circles are the coordinate positions output by the algorithm model when the interactive object looks at the sampling points. It can be seen that the figure contains four sets of output results. After optimization, the algorithm has greatly improved the eye tracking effect at the edge of the display area.

[0067] In the above-mentioned embodiments, the architecture is based on a camera device and a light source. In the annotation of the Yolov5 algorithm, after the eye position image is enlarged, the pupil position and the spot position are respectively annotated, and finally, the pupil position, the spot position and the head posture rotation information are fitted. The algorithm model in these preferred embodiments introduces a factor for considering the change of the head posture of the interactive object in the sight tracking algorithm, so that the main view position coordinates obtained according to the pupil position and the spot position in the known model are further corrected, especially, the positioning accuracy at the extreme positions such as the center and edge of the display area is significantly improved. Of course, in the process of training and learning the algorithm model, the preferred embodiment of the present invention still notes that due to the obvious difference between the size of the pupil position detection frame and the spot position detection frame, there may be accuracy problems in the process of calculating the distance between the pupil and the spot position center, and this error is difficult to be estimated and completely avoided; and when the scene light intensity is poor, or the display area content is a dark theme, etc., the aforementioned error may be further amplified for the spot position that is stretched and deformed and blurred after the sight line moves significantly. Although this error does not have an obvious impact on the final positioning result even after amplification, in order to ensure the performance of the algorithm model of the preferred embodiment of the present invention in some special scenarios, the preferred embodiment of the present invention has made further attempts.

[0068] Since there is an inevitable size difference between the pupil and the light spot, a light source device can be added. Two light sources can form two light spots at the eye position. When the user looks at the edge of the screen, the light spot on one side may appear too blurred. At this time, the light spot on the other side comes into play. At the same time, adjust the positions of the two light source devices to avoid overlapping of the light spot positions. Algorithmically, the original step of determining the center point loss by the pupil and light spot positions is improved to first find the center of the light spot position based on the coordinates of the two light spot positions, and then calculate the loss with the pupil position. Obviously, those skilled in the art know how to improve the training of the algorithm model when they know the scheme of a light source device and a camera device. Therefore, the specific implementation of the two light source devices (infrared lamps) scheme will not be repeated.

[0069] The above embodiments only express several implementation methods of the present invention, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present invention. It should be pointed out that, for a person of ordinary skill in the art, several modifications and improvements can be made without departing from the concept of the present invention, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention patent shall be subject to the attached claims.

Claims

1. A human-computer interaction sight tracking method, the method is used to track and locate the sight movement of an interactive object in a display area, the method comprising the following steps: The step of configuring at least one light source device and one camera device in the display area, the light source device generates illumination when in operation, and the camera device photographs the interactive object under illumination to obtain image data of the interactive object; Constructing a first reference coordinate system with two mutually perpendicular directions in the image data, intercepting the eye region of the interactive object from the image data according to a strategy of enlarging and then intercepting or intercepting and then enlarging, and determining the position coordinates of the pupil and the light spot under the eye region relative to the first reference coordinate system; Representing the head posture of the interactive object in the image data, and obtaining a three-dimensional rotation matrix representing the head posture angle; Configuring a plurality of sampling points covering the display area, each sampling point samples the head posture of the interactive object at least once, and then obtaining a position difference between a pupil position and a light spot position, and obtaining fitting parameters between the pupil position, the light spot position, the three-dimensional rotation matrix, and the gaze position coordinates of the interactive object in the display area according to the position difference and the three-dimensional rotation matrix; According to the fitting parameters, the pupil position and the spot position of the current interactive object are fitted with the three-dimensional rotation matrix to obtain the coordinates of the gaze position of the interactive object in the display area.

2. The human-computer interactive sight tracking method according to claim 1, wherein: The step of determining the position coordinates of the pupil and the light spot in the eye position area relative to the first reference coordinate system is specifically: The portrait data in the data training set are annotated by using a priori annotation method, and the first to third detection frames corresponding to the eye position, the pupil position and the light spot position are annotated respectively, and the first to third detection frames are represented by the combination of the coordinates in the first reference coordinate system and the width and height of the frame; The representation of the first to third detection boxes is used as input to the algorithm model, and a loss function is constructed to calculate the loss of the prediction results of the algorithm model and the contents of the prior annotations, and finally a weight file of the neural network parameters is obtained; The algorithm model is enabled to analyze the image data according to the parameter weight file to output a detection result.

3. The human-computer interactive sight tracking method according to claim 2, wherein: The loss function is: in, is an indicator factor indicating whether the neural network output matches the prior annotation, B is the number of true annotation boxes, is the coordination parameter of the coordinate center error weight, λ noobj is a preset constant. is the confidence of the algorithm output, is the category to which the real box belongs, The category of the box predicted by the algorithm.

4. The human-computer interactive sight tracking method according to claim 2, wherein: Before representing the head posture of the interactive object, the method further includes the steps of training the head posture algorithm based on the head posture data in the data training set and obtaining the head posture network weight file, and also includes: The head posture algorithm constructs a second reference coordinate system for representing the head posture angle from the image data to output a three-dimensional vector representing the rotation angle of the interactive object's head posture in each direction of the second reference coordinate system; The three-dimensional vector is reversely mapped to obtain the three-dimensional rotation matrix representing the head posture.

5. The human-computer interactive sight tracking method according to claim 4, wherein: It also includes using a rotation matrix to represent the rotation of the head posture of the interactive object.

6. The human-computer interactive sight tracking method according to claim 5, wherein: The specific steps of using rotation matrix for representation are: Simplify the rotation matrix using Gram-Schmidt orthogonalization or singular value decomposition to obtain two three-dimensional vectors; Reverse mapping the parameters contained in the two three-dimensional vectors to obtain the three-dimensional rotation matrix; The three-dimensional rotation matrix is ​​converted into Euler angles, and the difference between the Euler angles and the pupil position and the center point of the light spot position is combined to form a first vector.

7. The human-computer interactive sight tracking method according to claim 6, wherein: The rotation matrix is ​​simplified to satisfy: Among them, a1, a2 and a3 are three-dimensional vectors respectively, which are combined to identify the rotation matrix of the current head posture of the interactive object before simplification. and represents the parameters simplified by orthogonalization or SVD decomposition; The reverse mapping step fGS satisfies: Among them, b1, b2 and b3 are the three-dimensional rotation matrices representing the head posture.

8. The human-computer interactive sight tracking method according to claim 6, wherein: The step of obtaining fitting parameters between the pupil position, the light spot position, the three-dimensional rotation matrix, and the gaze position coordinates of the interactive object in the display area according to the position difference and the three-dimensional rotation matrix is ​​specifically as follows: Obtaining the coordinates of each sampling point corresponding to the first vector, which are defined as calibration coordinates; In two directions of the coordinate system of the display area, the fitting parameters are obtained by using the least square method.

Citation Information

Cited By

  • Eye movement tracking method and system based on multi-modal fusion

    CN121392947A

  • Method, device and equipment for detecting driver state in driving process

    CN122035004A