Eye movement tracking dynamic target identification method and device based on three-dimensional fixation point
By constructing a gaze-time series and deep learning model, combining template image recognition and a single-object tracker, the problems of large amount of operations and target loss of the three-dimensional gaze recognition algorithm in dynamic target recognition are solved, and dynamic target recognition with high accuracy and real-time performance are achieved.
Patent Information
- Application Number
- CN202510500113.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-08-22
AI Technical Summary
The existing three-dimensional gaze-based recognition algorithm has a large amount of calculation when identifying dynamic targets, which is prone to lose targets, and there is a lag between the user's gaze-based and expected recognition targets, resulting in poor recognition accuracy and real-time performance.
Using a three-dimensional eye movement tracking method based on gaze points, combining eye movement devices and deep learning models, we can filter high-density gaze points by building gaze points-time series, and use template image recognition models and single-object trackers to achieve real-time tracking and recognition of dynamic goals.
It improves the accuracy and real-timeness of dynamic target recognition, reduces target loss, reduces server pressure, and meets user expectations to focus on target recognition.
Smart Images

Figure CN120526469A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision and deep learning, and specifically relates to a method and device for dynamic target recognition based on three-dimensional gaze points and eye tracking.
[0002] Background Methods
[0003] Most current 3D gaze-based recognition algorithms require every frame of an image to identify an object, significantly increasing the computational load and often leading to target loss when faced with fast-moving objects. This not only affects recognition accuracy but also severely limits the system's real-time performance. Furthermore, due to the time lag between the user's gaze point and the intended target, relying solely on each 3D gaze-based object recognition can result in numerous false positives, reducing system reliability.
[0004] Integrating a single-target tracker into eye tracking can better simulate real-world user attention models and accurately reflect the user's gaze target. Applying this parameter to dynamic target recognition better reflects the user's actual usage experience, significantly improving target recognition accuracy and real-time performance, especially in dynamic environments, and facilitating more stable identification and tracking of moving targets. Summary of the Invention
[0005] In response to the above-mentioned problems, the present invention proposes a method and device for dynamic eye tracking recognition based on three-dimensional gaze points to solve the problem of inconsistency between the user's expected gaze target and the recognition target, as well as the server congestion problem caused by repeated uploading of recognition.
[0006] To achieve the above-mentioned purpose, the specific method scheme of the present invention is as follows: A three-dimensional gaze point-based eye tracking dynamic recognition method comprises the following steps:
[0007] S1 uses eye-tracking equipment to obtain the user's three-dimensional gaze point data, constructs a gaze point-time series, and filters out the user's high-density three-dimensional gaze points;
[0008] S2 establishes a spatial coordinate system based on the eye tracking device, with the center point of the binocular camera of the user's eye tracking device as the origin, the user's horizontal direction as the X-axis, and the right side as the positive direction of the X-axis; the vertical direction as the Y-axis, and the upward direction as the positive direction of the Y-axis; the front-back direction as the Z-axis, and the front direction as the positive direction of the Z-axis;
[0009] The position of the center point of the user's eyes in S3 is E, and the three-dimensional gaze point data is (x 注视 ,y 注视 ,z 注视 ), the position of the user's eye center and the three-dimensional gaze point data are substituted into the spatial coordinate system based on the eye movement device, and the sight line equation is established to obtain the user's gaze line function L(x视线 ,y 视线 ,z 视线 ), where L(x 视线 ,y 视线 ,z 视线 ) is the starting point of the user's eye center point E, and the end point is the three-dimensional gaze point (x 注视 ,y 注视 ,z 注视 );
[0010] S4 will look at the line of sight function L(x 视线 ,y 视线 ,z 视线 ) is mapped to the foreground camera coordinate system to obtain the two-dimensional coordinates (x', y') of the user's gaze point;
[0011] S5 selects a rectangular area with a set length and width of W*H pixels with the two-dimensional coordinates (x', y') of the gaze point as the center, establishes the initial coordinates of the recognition area (x', y', W / 2, H / 2), where x', y' represent the two-dimensional coordinates of the gaze point, W / 2, H / 2 represent the distance from the border of the rectangular area to the center, and sets the image of the W*H pixel rectangular area as a template image for subsequent tracking;
[0012] S6 uses the Pytorch framework to load the weight file pre-trained in the Yolov8 framework, initializes the template image recognition model, and uses the template image model to extract deep features of the template image. The deep features include low-level features, mid-level features, high-level features, and global features. The global features aggregate the features of the first three layers to generate semantically rich contextual information for subsequent target classification and definition;
[0013] S7 inputs the extracted deep features of the template image into the head network of the template image recognition model for classification. The head network calculates the probability distribution of each category and obtains the result with the highest confidence as the recognition result of the template image;
[0014] S8 sends the recognition result of the template image to the eye tracking device, which records the result received this time and compares the current result with the previous result. If the result changes, the recognition result is notified to the user; if there is no change, no operation is performed;
[0015] S9 initializes a single-target tracker based on the Siamese framework. It extracts each foreground frame from the 30FPS foreground video stream transmitted by the eye tracking device and numbers them in ascending order of time. The initial coordinates of the recognition region (x', y', W / 2, H / 2) and the template image are input into the template branch of the single-target tracker. The foreground frame with the smallest number is input into the search branch of the single-target tracker, and the smallest foreground frame is deleted.
[0016] S10 uses the HRnet network to extract features from the template branch and search branch of the single target tracker, and performs deep cross-correlation on the obtained template branch feature map and search branch feature map to obtain a similarity response map;
[0017] S11 generates multiple anchor points in the area with higher confidence in the similarity response map, performs regression and classification tasks on each anchor point, and generates tracking results;
[0018] S12 marks the tracking result in the video frame of the foreground image, completing the tracking process of the target image in one frame of the foreground image, and repeating steps S9 to S12 in the foreground video stream to achieve the tracking effect of the target in the foreground video;
[0019] While tracking the target in the foreground video, S13 monitors whether the two-dimensional coordinates (x', y') of the user's gaze point deviates from the tracking candidate frame for a long time. If the two-dimensional coordinates (x', y') of the gaze point exceed the set time threshold and do not enter the candidate frame area, it is determined that the user has lost attention on the target, and returns to step S5 to reselect the initial coordinates of the recognition area and the template image; if there is no deviation within the time threshold, the previous result is maintained.
[0020] Furthermore, the above step S1 includes the following contents:
[0021] S1.1 Organize the 3D gaze point data into time series according to timestamps;
[0022] S1.2 uses a density clustering algorithm to filter out high-density gaze points within a time period, and combines time information to establish a set of gaze point-time series (x, y, z, t), where x, y, z represent the high-density three-dimensional gaze points of the filtered users, and t represents the timestamp.
[0023] Furthermore, the above step S1.2 includes:
[0024] S1.2.1 Define the space-time distance calculation formula,
[0025] α represents the time weight coefficient;
[0026] S1.2.2 Use the DBSCAN algorithm to process the time series. Based on the temporal distribution of 3D fixation points, determine the ε and MinPts values. Compare the number of eligible points in the neighborhood of each data point with MinPts. ε is the neighborhood radius, which defines the range of the neighborhood around a point. MinPts is the minimum number of neighboring points. If the number of points around a point is greater than or equal to MinPts, the point is considered a core point.
[0027] S1.2.3 If a point is confirmed to be a core point and there are other points in its neighborhood, these points belong to the same cluster; clusters containing more than the set number threshold are selected, and the points in the cluster represent the selected user's high-density three-dimensional gaze points.
[0028] Furthermore, the above step S4 includes the following contents:
[0029] The S4.1 foreground camera coordinate system has the geometric center of the lens of the external camera mounted on the eye tracker as its origin, and the directions of the X, Y, and Z axes are consistent with the spatial coordinate system of the eye tracking device;
[0030] S4.2 The user's gaze function L(x 视线 ,y 视线 ,z 视线 ) is converted with the internal parameters of the foreground camera, and the gaze function in the eye movement device coordinate system is converted to the foreground camera coordinate system using the view transformation matrix to obtain the foreground gaze function K(x 前景 ,y 前景 ,z 前景 ), the conversion formula is K(x,y,z)=R*L(x,y,z)+T,
[0031] Where R is a 3×3 rotation matrix, representing the rotation of the eye tracking device coordinate system relative to the foreground camera coordinate system, and T is a 3×1 translation vector, representing the position of the eye tracking device origin in the foreground camera coordinate system;
[0032] S4.3 In the foreground camera coordinate system, according to the foreground sight function K(x 前景 ,y 前景 ,z 前景 ), assuming that the foreground image curtain is parallel to the X-axis and Y-axis in the foreground camera coordinate system, and the Z-axis coordinate is g, the foreground sight function K(x 前景 ,y 前景 ,z 前景 ) in z 前景 When the X-axis and Y-axis coordinates are equal to g, they are the two-dimensional coordinates (x', y') of the user's three-dimensional gaze point on the foreground image.
[0033] Furthermore, the above step S7 includes the following contents:
[0034] S7.1 passes the extracted deep features to the head network of the template image recognition model, and uses PAnet to further aggregate and optimize multi-layer features to ensure that the model takes into account both global and local features;
[0035] S7.2 outputs the confidence score of each category based on the multi-layer features after aggregation according to the classification loss function of yolov8;
[0036] S7.3 selects the category with the highest confidence as the final template image recognition result based on the confidence scores of all categories.
[0037] Furthermore, the above step S10 includes the following contents:
[0038] The template branch and search branch of S10.1 use a feature extraction network structure with shared weights to ensure that the template branch and search branch are matched in the same feature space. The backbone network of the high-resolution feature extraction network HRnet with the same structure is used to extract features from the template branch and search branch images respectively to retain more detailed information.
[0039] S10.2 After completing the feature extraction, the obtained template branch feature map and the search branch feature map are subjected to a deep cross-correlation operation. The deep cross-correlation operation is to perform a convolution operation on the template branch feature map and the search branch feature map channel by channel in the channel dimension to obtain a two-dimensional similarity response map, in which the value of each position represents the similarity between the template and the corresponding position of the search area.
[0040] Furthermore, the above step S11 includes the following contents:
[0041] S11.1 sets a threshold based on the percentage of the maximum value based on the obtained similarity response map, marking the area above 80% of the maximum value as a high-confidence area. Multiple anchor points are generated in the high-confidence area of the similarity response map. These anchor points are centered in the high-confidence area and each has a variety of scales and aspect ratios to adapt to the shape and size of the target. These anchor points serve as reference frames for target detection and provide a basis for subsequent regression and classification tasks.
[0042] S11.2 performs regression and classification tasks for each anchor point. The regression task calculates the one with the highest regression score among multiple scales and aspect ratios based on the regression loss function as the size and aspect ratio of the anchor box, making it closer to the target; the classification branch generates a confidence score for each anchor point based on the classification loss function, which indicates the probability of the target contained in the anchor point. The anchor point with the highest score is found based on the score of the classification branch, and its candidate box is used as the tracking result.
[0043] The present invention also provides an eye tracking dynamic recognition device based on three-dimensional gaze points, including an eye state detection device, a two-dimensional gaze point mapping device, a target recognition device, a single target tracking device, a gaze deviation monitoring device, and a recognition result broadcasting device. The eye state detection device screens out three-dimensional gaze points in a staring state according to the three-dimensional gaze points provided by the eye movement device, the gaze time, and the distribution of the gaze points; the two-dimensional gaze point mapping device is responsible for converting the screened three-dimensional gaze points into corresponding two-dimensional coordinates and mapping them to a foreground image corresponding to their time; the target recognition device is responsible for identifying the image of the user's gaze point area through a local recognition library or an online recognition interface to identify the target and extract it as a template image; the single target tracking device is used to track the target and track the target in the foreground image through the extracted template image; the gaze deviation monitoring device is used to combine the information of the tracking target with the user's gaze point information to determine whether the gaze area deviates from the tracking target for a long time; the recognition result broadcasting device is used to broadcast the recognition result to the user through voice and provide a voice broadcast of routine operation prompts.
[0044] Compared with existing methods, the results obtained by this invention are more consistent with the target the user is looking at, effectively reducing the loss of the target due to deformation or high-speed movement. Even if the target is temporarily lost, it can still be re-acquired based on the user's gaze point. Furthermore, this method does not require repeated upload and recognition of the same target, reducing communication volume with the server and reducing server pressure. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 This is a flow chart of the eye tracking dynamic target method based on three-dimensional gaze points.
[0046] Figure 2 Schematic diagram of the structure of the eye tracking dynamic target device based on three-dimensional gaze point.
[0047] Figure 3 This is the flow chart of the single target tracking device. DETAILED DESCRIPTION
[0048] The present invention is further described below with reference to the accompanying drawings and specific embodiments. It should be pointed out that the method scheme and design principle of the present invention are described in detail below only with an optimized method scheme, but the protection scope of the present invention is not limited to this.
[0049] The embodiments are preferred implementations of the present invention, but the present invention is not limited to the above implementations. Any obvious improvements, substitutions or modifications that can be made by method personnel in this field without departing from the essence of the present invention are within the scope of protection of the present invention.
[0050] The present invention is an eye tracking dynamic recognition method based on three-dimensional gaze point. Figure 1 As shown, the following steps are included:
[0051] S1 uses eye-tracking equipment to obtain the user's three-dimensional gaze point data, constructs a gaze point-time series, and filters out the user's high-density three-dimensional gaze points;
[0052] As a preferred embodiment of the present invention, step S1 includes the following contents:
[0053] S1.1 Organize the 3D gaze point data into time series according to timestamps;
[0054] S1.2 uses a density clustering algorithm to filter out high-density gaze points within a time period, and combines time information to establish a set of gaze point-time series (x, y, z, t), where x, y, z represent the filtered high-density three-dimensional gaze points of the user, and t represents the timestamp.
[0055] As a preferred embodiment of the present invention, step S1.2 includes:
[0056] S1.2.1 Define the space-time distance calculation formula,
[0057] α represents the time weight coefficient;
[0058] S1.2.2 Use the DBSCAN algorithm to process the time series. Based on the temporal distribution of 3D fixation points, determine the ε and MinPts values. Compare the number of eligible points in the neighborhood of each data point with MinPts. ε is the neighborhood radius, which defines the range of the neighborhood around a point. MinPts is the minimum number of neighboring points. If the number of points around a point is greater than or equal to MinPts, the point is considered a core point.
[0059] S1.2.3 If a point is confirmed to be a core point and there are other points in its neighborhood, these points belong to the same cluster; clusters containing more than the set number threshold are selected, and the points in the cluster represent the selected user's high-density three-dimensional gaze points.
[0060] S2) establishing a spatial coordinate system based on the eye tracking device, with the center point of the binocular camera of the user's eye tracking device as the origin, the user's horizontal direction as the X-axis, with the right side as the positive direction of the X-axis; the vertical direction as the Y-axis, with the top as the positive direction of the Y-axis; and the front-back direction as the Z-axis, with the front direction as the positive direction of the Z-axis;
[0061] S3) The position of the center point of the user's eyes is E, and the three-dimensional gaze point data is (x 注视 ,y 注视 ,z 注视), the position of the user's eye center and the three-dimensional gaze point data are substituted into the spatial coordinate system based on the eye movement device, and the sight line equation is established to obtain the user's gaze line function L(x 视线 ,y 视线 ,z 视线 ), where L(x 视线 ,y 视线 ,z 视线 ) is the starting point of the user's eye center point E, and the end point is the three-dimensional gaze point (x 注视 ,y 注视 ,z 注视 );
[0062] S4 converts the gaze point line of sight function L(x 视线 ,y 视线 ,z 视线 ) is mapped to the foreground camera coordinate system to obtain the two-dimensional coordinates (x', y') of the user's gaze point.
[0063] As a preferred embodiment of the present invention, step S4 includes the following contents:
[0064] The S4.1 foreground camera coordinate system has the geometric center of the lens of the external camera mounted on the eye tracker as its origin, and the directions of the X, Y, and Z axes are consistent with the spatial coordinate system of the eye tracking device;
[0065] S4.2 The user's gaze function L(x 视线 ,y 视线 ,z 视线 ) is converted with the internal parameters of the foreground camera, and the gaze function in the eye movement device coordinate system is converted to the foreground camera coordinate system using the view transformation matrix to obtain the foreground gaze function K(x 前景 ,y 前景 ,z 前景 ), the conversion formula is K(x,y,z)=R*L(x,y,z)+T,
[0066] Where R is a 3×3 rotation matrix, representing the rotation of the eye tracking device coordinate system relative to the foreground camera coordinate system, and T is a 3×1 translation vector, representing the position of the eye tracking device origin in the foreground camera coordinate system;
[0067] S4.3 In the foreground camera coordinate system, according to the foreground sight function K(x 前景 ,y 前景 ,z 前景 ), assuming that the foreground image curtain is parallel to the X-axis and Y-axis in the foreground camera coordinate system, and the Z-axis coordinate is g, the foreground sight function K(x 前景 ,y 前景 ,z 前景 ) in z 前景When the X-axis and Y-axis coordinates are equal to g, they are the two-dimensional coordinates (x', y') of the user's three-dimensional gaze point on the foreground image.
[0068] S5 selects a rectangular area with a set length and width of W*H pixels with the two-dimensional coordinates (x', y') of the gaze point as the center, establishes the initial coordinates of the recognition area (x', y', W / 2, H / 2), where x', y' represent the two-dimensional coordinates of the gaze point, W / 2, H / 2 represent the distance from the border of the rectangular area to the center, and sets the image of the W*H pixel rectangular area as a template image for subsequent tracking;
[0069] S6 uses the Pytorch framework to load the weight file pre-trained in the Yolov8 framework, initializes the template image recognition model, and uses the template image model to extract the deep features of the template image. The deep features include low-level features, mid-level features, high-level features, and global features. The global features aggregate the features of the first three layers to generate semantically rich contextual information for subsequent target classification and definition.
[0070] As a preferred embodiment of the present invention, the present invention uses the Pytorch framework to load the weight file that has been pre-trained in the Yolov8 framework to initialize the template image model. The weight file mainly contains the following four contents:
[0071] S6.1 Model architecture information, including the specific number of layers in the network and the parameters of each layer, including the number of channels, stride, and convolution kernel size.
[0072] S6.2 loss function parameters include bounding box regression loss, which is used to predict the coordinates and size of the bounding box; object confidence loss, which determines whether an object exists at a certain location; and classification loss, which is responsible for the classification task of the detection category.
[0073] S6.3 Detection head parameters, used for bounding box prediction, including the weights of the classification head, confidence head, and regression head.
[0074] S6.4 Training status information, including training hyperparameters such as learning rate, batch size, and weight decay.
[0075] As a preferred embodiment of the present invention, the template image is input into the template image model for preprocessing, the image is proportionally enlarged, and it is adjusted to (640, 640) and filled with black on the shorter side to make up for the size. The enlarged image is normalized, and the channel order is adjusted to (3, 640, 640) through dimensional conversion to meet the model input requirements. The preprocessed image is input into the template image recognition model, and the deep features are extracted through the backbone network of the template image recognition model. The deep features include low-level features, mid-level features, high-level features and global features. The low-level features have a high resolution and retain rich edge and texture information. The mid-level features are used to extract the contour, shape and local structure information of the object. The high-level features are used to identify the category of the object and background information. The global features aggregate the first three layers of features to generate semantically rich contextual information for subsequent target classification and definition.
[0076] S7 inputs the extracted deep features of the template image into the head network of the template image recognition model for classification. The head network calculates the probability distribution of each category and obtains the result with the highest confidence as the recognition result of the template image.
[0077] As a preferred embodiment of the present invention, step S7 includes the following contents:
[0078] S7.1 passes the extracted deep features to the head network of the template image recognition model, and uses PAnet to further aggregate and optimize multi-layer features to ensure that the model takes into account both global and local features;
[0079] S7.2 outputs the confidence score of each category based on the multi-layer features after aggregation according to the classification loss function of yolov8;
[0080] S7.3 selects the category with the highest confidence as the final template image recognition result based on the confidence scores of all categories.
[0081] S8 sends the recognition result of the template image to the eye tracking device, which records the result received this time and compares the current result with the previous result. If the result changes, the recognition result is notified to the user; if there is no change, no operation is performed;
[0082] S9 initializes a single-target tracker based on the Siamese framework. It extracts each foreground frame from the 30FPS foreground video stream transmitted by the eye tracking device and numbers them in ascending order of time. The initial coordinates of the recognition region (x', y', W / 2, H / 2) and the template image are input into the template branch of the single-target tracker. The foreground frame with the smallest number is input into the search branch of the single-target tracker, and the smallest foreground frame is deleted.
[0083] S10 uses the HRnet network to extract features from the template branch and search branch of the single target tracker, and performs deep cross-correlation on the obtained template branch feature map and search branch feature map to obtain a similarity response map.
[0084] As a preferred embodiment of the present invention, step S10 includes the following contents:
[0085] The template branch and search branch of S10.1 use a feature extraction network structure with shared weights to ensure that the template branch and search branch are matched in the same feature space. The backbone network of the high-resolution feature extraction network HRnet with the same structure is used to extract features from the template branch and search branch images respectively to retain more detailed information.
[0086] S10.2 After completing the feature extraction, the obtained template branch feature map and the search branch feature map are subjected to a deep cross-correlation operation. The deep cross-correlation operation is to perform a convolution operation on the template branch feature map and the search branch feature map channel by channel in the channel dimension to obtain a two-dimensional similarity response map, in which the value of each position represents the similarity between the template and the corresponding position of the search area.
[0087] S11 generates multiple anchor points in the area with higher confidence in the similarity response map, performs regression and classification tasks on each anchor point, and generates tracking results;
[0088] As a preferred embodiment of the present invention, step S11 includes the following contents:
[0089] S11.1 sets a threshold based on the percentage of the maximum value based on the obtained similarity response map, marking the area above 80% of the maximum value as a high-confidence area. Multiple anchor points are generated in the high-confidence area of the similarity response map. These anchor points are centered in the high-confidence area and each has a variety of scales and aspect ratios to adapt to the shape and size of the target. These anchor points serve as reference frames for target detection and provide a basis for subsequent regression and classification tasks.
[0090] S11.2 performs regression and classification tasks for each anchor point. The regression task calculates the one with the highest regression score among multiple scales and aspect ratios based on the regression loss function as the size and aspect ratio of the anchor box, making it closer to the target; the classification branch generates a confidence score for each anchor point based on the classification loss function, which indicates the probability of the target contained in the anchor point. The anchor point with the highest score is found based on the score of the classification branch, and its candidate box is used as the tracking result.
[0091] S12 marks the tracking result in the video frame of the foreground image, completing the tracking process of the target image in one frame of the foreground image, and repeating S9 to S12 in the foreground video stream to achieve the tracking effect of the target in the foreground video;
[0092] While tracking the target in the foreground video, S13 monitors whether the two-dimensional coordinates (x', y') of the user's gaze point deviates from the tracking candidate frame for a long time. If the two-dimensional coordinates (x', y') of the gaze point exceed the set time threshold and do not enter the candidate frame area, it is determined that the user has lost attention on the target, and returns to step S5 to reselect the initial coordinates of the recognition area and the template image; if there is no deviation within the time threshold, the previous result is maintained.
[0093] The present invention also provides an eye tracking dynamic recognition device based on three-dimensional gaze points, including an eye state detection device, a two-dimensional gaze point mapping device, a target recognition device, a single target tracking device, a gaze deviation monitoring device, and a recognition result broadcasting device. The eye state detection device screens out three-dimensional gaze points in a staring state according to the three-dimensional gaze points provided by the eye movement device, the gaze time, and the distribution of the gaze points; the two-dimensional gaze point mapping device is responsible for converting the screened three-dimensional gaze points into corresponding two-dimensional coordinates and mapping them to a foreground image corresponding to their time; the target recognition device is responsible for identifying the image of the user's gaze point area through a local recognition library or an online recognition interface to identify the target and extract it as a template image; the single target tracking device is used to track the target and track the target in the foreground image through the extracted template image; the gaze deviation monitoring device is used to combine the information of the tracking target with the user's gaze point information to determine whether the gaze area deviates from the tracking target for a long time; the recognition result broadcasting device is used to broadcast the recognition result to the user through voice and provide a voice broadcast of routine operation prompts.
[0094] As a preferred embodiment of the present invention, Figure 3 As shown, the implementation of the single target tracking device of the present invention includes the following contents:
[0095] 1) Load the improved Siamese model and input the initial coordinates of the recognition region (x', y', W / 2, H / 2) and the template image into the template branch of the single object tracker. Input the foreground video frame with the smallest number into the search branch of the single object tracker and delete the foreground video frame with the smallest number.
[0096] 2) The system integrates the HRnet network to extract features from the template image and the search image respectively to retain more high-resolution detail information.
[0097] 3) Perform a deep cross-correlation operation on the obtained template branch feature map and the search branch feature map to generate a two-dimensional similarity response map, where the value of each position represents the similarity between the template and the corresponding position in the search area.
[0098] 4) Based on the obtained similarity response map, high confidence areas are screened out according to the set threshold, and multiple anchor points are generated in the high confidence areas for subsequent target positioning.
[0099] 5) For each anchor point, regression and classification tasks are performed separately. The regression task determines the size and aspect ratio of the anchor box. The classification task finds the anchor point with the highest score and uses its candidate box as the tracking result.
[0100] 6) Output the final result, continue to input the foreground video frame with the smallest number into the search branch of the single target tracker, and repeat the process from step 1 to step 6.
Claims
1. A method for dynamic target recognition based on eye tracking of three-dimensional gaze points, characterized in that: The steps include: S1 uses eye-tracking equipment to obtain the user's three-dimensional gaze point data, constructs a gaze point-time series, and filters out the user's high-density three-dimensional gaze points; S2 establishes a spatial coordinate system based on the eye tracking device, with the center point of the binocular camera of the user's eye tracking device as the origin, the user's horizontal direction as the X-axis, and the right side as the positive direction of the X-axis; the vertical direction as the Y-axis, and the upward direction as the positive direction of the Y-axis; the front-back direction as the Z-axis, and the front direction as the positive direction of the Z-axis; The position of the center point of the user's eyes in S3 is E, and the three-dimensional gaze point data is (x 注视 ,y 注视 ,z 注视 ), the position of the user's eye center and the three-dimensional gaze point data are substituted into the spatial coordinate system based on the eye movement device, and the sight line equation is established to obtain the user's gaze line function L(x 视线 ,y 视线 ,z 视线 ), where L(x 视线 ,y 视线 ,z 视线 ) is the starting point of the user's eye center point E, and the end point is the three-dimensional gaze point (x 注视 ,y 注视 ,z 注视 ); S4 converts the gaze point line of sight function L(x 视线 ,y 视线 ,z 视线 ) is mapped to the foreground camera coordinate system to obtain the two-dimensional coordinates (x', y') of the user's gaze point; S5 selects a rectangular area with a set length and width of W*H pixels with the two-dimensional coordinates (x', y') of the gaze point as the center, establishes the initial coordinates of the recognition area (x', y', W / 2, H / 2), where x', y' represent the two-dimensional coordinates of the gaze point, W / 2, H / 2 represent the distance from the border of the rectangular area to the center, and sets the image of the W*H pixel rectangular area as a template image for subsequent tracking; S6 uses the Pytorch framework to load the weight file pre-trained in the Yolov8 framework, initializes the template image recognition model, and uses the template image model to extract deep features of the template image. The deep features include low-level features, mid-level features, high-level features, and global features. The global features aggregate the features of the first three layers to generate semantically rich contextual information for subsequent target classification and definition; S7 inputs the extracted deep features of the template image into the head network of the template image recognition model for classification. The head network calculates the probability distribution of each category and obtains the result with the highest confidence as the recognition result of the template image; S8 sends the recognition result of the template image to the eye tracking device, which records the result received this time and compares the current result with the previous result. If the result changes, the recognition result is notified to the user; if there is no change, no operation is performed; S9 initializes a single-target tracker based on the Siamese framework. It extracts each foreground video frame from the 30FPS foreground video stream transmitted by the eye tracking device and numbers them in ascending order of time. It inputs the initial coordinates (x', y', W / 2, H / 2) of the recognition area and the template image into the template branch of the single-target tracker. It inputs the foreground video frame with the smallest number into the search branch of the single-target tracker and deletes the foreground video frame with the smallest number. S10 uses the HRnet network to extract features from the template branch and search branch of the single target tracker, and performs deep cross-correlation on the obtained template branch feature map and search branch feature map to obtain a similarity response map; S11 generates multiple anchor points in the area with higher confidence in the similarity response map, performs regression and classification tasks on each anchor point, and generates tracking results; S12 marks the tracking result in the video frame of the foreground image, completing the tracking process of the target image in one frame of the foreground image, and repeating steps S9 to S12 in the foreground video stream to achieve the tracking effect of the target in the foreground video; While tracking the target in the foreground video, S13 monitors whether the two-dimensional coordinates (x', y') of the user's gaze point deviate from the tracking candidate frame for a long time. If the two-dimensional coordinates (x', y') of the gaze point exceed the set time threshold and do not enter the candidate frame area, it is determined that the user has lost attention on the target, and returns to S5 to reselect the initial coordinates of the recognition area and the template image; if there is no deviation within the time threshold, the previous result is maintained.
2. The method for dynamic target recognition based on three-dimensional gaze point eye tracking according to claim 1, characterized in that: The step S1 includes the following contents: S1.1 Organize the 3D gaze point data into time series according to timestamps; S1.2 uses a density clustering algorithm to filter out high-density gaze points within a time period, and combines time information to establish a set of gaze point-time series (x, y, z, t), where x, y, z represent the high-density three-dimensional gaze points of the filtered user, and t represents the timestamp.
3. The method for dynamic target recognition based on three-dimensional gaze point eye tracking according to claim 2, characterized in that: The step S1.2 includes: S1.2.1 Define the space-time distance calculation formula, α represents the time weight coefficient; S1.2.2 Use the DBSCAN algorithm to process the time series. Based on the temporal distribution of 3D fixation points, determine the ε and MinPts values. Compare the number of eligible points in the neighborhood of each data point with MinPts. ε is the neighborhood radius, which defines the range of the neighborhood around a point. MinPts is the minimum number of neighboring points. If the number of points around a point is greater than or equal to MinPts, the point is considered a core point. S1.2.3 If a point is confirmed to be a core point and there are other points in its neighborhood, these points belong to the same cluster; clusters containing more than the set number threshold are selected, and the points in the cluster represent the selected user's high-density three-dimensional gaze points.
4. The method for dynamic target recognition based on three-dimensional gaze point eye tracking according to claim 1, characterized in that: The step S4 includes the following contents: The S4.1 foreground camera coordinate system has the geometric center of the lens of the external camera mounted on the eye tracker as its origin, and the directions of the X, Y, and Z axes are consistent with the spatial coordinate system of the eye tracking device; S4.2 The user's gaze function L(x 视线 ,y 视线 ,z 视线 ) is converted with the internal parameters of the foreground camera, and the gaze function in the eye movement device coordinate system is converted to the foreground camera coordinate system using the view transformation matrix to obtain the foreground gaze function K(x 前景 ,y 前景 ,z 前景 ), the conversion formula is K(x,y,z)=R*L(x,y,z)+T, Where R is a 3×3 rotation matrix, representing the rotation of the eye tracking device coordinate system relative to the foreground camera coordinate system, and T is a 3×1 translation vector, representing the position of the eye tracking device origin in the foreground camera coordinate system; S4.3 In the foreground camera coordinate system, according to the foreground sight function K(x 前景 ,y 前景 ,z 前景 ), assuming that the foreground image curtain is parallel to the X-axis and Y-axis in the foreground camera coordinate system, and the Z-axis coordinate is g, the foreground sight function K(x 前景 ,y 前景 ,z 前景 ) in z 前景 When the X-axis and Y-axis coordinates are equal to g, they are the two-dimensional coordinates (x', y') of the user's three-dimensional gaze point on the foreground image.
5. The method for dynamic target recognition based on three-dimensional gaze point eye tracking according to claim 1, characterized in that: The step S7 includes the following contents: S7.1 passes the extracted deep features to the head network of the template image recognition model, and uses PAnet to further aggregate and optimize multi-layer features to ensure that the model takes into account both global and local features; S7.2 outputs the confidence score of each category based on the multi-layer features after aggregation according to the classification loss function of yolov8; S7.3 selects the category with the highest confidence as the final template image recognition result based on the confidence scores of all categories.
6. The method for dynamic target recognition based on three-dimensional gaze point eye tracking according to claim 1, characterized in that: The step S10 includes the following contents: The template branch and search branch of S10.1 use a feature extraction network structure with shared weights to ensure that the template branch and search branch are matched in the same feature space. The backbone network of the high-resolution feature extraction network HRnet with the same structure is used to extract features from the template branch and search branch images respectively to retain more detailed information. S10.2 After completing the feature extraction, the obtained template branch feature map and the search branch feature map are subjected to a deep cross-correlation operation. The deep cross-correlation operation is to perform a convolution operation on the template branch feature map and the search branch feature map channel by channel in the channel dimension to obtain a two-dimensional similarity response map, in which the value of each position represents the similarity between the template and the corresponding position of the search area.
7. The method for dynamic target recognition based on three-dimensional gaze point eye tracking according to claim 1, characterized in that: The step S11 includes the following contents: S11.1 sets a threshold based on the percentage of the maximum value based on the obtained similarity response map, marking the area above 80% of the maximum value as a high-confidence area. Multiple anchor points are generated in the high-confidence area of the similarity response map. These anchor points are centered in the high-confidence area and each has a variety of scales and aspect ratios to adapt to the shape and size of the target. These anchor points serve as reference frames for target detection and provide a basis for subsequent regression and classification tasks. S11.2 performs regression and classification tasks for each anchor point. The regression task calculates the one with the highest regression score among multiple scales and aspect ratios based on the regression loss function as the size and aspect ratio of the anchor box, making it closer to the target; the classification branch generates a confidence score for each anchor point based on the classification loss function, which indicates the probability of the target contained in the anchor point. The anchor point with the highest score is found based on the score of the classification branch, and its candidate box is used as the tracking result.
8. A dynamic target recognition device based on eye tracking of three-dimensional gaze points, characterized in that: The system comprises an eye state detection device, a two-dimensional gaze point mapping device, a target recognition device, a single target tracking device, a gaze deviation monitoring device, and a recognition result broadcasting device. The eye state detection device selects three-dimensional gaze points in a gaze state according to the three-dimensional gaze points provided by the eye movement device, the gaze time, and the distribution of the gaze points; the two-dimensional gaze point mapping device is responsible for converting the selected three-dimensional gaze points into corresponding two-dimensional coordinates and mapping them to the foreground image corresponding to the time; the target recognition device is responsible for identifying the target in the image of the user's gaze point area through a local recognition library or an online recognition interface, and extracting it as a template image; The single target tracking device is used to track the target in the foreground image through the extracted template image; The gaze deviation monitoring device is used to combine the information of the tracking target and the user's gaze point information to determine whether the gaze area deviates from the tracking target for a long time; the recognition result broadcasting device is used to broadcast the recognition result to the user through voice and provide regular operation prompt voice broadcast.
Citation Information
Cited By
Focus depth determination method and related device
CN120807743A