Human factor intelligent user gaze analysis method, device and system, edge computing device

By collecting visual and eye-tracking data through a head-mounted device, identifying and displaying target eye-tracking data, the system solves the problem of low efficiency in traditional eye-tracking systems, and enables real-time and efficient analysis of user visual behavior and in-depth understanding of multi-user behavior.

CN119148860BActive Publication Date: 2026-04-17KINGFAR INTERNATIONAL INC
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
KINGFAR INTERNATIONAL INC
Filing Date
2024-11-11
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Traditional eye-tracking systems are inefficient due to manual annotation and cannot effectively analyze user visual behavior.

Method used

Visual and eye-tracking data are collected by a head-mounted device to identify target objects and display their eye-tracking data in real time. Two-dimensional labeling and feature extraction technologies are used for precise positioning and mapping.

Benefits of technology

It enables real-time analysis of user visual behavior, improves efficiency and enhances user interaction experience, and supports multi-user behavior analysis and interface optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119148860B_ABST
    Figure CN119148860B_ABST
Patent Text Reader

Abstract

This application discloses a human-centric intelligent user gaze analysis method, apparatus, system, and edge computing device, belonging to the field of computer vision technology. The method includes: acquiring visual data of the user's field of vision via a camera of a head-mounted device, and acquiring eye movement data of the user within the field of vision via an eye tracker of the head-mounted device; identifying target objects in the visual data; determining target eye movement data associated with gaze on the target object within the eye movement data; and sending the target eye movement data to a target screen so that the target eye movement data is correspondingly displayed on the target object shown on the target screen. The embodiments of this application improve the efficiency of user visual behavior analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of computer vision technology, and in particular relates to a human factors intelligent user gaze analysis method, device and system, and edge computing device. Background Technology

[0002] Eye-tracking technology is widely used in fields such as psychology, medicine, advertising analysis, and autonomous driving. It tracks users' eye movements to understand their attention distribution, decision-making behavior, and reaction speed.

[0003] Traditional eye-tracking systems capture the movement of a user's eyes using cameras mounted on a monitor or device, analyzing the user's visual focus. However, this technology relies heavily on post-processing data analysis, involving manual annotation by personnel to map eye-tracking information onto the target screen, resulting in low efficiency in analyzing user visual behavior. Summary of the Invention

[0004] This application aims to address at least one of the technical problems existing in the prior art. To this end, this application proposes a human factors intelligent user gaze analysis method, apparatus and system, and edge computing device to improve the efficiency of user visual behavior analysis.

[0005] Firstly, this application provides a human-centric intelligent user gaze analysis method, including:

[0006] The head-mounted device collects visual data of the user's field of vision via its camera and eye-tracking device collects eye-tracking data of the user within that field of vision via its eye tracker.

[0007] Identify the target object in the visual data;

[0008] Identify the target eye movement data associated with the gaze of the target object from the eye movement data;

[0009] The target eye-tracking data is sent to the target screen so that the target eye-tracking data is displayed on the target object on the target screen.

[0010] According to the human-centric intelligent user gaze analysis method of this application, visual data of the user's field of vision is collected via a camera of a head-mounted device, and eye movement data of the user within the field of vision is collected via an eye tracker of the same head-mounted device. The method identifies target objects in the visual data; determines target eye movement data associated with the gaze of the target object; and sends the target eye movement data to a target screen so that the target eye movement data is displayed on the target object on the target screen. This embodiment of the application achieves effective identification of the target object being gazed upon by the user by collecting visual data and eye movement data, and displays the user's target eye movement data for the target object on the target screen. This application eliminates the need for manual annotation of the user's target gaze data regarding the target object, allowing for real-time display of the user's target eye movement data on the target screen for real-time analysis of user visual behavior. This enhances the user's interactive experience and improves the efficiency of user visual behavior analysis.

[0011] According to one embodiment of this application, the method further includes:

[0012] Determine the two-dimensional identifier to be used based on the scenario;

[0013] The two-dimensional markers are deployed on the key points of the target object, and the correspondence between the target object, key points, and two-dimensional markers is recorded;

[0014] The identification of target objects in the visual data includes:

[0015] The position of the target object in the visual data is located based on the position of the two-dimensional identifier in the visual data and the corresponding relationship.

[0016] In this embodiment, by determining appropriate two-dimensional markers according to the needs of the scenario, the selection of two-dimensional markers can be changed according to the scenario. This ensures that the two-dimensional markers have sufficient recognizability and information capacity in the visual data under different scenarios. Furthermore, two-dimensional markers are deployed on key points of the target object to record the corresponding relationships, providing necessary reference information for subsequent positioning. As a result, the position of the target object can be accurately calculated based on the position of the two-dimensional markers in the visual data, thus improving the accuracy of positioning.

[0017] According to one embodiment of this application, determining the two-dimensional identifier to be used based on the scenario includes:

[0018] The size of the white space in the two-dimensional logo is determined according to the scenario; wherein, the white space is the area between the edge of the two-dimensional logo and the background area of ​​the two-dimensional logo;

[0019] The two-dimensional identifier to be used is determined based on the size of the blank area.

[0020] In this embodiment, by determining the appropriate size of the blank area according to the scenario, the risk of misidentification can be reduced because the blank area provides additional visual buffer, helps the two-dimensional sign to be distinguished from the surrounding environment, and improves the accuracy of two-dimensional sign recognition.

[0021] According to one embodiment of this application, determining the two-dimensional identifier to be used based on the scenario includes:

[0022] Determine the area ratio of the background region to the two-dimensional label in the two-dimensional label based on the scenario described;

[0023] The two-dimensional identifier to be used is determined based on the area ratio.

[0024] In this embodiment, by precisely controlling the area ratio of the background region to the marker itself in the 2D marker, the marker is neither too prominent, creating an overly strong contrast with the background, nor too concealed, making it difficult to detect. This improves the efficiency of the algorithm when processing images, as it can lock onto the target marker more quickly and reduce the processing of irrelevant background information. Furthermore, an appropriate area ratio also enhances the marker's stability under different lighting and viewing angle conditions, enabling reliable target localization even in varying environmental conditions.

[0025] According to one embodiment of this application, the target object includes multiple key points, and deploying the two-dimensional markers on the key points of the target object includes:

[0026] Select multiple non-collinear key points on the target object; deploy the two-dimensional markers on the multiple non-collinear key points respectively, and record the correspondence between the target object, key points and two-dimensional markers;

[0027] Alternatively, select multiple key points from the target object to form a polygon; deploy the two-dimensional markers on the multiple key points forming the polygon, and record the correspondence between the target object, key points, and two-dimensional markers.

[0028] In this embodiment, by selecting multiple non-collinear key points on the target object and deploying two-dimensional markers on each, the geometric distribution advantage of the non-collinear key points is utilized. Even if some markers are unrecognizable due to occlusion or damage, the target object can still be located using other visible markers, thereby reducing the impact of a single marker failure on the overall positioning accuracy. By utilizing the geometric properties of polygons to provide a stable reference frame for positioning, with the vertices of the polygon (i.e., key points) serving as the deployment locations for two-dimensional markers, even if some markers are temporarily unrecognizable due to environmental factors (such as occlusion or changes in lighting), effective positioning can still be achieved based on other visible markers, improving positioning accuracy.

[0029] According to one embodiment of this application, locating the target object based on the position of the two-dimensional identifier in the visual data and the correspondence includes:

[0030] Feature extraction is performed on the visual data; based on the extracted features, the two-dimensional identifiers included in the visual data and their positions in the visual data are detected; based on the correspondence and the detection results of the two-dimensional identifiers and their positions in the visual data, the target object in the visual data is located.

[0031] Alternatively, in response to detecting that the number of two-dimensional icons in the visual data is less than the number of deployed two-dimensional icons, the positions of the undetected two-dimensional icons are fitted based on the geometric positional relationships of the key points;

[0032] The target object is located based on the position of the two-dimensional markers detected in the visual data, the position of the undetected two-dimensional markers obtained by fitting, and the corresponding relationship.

[0033] In this embodiment, feature extraction from visual data allows for the identification and detection of two-dimensional (2D) markers within the visual data. By combining these markers with their corresponding relationships, the target object can be accurately located. Since the number of detected 2D markers in the visual data is less than the number of deployed 2D markers, the location information of the detected 2D markers and the geometric relationships of key points can be used to predict the possible locations of undetected 2D markers through fitting. This allows for the reconstruction of the 2D marker layout even in the absence of markers, enabling effective target object localization and improving accuracy.

[0034] According to one embodiment of this application, the target screen includes a field of view partition and a target object partition;

[0035] The visual field partitions show the positions of the visual data and the eye movement data within the visual data;

[0036] The target object partition displays the location of the target object and the target eye-tracking data on the target object.

[0037] In this embodiment, by dividing the target screen into a visual field partition and a target object partition, the visual field partition allows for a visual display of the relative position of the user's eye movement data within their entire visual field, which helps in analyzing the distribution of the user's visual focus and visual path. The target object partition further focuses on the specific target the user is looking at. By displaying the specific position of the target eye movement data on the target object, it is possible to more accurately understand the user's attention allocation and attention details to the target object. This not only enhances the visual analysis of the user's gaze behavior but also provides guidance for designing user interfaces, optimizing ad layouts, and improving user experience.

[0038] According to one embodiment of this application, sending the target eye-tracking data to a target screen to display the target eye-tracking data correspondingly on the target object displayed on the target screen includes:

[0039] Acquire the positional information of the target eye-tracking data on the target object;

[0040] The visual data is identified to obtain the position information of each marker pre-marked on the target object;

[0041] The conversion relationship is determined based on the location information of each marker;

[0042] Substituting the position information of the target eye-tracking data on the target object into the transformation relationship, the target coordinates of the target eye-tracking data on the target object in the target screen are obtained.

[0043] In this embodiment, by recognizing visual data, the positional information of each identified marker helps to locate the specific position of the target object and the marker area containing the target eye-tracking data. Obtaining the positional information of the target eye-tracking data helps to accurately map the target eye-tracking data onto the target screen. The positional information of the markers can establish a mapping relationship between the target object and the target screen, which helps to ensure that the target object and the target screen can be accurately aligned during the mapping process. Different marker layouts can adapt to different screen sizes and shapes. After determining the conversion relationship, the positional information of the target eye-tracking data can be mapped onto the target screen in real time, supporting interaction with the target object and seeing the corresponding response on the target screen, thus improving the real-time performance of target point mapping.

[0044] According to one embodiment of this application, determining the conversion relationship based on the location information of each identifier includes:

[0045] Based on the location information of each marker, a first set of parameters is determined for calculating the abscissa of the target coordinates, a second set of parameters is determined for calculating the ordinate of the target coordinates, and a third set of parameters is determined for calculating the homogeneous coordinate normalization factor.

[0046] The expression for the first intermediate variable is determined based on the third parameter set;

[0047] Based on the expression of the first intermediate variable and the first parameter set, a first expression for calculating the x-coordinate of the target coordinate is obtained;

[0048] Based on the expression of the first intermediate variable and the second parameter set, a second expression is obtained for calculating the ordinate of the target coordinates. The first expression and the second expression constitute the transformation relationship.

[0049] In this embodiment, the parameter sets of normalization factors for the horizontal, vertical, and homogeneous coordinates are determined respectively, which can more accurately describe the mapping relationship from the target object to the target screen. Introducing intermediate variables can simplify the final transformation relationship expression, making the calculation process more efficient. By combining the horizontal and vertical coordinate expressions, a complete transformation relationship is formed, ensuring the consistency of the target eye-tracking data position between the target object and the target screen. After the transformation relationship is determined, it can be reused multiple times, improving the real-time performance and reusability of the system.

[0050] According to one embodiment of this application, the number of markers pre-marked on the target object is four, and the corresponding position information is a first coordinate, a second coordinate, a third coordinate, and a fourth coordinate, respectively;

[0051] The step of determining a first set of parameters for calculating the abscissa of the target coordinates, a second set of parameters for calculating the ordinate of the target coordinates, and a third set of parameters for calculating the homogeneous coordinate normalization factor based on the position information of each marker includes:

[0052] Calculate the second intermediate variable based on the second coordinate, the third coordinate, and the fourth coordinate;

[0053] Calculate the third intermediate variable based on the second intermediate variable, the first coordinate, the second coordinate, and the fourth coordinate;

[0054] The fourth intermediate variable is calculated based on the second intermediate variable, the first coordinate, the third coordinate, and the fourth coordinate;

[0055] Based on the first coordinate, the second coordinate, the third coordinate, the third intermediate variable, and the fourth intermediate variable, determine the first parameter set and the second parameter set;

[0056] The third parameter set is calculated based on the third intermediate variable and the fourth intermediate variable.

[0057] In this embodiment, a stable framework can be formed by four markers, reducing mapping errors caused by differences in the shape, size, resolution, and other factors of the target objects. It can adapt to different types of configurations and has good versatility. The calculation process is based on explicit mathematical expressions and intermediate variables, making it easy to implement.

[0058] According to one embodiment of this application, determining the first parameter set and the second parameter set based on the first coordinate, the second coordinate, the third coordinate, the third intermediate variable, and the fourth intermediate variable includes:

[0059] The product of the difference between the abscissas of the third coordinate and the first coordinate and the third intermediate variable is used as the first parameter; the product of the difference between the abscissas of the second coordinate and the first coordinate and the fourth intermediate variable is used as the second parameter; and the abscissa of the first coordinate is used as the third parameter. The first parameter, the second parameter, and the third parameter constitute the first parameter set.

[0060] The product of the difference between the ordinates of the third coordinate and the first coordinate and the third intermediate variable is used as the fourth parameter, the product of the difference between the ordinates of the second coordinate and the first coordinate and the fourth intermediate variable is used as the fifth parameter, and the ordinate of the first coordinate is used as the sixth parameter. The fourth parameter, the fifth parameter, and the sixth parameter constitute the second parameter set.

[0061] In this embodiment, by calculating parameters based on the position information of four markers and through specific mathematical operations such as product and difference, accurate mapping results can be provided. The difference operation helps to capture the relative positional relationship between different markers. Using multiple markers as reference points and calculating parameter sets based on these reference points can improve the stability of the system. Even if there is an error in the position detection of one of the markers, it can be corrected by the information of other markers, thereby reducing the impact of the error on the mapping results.

[0062] According to one embodiment of this application, the step of identifying the visual data to obtain the position information of each marker pre-marked on the target object includes:

[0063] Identify the identification area of ​​each pre-marked object on the target object from the visual data;

[0064] Determine the coordinates of the center point of each marked area in the coordinate system established based on the visual data to obtain the position information of each marker;

[0065] Determining the center point coordinates of each identified area includes:

[0066] For each identified region in the visual data, determine the coordinates of each point in the outline of the identified region, and calculate the average coordinates of each point in the outline of the identified region to obtain the coordinates of the center point of the identified region.

[0067] In this embodiment, by utilizing an image-based coordinate system, the marker can be accurately located, ensuring the precision of the location information. Identifying the marker's area and determining the center point coordinates accurately reflects the marker's actual position in the image. By calculating the average coordinates of all points on the contour as the center point coordinates, the overall shape and size of the marker area are considered, resulting in a more accurate reflection of the marker's actual position. Compared to methods that only consider some points or specific points on the contour, the average coordinate method reduces errors and improves the accuracy of location information.

[0068] According to one embodiment of this application, sending the target eye-tracking data to a target screen to display the target eye-tracking data correspondingly on the target object displayed on the target screen includes:

[0069] The first video frame in which the visual data is acquired includes a first image block containing the target eye-tracking data;

[0070] In response to the first image block satisfying a preset condition, the coordinate mapping relationship between the first video frame and the second video frame in the target screen is calculated based on the first image block;

[0071] The target eye-tracking data is mapped to the second video frame according to the coordinate mapping relationship.

[0072] In this embodiment, by acquiring image blocks containing target eye-tracking data, the coordinate mapping relationship between the first video frame and the second image frame is calculated when the image blocks meet preset conditions. This makes the calculation process focus more on the region in the image related to the target eye-tracking data, rather than the entire first video frame, reducing the impact of environmental changes on the calculation process, improving the accuracy of coordinate mapping relationship calculation, and thus improving the accuracy of eye tracking.

[0073] According to one embodiment of this application, the first video frame in which the visual data is acquired includes a first image block containing the target eye-tracking data, comprising:

[0074] Determine the position of the target eye-tracking data in the first video frame;

[0075] The region within a first preset range centered on the position of the target eye-tracking data in the first video frame is defined as the first image block.

[0076] In this embodiment, by accurately identifying the location of the target eye-tracking data and determining the first image block based on the location of the target eye-tracking data, irrelevant areas in the processed image can be reduced, allowing the calculation process to focus on the area near the target eye-tracking data, thereby reducing invalid calculations in irrelevant areas and improving calculation efficiency and accuracy.

[0077] According to one embodiment of this application,

[0078] The preset conditions include: the number of matching points in the image block is greater than or equal to a preset number, and / or multiple matching points in the image block are not collinear; and / or,

[0079] The method further includes: in response to the first image block not meeting a preset condition, acquiring a second image block in the first video frame that includes the target eye-tracking data; and / or,

[0080] The method further includes: in response to the second image block satisfying a preset condition, calculating the coordinate mapping relationship between the first video frame and the second video frame in the target screen based on the second image block.

[0081] In this embodiment, by setting a preset condition requiring the number of matching points to be greater than or equal to a preset number, and these matching points to be non-collinear, there are enough matching points to calculate the coordinate mapping relationship. The condition that multiple matching points are non-collinear allows the matching points to provide enough geometric information to solve the coordinate mapping relationship between the first video frame and the second video frame, further improving the accuracy of target eye-tracking data mapping.

[0082] In this embodiment, by acquiring a second image block containing the target eye-tracking data in the first video frame when the first image block does not meet the preset conditions, even if a sufficient number or appropriately distributed matching points are not found within the range of the first image block, the possibility of finding effective matching points can be increased by adjusting the search area, thereby enhancing the flexibility and success rate of the target eye-tracking data mapping process.

[0083] According to one embodiment of this application, obtaining a second image block in the first video frame that includes the target eye-tracking data includes:

[0084] The second image block is obtained based on the position of the first image block in the first video frame; or...

[0085] The position of the target eye-tracking data in the first video frame is determined; the area within a second preset range centered on the position of the target eye-tracking data in the first video frame is determined as the second image block; wherein the second preset range is larger than the first preset range.

[0086] In this embodiment, by obtaining the second image block based on the position of the first image block in the first video frame, the new search range can be redefined with the first image block as a reference, thereby improving the efficiency of the target eye-tracking data mapping calculation process.

[0087] In this embodiment, by expanding the first image block into a second image block, even if a sufficient number or appropriately distributed number of matching points are not found within the initial search range, the possibility of finding effective matching points can be increased by expanding the search area, thereby enhancing the flexibility and success rate of the target eye-tracking data mapping process.

[0088] According to one embodiment of this application, before the first image block containing the target eye-tracking data is included in the first video frame of the visual data acquisition, the process includes:

[0089] Feature point matching is performed on the first video frame and the second video frame in the target screen. In response to the failure of matching between the first video frame and the second video frame, a new video frame is selected from the visual data to replace the first video frame and the second video frame for feature point matching.

[0090] In this embodiment, when feature point matching fails to establish an accurate correspondence between the first video frame and the second video frame, a new video frame can be flexibly selected from the continuous video stream as the first video frame, reducing the risk of interruption of the entire eye-tracking process due to a single matching failure and improving the continuity of the system.

[0091] According to one embodiment of this application, the step of reselecting video frames from the visual data to replace the first video frame and the second video frame for feature point matching includes:

[0092] The video frame adjacent to the first video frame in the visual data, or other video frames in the visual data, are identified as the reselected video frame.

[0093] In this embodiment, when feature point matching between the first video frame and the second video frame fails, a video frame adjacent to the current first video frame is selected from the target video stream as a replacement to re-perform feature point matching. This ensures that the re-selected video frame maintains temporal continuity, or other video frames are selected as replacements, thereby improving the stability and accuracy of the eye-tracking process.

[0094] According to one embodiment of this application, the feature point matching of the first video frame and the second video frame in the target screen includes: inputting the first video frame and the second video frame into a preset neural network model so that the neural network model can extract feature points of the first video frame and the second video frame, and perform matching based on the extracted feature points;

[0095] or,

[0096] The step of performing feature point matching on the first video frame and the second video frame in the target screen includes: preprocessing the first video frame and the second video frame; the preprocessing includes removing feature points within a preset range of the boundary of the first video frame and removing feature points within a preset range of the boundary of the second video frame; performing feature point matching on the preprocessed first video frame and the second video frame; and / or inputting the preprocessed first video frame and the second video frame into a preset neural network model so that the neural network model can extract feature points from the preprocessed first video frame and the second video frame and perform matching based on the extracted feature points.

[0097] In this embodiment, by using a preset neural network model to match feature points of the first and second video frames, the powerful feature extraction capabilities of deep learning can be utilized to improve the accuracy and efficiency of the matching.

[0098] In this embodiment, by removing feature points within a preset range of the image boundary, the negative impact of noise or incompleteness that may exist in the image boundary region on the matching accuracy can be reduced, allowing more focus to be placed on more stable and information-rich areas in the image, thereby improving the accuracy of the matching.

[0099] According to one embodiment of this application, the method further includes:

[0100] Store the feature points of the second video frame extracted by the neural network model;

[0101] Reselecting video frames from the visual data to replace the first and second video frames for feature point matching includes:

[0102] The replaced first video frame is input into a preset neural network model so that the neural network model can extract the feature points of the first video frame and match them with the feature points of the stored second video frame.

[0103] In this embodiment, features of the second video frame are extracted and stored using a neural network model. When other video frame images are matched with the second video frame, the feature points of the second video frame can be called for feature matching, eliminating the need to repeatedly extract the features of the second video frame and improving processing efficiency.

[0104] According to one embodiment of this application, the step of calculating the coordinate mapping relationship between the first video frame and the second video frame in the target screen based on the first image block includes:

[0105] Obtain the first position coordinates of the matching point in the first image block in the first video frame and the second position coordinates of the matching point in the second video frame;

[0106] The homography matrix corresponding to the matching point is calculated based on the first position coordinates and the second position coordinates; wherein the homography matrix represents the coordinate mapping relationship.

[0107] In this embodiment, by calculating the corresponding position coordinates of the matching points in the first and second video frames, a mathematical model can be established to describe the spatial relationship between these matching points, and a homography matrix can be determined to represent the coordinate mapping relationship. Using the homography matrix to represent the coordinate mapping relationship can adapt to complex image transformations, thereby improving the accuracy and robustness of target eye-tracking data mapping.

[0108] According to one embodiment of this application, the method further includes:

[0109] Acquire target eye-tracking data associated with multiple users' gazes at the same target object;

[0110] Multi-user fixation behavior is analyzed based on target eye-tracking data from multiple users.

[0111] In this embodiment, the ability to analyze the behavior of multiple users is improved by analyzing the target eye-tracking data of multiple users.

[0112] According to one embodiment of this application, the step of analyzing multi-user gaze behavior based on target eye-tracking data corresponding to multiple users includes:

[0113] Multiple user-specific eye-tracking data are sent to the target screen to overlay the multiple user-specific eye-tracking data onto the target object displayed on the target screen.

[0114] In this embodiment, by overlaying and displaying the target eye-tracking data of multiple users on the target screen, researchers and designers can intuitively observe and compare the common fixation points and fixation patterns of different users when looking at the same target object, thereby identifying the common interest areas or hot spots of the user group.

[0115] According to one embodiment of this application, the step of analyzing multi-user gaze behavior based on target eye-tracking data corresponding to multiple users includes:

[0116] Based on the target eye movement data of multiple users, analyze the eye movement trajectory and / or eye movement point heatmap of multiple users to obtain the primary viewing position of multiple users and the habitual operation process of multiple users when performing preset operations;

[0117] The scene is optimized based on the primary viewing position of multiple users and their habitual operation flow when performing preset operations.

[0118] In this embodiment, by analyzing the eye movement trajectories and eye movement heatmaps of multiple users, it is possible to reveal the user's primary viewing position and habitual operation process, thereby achieving a deep understanding of the user's behavior patterns. This embodiment provides a quantitative means to evaluate and optimize the user's interaction experience with products or services. By identifying common gaze points and gaze patterns, designers can optimize the interface layout so that key information and functions conform to the user's natural eye movement and operating habits, thereby improving user satisfaction and operating efficiency.

[0119] Secondly, this application provides a human-centric intelligent user gaze analysis device, comprising:

[0120] The acquisition module is used to acquire visual data of the user's field of vision via the camera of the head-mounted device, and to acquire eye movement data of the user within the field of vision via the eye tracker of the head-mounted device.

[0121] The recognition module is used to identify target objects in the visual data;

[0122] The determination module is used to determine the target eye movement data associated with the gaze of the target object in the eye movement data;

[0123] The display module is used to send the target eye-tracking data to the target screen so that the target eye-tracking data is displayed on the target object on the target screen.

[0124] According to the human-centric intelligent user gaze analysis device of this application, visual data of the user's field of vision is collected via a camera of a head-mounted device, and eye movement data of the user within the field of vision is collected via an eye tracker of the same head-mounted device. The device identifies target objects within the visual data; determines target eye movement data associated with the gaze of the target object; and sends the target eye movement data to a target screen so that the target eye movement data is displayed on the target object on the target screen. This embodiment of the application achieves effective identification of the target object being gazed upon by the user by collecting visual data and eye movement data, and displays the user's target eye movement data for the target object on the target screen. This application eliminates the need for manual annotation of the user's target gaze data regarding the target object, allowing for real-time display of the user's target eye movement data on the target screen for real-time analysis of user visual behavior. This enhances the user's interactive experience and improves the efficiency of user visual behavior analysis.

[0125] Thirdly, this application provides an edge computing device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the human factors intelligent user gaze analysis method as described in the first aspect above.

[0126] According to one embodiment of this application, the edge computing device includes a head-mounted edge computing device; the head-mounted edge computing device includes a camera and an eye tracker.

[0127] Fourthly, this application provides a human-centric intelligent user gaze analysis system, comprising: a target screen and an edge computing device as described in the third aspect above.

[0128] Fifthly, this application provides a computer-readable storage medium storing a program or instructions that, when executed by a processor, implement the human factors intelligence user gaze analysis method as described in the first aspect above.

[0129] In a sixth aspect, this application provides a chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the human factors intelligent user gaze analysis method as described in the first aspect above.

[0130] In a seventh aspect, this application provides a computer program product, including a computer program that, when executed by a processor, implements the human factors intelligence user gaze analysis method as described in the first aspect above.

[0131] The above-described one or more technical solutions in the embodiments of this application have at least one of the following technical effects:

[0132] According to the human-centric intelligent user gaze analysis method of this application, visual data of the user's field of vision is collected via a camera of a head-mounted device, and eye movement data of the user within the field of vision is collected via an eye tracker of the same head-mounted device. The method identifies target objects in the visual data; determines target eye movement data associated with the gaze of the target object; and sends the target eye movement data to a target screen so that the target eye movement data is displayed on the target object on the target screen. This embodiment of the application achieves effective identification of the target object being gazed upon by the user by collecting visual data and eye movement data, and displays the user's target eye movement data for the target object on the target screen. This application eliminates the need for manual annotation of the user's target gaze data regarding the target object, allowing for real-time display of the user's target eye movement data on the target screen for real-time analysis of user visual behavior. This enhances the user's interactive experience and improves the efficiency of user visual behavior analysis.

[0133] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0134] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0135] Figure 1 This is a flowchart illustrating the human factors intelligent user gaze analysis method provided in the embodiments of this application;

[0136] Figure 2 This is a schematic diagram of the two-dimensional identifier provided in the embodiments of this application;

[0137] Figure 3 This is a schematic diagram of the field of view partitioning and target object partitioning provided in the embodiments of this application;

[0138] Figure 4 This is a schematic diagram of the target eye-tracking data mapping process provided in an embodiment of this application;

[0139] Figure 5 This is a schematic diagram of the structure of the human factors intelligent user gaze analysis device provided in the embodiments of this application;

[0140] Figure 6 This is a schematic diagram of the structure of the edge computing device provided in the embodiments of this application. Detailed Implementation

[0141] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0142] The terms "first," "second," etc., used in this application's specification are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class, without limiting the number of objects; for example, a first object can be one or more. Furthermore, in the specification, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects have an "or" relationship.

[0143] It should be noted that, in the optional embodiments of this application, the data related to object information, when applied to specific products or technologies, requires the permission or consent of the object. Furthermore, the collection, use, and processing of this data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. In other words, if the embodiments of this application involve data related to an object, it must be obtained with the object's authorization and consent, the authorization and consent of relevant departments, and in accordance with the relevant laws, regulations, and standards of the country and region. If the embodiments involve personal information, the acquisition of all personal information requires the individual's consent. If sensitive information is involved, the separate consent of the information subject is required. The embodiments also need to be implemented with the object's authorization and consent.

[0144] The human-centric intelligent user gaze analysis method, apparatus, system, and edge computing device provided in this application will be described in detail below with reference to the accompanying drawings and through specific embodiments and application scenarios.

[0145] Among them, the human factors intelligent user gaze analysis method can be applied to the terminal, specifically executed by the hardware or software in the terminal.

[0146] The terminal includes, but is not limited to, portable communication devices such as mobile phones or tablets with touch-sensitive surfaces (e.g., touchscreen displays and / or touchpads). It should also be understood that, in some embodiments, the terminal may not be a portable communication device, but rather a desktop computer with touch-sensitive surfaces (e.g., touchscreen displays and / or touchpads).

[0147] The following embodiments describe a terminal including a display and a touch-sensitive surface. However, it should be understood that the terminal may include one or more other physical user interface devices such as a physical keyboard, mouse, and joystick.

[0148] The human-centric intelligent user gaze analysis method provided in this application embodiment can be executed by an edge computing device or a functional module or entity within an edge computing device that can implement the human-centric intelligent user gaze analysis method. The edge computing device mentioned in this application embodiment can be a portable device (such as a mobile phone or tablet computer), a wearable device (such as a smartwatch or smart bracelet), an eye tracker, an AR (Augmented Reality) / VR (Virtual Reality) / XR (Extended Reality) / MR (Mixed Reality) device, an edge server, an edge gateway, an in-vehicle device, a cloud computing device, etc. The human-centric intelligent user gaze analysis method provided in this application embodiment will be described below using an edge computing device as the execution subject as an example.

[0149] like Figure 1 As shown, the human-based intelligent user gaze analysis method includes steps 110, 120, 130, and 140.

[0150] Step 110: Visual data of the user's field of vision is collected via the camera of the head-mounted device, and eye movement data of the user within the field of vision is collected via the eye tracker of the head-mounted device.

[0151] In this embodiment, visual data of the user's field of vision is collected via the camera of the head-mounted device. This visual data corresponds to what the user's eyes see, i.e., what is within the user's field of vision. The visual data can be images or videos. Eye-tracking data can be data related to the user's eye gaze, such as eye gaze point, gaze direction, and gaze trajectory.

[0152] In this embodiment, eye-tracking data can be data related to the user's eye movements captured by an eye tracker in a head-mounted device, while visual data can be data obtained by capturing objects within the user's field of vision through a camera in a head-mounted device. For example, the eye tracker can be a wearable eye tracker, a screen-based eye tracker, or a head-mounted eye tracker. The eye tracker can capture the real-time state of the eyes, such as pupil size and corneal reflection, to obtain eye-tracking data, while the camera can capture the real-time state of objects within the field of vision, thus obtaining visual field data.

[0153] In one scenario, visual and eye-tracking data can be collected using a head-mounted device. Specifically, the head-mounted device can be equipped with a camera and an eye tracker. The camera is responsible for collecting image or video data within the user's field of vision, while the eye tracker simultaneously collects eye-tracking data such as the user's fixation point, eye movement trajectory, and fixation duration within the field of vision.

[0154] Step 120: Identify the target object in the visual data.

[0155] In this application embodiment, the human-centric intelligent user gaze analysis method provided can be applied to various scenarios, such as user gaze analysis in augmented reality, industrial automation, robot navigation, driving simulators, advertising and marketing, education and entertainment, and user experience scenarios. Of course, it can also be applied to other scenarios, and this application embodiment does not limit its application to these. Depending on the scenario, the target object can be different objects. For example, in a driving simulator scenario, the target object can be the screen; in a multi-turn robotic arm recognition scenario, the target object can be the robotic arm; in a robot navigation scenario, the target object can be a turning point, intersection, or specific task area on the navigation path; and in an advertising and marketing scenario, the target object can be a product or an advertisement for a product. This application embodiment does not limit the target object.

[0156] The target object is the object of gaze analysis. For example, in a simulated driving scenario, the target object is the central control screen or different areas of the central control screen. Users can perform some operations by gazing at different areas of the screen, such as turning on the air conditioner, playing music, or adjusting the driving mode.

[0157] In this embodiment, conventional image processing techniques can be used to identify target objects from visual data. For example, visual data can be prepared through preprocessing steps such as grayscale conversion and filtering to remove noise. Then, image features of the visual data can be extracted using algorithms such as edge detection and corner detection. Image segmentation techniques such as thresholding or region growing can be used to separate the target object from the background. Finally, feature matching and classical classifiers, such as Support Vector Machine (SVM) or Adaboost, can be used to identify and classify the target object, thereby identifying the target object in the visual data and determining its position in the visual field. Alternatively, a recognition model can be pre-trained using deep learning algorithms and then deployed in practical applications. By inputting visual data into the recognition model, the target object in the visual data and its position in the visual field can be obtained from the model's output. Of course, target objects in visual data can also be identified in any other way, and this embodiment does not limit this method.

[0158] Step 130: Identify the target eye movement data associated with the fixation of the target object in the eye movement data.

[0159] In this embodiment, an eye tracker can capture the position and movement of the user's eyes. By utilizing the reflection generated when light shines into the eyes, the reflection features of the pupil and cornea can be extracted. Through analysis of the pupil position and corneal reflection points, the direction of the gaze can be calculated. Using geometric principles, combined with eye structure and light paths, the specific location of the eye movement data in the visual data can be deduced.

[0160] After identifying the target object in the visual data, the position of the target object in the visual data can be determined. Based on the position of the eye movement data in the visual data and the position of the target object in the visual data, it can be determined whether the eye movement data is on the target object. If the eye movement data is on the target object, then the eye movement data can be identified as the target eye movement data associated with the gaze of the target object.

[0161] Step 140: Send the target eye-tracking data to the target screen so that the target eye-tracking data is displayed on the target object on the target screen.

[0162] In the embodiments of this application, the target screen can be any device with display function. For example, the target screen can be a computer screen, a mobile phone screen, a central control screen, or a screen on a head-mounted device. The target screen can also be other screens. The embodiments of this application are not limited to any particular type.

[0163] In this embodiment, the target screen may display a target object. It should be noted that the target object displayed on the target screen can be a target object photographed from a different viewpoint than the visual data or a target object photographed from the same viewpoint. The target object displayed on the target screen is not displayed after receiving the target eye-tracking data, but is displayed beforehand. One purpose of this embodiment is to map the user's visual data to the target object's eye-tracking data onto the same target object displayed on the target screen, thereby analyzing the user's gaze behavior towards the target object.

[0164] In this embodiment, target eye-tracking data can be sent to a target screen, which then displays the target eye-tracking data on the target object it is showing. This maps the user's visual data to the eye-tracking data of the target object onto the same target object displayed on the target screen, facilitating the analysis of the user's gaze behavior. For example, the position of the target eye-tracking data on the target object can be represented by markers (such as dots, crosshairs, etc.).

[0165] Of course, in order to further clearly display the target eye-tracking data and the position of the target object to the user or observer, information such as fixation time, fixation frequency, and fixation trajectory can also be displayed on the target screen.

[0166] In some embodiments, the target object displayed on the target screen may be a target object identified from visual data. The content displayed on the target screen may be the same as or partially the same as the visual data. For example, the content displayed on the target screen may be a close-up of the target object, that is, the content displayed on the target screen is mainly focused on the target object. The difference between the content displayed on the target screen and the visual data may be that the viewpoint of the target object is different, or that other content including the target object is displayed.

[0167] The target object can be different in different application scenarios. For example, in a driving simulation scenario, the target object can be a driving screen; in a multi-turn robotic arm recognition scenario, the target object can be a robotic arm; in a robot navigation scenario, the target object can be a turning point, intersection, or specific task area on the navigation path. In other application scenarios, the target object can also be other items, advertisements, etc. This application embodiment does not limit the target object.

[0168] In simulated driving scenarios, for single-user analysis, the central control screen can be used as the target screen. The user's eye-tracking data for the target object can be mapped onto the central control screen for display and analysis. Alternatively, the user's eye-tracking data for the target object can be mapped onto other videos containing the target object for display and analysis. For multi-user analysis, since the visual and eye-tracking data collected by different users are not entirely the same, the target screen can focus on displaying the eye-tracking data for different users regarding the target object, enabling multi-user gaze behavior analysis of the target object.

[0169] According to the human-centric intelligent user gaze analysis method of this application, visual data of the user's field of vision is collected via a camera of a head-mounted device, and eye movement data of the user within the field of vision is collected via an eye tracker of the head-mounted device. The method identifies target objects in the visual data; determines target eye movement data associated with the gaze of the target object; and sends the target eye movement data to a target screen so that the target eye movement data is displayed on the target object on the target screen. This application's embodiment achieves effective identification of the target object being gazed upon by the user by collecting visual data and eye movement data, and displays the user's target eye movement data for the target object on the target screen. This application eliminates the need for manual annotation of the user's target gaze data for the target object, and can display the user's target eye movement data for the target object on the target screen in real time, enabling real-time analysis of user visual behavior on the target screen, enhancing the user's interactive experience, and improving the efficiency of user visual behavior analysis.

[0170] In some embodiments, the method further includes:

[0171] Determine the two-dimensional identifier to be used based on the scenario;

[0172] Two-dimensional markers are deployed on key points of the target object, and the correspondence between the target object, key points, and two-dimensional markers is recorded;

[0173] Identifying target objects in visual data, including:

[0174] The location of the target object in the visual data is determined by the position and correspondence of the two-dimensional markers in the visual data.

[0175] In this embodiment, the two-dimensional identifier is a pattern created in a two-dimensional space using a specific encoding method. This pattern can be used to store and transmit information. Two-dimensional identifiers are typically composed of black and white squares or dot matrices, and can be recognized and decoded by scanning devices such as smartphones or scanners to retrieve the information stored within them. The design of two-dimensional identifiers allows for the storage of more data in a limited space, providing higher data density and information capacity compared to traditional one-dimensional barcodes. In this embodiment, the two-dimensional identifier can be a QR code, a data matrix, an Aztec Code, or a two-dimensional identifier designed according to other encoding methods; this application embodiment does not limit this.

[0176] In this embodiment, the two-dimensional sign to be used can be determined based on the scenario. For example, different scenarios typically correspond to different environments, and the two-dimensional sign to be used can be determined based on the special effects of different scenarios. Some scenarios may correspond to outdoor environments with strong or complex lighting, in which case two-dimensional signs with high contrast and good reflectivity can be selected; some scenarios may correspond to cluttered background environments, in which case two-dimensional signs with unique and easily distinguishable patterns can be selected; some scenarios may correspond to dynamic environments, in which case signs that can be quickly identified and are not sensitive to blur and motion can be selected, and so on.

[0177] In some embodiments, the type of two-dimensional identifier can be selected to match the scenario. In this embodiment, the type of two-dimensional identifier can represent different kinds of two-dimensional identifiers, such as QR codes, data matrices, Aztec Codes, or two-dimensional identifiers designed according to other encoding methods.

[0178] The type of 2D identifier can also identify different representations of the same type of 2D identifier. For example, if the 2D identifier is an ArUco (Augmented Reality University of Cordoba) code, then the type of the 2D identifier can be an ArUco code marking pattern. For example, the ArUco library provides a variety of different marking patterns, which are mainly distinguished by the number of squares and the number of sub-squares in each square. Among them, the ArUco code is a binary square reference mark. The ArUco code consists of a wide black background and an inner white binary matrix. Its size and the size of the inner matrix determine the identifier of the mark, that is, each ArUco code has a unique ID. In addition to ID identification, the ArUco code can also be used to accurately estimate the position and orientation of the camera relative to the mark.

[0179] Key points of an object can be those locations on the object that possess unique geometric features, are easily identifiable, and remain relatively consistent across different viewpoints. Key points can be corners, edge intersections, center points, or other significant geometric locations on the object. For example, in a multi-screen recognition scenario of a driving simulator, key points could be the corners of the screens; in a multi-turn robotic arm recognition scenario, key points could be the joints on the robotic arm. The selection of key points is crucial for subsequent recognition and localization because key points provide a stable positional reference frame.

[0180] In this embodiment, after selecting key points, two-dimensional markers can be deployed on the key points, for example, by pasting two-dimensional markers onto the key points.

[0181] In this embodiment, a database or data structure can be pre-created to store the correspondence between target objects, key points, and 2D identifiers. For each key point, its specific location on the target object is recorded, such as coordinate information. Then, each 2D identifier is associated with its specific location on the key point, and this correspondence is recorded. For example, the correspondence could be: target object A - key point a1 - 2D identifier a11, target object A - key point a2 - 2D identifier a21, target object B - key point b1 - 2D identifier b11, etc.

[0182] In this embodiment, computer vision algorithms, such as edge detection, feature matching, or pattern recognition, can be used to identify specific patterns and structures of two-dimensional identifiers, thereby detecting two-dimensional identifiers in visual data. Once a two-dimensional identifier is detected, it can be decoded to obtain information from it. This information can be encoded data, such as numbers or text. It should be noted that the corresponding decoding method can be selected based on the type of the two-dimensional identifier.

[0183] Taking the detection of the position of a 2D label in visual data using an edge detection algorithm as an example, in this example, the visual data is image data. After acquiring the image data, the gradient of each pixel in the image can be calculated using an edge detection algorithm. For example, filters (such as the Sobel operator or Canny operator) can be applied to the image to calculate the gradient of each pixel. Then, the pixels are sorted according to the gradient magnitude. The algorithm can cluster pixels with similar gradient directions and magnitudes, and calculate the optimal straight line using the least squares method. These straight lines represent the edges in the image. Different contours can be identified through the edges in the image. If a contour is similar to the contour of a 2D label, the position of that contour can be determined as the position of the 2D label in the image.

[0184] After obtaining the decoding information of the 2D identifier, the corresponding relationship of the 2D identifier can be found in the database based on the decoding information, thereby determining the associated target object and key points. The position of the 2D identifier in the visual data is the position of the key point corresponding to the 2D identifier. Furthermore, the target object can be located based on the position of the key point.

[0185] In this embodiment, by determining appropriate two-dimensional markers according to the needs of the scenario, the selection of two-dimensional markers can be changed according to the scenario. This ensures that the two-dimensional markers have sufficient recognizability and information capacity in the visual data under different scenarios. Furthermore, two-dimensional markers are deployed on key points of the target object to record the corresponding relationships, providing necessary reference information for subsequent positioning. As a result, the position of the target object can be accurately calculated based on the position of the two-dimensional markers in the visual data, thus improving the accuracy of positioning.

[0186] In some embodiments, determining the two-dimensional identifier to be used based on the scenario includes:

[0187] The size of the white space in the 2D logo is determined based on the scenario; the white space is the area between the edge of the 2D logo and its background area.

[0188] The size of the blank area determines the two-dimensional identifier to be used.

[0189] In this embodiment, the two-dimensional identifier may include a background area and a blank area. For example... Figure 2 As shown, the background area is where the encoding is located. The background area can include a black background and a white binary matrix inside the black background. Different white binary matrices indicate different information contained in the two-dimensional identifier. The blank area is the area between the edge of the two-dimensional identifier and the edge of the background area. The blank area is a blank area without encoding. The blank area is crucial for the recognition of two-dimensional identifiers because it provides the necessary visual buffer to help the scanning device determine the boundary of the two-dimensional identifier.

[0190] In this embodiment, the size of the white space can be determined according to the characteristics of the scene. A larger white space indicates a greater distance between the edge of the 2D sign and the edge of the background area. For example, in strong light or reflective environments, a larger white space helps the scanning device better identify the boundary of the 2D sign; in close-range recognition scenarios, a smaller white space helps improve the visibility and recognition rate of the sign. Tests were conducted in a simulated driving scenario. The tests revealed that when the distance between the outer edge of the white space and the edge of the background area was 0.1 cm, robustness was poor in different driving scenarios. Increasing the distance to 0.5 cm significantly improved the recognition rate. Finally, tests showed that the highest recognition accuracy was achieved when the distance between the outer edge of the white space and the edge of the background area was 0.5-1 cm.

[0191] In this embodiment, two-dimensional markers for different scenes and different blank areas can be tested and calibrated in advance to determine the blank areas with good positioning performance for each scene and establish a correlation. After determining the scene, the size of the blank area corresponding to the scene can be determined based on the correlation, thereby determining the two-dimensional marker to be used.

[0192] In this embodiment, by determining the appropriate size of the blank area according to the scenario, the risk of misidentification can be reduced because the blank area provides additional visual buffer, helps the two-dimensional sign to be distinguished from the surrounding environment, and improves the accuracy of two-dimensional sign recognition.

[0193] In some embodiments, determining the two-dimensional identifier to be used based on the scenario includes:

[0194] Determine the area ratio of the background region to the area of ​​the two-dimensional sign based on the scenario;

[0195] The two-dimensional identifier to be used is determined based on the area ratio.

[0196] In this embodiment, the characteristics of the usage scenario can be analyzed in detail, considering factors such as lighting conditions, the distance to the target object, background complexity, and dynamic changes. Based on the characteristics of the scenario, the area ratio of the background region to the label itself in the 2D label is determined. The area ratio is crucial for the visibility and ease of recognition of the 2D label. For example, in complex backgrounds, the area of ​​the background region can be increased to make it easier to distinguish the 2D label from the background in the image. After determining the area ratio, the 2D label can be designed so that it can be accurately recognized in different environments. Specifically, during the design process, the size, color, and pattern of the label can be adjusted to ensure its performance in specific scenarios.

[0197] In this embodiment, 2D markers with different scene ratios can be tested and calibrated in advance to determine the area ratio with good positioning effect for each scene and establish a correlation. After determining the scene, the area ratio corresponding to the scene can be determined based on the correlation, thereby determining the 2D marker to be used. Through testing on a simulated driving scene, it was found that the highest recognition accuracy was achieved when the area ratio of the background area to the 2D marker was 0.51-0.73.

[0198] In this embodiment, by precisely controlling the area ratio of the background region to the marker itself in the 2D marker, the marker is neither too prominent, creating an overly strong contrast with the background, nor too concealed, making it difficult to detect. This improves the efficiency of the algorithm when processing images, as it can lock onto the target marker more quickly and reduce the processing of irrelevant background information. Furthermore, an appropriate area ratio also enhances the marker's stability under different lighting and viewing angle conditions, enabling reliable target localization even in varying environmental conditions.

[0199] In some embodiments, the target object includes multiple key points, and two-dimensional markers are deployed on the key points of the target object, including:

[0200] Select multiple non-collinear key points on the target object; deploy two-dimensional markers on the multiple non-collinear key points respectively, and record the correspondence between the target object, key points and two-dimensional markers;

[0201] Alternatively, select multiple key points from the target object to form a polygon; deploy two-dimensional markers on the multiple key points forming the polygon, and record the correspondence between the target object, key points, and two-dimensional markers.

[0202] In this embodiment, the target object includes multiple key points, and the arrangement of these key points on the target object conforms to a preset positional relationship. For example, the preset positional relationship can be symmetrical, grid-like, or any other geometric layout that facilitates positioning.

[0203] In this embodiment, the selected key points are not completely collinear. For example, if three key points are selected, these three key points are not on a straight line; if four key points are selected, these four key points are not on a straight line.

[0204] In this embodiment, the polygon can be determined based on the number of key points. For example, if the number of key points is 3, the polygon can be a triangle; if the number of key points is 4, the polygon can be a quadrilateral. Depending on the number of key points, the polygon can also be any other polygon, such as a pentagon or a hexagon, etc. This embodiment of the application does not limit this.

[0205] In this embodiment, by selecting multiple non-collinear key points on the target object and deploying two-dimensional markers on each, the geometric distribution advantage of the non-collinear key points is utilized. Even if some markers are unrecognizable due to occlusion or damage, the target object can still be located using other visible markers, thereby reducing the impact of a single marker failure on the overall positioning accuracy. By utilizing the geometric properties of polygons to provide a stable reference frame for positioning, with the vertices of the polygon (i.e., key points) serving as the deployment locations for two-dimensional markers, even if some markers are temporarily unrecognizable due to environmental factors (such as occlusion or changes in lighting), effective positioning can still be achieved based on other visible markers, improving positioning accuracy.

[0206] In some embodiments, locating a target object based on the position and correspondence of two-dimensional markers in visual data includes:

[0207] Feature extraction is performed on the visual data; based on the extracted features, the two-dimensional icons included in the visual data and their positions in the visual data are detected; based on the correspondence and the detection results of the two-dimensional icons and their positions in the visual data, the target objects in the visual data are located.

[0208] Alternatively, in response to the detection that the number of 2D markers in the visual data is less than the number of deployed 2D markers, the positions of the undetected 2D markers are fitted based on the geometric positional relationships of the key points; the target object is located based on the positions of the detected 2D markers in the visual data, the fitted positions of the undetected 2D markers, and their corresponding relationships.

[0209] In this embodiment, computer vision algorithms can be used to identify features in the image. For example, a scale-invariant feature transform algorithm can be used, which extracts features by finding local feature points in the image and calculating descriptors. These feature points are typically corner points, edge points, or other significant image features. The algorithm describes the image patches around these feature points to generate a feature vector, which can be used to match the same feature points in different images.

[0210] After feature extraction, algorithms can be used to find specific patterns and structures, such as feature matching or pattern recognition of feature points, to identify the outline of the two-dimensional sign. For example, edge detection algorithms can be used to identify edges in an image, detect the outline of the two-dimensional sign, and thus determine the position of the two-dimensional sign in the visual data.

[0211] After detecting the location of the two-dimensional marker, the marker can be decoded to obtain the encoded data in the marker. The encoded data can include information such as numbers and text, and then the target object in the visual data can be located by combining the corresponding relationship.

[0212] In this embodiment, if the number of 2D markers detected in the visual data is less than the number of deployed 2D markers, it may be due to some 2D markers being occluded or the shooting conditions being unsatisfactory. In this case, the positions of the undetected 2D markers in the visual data can be fitted by combining the positions of the detected 2D markers in the visual data with the geometric positional relationships of the key points. It should be noted that since the 2D markers are deployed on key points, the geometric positional relationships of the key points are the same as the geometric positional relationships of the deployed 2D markers.

[0213] Specifically, the location of undetected 2D markers can be fitted using any method. For example, since the arrangement of 2D markers on the target object conforms to the geometric positional relationship of key points, when the locations of some 2D markers are known, the possible locations of undetected 2D markers can be fitted using linear fitting, polynomial fitting, geometric fitting, etc. Alternatively, machine learning algorithms can be used to learn the geometric positional relationship of key points and predict the location of undetected 2D markers based on the locations of already detected 2D markers.

[0214] For example, if three 2D markers are deployed at three key points of a target object, and the geometric relationship between the key points is a triangle, then if any two 2D markers are detected, the position of the undetected 2D markers can be fitted based on the geometric relationship of the triangle. Similarly, if four 2D markers are deployed at four key points of the target object, and the geometric relationship between the key points is a quadrilateral, then if any three 2D markers are detected, the position of the undetected 2D markers can be fitted based on the geometric relationship of the quadrilateral. Likewise, if the number of 2D markers and the geometric relationship between the key points are other than normal, the position of the undetected 2D markers can be fitted based on the positions of the detected 2D markers in the visual data and the geometric relationship between the key points.

[0215] After fitting the positions of the undetected 2D markers, the position of the target object in the visual data can be determined by combining the positions of the detected 2D markers in the visual data with the correspondence between the target object, key points and 2D markers.

[0216] In this embodiment, feature extraction from visual data allows for the identification and detection of two-dimensional (2D) markers within the visual data. By combining these markers with their corresponding relationships, the target object can be accurately located. Since the number of detected 2D markers in the visual data is less than the number of deployed 2D markers, the location information of the detected 2D markers and the geometric relationships of key points can be used to predict the possible locations of undetected 2D markers through fitting. This allows for the reconstruction of the 2D marker layout even in the absence of markers, enabling effective target object localization and improving accuracy.

[0217] In some embodiments, the target screen includes a field of view partition and a target object partition;

[0218] Visual field partitioning shows the location of visual data and eye-tracking data within the visual data;

[0219] The target object is displayed in a partitioned manner, showing the location of the target object and the target eye-tracking data on the target object.

[0220] In this embodiment, such as Figure 3 As shown, the target screen can include a visual field partition and a target object partition. The visual field partition displays the position of the user's visual data and eye-tracking data within the visual data. The target object partition displays the position of the target object and its corresponding eye-tracking data.

[0221] In this embodiment, by dividing the target screen into a visual field partition and a target object partition, the visual field partition allows for a visual display of the relative position of the user's eye movement data within their entire visual field, which helps in analyzing the distribution of the user's visual focus and visual path. The target object partition further focuses on the specific target the user is looking at. By displaying the specific position of the target eye movement data on the target object, it is possible to more accurately understand the user's attention allocation and attention details to the target object. This not only enhances the visual analysis of the user's gaze behavior but also provides guidance for designing user interfaces, optimizing ad layouts, and improving user experience.

[0222] In implementing the embodiments of this application, the inventors discovered that in multi-screen systems, devices with image / video acquisition capabilities (such as cameras or devices with cameras, which may be referred to as visual signal acquisition devices) may have misaligned shooting angles relative to the target object during image / video acquisition. For example, the positioning of the visual signal acquisition device may not be directly in front of the target object, or the shooting angle of the visual signal acquisition device may move or change during image / video acquisition. This causes the image / video acquired by the visual signal acquisition device to show a distorted shape of the target object, which should be of a standard shape. The distortion of the image makes it difficult to accurately obtain the operation information received on the target object, such as the user's gaze point, touch point, mouse click position, etc. Therefore, the eye-tracking data related to the operation information in the distorted sub-image of the target object in the visual data acquired by the visual signal acquisition device can be mapped onto a non-distorted standard / normal shaped target screen. Through point mapping, the eye-tracking data on the target object in the acquired image or video can be displayed on the target screen, thereby accurately obtaining the operation information received on the target object.

[0223] In some embodiments, sending target eye-tracking data to a target screen to display the target eye-tracking data correspondingly on a target object displayed on the target screen includes:

[0224] Acquire the location information of the target's eye-tracking data on the target object;

[0225] Visual data is identified to obtain the position information of each marker pre-marked on the target object;

[0226] The conversion relationship is determined based on the location information of each marker;

[0227] By substituting the position information of the target eye-tracking data on the target object into the transformation relationship, the target coordinates of the target eye-tracking data on the target object in the target screen are obtained.

[0228] In this embodiment, referring to the description of step 130, the specific position of the eye movement data in the visual data can be calculated. If the eye movement data falls on the target object, the position information of the eye movement data on the target object can be further determined based on the specific position of the eye movement data in the visual data, that is, the position information of the target eye movement data on the target object can be determined.

[0229] In this embodiment, the markers can be the aforementioned two-dimensional markers. The two-dimensional markers can be deployed on key points of the target object. By recognizing the visual data, the position information of each marker on the target object can be obtained.

[0230] For example, the target object is a screen, and the markers are deployed at the four corners of the screen, forming a valid quadrilateral. The position information of the markers is their coordinates. We can calculate the distance of each coordinate to the origin of the coordinate system, take the point with the smallest distance to the origin as the starting point, calculate the vectors from the starting point to the other three points, and determine the angle between each vector and the x-coordinate of the coordinate system. Represent these three points in descending order of their respective angles as follows: , and The starting point is represented as The transformation relationship can be determined based on the coordinates of the four points obtained. The transformation relationship is used to map the eye-tracking data of the target on the twisted rectangle corresponding to the target object to the target screen.

[0231] Specifically, the location information of the target eye-tracking data can be represented as The target coordinates can be represented as The transformation relation is converted into an expression that includes the location information of the target eye-tracking data. Substituting the location information of the target eye-tracking data into the expression yields the target coordinates, thus enabling real-time mapping of the target eye-tracking data.

[0232] In this embodiment, by recognizing visual data, the positional information of each identified marker helps to locate the specific position of the target object and the marker area containing the target eye-tracking data. Obtaining the positional information of the target eye-tracking data helps to accurately map the target eye-tracking data onto the target screen. The positional information of the markers can establish a mapping relationship between the target object and the target screen, which helps to ensure that the target object and the target screen can be accurately aligned during the mapping process. Different marker layouts can adapt to different screen sizes and shapes. After determining the conversion relationship, the positional information of the target eye-tracking data can be mapped onto the target screen in real time, supporting interaction with the target object and seeing the corresponding response on the target screen, thus improving the real-time performance of target point mapping.

[0233] In some embodiments, determining the conversion relationship based on the location information of each identifier includes:

[0234] Based on the location information of each marker, determine the first set of parameters for calculating the x-coordinate of the target coordinate, the second set of parameters for calculating the y-coordinate of the target coordinate, and the third set of parameters for calculating the homogeneous coordinate normalization factor.

[0235] The expression for the first intermediate variable is determined based on the third parameter set;

[0236] Based on the expression of the first intermediate variable and the first parameter set, the first expression for calculating the x-coordinate of the target coordinate is obtained;

[0237] Based on the expression of the first intermediate variable and the second parameter set, a second expression for calculating the ordinate of the target coordinate is obtained. The first and second expressions constitute a transformation relationship.

[0238] In this embodiment, it can be based on , and Calculate the second intermediate variable, which can be represented as: ,based on , , and It is possible to calculate the third intermediate variable, which can be represented as: ,based on , , and It is possible to calculate the fourth intermediate variable, which can be represented as: .

[0239] Furthermore, based on the second intermediate variable Third intermediate variable and the fourth intermediate variable It can calculate the first parameter set, the second parameter set, and the third parameter set.

[0240] In some embodiments, the expression for the first intermediate variable is:

[0241]

[0242] in, and For parameters in the third parameter set, Location information of the target point;

[0243] The first expression is:

[0244]

[0245] in, The x-coordinate of the target coordinates. , and For the parameters in the first parameter set, As the first intermediate variable, Location information for the target eye-tracking data;

[0246] The second expression is:

[0247]

[0248] in, The ordinate of the target coordinates. , and For parameters in the second parameter set, As the first intermediate variable, Location information for the target eye-tracking data.

[0249] In this embodiment, the parameter sets of normalization factors for the horizontal, vertical, and homogeneous coordinates are determined respectively, which can more accurately describe the mapping relationship from the target object to the target screen. Introducing intermediate variables can simplify the final transformation relationship expression, making the calculation process more efficient. By combining the horizontal and vertical coordinate expressions, a complete transformation relationship is formed, ensuring the consistency of the target eye-tracking data position between the target object and the target screen. After the transformation relationship is determined, it can be reused multiple times, improving the real-time performance and reusability of the system.

[0250] In some embodiments, the number of markers pre-marked on the target object is four, and the corresponding position information is the first coordinate, the second coordinate, the third coordinate, and the fourth coordinate, respectively;

[0251] Based on the location information of each marker, a first set of parameters is determined for calculating the x-coordinate of the target coordinates, a second set of parameters is determined for calculating the y-coordinate of the target coordinates, and a third set of parameters is determined for calculating the homogeneous coordinate normalization factor, including:

[0252] Calculate the second intermediate variable based on the second, third, and fourth coordinates;

[0253] Calculate the third intermediate variable based on the second intermediate variable, the first coordinate, the second coordinate, and the fourth coordinate;

[0254] Calculate the fourth intermediate variable based on the second intermediate variable, the first coordinate, the third coordinate, and the fourth coordinate;

[0255] Based on the first coordinate, the second coordinate, the third coordinate, the third intermediate variable, and the fourth intermediate variable, determine the first parameter set and the second parameter set;

[0256] The third parameter set is calculated based on the third and fourth intermediate variables.

[0257] In this embodiment, the first coordinate is represented as The second coordinate is represented as The third coordinate is represented as The fourth coordinate is represented as It can be based on , and Calculate the second intermediate variable, whose expression is:

[0258]

[0259] Furthermore, it is possible to base it on the second intermediate variable. Calculate the third intermediate variable and the fourth intermediate variable Third intermediate variable The expression is:

[0260]

[0261] Fourth intermediate variable The expression is:

[0262]

[0263] The third parameter set includes parameters and parameters , , .

[0264] In this embodiment, a stable framework can be formed by four markers, reducing mapping errors caused by differences in the shape, size, resolution, and other factors of the target objects. It can adapt to different types of configurations and has good versatility. The calculation process is based on explicit mathematical expressions and intermediate variables, making it easy to implement.

[0265] In some embodiments, determining the first parameter set and the second parameter set based on the first coordinate, the second coordinate, the third coordinate, the third intermediate variable, and the fourth intermediate variable includes:

[0266] The product of the difference between the x-coordinates of the third coordinate and the first coordinate and the third intermediate variable is used as the first parameter; the product of the difference between the x-coordinates of the second coordinate and the first coordinate and the fourth intermediate variable is used as the second parameter; and the x-coordinate of the first coordinate is used as the third parameter. The first parameter, the second parameter, and the third parameter constitute the first parameter set.

[0267] The product of the difference between the ordinates of the third and first coordinates and the third intermediate variable is used as the fourth parameter, the product of the difference between the ordinates of the second and first coordinates and the fourth intermediate variable is used as the fifth parameter, and the ordinate of the first coordinate is used as the sixth parameter. The fourth, fifth, and sixth parameters constitute the second parameter set.

[0268] The first parameter set includes parameters A, B, and C; the second parameter set includes parameters D, E, and F; and the third parameter set includes parameters G and H. , , , , , .in, F represents the offset of the target point translation, and G and H are used to calculate the normalization factor of the homogeneous coordinates. The normalization factor is used to maintain the correctness of the perspective transformation.

[0269] In this embodiment, by calculating parameters based on the position information of four markers and through specific mathematical operations such as product and difference, accurate mapping results can be provided. The difference operation helps to capture the relative positional relationship between different markers. Using multiple markers as reference points and calculating parameter sets based on these reference points can improve the stability of the system. Even if there is an error in the position detection of one of the markers, it can be corrected by the information of other markers, thereby reducing the impact of the error on the mapping results.

[0270] In some embodiments, visual data is identified to obtain position information of each marker pre-marked on the target object, including:

[0271] Identify the marked area of ​​each pre-marked object on the target object from visual data;

[0272] Determine the coordinates of the center point of each marked area in a coordinate system established based on visual data to obtain the location information of each marker.

[0273] Determine the coordinates of the center point of each marked area, including:

[0274] For each labeled region in the visual data, determine the coordinates of each point in the outline of the labeled region, and calculate the average coordinates of each point in the outline of the labeled region to obtain the coordinates of the center point of the labeled region.

[0275] In this embodiment, the lower right corner of the visual data can be set as the origin of the coordinate system, with the positive x-axis pointing to the right and the positive y-axis pointing upwards. Image processing techniques such as edge detection, threshold segmentation, and template matching are used to identify the outline of the marked region of each pre-marked object on the target object from the visual data.

[0276] In this embodiment, for any identified region in the visual data, the outline of the identified region is represented by the coordinates of a series of points. By traversing each point in the outline, the number of points in the outline is determined, and the average abscissa and average ordinate of each point are calculated. The average abscissa is the ratio of the sum of the abscissas of each point to the number of points, and the average ordinate is the ratio of the sum of the ordinates of each point to the number of points. Through the above steps, the coordinates of the center point of each identified region can be obtained.

[0277] In this embodiment, by utilizing an image-based coordinate system, the marker can be accurately located, ensuring the precision of the location information. Identifying the marker's area and determining the center point coordinates accurately reflects the marker's actual position in the image. By calculating the average coordinates of all points on the contour as the center point coordinates, the overall shape and size of the marker area are considered, resulting in a more accurate reflection of the marker's actual position. Compared to methods that only consider some points or specific points on the contour, the average coordinate method reduces errors and improves the accuracy of location information.

[0278] In some embodiments, sending target eye-tracking data to a target screen to display the target eye-tracking data correspondingly on a target object displayed on the target screen includes:

[0279] The first video frame in which visual data is acquired includes a first image block containing target eye-tracking data;

[0280] In response to the first image block meeting the preset conditions, the coordinate mapping relationship between the first video frame and the second video frame in the target screen is calculated based on the first image block;

[0281] The target eye-tracking data is mapped to the second video frame based on the coordinate mapping relationship.

[0282] In this embodiment, the first image block is a local region of the first video frame, and the local region includes the target eye movement data.

[0283] In this embodiment, the image the user is viewing or interacting with is such as content on a computer screen, a book page, or a real-world field of view. Through coordinate mapping, the target eye-tracking data can be mapped onto video frames of the target screen, thereby enabling in-depth analysis and understanding of the user's visual behavior.

[0284] When the first image block meets preset conditions, such as the quality of the first image block and the information it contains meeting certain conditions, the coordinate mapping relationship between the first video frame and the second video frame can be calculated through image processing and computer vision algorithms.

[0285] In this embodiment, as described above, the position of the target eye-tracking data in the visual data can be calculated, thereby obtaining the position of the target eye-tracking data in the first video frame of the visual data. After obtaining the coordinate mapping relationship, calculations can be performed based on the coordinate mapping relationship and the position of the target eye-tracking data in the first video frame to obtain the position of the target eye-tracking data in the second video frame of the target screen.

[0286] In this embodiment, by acquiring image blocks containing target eye-tracking data, the coordinate mapping relationship between the first video frame and the second image frame is calculated when the image blocks meet preset conditions. This makes the calculation process focus more on the region in the image related to the target eye-tracking data, rather than the entire first video frame, reducing the impact of environmental changes on the calculation process, improving the accuracy of coordinate mapping relationship calculation, and thus improving the accuracy of eye tracking.

[0287] In some embodiments, the first video frame from which visual data is acquired includes a first image block containing target eye-tracking data, comprising:

[0288] Determine the position of the target eye-tracking data in the first video frame;

[0289] The region within a first preset range centered on the position of the target eye-tracking data in the first video frame is defined as the first image block.

[0290] In this embodiment, after determining the position of the target eye-tracking data in the first video frame, a first image block can be determined based on the position of the target eye-tracking data. For example, the first image block is a region within a preset range in the first video frame that contains the position of the target eye-tracking data; the first image block can be a circular region, a square region, etc.

[0291] In some embodiments, a region within a first preset range centered on the position of the target eye-tracking data in the first video frame can be defined as the first image block. The first preset range can be a circular region, a square region, etc., and the first preset range is smaller than the first video frame. For example, if the resolution of the first video frame is W*H, then the first preset range can be a region centered on the target eye-tracking data with a size of 0.1W and 0.1H.

[0292] In this embodiment, by accurately identifying the location of the target eye-tracking data and determining the first image block based on the location of the target eye-tracking data, irrelevant areas in the processed image can be reduced, allowing the calculation process to focus on the area near the target eye-tracking data, thereby reducing invalid calculations in irrelevant areas and improving calculation efficiency and accuracy.

[0293] In some embodiments, the preset conditions include: the number of matching points in the image block is greater than or equal to a preset number, and / or multiple matching points in the image block are not collinear.

[0294] In this embodiment, by setting a preset condition requiring the number of matching points to be greater than or equal to a preset number, and these matching points to be non-collinear, there are enough matching points to calculate the coordinate mapping relationship. The condition that multiple matching points are non-collinear allows the matching points to provide enough geometric information to solve the coordinate mapping relationship between the first video frame and the second video frame, further improving the accuracy of target eye-tracking data mapping.

[0295] In some embodiments, the method further includes: in response to a first image block not meeting a preset condition, acquiring a second image block in the first video frame that includes target eye-tracking data.

[0296] The method further includes: in response to the second image block meeting a preset condition, calculating the coordinate mapping relationship between the first video frame and the second video frame in the target screen based on the second image block.

[0297] In this embodiment, if the first image block does not meet the preset conditions, a second image block can be redefined. The second image block also includes target eye-tracking data. The second image block and the first image block are different local regions in the first video frame. The second image block and the first image block can be image blocks of the same shape or image blocks of different shapes. The second image block can be an image block obtained by translating the first image block or an image block obtained by magnifying the first image block.

[0298] In this embodiment, if the second image block meets the preset conditions, the coordinate mapping relationship between the first video frame and the second video frame can be calculated based on the second image block.

[0299] In this embodiment, by acquiring a second image block containing target eye-tracking data in the first video frame when the first image block does not meet the preset conditions, even if a sufficient number or appropriately distributed matching points are not found within the range of the first image block, the possibility of finding effective matching points can be increased by adjusting the search area, thereby enhancing the flexibility and success rate of the target eye-tracking data mapping process.

[0300] In some embodiments, obtaining a second image block in a first video frame that includes target eye-tracking data includes:

[0301] The second image block is obtained based on the position of the first image block in the first video frame.

[0302] In this embodiment, the second image block can be an image block obtained by translating the first image block, or an image block obtained by magnifying the first image block. The second image block can be determined by the position of the first image block.

[0303] In this embodiment, by obtaining the second image block based on the position of the first image block in the first video frame, the new search range can be redefined using the first image block as a reference, thereby improving the efficiency of the target eye-tracking data mapping calculation process.

[0304] In some embodiments, obtaining a second image block in a first video frame that includes target eye-tracking data includes:

[0305] The position of the target eye-tracking data in the first video frame is determined; the area within a second preset range centered on the position of the target eye-tracking data in the first video frame is determined as the second image block; wherein, the second preset range is larger than the first preset range.

[0306] In this embodiment, the first preset range can be expanded to a second preset range to obtain a second image block, thereby expanding the search range to find enough matching points and improving the accuracy of calculating coordinate mapping relationships. For example, the first preset range can be expanded to 1.1 times, 1.2 times, 1.5 times, or 2 times the original, or expanded to other multiples of the original, to obtain the second preset range.

[0307] In this embodiment, by expanding the first image block into a second image block, even if a sufficient number or appropriately distributed number of matching points are not found within the initial search range, the possibility of finding effective matching points can be increased by expanding the search area, thereby enhancing the flexibility and success rate of the target eye-tracking data mapping process.

[0308] In some embodiments, prior to including a first image block containing target eye-tracking data in a first video frame for acquiring visual data, the method includes:

[0309] Feature point matching is performed on the first video frame and the second video frame in the target screen. In response to the failure of matching between the first video frame and the second video frame, a new video frame is selected from the visual data to replace the first video frame and the second video frame for feature point matching.

[0310] In this embodiment, feature points are unique and identifiable points in an image, such as bright spots, freckles, corner points, or other significant visual features. These feature points are equivalent to visual landmarks in the image, and are easily identifiable and distinguishable parts of the image. Matching points refer to identical or similar feature points between two or more images. In tasks such as image registration, stereo vision, image stitching, or object recognition, matching points are used to establish correspondences between different images. It should be noted that feature points and matching points can be a single pixel or a region comprising multiple pixels.

[0311] Feature point matching is the process of extracting feature points from two or more images and identifying whether these feature points are matching points. Specifically, feature extraction algorithms, such as scale-invariant feature transform and accelerated robust feature extraction, can be used to extract feature points from a first video frame and a second video frame, respectively, and generate descriptors for the feature points. These descriptors can be quantized representations of the neighborhood surrounding the feature point, describing its characteristics and enabling it to be recognized and matched in different images. Then, points in the second video frame that match the feature points in the first video frame are searched. For example, the descriptors of the feature points in the first and second video frames can be compared, and the similarity of the feature points can be evaluated using distance metrics such as Euclidean distance or Hamming distance. Feature points with a similarity greater than a preset threshold are identified as matching points.

[0312] In this embodiment, matching points are obtained by feature point matching, which can establish the spatial relationship between the first video frame and the second video frame, providing an accurate reference benchmark for subsequent image segmentation. Moreover, it enables the target eye-tracking data to be effectively located and mapped to the corresponding position in the second video frame. Furthermore, the determination of matching points also helps to optimize the allocation of computing resources, as these matching points can be focused on without redundant processing of the entire image, thus improving the efficiency and accuracy of eye tracking.

[0313] When the first video frame and the second video frame fail to match, it means that the similarity between the feature points in the first video frame and the feature points in the second video frame does not reach the preset threshold. Other video frames in the visual data can be selected to replace the first video frame and used as the new first video frame and second video frame for matching.

[0314] In this embodiment, when feature point matching fails to establish an accurate correspondence between the first video frame and the second video frame, a new video frame can be flexibly selected from the continuous video stream as the first video frame, reducing the risk of interruption of the entire eye-tracking process due to a single matching failure and improving the continuity of the system.

[0315] In some embodiments, reselecting video frames from visual data to replace the first and second video frames for feature point matching includes:

[0316] The video frame adjacent to the first video frame in the visual data, or other video frames in the visual data, are identified as the reselected video frame.

[0317] In this embodiment, video frames adjacent to the first video frame in the visual data or other video frames in the visual data can be identified as reselected video frames and used as new first video frames to continue matching with the second video frame.

[0318] In this embodiment, when feature point matching between the first video frame and the second video frame fails, a video frame adjacent to the current first video frame is selected from the target video stream as a replacement to re-perform feature point matching. This ensures that the re-selected video frame maintains temporal continuity, or other video frames are selected as replacements, thereby improving the stability and accuracy of the eye-tracking process.

[0319] In some embodiments, feature point matching between a first video frame and a second video frame in a target screen includes: inputting the first video frame and the second video frame into a preset neural network model so that the neural network model can extract feature points from the first video frame and the second video frame, and perform matching based on the extracted feature points.

[0320] In this embodiment, a neural network model can be pre-trained so that the neural network model can identify and process feature points in the image. For example, the neural network in the neural model can analyze the first and second video frames of the input through its multi-layer structure, extract key visual features, including texture, color, shape, edge, etc., and then calculate the similarity of feature points to output the final matching point.

[0321] In this embodiment, by using a preset neural network model to match feature points of the first and second video frames, the powerful feature extraction capabilities of deep learning can be utilized to improve the accuracy and efficiency of the matching.

[0322] In some embodiments, feature point matching is performed on a first video frame and a second video frame in a target screen, including: preprocessing the first video frame and the second video frame; the preprocessing includes removing feature points within a preset range of the boundary of the first video frame and removing feature points within a preset range of the boundary of the second video frame; performing feature point matching on the preprocessed first video frame and the second video frame; and / or inputting the preprocessed first video frame and the second video frame into a preset neural network model so that the neural network model can extract feature points from the preprocessed first video frame and the second video frame and perform matching based on the extracted feature points.

[0323] In some embodiments, the preprocessed first video frame and the second video frame can be input into a preset neural network model so that the neural network model can extract feature points from the preprocessed first video frame and the second video frame and perform matching based on the extracted feature points.

[0324] In this embodiment, by removing feature points within a preset range of the image boundary, the negative impact of noise or incompleteness that may exist in the image boundary region on the matching accuracy can be reduced, allowing more focus to be placed on more stable and information-rich areas in the image, thereby improving the accuracy of the matching.

[0325] In some embodiments, the method further includes:

[0326] Store the feature points of the second video frame extracted by the neural network model;

[0327] Reselecting video frames from the visual data to replace the first and second video frames for feature point matching includes:

[0328] The replaced first video frame is input into a preset neural network model so that the neural network model can extract the feature points of the first video frame and match them with the feature points of the stored second video frame.

[0329] In this embodiment, since the first video frame and the second video frame may fail to match, if the feature points of the second video frame are extracted every time a new video frame is selected as the first video frame and the second video frame for matching after each matching failure, it will waste computing resources. Therefore, the feature points of the second video frame can be stored so that after a matching failure, the feature points of the new first video frame can be extracted and the stored feature points of the second video frame can be reused without repeatedly extracting the feature points of the second video frame.

[0330] In this embodiment, features of the second video frame are extracted and stored using a neural network model. When other video frame images are matched with the second video frame, the feature points of the second video frame can be called for feature matching, eliminating the need to repeatedly extract the features of the second video frame and improving processing efficiency.

[0331] In some embodiments, calculating the coordinate mapping relationship between the first video frame and the second video frame in the target screen based on the first image block includes:

[0332] Obtain the first position coordinates of the matching point in the first video frame and the second position coordinates of the matching point in the second video frame in the first image block;

[0333] The homography matrix corresponding to the matching point is calculated based on the first and second position coordinates; where the homography matrix represents the coordinate mapping relationship.

[0334] A homography matrix describes the projection relationship between different planes. This projection relationship can be used to map points on one plane to corresponding points on another plane, typically used in scenarios such as image matching, stereo vision, and image stitching. In this embodiment, a homography matrix can be used to represent the coordinate mapping relationship, transforming points in the first video frame to the coordinate system of the second video frame, so that the two images can be aligned. For example, For points in the first video frame, For a point in the second video frame, then and The coordinate mapping relationship can be expressed as: .in This represents the homography matrix.

[0335] In this embodiment, the first position coordinates of the matching point in the first image block in the first video frame and the second position coordinates of the matching point in the second video frame can be obtained. For example, there are four matching points in the first image block, namely matching point 1, matching point 2, matching point 3, and matching point 4. The coordinates of matching point 1 in the first video frame are... The coordinates in the second video frame are The coordinates of matching point 2 in the first video frame are: The coordinates in the second video frame are The coordinates of matching point 3 in the first video frame are: The coordinates in the second video frame are The coordinates of matching point 4 in the first video frame are: The coordinates in the second video frame are Then, using the position coordinates of these matching points in different images, a system of linear equations is constructed to calculate the homography matrix. :

[0336]

[0337]

[0338] in, Homography matrix elements, .

[0339] Then, the above equations are solved using the least squares method or other numerical methods, such as direct linear transformation, to obtain the homography matrix. elements The value of can be determined. It can be seen that if the number of matching points is greater than 4, and these matching points are not collinear, the homography matrix can be obtained. The solution is unique; therefore, the preset conditions can be determined based on the number and / or distribution of the matching points. For example, the preset conditions can be set to have a number of matching points greater than or equal to 4, and these matching points are not collinear. Of course, the above calculation process is based on a number of matching points of 4. When the number of matching points is other values, the above calculation process can be adaptively modified according to the number of matching points.

[0340] like Figure 4 As shown, the pentagram in the second video frame represents the location of the target eye-tracking data. After calculating the coordinate mapping relationship, the position of the target eye-tracking data in the second video frame can be obtained by calculating the location of the target eye-tracking data with the coordinate mapping relationship.

[0341] In this embodiment, by calculating the corresponding position coordinates of the matching points in the first and second video frames, a mathematical model can be established to describe the spatial relationship between these matching points, and a homography matrix can be determined to represent the coordinate mapping relationship. Using the homography matrix to represent the coordinate mapping relationship can adapt to complex image transformations, thereby improving the accuracy and robustness of target eye-tracking data mapping.

[0342] In some embodiments, the method further includes:

[0343] Acquire target eye-tracking data associated with multiple users' gazes at the same target object;

[0344] Multi-user fixation behavior is analyzed based on target eye-tracking data from multiple users.

[0345] In this embodiment, target eye-tracking data associated with the gaze of multiple users toward the same target object can be acquired, and multi-user gaze behavior can be analyzed based on this target eye-tracking data. For example, the concentrated location of each user's gaze on the target object can be analyzed based on this target eye-tracking data, and the location where each user gazes on the target object for the longest time can be analyzed, thereby analyzing the location that most users are more interested in.

[0346] In this embodiment, the ability to analyze the behavior of multiple users is improved by analyzing the target eye-tracking data of multiple users.

[0347] In some embodiments, target eye-tracking data of a single user can also be analyzed to analyze the user's gaze behavior. For example, the user's target eye-tracking data can be used to analyze the position on the target object where the user gazes for the longest time and the position with the highest gaze frequency, thereby analyzing the position on the target object that is most attractive to the user.

[0348] In some embodiments, multi-user gaze behavior is analyzed based on target eye-tracking data corresponding to multiple users, including:

[0349] Multiple user-specific eye-tracking data are sent to the target screen to overlay the target object displayed on the target screen with the user-specific eye-tracking data.

[0350] In this embodiment, multiple target eye-tracking data of different users can be superimposed and displayed on the target screen. For example, multiple visual data of different users and the position of eye-tracking data of different users in the visual data can be displayed in the visual field partition. Target objects can be displayed in the target object partition, and the positions of target eye-tracking data of multiple users on the target object can be superimposed and displayed on the target object.

[0351] In this embodiment, by overlaying and displaying the target eye-tracking data of multiple users on the target screen, researchers and designers can intuitively observe and compare the common fixation points and fixation patterns of different users when looking at the same target object, thereby identifying the common interest areas or hot spots of the user group.

[0352] In some embodiments, multi-user gaze behavior is analyzed based on target eye-tracking data corresponding to multiple users, including:

[0353] Based on the target eye movement data of multiple users, analyze the eye movement trajectory and / or eye movement point heatmap of multiple users to obtain the primary viewing position of multiple users and the habitual operation process of multiple users when performing preset operations;

[0354] The scene is optimized based on the primary viewing position of multiple users and their habitual operation flow when performing preset operations.

[0355] In this embodiment, the user's eye movement trajectory can be analyzed using target eye movement data to identify the user's viewing patterns and order, and to find the primary viewing position for different users. By collecting data such as eye movement points and fixation duration from different users, an eye movement point heatmap is generated, and color changes are used to display the concentrated areas of user fixation. For example, warm colors such as red and yellow can be used to represent areas of dense fixation or long duration.

[0356] By analyzing eye-tracking patterns and heatmaps to understand users' habitual workflows when performing preset actions, the design can be optimized based on the user's primary viewing position and habitual workflow. For example, in user interface design, important information or controls can be placed in the user's primary viewing position to simplify the user's operation process and improve the user experience.

[0357] In this embodiment, by analyzing the eye movement trajectories and eye movement heatmaps of multiple users, it is possible to reveal the user's primary viewing position and habitual operation process, thereby achieving a deep understanding of the user's behavior patterns. This embodiment provides a quantitative means to evaluate and optimize the user's interaction experience with products or services. By identifying common gaze points and gaze patterns, designers can optimize the interface layout so that key information and functions conform to the user's natural eye movement and operating habits, thereby improving user satisfaction and operating efficiency.

[0358] The human factors intelligent user gaze analysis method provided in this application can be executed by a human factors intelligent user gaze analysis device. This application uses the execution of the human factors intelligent user gaze analysis method by a human factors intelligent user gaze analysis device as an example to illustrate the human factors intelligent user gaze analysis device provided in this application.

[0359] This application also provides a human factors intelligent user gaze analysis device.

[0360] like Figure 5 As shown, the human gaze analysis device includes:

[0361] The acquisition module 510 is used to acquire visual data of the user's field of vision via the camera of the head-mounted device, and to acquire eye movement data of the user within the field of vision via the eye tracker of the head-mounted device.

[0362] The recognition module 520 is used to identify target objects in visual data;

[0363] Module 530 is used to determine the target eye movement data associated with the fixation of the target object in the eye movement data;

[0364] The display module 540 is used to send target eye-tracking data to the target screen so that the target eye-tracking data is displayed on the target object on the target screen.

[0365] According to the human-centric intelligent user gaze analysis device of this application, visual data of the user's field of vision is collected via a camera of a head-mounted device, and eye movement data of the user within the field of vision is collected via an eye tracker of the head-mounted device; target objects are identified in the visual data; target eye movement data associated with the gaze of the target object is determined in the eye movement data; and the target eye movement data is sent to a target screen so that the target eye movement data is displayed on the target object on the target screen. This application, by collecting visual data and eye movement data, achieves effective identification of the target object being gazed at by the user, and displays the user's target eye movement data for the target object on the target screen. This application eliminates the need for manual annotation of the user's target gaze data for the target object, and can display the user's target eye movement data for the target object on the target screen in real time, enabling real-time analysis of the user's visual behavior on the target screen, enhancing the user's interactive experience and improving the efficiency of user visual behavior analysis.

[0366] The human-centric intelligent user gaze analysis device in this application embodiment can be an edge computing device or a component within an edge computing device, such as an integrated circuit or chip. The edge computing device can be a terminal or other devices besides a terminal. For example, the edge computing device can be a mobile phone, tablet computer, laptop computer, eye tracker, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the device. In an optional example, deploying the human factors intelligence user gaze analysis device as an edge computing device or component on the edge deployment side can migrate at least some computing tasks to the edge device, leveraging the advantages of edge computing to improve data processing efficiency, reduce latency and network resource consumption and dependence, and also benefit data security and privacy protection.

[0367] The human-centric intelligent user gaze analysis device in this application embodiment can be a device with an operating system. This operating system can be a Microsoft (Windows) operating system, an Android operating system, an iOS operating system, or other possible operating systems; this application embodiment does not specifically limit it.

[0368] In some embodiments, such as Figure 6 As shown, this application embodiment also provides an edge computing device 600, including a processor 601, a memory 602, and a computer program stored on the memory 602 and executable on the processor 601. When the program is executed by the processor 601, it implements the various processes of the above-described human factors intelligent user gaze analysis method embodiment and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0369] It should be noted that the edge computing devices in this application embodiment include the aforementioned mobile edge computing devices and non-mobile edge computing devices.

[0370] In some embodiments, the edge computing device may be a head-mounted edge computing device; the head-mounted edge computing device may include a camera and an eye tracker.

[0371] In some embodiments, this application also provides a human factors intelligent user gaze analysis system, including: a target screen and the aforementioned edge computing device.

[0372] This application also provides a computer-readable storage medium on which a program or instruction is stored. When the program or instruction is executed by a processor, it implements the various processes of the above-described human factors intelligent user gaze analysis method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0373] The processor is the processor in the edge computing device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0374] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described human factors intelligent user gaze analysis method.

[0375] The processor is the processor in the edge computing device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0376] This application also provides a chip, which includes a processor and a communication interface. The communication interface and the processor are coupled. The processor is used to run programs or instructions to implement the various processes of the above-described human factors intelligent user gaze analysis method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0377] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0378] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0379] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0380] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other modifications under the guidance of this application without departing from its spirit, and all of these modifications are within the scope of protection of this application.

[0381] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

Claims

1. A human factors intelligent user gaze analysis method, characterized in that, include: The device collects visual data of the user wearing the head-mounted device within the user's field of vision via its camera, and eye movement data of the user within the field of vision via the eye tracker of the head-mounted device; the visual data includes data obtained by the camera of the head-mounted device capturing objects within the user's field of vision; the eye movement data includes data related to the user's eye movements captured by the eye tracker of the head-mounted device. Identify target objects in the visual data within the user's field of vision; Determine the target eye movement data associated with the fixation on the target object from the eye movement data within the user's field of vision; The target eye-tracking data is located on the target object within the visual data; Sending the target eye-tracking data to the target screen, so as to display the target eye-tracking data on the target object displayed on the target screen, includes: The position information of the target eye-tracking data on the target object is obtained; the visual data is identified to obtain the position information of each marker pre-marked on the target object; a transformation relationship is determined based on the position information of each marker; the position information of the target eye-tracking data on the target object is substituted into the transformation relationship to obtain the target coordinates of the target eye-tracking data on the target object in the target screen. Alternatively, the first video frame for acquiring the visual data includes a first image block containing the target eye-tracking data; in response to the first image block satisfying a preset condition, the coordinate mapping relationship between the first video frame and a second video frame in the target screen is calculated based on the first image block; and the target eye-tracking data is mapped to the second video frame based on the coordinate mapping relationship.

2. The method according to claim 1, characterized in that, The method further includes: Determine the two-dimensional identifier to be used based on the scenario; The two-dimensional markers are deployed on the key points of the target object, and the correspondence between the target object, key points, and two-dimensional markers is recorded; The identification of target objects in the visual data includes: The position of the target object in the visual data is located based on the position of the two-dimensional identifier in the visual data and the corresponding relationship.

3. The method according to claim 2, characterized in that, The step of determining the two-dimensional identifier to be used based on the scenario includes: The size of the white space in the two-dimensional logo is determined according to the scenario; wherein, the white space is the area between the edge of the two-dimensional logo and the background area of ​​the two-dimensional logo; The two-dimensional identifier to be used is determined based on the size of the blank area.

4. The method according to claim 2, characterized in that, The step of determining the two-dimensional identifier to be used based on the scenario includes: Determine the area ratio of the background region to the two-dimensional label in the two-dimensional label based on the scenario described; The two-dimensional identifier to be used is determined based on the area ratio.

5. The method according to claim 2, characterized in that, The target object includes multiple key points, and deploying the two-dimensional markers on the key points of the target object includes: Select multiple non-collinear key points on the target object; deploy the two-dimensional markers on the multiple non-collinear key points respectively, and record the correspondence between the target object, key points and two-dimensional markers; Alternatively, select multiple key points from the target object to form a polygon; deploy the two-dimensional markers on the multiple key points forming the polygon, and record the correspondence between the target object, key points, and two-dimensional markers.

6. The method according to claim 2, characterized in that, The step of locating the target object based on the position of the two-dimensional marker in the visual data and the corresponding relationship includes: Feature extraction is performed on the visual data; based on the extracted features, the two-dimensional identifiers included in the visual data and their positions in the visual data are detected; based on the correspondence and the detection results of the two-dimensional identifiers and their positions in the visual data, the target object in the visual data is located. Alternatively, in response to detecting that the number of two-dimensional icons in the visual data is less than the number of deployed two-dimensional icons, the positions of the undetected two-dimensional icons are fitted based on the geometric positional relationships of the key points; The target object is located based on the position of the two-dimensional markers detected in the visual data, the position of the undetected two-dimensional markers obtained by fitting, and the corresponding relationship.

7. The method according to claim 1, characterized in that, The target screen includes a field-of-view partition and a target object partition; The visual field partitions show the positions of the visual data and the eye movement data within the visual data; The target object partition displays the location of the target object and the target eye-tracking data on the target object.

8. The method according to claim 1, characterized in that, The step of determining the conversion relationship based on the location information of each marker includes: Based on the location information of each marker, a first set of parameters is determined for calculating the abscissa of the target coordinates, a second set of parameters is determined for calculating the ordinate of the target coordinates, and a third set of parameters is determined for calculating the homogeneous coordinate normalization factor. The expression for the first intermediate variable is determined based on the third parameter set; Based on the expression of the first intermediate variable and the first parameter set, a first expression for calculating the x-coordinate of the target coordinate is obtained; Based on the expression of the first intermediate variable and the second parameter set, a second expression is obtained for calculating the ordinate of the target coordinates. The first expression and the second expression constitute the transformation relationship.

9. The method according to claim 8, characterized in that, The number of markers pre-marked on the target object is four, and the corresponding position information is the first coordinate, the second coordinate, the third coordinate, and the fourth coordinate, respectively; The step of determining a first set of parameters for calculating the abscissa of the target coordinates, a second set of parameters for calculating the ordinate of the target coordinates, and a third set of parameters for calculating the homogeneous coordinate normalization factor based on the position information of each marker includes: Calculate the second intermediate variable based on the second coordinate, the third coordinate, and the fourth coordinate; Calculate the third intermediate variable based on the second intermediate variable, the first coordinate, the second coordinate, and the fourth coordinate; The fourth intermediate variable is calculated based on the second intermediate variable, the first coordinate, the third coordinate, and the fourth coordinate; Based on the first coordinate, the second coordinate, the third coordinate, the third intermediate variable, and the fourth intermediate variable, determine the first parameter set and the second parameter set; The third parameter set is calculated based on the third intermediate variable and the fourth intermediate variable.

10. The method according to claim 9, characterized in that, The step of determining the first parameter set and the second parameter set based on the first coordinate, the second coordinate, the third coordinate, the third intermediate variable, and the fourth intermediate variable includes: The product of the difference between the abscissas of the third coordinate and the first coordinate and the third intermediate variable is used as the first parameter; the product of the difference between the abscissas of the second coordinate and the first coordinate and the fourth intermediate variable is used as the second parameter; and the abscissa of the first coordinate is used as the third parameter. The first parameter, the second parameter, and the third parameter constitute the first parameter set. The product of the difference between the ordinates of the third coordinate and the first coordinate and the third intermediate variable is used as the fourth parameter, the product of the difference between the ordinates of the second coordinate and the first coordinate and the fourth intermediate variable is used as the fifth parameter, and the ordinate of the first coordinate is used as the sixth parameter. The fourth parameter, the fifth parameter, and the sixth parameter constitute the second parameter set.

11. The method according to claim 1, characterized in that, The step of identifying the visual data to obtain the position information of each marker pre-marked on the target object includes: Identify the identification area of ​​each pre-marked object on the target object from the visual data; Determine the coordinates of the center point of each marked area in the coordinate system established based on the visual data to obtain the position information of each marker; Determining the center point coordinates of each identified area includes: For each identified region in the visual data, determine the coordinates of each point in the outline of the identified region, and calculate the average coordinates of each point in the outline of the identified region to obtain the coordinates of the center point of the identified region.

12. The method according to claim 11, characterized in that, The first video frame in which the visual data is acquired includes a first image block containing the target eye-tracking data, comprising: Determine the position of the target eye-tracking data in the first video frame; The region within a first preset range centered on the position of the target eye-tracking data in the first video frame is defined as the first image block.

13. The method according to claim 1, characterized in that, The preset conditions include: the number of matching points in the image block is greater than or equal to a preset number, and / or multiple matching points in the image block are not collinear; and / or, The method further includes: in response to the first image block not meeting a preset condition, acquiring a second image block in the first video frame that includes the target eye-tracking data; and / or, The method further includes: in response to the second image block satisfying a preset condition, calculating the coordinate mapping relationship between the first video frame and the second video frame in the target screen based on the second image block.

14. The method according to claim 13, characterized in that, The step of acquiring a second image block in the first video frame that includes the target eye-tracking data includes: The second image block is obtained based on the position of the first image block in the first video frame; or... The position of the target eye-tracking data in the first video frame is determined; the area within a second preset range centered on the position of the target eye-tracking data in the first video frame is determined as the second image block; wherein the second preset range is larger than the first preset range.

15. The method according to claim 1, characterized in that, Before the first image block containing the target eye-tracking data is included in the first video frame in which the visual data is acquired, the process includes: Feature point matching is performed on the first video frame and the second video frame in the target screen. In response to the failure of matching between the first video frame and the second video frame, a new video frame is selected from the visual data to replace the first video frame and the second video frame for feature point matching.

16. The method according to claim 15, characterized in that, The step of reselecting video frames from the visual data to replace the first and second video frames for feature point matching includes: The video frame adjacent to the first video frame in the visual data, or other video frames in the visual data, are identified as the reselected video frame.

17. The method according to claim 15, characterized in that, The step of performing feature point matching on the first video frame and the second video frame in the target screen includes: inputting the first video frame and the second video frame into a preset neural network model so that the neural network model can extract feature points of the first video frame and the second video frame, and perform matching based on the extracted feature points; or, The step of performing feature point matching on the first video frame and the second video frame in the target screen includes: preprocessing the first video frame and the second video frame; the preprocessing includes removing feature points within a preset range of the boundary of the first video frame and removing feature points within a preset range of the boundary of the second video frame; performing feature point matching on the preprocessed first video frame and the second video frame; and / or inputting the preprocessed first video frame and the second video frame into a preset neural network model so that the neural network model can extract feature points from the preprocessed first video frame and the second video frame and perform matching based on the extracted feature points.

18. The method according to claim 17, characterized in that, The method further includes: Store the feature points of the second video frame extracted by the neural network model; Reselecting video frames from the visual data to replace the first and second video frames for feature point matching includes: The replaced first video frame is input into a preset neural network model so that the neural network model can extract the feature points of the first video frame and match them with the feature points of the stored second video frame.

19. The method according to claim 1, characterized in that, The step of calculating the coordinate mapping relationship between the first video frame and the second video frame in the target screen based on the first image block includes: Obtain the first position coordinates of the matching point in the first image block in the first video frame and the second position coordinates of the matching point in the second video frame; The homography matrix corresponding to the matching point is calculated based on the first position coordinates and the second position coordinates; wherein the homography matrix represents the coordinate mapping relationship.

20. The method according to claim 1, characterized in that, The method further includes: Acquire target eye-tracking data associated with multiple users' gazes at the same target object; Multi-user fixation behavior is analyzed based on target eye-tracking data from multiple users.

21. The method according to claim 20, characterized in that, The analysis of multi-user fixation behavior based on target eye-tracking data from multiple users includes: Multiple user-specific eye-tracking data are sent to the target screen to overlay the multiple user-specific eye-tracking data onto the target object displayed on the target screen.

22. The method according to claim 20, characterized in that, The analysis of multi-user fixation behavior based on target eye-tracking data from multiple users includes: Based on the target eye movement data of multiple users, analyze the eye movement trajectory and / or eye movement point heatmap of multiple users to obtain the primary viewing position of multiple users and the habitual operation process of multiple users when performing preset operations; The scene is optimized based on the primary viewing position of multiple users and their habitual operation flow when performing preset operations.

23. A human factors intelligent user gaze analysis device, characterized in that, include: The acquisition module is used to acquire visual data of the user's field of vision via the camera of the head-mounted device, and to acquire eye movement data of the user within the field of vision via the eye tracker of the head-mounted device; the visual data includes data obtained by the camera of the head-mounted device capturing objects within the user's field of vision; the eye movement data includes data related to the user's eye movements captured by the eye tracker of the head-mounted device. The recognition module is used to identify target objects in the visual data within the user's field of vision. The determination module is used to determine the target eye movement data associated with the gaze of the target object in the eye movement data of the user's field of vision; The target eye-tracking data is located on the target object within the visual data; A display module, configured to send the target eye-tracking data to a target screen, so as to display the target eye-tracking data on a target object displayed on the target screen, includes: The position information of the target eye-tracking data on the target object is obtained; the visual data is identified to obtain the position information of each marker pre-marked on the target object; a transformation relationship is determined based on the position information of each marker; the position information of the target eye-tracking data on the target object is substituted into the transformation relationship to obtain the target coordinates of the target eye-tracking data on the target object in the target screen. Alternatively, the first video frame for acquiring the visual data includes a first image block containing the target eye-tracking data; in response to the first image block satisfying a preset condition, the coordinate mapping relationship between the first video frame and a second video frame in the target screen is calculated based on the first image block; and the target eye-tracking data is mapped to the second video frame based on the coordinate mapping relationship.

24. An edge computing device, characterized in that, It includes a processor and a memory, the memory storing programs or instructions that can run on the processor, the programs or instructions being executed by the processor to implement the human factors intelligence user gaze analysis method as described in any one of claims 1 to 22.

25. The edge computing device according to claim 24, characterized in that, The edge computing device includes a head-mounted edge computing device; the head-mounted edge computing device includes a camera and an eye tracker.

26. A human factors intelligent user gaze analysis system, characterized in that, include: The target screen, or the edge computing device as described in claim 24 or 25.

27. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program or instructions that, when executed by a processor, implement the human factors intelligence user gaze analysis method as described in any one of claims 1 to 22.

Citation Information

Patent Citations

  • Target predictive thinking evaluation and training method and system

    CN111401721A

  • Multi-selectivity attention evaluation and training method and system

    CN112086196A

  • Eye movement target acquisition device, method and system based on display simulator

    CN115826766A

  • Head-mounted eye movement data calibration method and system applied to GRAIL system

    CN118902378A