An adaptive interest region processing method and device based on an eye movement test scene
Patent Information
- Application Number
- CN202611345040.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-09-01
- Publication Date
- 2026-09-29
AI Technical Summary
[0005]针对现有AOI处理策略中算力分配与行为学信息有效性不匹配的技术不足,本申请提供了一种基于眼动测试场景的自适应兴趣区域处理方法及装置
1,通过在测试运行前对静态类型兴趣区域的轮廓边界执行顶点精简处理,使命中判定所需遍历的边界点数量从正比于原始采样点数量降低至正比于精简后顶点数量,从而降低静态场景下每帧命中判定的计算开销;
Smart Images

Figure CN122841930A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of eye-tracking data processing technology, specifically to an adaptive processing method and apparatus that differentiates the calculation and update frequency of the boundary of the region of interest based on the type of region of interest and the user's eye movement state in an eye-tracking test scenario. Background Technology
[0002] Eye-tracking technology is widely used in fields such as psychological research, user experience testing, and advertising effectiveness evaluation. Its core processing flow includes raw gaze coordinate acquisition, fixation and saccade classification, and Area of Interest (AOI) analysis. AOI analysis determines whether the user's eye movement data points fall within a predefined area and uses this to calculate behavioral indicators such as dwell time and number of visits.
[0003] However, existing AOI processing strategies use a uniform processing method without distinguishing between target region type and user eye movement state. This results in a mismatch between computing power and the effectiveness of information acquisition when the target region outline remains fixed in static AOI scenarios, and the system still performs full traversal hit determination on the high-density original edge point set. In dynamic AOI scenarios, when the user is in a saccade state, the human eye enters a visual inhibition period and cannot effectively acquire visual information, but the system still performs boundary updates and hit determination of dynamic AOI at full frame rate, generating a large amount of redundant computation without behavioral significance.
[0004] The two types of problems mentioned above together cause a mismatch between computing resources and the effectiveness of information acquisition, making it difficult to achieve smooth real-time analysis in resource-constrained testing environments. Summary of the Invention
[0005] To address the technical shortcomings of existing AOI processing strategies that mismatch between computational power allocation and the effectiveness of behavioral information, this application provides an adaptive region of interest (ROI) processing method and apparatus based on eye-tracking testing scenarios. This method simplifies the contour vertices of static AOIs according to ROI type and performs adaptive frequency updates on dynamic AOIs based on eye-tracking states. This ensures the effectiveness of data behavioral analysis while matching the computational cost of hit determination and the computational frequency of ROI updates with the effectiveness of information acquisition in the current scene.
[0006] Specifically, this application provides the following technical solutions: This application provides an adaptive region of interest processing method based on an eye-tracking test scenario, including: Obtain the type label of each region of interest in the eye-tracking test scenario, wherein the type label indicates whether the region of interest belongs to a static type or a dynamic type; For the first region of interest marked as static, vertex simplification is performed on the contour boundary of the first region of interest before the test run to obtain a simplified contour, which remains unchanged during the test run. During the test run, each frame of eye movement data points was classified into eye movement states to obtain a first state label or a second state label. The first state label corresponds to the fixation state, and the second state label corresponds to the saccade state. Based on the type label of each region of interest and the status label of the current frame data point, differentiated region of interest hit determination is performed for each frame of eye-tracking data points: For the first region of interest, a hit determination is performed based on the simplified outline; For the second region of interest marked as dynamic, when the current frame status label is the first status label, boundary update and hit determination are performed at the first frequency; when the current frame status label is the second status label, boundary update and hit determination are performed at the second frequency. During the execution frame interval of the second frequency, the line-of-sight coordinates are recorded frame by frame at the first frequency, and hit determination is performed and recorded using the buffered boundary. The second frequency is lower than the first frequency, and it is restored to the first frequency immediately when the status label is switched from the second status label to the first status label.
[0007] Optionally, performing vertex simplification on the contour boundary includes: Edge detection is performed on the image region corresponding to the region of interest to obtain the original edge point set, and the contour boundary includes the original edge point set; Vertex simplification is performed on the original edge point set based on the first distance parameter, and redundant vertices within each arc segment whose vertical offset distance does not exceed the first distance parameter are deleted to obtain the simplified contour.
[0008] Optionally, performing vertex simplification on the original set of edge points based on the first distance parameter includes: Select the two farthest vertices on the closed contour as the initial dividing endpoints, and divide the closed contour into two arc segments; For each arc segment, calculate the perpendicular distance from each vertex on the arc segment to the line connecting the two endpoints of the arc segment. If the maximum perpendicular distance is greater than the first distance parameter, then the arc segment is recursively divided and calculated using that vertex as the new dividing endpoint. If the maximum perpendicular distance is not greater than the first distance parameter, then only the two endpoints of the arc segment are retained and the remaining vertices in the arc segment are deleted. The simplified results of each arc segment are combined to obtain the refined outline.
[0009] Optionally, the method further includes: The simplified contour generated based on the current first distance parameter is superimposed on the original contour corresponding to the original edge point set for display. In response to the adjustment operation, the first distance parameter is updated, and vertex simplification is re-performed based on the updated first distance parameter to generate an updated simplified outline, and the overlay display is refreshed; In response to the confirmation operation, the current first distance parameter and the corresponding simplified contour are determined as the final boundary configuration of the region of interest.
[0010] Optionally, during the test run, the eye movement state classification of each frame of eye movement data points includes: Calculate the displacement velocity between the current frame eye-tracking data point and the previous frame sampling point; The displacement velocity is compared with a first velocity parameter. If the displacement velocity does not exceed the first velocity parameter and the cumulative duration of the displacement velocity does not exceed the first velocity parameter reaches a first duration threshold, then the current frame data point is assigned a first state label; otherwise, the current frame data point is assigned a second state label.
[0011] Optionally, the hit determination step for regions of interest marked as dynamic type further includes: The second frequency is determined based on the first frequency and a preset first coefficient, wherein the first coefficient is a positive integer. During the frame interval when boundary updates and hit determination are performed at the second frequency, the gaze coordinates of each frame's eye-tracking data points are still recorded at the first frequency, and hit determination is performed using the boundary determined during the most recent boundary update.
[0012] Optionally, the method further includes: A ray is emitted from the current eye-tracking data point along a preset direction; Count the number of intersections between the ray and each edge segment of the simplified contour; If the number of crossovers is odd, the current eye-tracking data point is determined to fall within the region of interest; if it is even, it is determined to fall outside the region of interest.
[0013] Optionally, obtaining the type label of each region of interest in the eye-tracking test scene includes: Obtain the test material sequence and type information of each material in the eye-tracking test scenario; Based on the type information, determine the type label of the region of interest contained in each material.
[0014] This application also provides an adaptive region of interest processing device based on an eye-tracking test scenario, including: The type tagging module is used to obtain the type tag of each region of interest in the eye-tracking test scenario. The type tag indicates whether the region of interest belongs to a static type or a dynamic type. The contour simplification module is used to perform vertex simplification processing on the contour boundary of the first region of interest, which is marked as static, before the test run to obtain a simplified contour, which remains unchanged during the test run. The eye-tracking classification module is used to classify the eye-tracking state of each frame of eye-tracking data points collected during the test run, and obtain a first state label or a second state label. The first state label corresponds to the fixation state, and the second state label corresponds to the saccade state. An adaptive processing module is used to perform differentiated region of interest (ROI) hit determination for each frame of eye-tracking data points based on the type label of each ROI and the status label of the current frame data points: for the first ROI, hit determination is performed based on the simplified contour; for the second ROI labeled as dynamic, boundary update and hit determination are performed at a first frequency when the first status label is used, and at a second frequency when the second status label is used, the boundary update and hit determination are performed at a second frequency, and during the frame interval of the second frequency, the gaze coordinates are recorded frame by frame at the first frequency and the hit determination is recorded with buffered boundaries. The second frequency is lower than the first frequency, and it is restored to the first frequency immediately when the status label is switched from the second status label to the first status label.
[0015] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.
[0016] Compared with the prior art, this application has the following beneficial effects: 1. By performing vertex simplification on the contour boundaries of static type interest regions before test run, the number of boundary points required for hit determination in the hit is reduced from proportional to the number of original sampling points to proportional to the number of simplified vertices, thereby reducing the computational overhead of hit determination per frame in static scenes. 2. By implementing adaptive frequency updates for dynamic type interest regions driven by eye-tracking state classification results, the execution frequency of boundary updates and hit determination is reduced during the saccade corresponding to the second state label, and redundant calculations without behavioral significance are eliminated by utilizing the physiological characteristics of visual inhibition of the human eye during the saccade period. 3. By instantly restoring to the first frequency when the status label switches from the second status label to the first status label, the integrity and accuracy of the gaze initiation time data are ensured; 4. The three mechanisms mentioned above work together to match computing power allocation with the effectiveness of information acquisition, enabling smooth real-time analysis in test environments with limited computing resources. Attached Figure Description
[0017] Figure 1 This is a schematic diagram of the overall process of the adaptive region of interest processing method provided in the embodiments of this application.
[0018] Figure 2 A simplified diagram of a 3-pixel recursive vertex provided in an embodiment of this application.
[0019] Figure 3 A simplified diagram of a 6-pixel recursive vertex provided in an embodiment of this application.
[0020] Figure 4 This is a schematic diagram illustrating the relationship between the number of vertices and the computational complexity of hit determination, provided in an embodiment of this application. Detailed Implementation
[0021] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with the accompanying drawings and specific embodiments. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0022] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.
[0023] Eye-tracking technology has developed mature data processing pipelines in fields such as psychological research, user experience testing, and advertising effectiveness evaluation. Its core components include raw gaze coordinate acquisition, fixation and saccade classification, and region of interest (ROI) detection. In practical engineering, the computing devices that support these processing pipelines are diverse, including desktop PCs, tablets, laptops, and embedded eye-tracking analyzers. The processed data fall into two main categories: static image-based test materials and dynamic video-based test materials.
[0024] Existing processing solutions employ a uniform processing strategy for regions of interest across different types of materials, without distinguishing between the target region type and the user's real-time eye movement state. For example, in static image scenes, the system performs a full traversal hit determination on the high-density original edge point set; in dynamic video scenes, even if the user is in a saccade state and the human eye has entered a visual inhibition period, the system still performs boundary updates and hit determination at the full frame rate.
[0025] The aforementioned two types of redundant computations lead to a mismatch between computing power and the effectiveness of information acquisition in resource-constrained testing environments. Therefore, this application provides an adaptive region of interest processing method and apparatus based on an eye-tracking testing scenario.
[0026] Example 1 The adaptive region of interest processing method provided in this embodiment, such as Figure 1 As shown, it includes the following steps: S100, obtain the type label of each region of interest in the eye-tracking test scene, wherein the type label indicates whether the region of interest belongs to static type or dynamic type.
[0027] Before the official test runs, the system preprocesses all test materials involved in the test to establish type tags for each region of interest. This step is performed during the configuration phase before the test starts, and the generated type tag information remains unchanged throughout the entire test run, without being dynamically updated as the test progresses. The stability of the type tags allows the system to directly retrieve the type tag from the table and enter the corresponding processing branch without needing to determine the type of the current region of interest frame by frame during the test run, thus minimizing the overhead of branch selection. Table lookup operation.
[0028] The system acquires the test material sequence and the type information of each material. The test material sequence is pre-arranged by the testers and maintains a fixed order throughout the test. The system obtains a complete list of materials during the configuration phase. The type information is explicitly given by the material file format or test configuration file, with the following correspondence: For materials whose type information indicates image content, their contained regions of interest are marked as static and recorded in the test configuration data structure; for materials whose type information indicates video content, their contained regions of interest are marked as dynamic, corresponding to the frequency adaptive processing path in the differentiated hit determination.
[0029] For example, a typical eye-tracking test sequence contains four segments: segment 1 is a static image (PNG format, containing three predefined regions of interest); segment 2 is a video clip (MP4 format, containing one dynamic target region of interest); segment 3 is a static image (JPG format, containing two predefined regions of interest); and segment 4 is a video clip (containing two dynamic target regions of interest). After the system completes type labeling, it records a total of five static type regions of interest and three dynamic type regions of interest. The type labeling information is written into the configuration structure and remains unchanged throughout the test, without requiring frame-by-frame re-evaluation.
[0030] In mixed scenarios (containing both image and video footage in the same test), the system maintains a type tag lookup table according to the presentation order of the footage. The key is the footage identifier, and the value is the set of type tags for each region of interest (ROI) under that footage. During test execution, when a footage switching event is triggered (such as switching from image footage to video footage), the system updates the currently active ROI set and connects each ROI of the new footage to the static or dynamic processing path according to the type tags in the lookup table, without interrupting the data acquisition pipeline. The perception of footage switching is time-stamped by the test control module. This timestamp is synchronously written to the eye-tracking event record sequence for use in subsequent data analysis phases to divide the data time periods for each footage.
[0031] In the test configuration data structure, each region of interest corresponds to one record, containing the following core fields: a unique identifier for the region of interest (used for indexing across multiple regions of interest), a type tag (static or dynamic type), an identifier for the associated material (pointing to the material entry to which this region of interest belongs), and type-related configuration parameters (a simplified outline of static type records). and the first distance parameter The dynamic type records the first coefficient. (and initial bounding box). The system loads the entire configuration data structure into memory at the start of the test run, and only performs memory-level read and write operations during the test, without accessing disk files during each frame processing.
[0032] In another implementation, type tags can also be established manually by the operator in the test configuration interface. This is suitable for scenarios where the source file format cannot be directly mapped to static or dynamic types (e.g., some animated GIF files are technically image formats but contain dynamic content). The system provides a source preview function, allowing operators to view each source file individually, manually select type tags, and manually overwrite and correct automatically tagged sources. The output data structure format is the same for both type tagging methods (automatic and manual tagging), and the subsequent execution logic is independent of the source method of the type tags.
[0033] S200, for the first region of interest marked as static, vertex simplification is performed on its contour boundary before the test run to obtain a simplified contour, which remains unchanged during the test run.
[0034] The contour boundary of the static type interest region (i.e., the first interest region) remains fixed during the test run, so contour optimization can be completed all at once before the test starts. Specifically, it includes the following sub-steps.
[0035] S210, perform edge detection on the image region corresponding to the region of interest to obtain the original edge point set, wherein the contour boundary includes the original edge point set.
[0036] The system performs grayscale processing on image materials and uses edge detection operators to extract the closed contours of the target region to obtain the original edge point set.
[0037] For example, for a still image with a resolution of 1920 pixels × 1080 pixels, when the distance tolerance parameter (i.e., the first distance parameter) ε = 0 (without performing any simplification), the original set of edge points output by the edge detection is... Include For target regions with relatively complex shapes, each vertex has a specific value. Typically, it consists of between 300 and 500 vertices. Original edge point set Each element in Two-dimensional pixel coordinates in the image coordinate system .
[0038] Grayscale conversion transforms a color image from RGB three-channel to a single-channel grayscale image. A weighted average is used here to reflect the differences in human eye perception of brightness across different color channels. In the formula, , , These represent the red, green, and blue channel values of the pixel (range 0 to 255). After grayscale conversion, the outline of the target region appears as a region of abrupt changes in grayscale values in the grayscale image, facilitating edge detection operator recognition. For typical test materials where the target region has uniform color and significant contrast with the background (such as product images on a white background), the edges of the target region in the grayscale image usually exhibit clear high gradient bands, resulting in high-quality edge detection output. The deviation between the extracted closed contour and the actual boundary of the target region typically does not exceed 2 to 3 pixels.
[0039] In an alternative implementation, a classic edge detection method based on gradient magnitude calculation and non-maximum suppression can be used. Edge connectivity is achieved by setting a dual threshold range of low and high thresholds, typically with the low-threshold to high-threshold ratio set between 1:2 and 1:3. The edge detection method used does not affect the subsequent vertex simplification execution logic. The core constraint here is: the output... It is a set of discrete coordinates arranged in order on a closed curve, with the vertex spacing not exceeding 1 pixel.
[0040] In one implementation, the specific execution process of edge detection is as follows: the system performs Gaussian smoothing on the input image to suppress high-frequency disturbances introduced by image acquisition noise. The standard deviation of the smoothing kernel is usually in the range of 0.8 to 2 pixels, and in this embodiment, it is 1 pixel.
[0041] After smoothing, the gradient magnitudes of the image in the horizontal and vertical directions are calculated. Pixels with gradient magnitudes exceeding a high threshold are taken as strong edge anchors, and pixels with gradient magnitudes in the range of low to high thresholds are taken as weak edge candidates. Through connectivity analysis, weak edge candidate points adjacent to strong edges are included in the edge set, and the remaining candidate points are discarded.
[0042] The final output closed contour is extracted from the edge set using a contour tracking algorithm, recording coordinates pixel by pixel along the closed edge to form an ordered sequence. Point set. For image material containing multiple independent target regions, the system independently extracts the closed contour of each target region and stores them as separate point sets. For further processing.
[0043] S220, perform vertex simplification on the original edge point set according to the first distance parameter, delete redundant vertices in each arc segment whose vertical offset distance does not exceed the first distance parameter, and obtain the simplified contour.
[0044] In this embodiment, vertex simplification is achieved through a recursive arc segment division and redundant point deletion mechanism, which compresses the number of vertices to the greatest extent possible while preserving the main features of the contour shape.
[0045] Specifically, in closed contours Above, select the two vertices with the greatest Euclidean distance. and As the initial segmentation endpoints, the closed contour is divided into two arc segments. and Among them, for For each vertex in the array, compute its connection to the connected vertex. and The perpendicular distance of the straight line segment; for Perform the same calculation. If, at this point, there exists a vertical distance within the arc segment greater than the first distance parameter... If the vertex is the same as the vertex in the equation, then the vertex with the largest vertical distance is selected. As the new dividing endpoint, the arc segment is then divided into two sub-arc segments, and the above operation is recursively performed on each sub-arc segment; if the perpendicular distance between all vertices in the arc segment is not greater than... If so, only the two endpoints of the arc segment are retained, and all intermediate vertices within the arc segment are deleted.
[0046] After the recursive process is completed, the simplified results of all arc segments are merged to obtain the simplified set of edge points. ,in .
[0047] For example, suppose Include Each vertex represents a target area that is approximately elliptical in shape, serving as a product display area. For example... Figure 2As shown, when the first distance parameter When taking 3 pixels, after performing the above recursive vertex simplification, Include The system retains the key curvature change points of the four arcs of the ellipse, while removing smooth transition points within each arc segment whose offset does not exceed 3 pixels. For example... Figure 3 As shown, when When increased to 6 pixels, The number of vertices was further reduced to about 25, but the elliptical arc segment exhibited obvious polygonal polyline characteristics, which had a significant impact on the accuracy of the region coverage for subsequent hit detection. The value of is typically in the range of 2 pixels to 8 pixels; in this embodiment, it is 1. Pixels are used as the default configuration. (The above...) The specific values will be determined manually later.
[0048] In another embodiment, the vertex simplification can also be achieved through a curvature-based resampling mechanism: calculating The local curvature of each vertex is calculated, and vertices with curvature exceeding a preset threshold are retained as anchor points, while low-curvature smooth segments between anchor points are deleted. Both implementation methods aim to reduce the number of vertices while preserving the contour shape features, and are suitable for target regions of different shape types: the recursive arc segment division method is more effective for straight segments and polygonal regions, while the curvature resampling method has smaller errors for circular and organic curve regions.
[0049] The recursive arc segment partitioning vertex simplification mechanism used in this embodiment is commonly known in the algorithm field as the Douglas-Peucker algorithm, whose core characteristic is its time complexity. (in To input the number of vertices, (For the number of output vertices), in the number of contour vertices It exhibits good real-time execution performance even with values in the hundreds. First distance parameter. The only control parameter of this algorithm directly determines the trade-off between the approximate accuracy of the simplified contour and the vertex compression rate.
[0050] For example, when the outline of the target region has many regular straight line edges, the same Choosing a higher vertex compression ratio is possible because redundant intermediate points on straight line segments can be completely removed; when the contour exhibits high-frequency curve variations, a relatively larger number of vertices are required to retain more points of curvature variation. The compression ratio is slightly low. Operators can flexibly adjust the parameters for target areas with different shape characteristics in the subsequent interactive parameter tuning interface. This ensures that the simplified outline maintains an acceptable degree of approximation in shape semantics.
[0051] It should be noted that the first distance parameter The relationship between the shape fidelity of the original contour and the shape fidelity can be quantitatively evaluated using the Hausdorff distance: the Hausdorff distance is defined as the distance between the original contour and the shape fidelity. With streamlined outline The maximum point set distance between them is numerically equal to Neutral The farthest point and The distance between them. When performing recursive arc segment division, the perpendicular distance from each deleted intermediate vertex to the corresponding arc endpoint does not exceed the distance between them. ,therefore arrive The one-way Hausdorff distance has an upper bound. .
[0052] In practice, Setting the physical size to 1 to 2 pixels on the display (approximately 0.3 mm to 0.6 mm) enables significant vertex compression while ensuring that the outline shape visually matches the height of the original outline, meeting the engineering requirements for the accuracy of the region of interest boundary in most eye-tracking testing scenarios.
[0053] For test scenarios where the shape complexity of the regions of interest varies greatly, the system supports setting independent parameters for each region of interest. Value: For example, use a larger value for areas with relatively regular shapes (such as rectangles or circles). (e.g., 5 to 8 pixels) to achieve a higher compression ratio; use smaller pixels for areas with irregular shapes (e.g., facial contours, irregularly shaped product areas). (e.g., 2 to 3 pixels) to retain more shape detail. This regional differentiation setting This strategy achieves a more refined trade-off between the computational cost of hit detection in the overall test scenario and the fidelity of the contour shape, which is superior to applying a single method uniformly to all regions. A worthwhile solution.
[0054] S230, the simplified contour generated based on the current first distance parameter is superimposed on the original contour corresponding to the original edge point set and displayed, and the first distance parameter is updated in response to the adjustment operation to regenerate the simplified contour with the updated first distance parameter; in response to the confirmation operation, the current first distance parameter and the corresponding simplified contour are determined as the final boundary configuration of the region of interest.
[0055] The system provides operators with an interactive parameter adjustment interface, which will... The corresponding original contour and the current Below The corresponding simplified outlines are displayed overlaid on the same image: The original outline is drawn with thin lines, the simplified outline is overlaid with thick lines, and the remaining vertices of the simplified outline are highlighted with markers. The interface also provides... The value adjustment control allows operators to adjust the value by dragging a slider or entering a specific value. The system re-executes S220 in real time and refreshes the overlay display of the simplified outline, allowing operators to intuitively observe the different... The trade-off between contour simplification and shape fidelity under different values.
[0056] After confirming that the current simplified outline has not undergone unacceptable deformation, the operator submits a confirmation operation, and the system will then update the current outline. Value and The final boundary configuration of this static type of interest region is written into the test configuration data structure. If a test contains multiple static type interest regions, the system executes S210 to S230 independently for each interest region, allowing different settings to be applied to regions with different shape complexities. The value is determined to achieve the optimal trade-off between contour simplification and shape fidelity in the overall test scenario.
[0057] After completing S200, each static type interest region holds a fixed During the test run, the output of the S200... No longer updated, the system is now... Perform subsequent hit detection for a single contour input, without revisiting it every frame. .
[0058] It should be noted that during the entire test configuration phase, the S200 processes all static type regions of interest in batches offline, without consuming real-time computing resources during the test run. A complete S200 process (including S210 edge detection, S220 vertex simplification, and S230 interactive parameter tuning) typically takes only a few seconds, far shorter than the overall configuration time during the test preparation phase. After processing, the system will assign each region of interest... Value and Serialization is written to the test configuration file, and the test runtime directly loads and reads the configuration file. This eliminates the need to re-execute edge detection and vertex simplification processes. This allows for the reuse of the same set of simplified outlines across multiple test runs of the same batch of static test materials, further reducing system initialization overhead.
[0059] In addition, the system provides a version management mechanism for the test configuration files generated by S200: each time S200 generates a new configuration, the system assigns a version identifier (including the generation timestamp and operator identifier) to the current configuration and associates it with the corresponding version identifier. value, Storing the data along with the original material identifier allows operators to select and load historical configuration versions in subsequent tests, eliminating the need to repeatedly perform edge detection and interactive parameter tuning on the same batch of materials. This version management mechanism is particularly important for longitudinal research scenarios that require maintaining consistency in region-of-interest boundaries across different test batches, ensuring that the same group of subjects faces identical region-of-interest boundary definitions across different test time periods, thus eliminating interference from boundary differences in cross-batch comparisons.
[0060] For example, in a test scenario containing 5 statically typed regions of interest, the number of vertices in the original edge point set of each region. The numbers are 320, 480, 215, 390, and 270 respectively. After processing with S200, the number of simplified contour vertices in each region is... The values are 38, 52, 28, 44, and 31, respectively, with an average compression ratio of approximately... In a test run at a 120Hz sampling rate and approximately 80 valid data points per frame, each frame needs to execute... The edge segment traversal count for each region of interest hit determination is 335 times on average (using...). The frequency dropped to an average of 38.6 times (using) The computational workload is reduced to approximately 11.5% of the original.
[0061] During the test run, the S300 classifies the eye movement state of each frame of eye movement data points collected to obtain a first state label or a second state label. The first state label corresponds to the fixation state, and the second state label corresponds to the saccade state.
[0062] After the test officially started, the system continuously acquired raw gaze coordinate data from the eye tracker at a fixed sampling frequency, and performed real-time eye movement state classification on the eye movement data points arriving in each frame, assigning them to one of two mutually exclusive states. In this embodiment, the first state label specifically corresponds to the fixation state, and the second state label specifically corresponds to the saccade state.
[0063] In this embodiment, the system's sampling frequency is determined by the connected eye tracker hardware, typically ranging from 60Hz to 600Hz, with 120Hz being a commonly used value. Taking 120Hz as an example, the sampling time interval... Milliseconds, meaning 120 frames of eye-tracking data are collected per second. Each frame contains a timestamp and horizontal viewing angle coordinates. and vertical view coordinates (In degrees). The origin of the viewing angle coordinates is usually defined as the center of the display screen, with positive to the right in the horizontal direction and positive upward in the vertical direction. The coordinate range is determined by the field of view of the eye tracker, with a typical value being horizontal. to ,vertical to Specifically, it includes the following sub-steps: S310, calculate the displacement velocity between the current frame eye-tracking data point and the previous frame sampling point.
[0064] The system uses sampling time intervals Using the basic time unit, velocity estimation is performed on the relative displacement between eye-tracking data points in two consecutive frames. Let the first... Frame view coordinates are , No. Frame view coordinates are Then the Euclidean displacement between the two frames for: In the formula, the coordinate unit is viewing angles to ensure that the velocity calculation result is independent of the physical size of the screen and the distance from the eye tracker to the screen. Displacement velocity for: In the formula, For the first The displacement velocity of a frame, measured in degrees per second; is the Euclidean displacement between frame t and frame t-1, in viewing angles; The time interval between adjacent sampled frames, in seconds, is determined by the fixed sampling frequency of the eye tracker. Decision, that is .
[0065] For example, if the system sampling frequency is 120Hz, then Seconds. If the first Frame coordinates are , No. Frame coordinates are ,but Corresponding displacement velocity .
[0066] S320, compare the displacement velocity with the first velocity parameter to obtain a first state label or a second state label.
[0067] The system will output displacement velocity Compared with the first velocity parameter (i.e., the velocity classification threshold) The comparison is made in conjunction with a first duration threshold (i.e., minimum fixation duration). Execution status determination. The determination logic is as follows: like And since the most recent satisfaction The cumulative duration of consecutive frames has reached If the condition is met, the current frame data point will be assigned the first state label; otherwise, the current frame data point will be assigned the second state label.
[0068] Among them, the first velocity parameter Usually in / s to Within the range of / s, this is taken / s. First duration threshold. Typically, the timeframe is between 50ms and 150ms; here, we take 60ms. At a 120Hz sampling rate, 60ms corresponds to approximately 7 to 8 consecutive low-speed samples.
[0069] Taking the aforementioned example data as an example: Therefore, the data point of the current frame is assigned a second state label (scanning state). If the change in gaze coordinates slows down in subsequent frames, and the condition is met for 8 consecutive frames... Starting from frame 8, the corresponding data point is assigned a first state label (gaze state). The first 7 frames did not reach this state. The second state label is still assigned, and so on.
[0070] In another implementation, eye-tracking state classification can also be achieved through a sequence recognition mechanism based on a Hidden Markov Model (HMM). This involves modeling fixation and saccades as implicit state sequences, using velocity sequences as observations, and employing the Viterbi algorithm to solve for the optimal state path. This mechanism exhibits better stability of the classification boundary than methods based on single-frame velocity thresholds, even in scenarios with short-term oscillations in the velocity sequence, at the cost of introducing frame-by-frame probability update calculations. Both mechanisms output a frame-by-frame sequence of first or second state labels, with identical interface formats; subsequent execution logic is independent of the classification mechanism employed.
[0071] First speed parameter and the first duration threshold Both parameters are configurable, and operators can adjust their values based on the characteristics of the test group (e.g., children's gaze duration is typically shorter than adults') and the requirements of the test scenario (e.g., the strictness of the definition of saccades differs between rapid browsing and detailed reading tests). For example, in an advertising effectiveness test scenario where rapid browsing is the primary behavioral pattern, It can be appropriately increased to / s to / s, to avoid misclassifying transitional slow eye movements as fixation; for text comprehension test scenarios that primarily involve close reading, It can be extended to 100ms to 150ms to filter out pseudo-foveations that are too short.
[0072] For state transition detection, the system maintains a continuously updated low-speed frame counter. : when hour ,when hour .when When a gaze is confirmed, the current frame is assigned a first state label, indicating a gaze is established; otherwise, a second state label is assigned. This counter is implemented as an integer increment, with a computational overhead of [missing information]. It is suitable for embedding into a real-time sampling and processing pipeline for frame-by-frame execution.
[0073] For example, suppose the sampling frequency is 120Hz ( ), , (correspond (Rounded up to 8 frames). If starting from the first... The following 8 consecutive frames meet the requirements ,but , No. The first state label is assigned at the start of the frame. If the first... Frame appearance ,but The second state label is assigned to this frame and subsequent frames before the threshold is met again by continuous low-speed counting, so as to realize the instantaneous switching from gaze state to saccade state.
[0074] It should be noted that the eye-tracking state classification mechanism used in S310 to S320 of this embodiment is commonly referred to in the field of eye-tracking research as the Identification by Velocity Threshold (I-VT) algorithm, which is one of the real-time classification methods with the lowest computational cost in engineering practice. The advantage of the I-VT algorithm is that the classification computation per frame is... (Requires only one subtraction, one division, and two comparisons), suitable for real-time operation on resource-constrained embedded eye-tracking analyzers; Algorithm parameters (first velocity parameter) and the first duration threshold The meaning is intuitive, making it easy for operators to manually adjust according to the characteristics of the scene. The main limitation of the I-VT algorithm is that it is sensitive to short-term speed jitter, and may misidentify eye tremors (minor physiological tremors of the eyeball during fixation) as brief saccades. This application introduces a first duration threshold. (Requires consecutive frame rates to be lower than) Continue to reach Only then was it identified as a gaze), effectively filtering out those with a duration shorter than [a certain value]. The instantaneous velocity fluctuations reduce misclassification caused by eye movement microsporia.
[0075] S400 performs differentiated region of interest (ROI) hit determination for each frame of eye-tracking data points based on the type label of each ROI and the status label of the current frame data points.
[0076] In this embodiment, the system uses the type label established in S100 and the status label output in real time in S300 as dual inputs to perform a branched hit determination strategy on each frame of eye-tracking data points.
[0077] The design of the differentiated hit determination strategy is based on two independent technical criteria: Firstly, the boundaries of static type interest regions have undergone vertex simplification using S200 during the test configuration phase, reducing the computational complexity of in-process determination from... Down to (in First, the reduction was effective throughout the test run and did not depend on eye movement. Second, the boundaries of dynamic interest regions need to be updated frame by frame from the video frame content. Boundary updates are the main computational overhead. However, the human visual system is in a state of visual inhibition during saccades, and the brain's processing efficiency for visual information input during saccades is greatly reduced. High-frequency boundary updates cannot provide additional behavioral information value during saccades. Therefore, reducing the frequency of boundary updates during saccades can significantly reduce computational overhead without affecting data validity.
[0078] From the perspective of the overall system's computational load distribution, the differentiated hit determination strategy here aligns computational resources with the effectiveness of information acquisition in two dimensions: the first dimension is the AOI type dimension, where the computational load for hit determination in static scenarios is reduced from... Down to This optimization remains effective throughout the entire test run and is independent of eye movement status. The second dimension is the eye movement status dimension. In dynamic scenes, the boundary update frequency remains at the first frequency during the fixation period (first state label, effective information acquisition period) and drops to the second frequency during the saccade period (second state label, visual inhibition period), so that the computational resources for boundary updates are allocated according to the effectiveness of information.
[0079] It's easy to understand that if the test scenario includes both statically and dynamically typed regions of interest, both optimizations will take effect simultaneously, and the computational savings will be the sum of the two. Specifically: S410, For regions of interest marked as static, perform a hit determination based on the simplified outline.
[0080] The static type region of interest (ROI) hit determination needs to be performed in every frame during the test run, with the processing frequency being the same as the eye tracker sampling frequency (first frequency). It is independent of the eye tracker state label of the current frame. Regardless of whether the current frame is labeled with the first or second state, the static type ROI hit determination uses a simplified outline. This is performed at full frame rate for boundary updates. This contrasts with the differentiated frequency strategy for dynamic regions of interest (ROIs): because the main overhead of dynamic ROIs lies in boundary updates (dependent on target tracking, with computational costs related to video resolution and the number of targets), frequency reduction is significantly beneficial; while for static ROIs, the boundaries remain unchanged, and the main overhead is the hit-determination traversal per frame, which is reduced by vertex simplification using S200. Down to Without further frequency reduction, full frame rate execution ensures accurate recording of data for each frame in static scenes.
[0081] For each first region of interest labeled as static, the system defines a simplified outline using S200. As boundary input, the coordinates of the eye-tracking data points in the current frame. The determination of whether the execution point is inside the polygon. This embodiment uses ray casting to perform the hit determination, that is, starting from the eye-tracking data points of the current frame. A ray is emitted along the positive horizontal direction from a point. The relationship between this ray and... Number of intersections of each edge segment Determine the hit result according to the odd / even rule of polygons: If If it is odd, then determine If the target falls within the region of interest, the result is a hit; if If it is even (including 0), then determine If the target falls outside the region of interest, the result is a miss.
[0082] For example, suppose Include For a given vertex, the ray casting method needs to determine whether each of the 42 edge segments intersects with the ray. For a given frame of eye-tracking data points... (Pixel coordinates), emitting a ray along the positive horizontal direction, and... The edge segments of lines 3, 17, and 31 intersect; the number of intersections... If the number is odd, the hit result is a hit, meaning that the data point of that frame falls within the region of interest.
[0083] like Figure 4 As shown, compared with the original edge point set ( Compared to (350 line segment intersection checks corresponding to each vertex), based on ( The computational cost of a single hit determination (42 checks per vertex) is reduced to... In a typical scenario with 10 static type interest regions and a sampling frequency of 120Hz, it needs to execute [the following] per second. In the second hit determination, the number of line segment intersection calculations per determination was reduced from 350 to 42, a reduction of approximately 88% in computational complexity.
[0084] In an alternative implementation, the determination of hits for static regions of interest can also be achieved through a two-stage filtering mechanism based on pre-computed bounding boxes. The first stage performs a coarse screening using axis-aligned bounding boxes with simplified outlines; data points falling outside the bounding boxes are directly determined as misses without traversing edge segments. The second stage performs precise ray casting determination on the data points that pass the coarse screening. This mechanism further reduces the average computational cost by terminating the computation early in scenarios where eye-tracking data points are scattered and multiple regions of interest do not overlap, through the early termination of the coarse screening stage.
[0085] Axis-aligned bounding box of two-stage filtering mechanism After S200 processing is completed, the calculation and cache are defined as follows: ,in and They are respectively The minimum and maximum x-coordinates of all vertices in the array. and Similarly, the logic for the bounding box coarse screening is: if or or or If the bounding box coarse screening fails, it is directly determined as a miss, and the subsequent ray projection steps are skipped. The computational cost of bounding box coarse screening is... (4 comparisons) far lower than the light projection (m-times of line segment intersection judgment) In typical scenarios where the area of interest accounts for a small proportion of the total screen area (e.g., 10% to 30%), approximately 70% to 90% of the data points can be terminated in advance during the coarse screening stage, further reducing the average calculation amount for hit judgment.
[0086] S420 performs adaptive frequency hit determination based on eye-tracking state labels for regions of interest marked as dynamic.
[0087] Processing of dynamic regions of interest requires simultaneous scheduling across two interrelated dimensions: boundary update (determining whether target tracking should be performed in the current frame to obtain the latest boundary) and hit determination (determining whether to perform and which version of the boundary should be used in the current frame). The scheduling of these two dimensions is driven by the same trigger signal, the current frame's state label, forming a coordinated mechanism. Under the first state label (gaze state), both dimensions are executed at the first frequency; under the second state label (scanning state), both dimensions are executed at a reduced frequency (second frequency), and hit determination for non-execution frames uses cached boundaries. This coordinated design avoids inconsistencies in the frequency of boundary updates and hit determination, ensuring that the hit determination result of the same frame always corresponds to the output of the same boundary update, preventing contradictory situations where a low-frequency updated boundary performs precise determination on high-frequency sampling points.
[0088] The target location and outline of the dynamic type of region of interest change continuously with the video frame content, requiring frame-by-frame boundary updates. This step dynamically switches between two execution frequencies based on the status labels output in real time by the S300, to match the effectiveness of information acquisition under the current eye-tracking state. The switching strategy between the two frequencies is as follows.
[0089] When the current frame state label is the first state label (gazing state), the system performs dynamic interest region boundary updates and hit determination at the first frequency (i.e., normal frequency, equal to the sampling frame rate): The system performs target tracking on the current frame image to obtain the updated target region boundary. It then uses this updated boundary to perform a hit determination on the current frame's eye-tracking data points, and writes the hit result and the current frame's boundary state into the eye-tracking event log. During fixation, the user is effectively acquiring visual information, and the high-frequency execution of boundary updates and hit determination ensures the accuracy of the data during this period.
[0090] When the current frame state label is the second state label (scanning state), the system performs the boundary update and hit determination at a second frequency; here, the second frequency is the first frequency. Divided by the first coefficient ,Right now The first coefficient takes a positive integer value between 2 and 8. In this embodiment, the first coefficient is 4.
[0091] For example, when the first frequency is 120Hz, the second frequency is 120 / 4=30Hz, that is, the boundary update and hit determination are performed only once in the first frame out of every 4 frames, and the boundary update is skipped in the other 3 frames.
[0092] It should be noted that during the frame interval when boundary updates and hit determination are performed at the second frequency, the system still records the gaze coordinates of each frame's eye-tracking data points at the first frequency, and performs hit determination recording on the data points of that frame using the boundary determined during the most recent boundary update, ensuring that the complete spatial trajectory of the saccade path is preserved and can be used for subsequent supplementary statistical analysis.
[0093] In addition, the first coefficient The larger the value, the lower the boundary update frequency during the scan, and the more significant the computational savings. However, the longer the boundary lag time when using historical boundary snapshots to perform hit determination (the maximum lag time is...). Taking a 120Hz sampling rate as an example, The maximum latency is approximately 8.3 milliseconds. The maximum latency is approximately 25 milliseconds. The maximum lag is approximately 58.3 milliseconds. Since the human eye is in a state of visual inhibition during saccades, a boundary lag of 25 to 58 milliseconds does not affect the behavioral validity of the data. The recommended value range is 2 to 8; in this embodiment, 4 is used. For scenarios where the target moves at a high speed (such as a fast-moving video target), the operator can select a smaller value. (e.g., 2 or 3) to control lag; for embedded eye trackers with extremely limited computing resources, a larger lag can be selected. (e.g., 6 or 8) to maximize the benefits of downclocking.
[0094] The lightweight hit detection and recording mechanism during the scan is as follows: During the non-execution frames in the second frequency frame interval, the system does not invoke target tracking processing, but directly reads the boundary snapshot cached during the most recent execution boundary update from memory. ;by and To input the hit determination, the results are written to the eye-tracking event log for that frame, and the log is marked as using cached boundaries instead of real-time boundaries. In subsequent data analysis, this marking can be used to distinguish between frames using real-time and cached boundaries. For scenarios requiring high-precision analysis, only data from frames using real-time boundaries can be selected. The entire lightweight hit determination process does not perform boundary updates. For boundaries represented by polygon vertex sets, the computational cost for hit determination per frame is... ( (The number of vertices used to cache the boundary) has a negligible overhead compared to the computational cost of tracking dynamic targets.
[0095] For example, suppose the sampling frequency is 120Hz ( (seconds), first coefficient The second frequency is 30Hz. During a 100ms scan period (corresponding to 12 frames of data), the system performs boundary updates and precise hit determination only in frames 1, 5, and 9. The remaining 9 frames use the most recently known boundary snapshot for lightweight hit recording. Compared to full frame rate execution, the number of boundary updates during this scan period is reduced from 12 to 3, and the computational load for target tracking is reduced by approximately 25%.
[0096] When the current frame's state label switches from the second state label to the first state label, the system immediately resumes from the second frequency to the first frequency without waiting for the counter to be reset or the buffer window to end. The current frame immediately performs boundary updates and hit determination at the first frequency. The immediate recovery mechanism ensures that the first frame data at the gaze initiation moment (i.e., the moment when the user begins to effectively acquire visual information) uses the latest accurate boundaries, rather than switching after a delay of several frames, thereby ensuring the temporal integrity and accuracy of the gaze period data.
[0097] In actual testing, the frequency of alternation between fixation and saccade states is relatively high: in typical eye-tracking recordings, there are 2 to 4 saccade-fixation transitions per second, with each saccade lasting approximately 20 to 200 milliseconds and each fixation lasting approximately 150 to 500 milliseconds. The instantaneous recovery mechanism precisely triggers frequency recovery at each state transition, with a significant cumulative effect: at a 120Hz sampling rate, if the test duration is 60 seconds and the fixation rate is approximately 70% (typical value), then approximately 5040 frames of the fixation period are executed at the first frequency in full, while only about 490 frames out of approximately 1960 frames of the saccade period undergo boundary updates. This saves approximately 1470 boundary update operations compared to the full frame rate, without sacrificing any accuracy of the fixation period data.
[0098] In test scenarios containing multiple dynamic types of regions of interest (ROIs), the eye-tracking state labels of each ROI share the output of S300. This means the current frame's state label is applicable to frequency switching judgments for all dynamic ROIs. For example, if the test scenario contains three dynamic ROIs, and the current frame's state label is the second state label (saccade state), then all three ROIs will perform boundary updates and hit determinations at the second frequency (i.e., 1 / 4 of the first frequency). The number of boundary updates saved by the three ROIs combined is three times the amount saved by a single ROI. This makes the overall computational savings of the frequency reduction strategy more significant the number of dynamic ROIs in the test scenario.
[0099] The timing guarantees for switching from a status label to instant recovery are as follows: Let the first... When the frame status label changes from the second status label to the first status label, the system processes the first frame. The recovery from the second frequency to the first frequency is triggered at frame rate. The frame performs a complete boundary update (with the latest video frame content as input) and hit determination to ensure that the gaze start frame uses the latest boundary based on the current frame content, rather than a boundary snapshot of a historical frame during the scan.
[0100] It should be noted that the instant recovery mechanism in this embodiment is different from the recovery method based on a fixed-period counter in the traditional downsampling strategy: the fixed-counter recovery method will wait until the next predetermined execution frame after the state switch before restoring the full frequency, which may result in the use of outdated boundaries for several frames after the gaze starts; while the instant recovery mechanism of this application uses the state label switching event as the trigger point to eliminate this delay.
[0101] In a test scenario containing multiple dynamically typed regions of interest (ROIs), the real-time recovery operations for each ROI are executed concurrently within the same frame: when the... When the frame status label changes from the second status label to the first status label, all dynamic type interest regions are in the first state. Frame-triggered boundary updates are computationally independent for each region of interest (ROI), allowing for parallel execution using multithreading. The hit determination for each ROI is also performed in the first frame. Frame complete.
[0102] In this embodiment, the concurrent execution strategy ensures that the processing latency of a scene containing multiple dynamically typed regions of interest (ROIs) does not increase linearly with the number of ROIs during state transition frames, thus meeting the throughput requirements for real-time processing. For a single-threaded execution environment (such as a resource-constrained embedded processor), the system serially updates the boundaries of each ROI according to a predetermined order of ROI identifiers, completing the processing of all ROIs within a single frame time budget.
[0103] This application also provides an adaptive region of interest processing device based on an eye-tracking test scenario, including: The type labeling module is used to obtain type labels for each region of interest in the eye-tracking test scene. The type label indicates whether the region of interest belongs to a static or dynamic type. During the test configuration phase, the type labeling module reads the test material sequence and the type information of each material, labels the regions of interest contained in the image content as static type, and labels the regions of interest contained in the video content as dynamic type based on the type information, and writes the type labels into the test configuration data structure for use by subsequent modules.
[0104] The contour simplification module performs vertex simplification on the contour boundaries of statically labeled regions of interest before the test run, obtaining a simplified contour that remains unchanged during the test run. The contour simplification module includes an edge detection unit that performs edge detection on image materials and outputs the original edge point set. Edge simplification unit, based on the first distance parameter right Perform recursive arc segment division and redundant point deletion, and output a simplified outline. The interactive parameter tuning unit will... and The display is overlaid on the interactive interface and refreshes in real time in response to the operator's parameter adjustment operations. and respond to the confirmation operation to the current and Write the final boundary configuration.
[0105] The eye-tracking classification module is used to classify the eye-tracking states of each frame of eye-tracking data points acquired during the test run, obtaining either a first state label or a second state label. The first state label corresponds to the fixation state, and the second state label corresponds to the saccade state. The eye-tracking classification module acquires the raw gaze coordinates from the eye tracker at a fixed sampling frequency, calculates the displacement velocity between adjacent frames, and compares the displacement velocity with a first velocity parameter. Compare and combine with the first duration threshold Execution status determination, output the first status label or the second status label frame by frame.
[0106] The adaptive processing module is used to perform differentiated interest region (ORM) hit determination for each frame of eye-tracking data points based on the type label of each ORM and the status label of the current frame data points. For ORMs labeled as static, the adaptive processing module uses the simplified outline... A ray casting method is used to determine the boundary. For regions of interest marked as dynamic, boundary updates and hit determination are performed at a first frequency when the first state label is displayed, and at a second frequency when the second state label is displayed. The second frequency is lower than the first frequency, and the frequency is restored to the first frequency immediately when the state label is switched from the second state label to the first state label.
[0107] The above modules can be implemented in software or a combination of software and hardware on general-purpose computing devices, including but not limited to desktop PC workstations, laptops, tablets, and embedded eye-tracking analyzers.
[0108] The interactive parameter tuning unit in the contour simplification module provides visual interactive functions: it overlays the simplified contour and the original contour on the same image canvas, distinguishes the two contour lines with different colors, and marks the retained vertices of the simplified contour with dots; it also provides a first distance parameter. The numerical input box and step adjustment button allow the operator to modify the value each time. Afterwards, the system re-executes S220 and refreshes the overlay display within 200 milliseconds (typical response time), supporting different... Quickly switch between and compare values. For test scenarios with multiple static type regions of interest, the interactive parameter tuning unit supports configuring each region individually and provides a "batch application" function to combine current values. One-click application to multiple areas with similar shapes improves configuration efficiency.
[0109] The data flow between the four modules is as follows: During the test configuration phase, the output of the type marking module (the set of type marks for each region of interest) is passed to the contour simplification module. The contour simplification module executes steps S210 to S230 sequentially for the static type regions of interest, simplifying the contours. and The test configuration data structure is written; during the test run, the eye-tracking classification module collects gaze coordinates for each frame from the eye tracker and outputs status labels. The status labels and frame coordinates are simultaneously passed to the adaptive processing module; the adaptive processing module reads type tags from the test configuration data structure and assigns static type interest regions... Full-frame-rate hit determination is performed for the boundary, and boundary updates and hit determination are performed for dynamic type interest regions using state label-driven frequency switching logic. The hit determination results of each interest region in each frame are written into the eye-tracking event recording sequence.
[0110] In addition, the adaptive processing module maintains a lightweight scheduling state machine to manage the frequency switching of dynamic type interest regions: the scheduling state machine takes the current frame state label (first state label or second state label) and the execution frame counter of each interest region as input, and outputs a decision signal for whether each interest region should perform boundary update and hit determination in this frame; when the state label is the first state label, the execution frame counter of all dynamic type interest regions is reset and boundary update is triggered; when the state label is the second state label, the execution frame counter is incremented, and the count reaches the limit. Boundary updates are triggered only in frames that are integer multiples of the boundary snapshot; for other frames, only a lightweight hit determination based on the most recent boundary snapshot is performed. The computational cost per frame of this scheduling state machine is constant and does not increase linearly with the number of regions of interest.
[0111] Example 2 The difference between this embodiment and Embodiment 1 is that, in step S200, when performing vertex simplification processing on static type regions of interest, a single first distance parameter is not uniformly applied to all regions of interest. Instead, it assigns differentiated first distance parameters based on the importance level of each region of interest, such as retaining higher contour accuracy in critical regions and obtaining a larger vertex compression ratio in secondary regions.
[0112] During the test configuration phase, operators assign importance levels to each static type of region of interest in the test materials. These importance levels are divided into two categories: Level 1 (critical areas) and Level 2 (background areas). Critical areas are those directly related to the core test task, where the hit detection error significantly impacts the data analysis conclusions; examples include key product function areas, facial expression areas, or text reading areas. Background areas are those used solely to record whether they attract background attention; their hit detection requires lower contour accuracy.
[0113] For critical areas marked as level one, the system will use the first distance parameter. Set to a smaller value to retain more contour details; for background areas labeled as level two, the system will use the first distance parameter. Set to a larger value to obtain a higher vertex compression ratio.
[0114] For example, Take 2 pixels, Take 6 pixels. In a test scene containing 3 key regions and 5 background regions, the simplified outline of the key regions. On average, about 50 vertices are retained, resulting in a simplified outline of the background area. On average, approximately 22 vertices are retained. When performing hit detection using ray casting, the critical region traverses approximately 50 line segments per detection, while the background region traverses approximately 22 line segments per detection. This is significantly less efficient than applying a uniform method to all regions. With a pixel-based (approximately 50 vertices on average) scheme, this embodiment reduces the computational load for background region hit determination to approximately 44%, further lowering the overall average computational load of the system.
[0115] Example 3 The difference between this embodiment and Embodiment 1 is that: when performing eye-tracking state classification in step S300, the first velocity parameter... Instead of using a preset fixed value, the value is adaptively determined based on the actual eye movement speed distribution of the current subject after the test begins, in order to eliminate the impact of individual speed differences among subjects on classification accuracy.
[0116] Before the test officially began, the system presented the participants with a 5-second gaze calibration video, instructing them to sequentially fixate on five fixed gaze points (arranged in a cross shape) on the screen. During the calibration process, the system collected the participants' eye movement data and calculated the displacement velocity sequence for each frame during the calibration period. The system performs bimodal analysis on the velocity sequences: the velocity sequences typically form a peak in the low-speed region (fixation) and a peak in the high-speed region (saccades). The system uses the velocity at the trough between the two peaks as the individual velocity threshold for the subject. .
[0117] For example, in the velocity sequence collected during the calibration phase for a normal adult subject, the center of the low-velocity peak is approximately 10. The high-speed peak center is approximately 200. The valley floor of the twin peaks is located at approximately 45. Therefore, the system will be set to 45. For a specific subject whose range of eye movement was limited due to surgery or eye disease, the peak value of high-speed saccades was only about 80. The valley floor of the twin peaks is approximately 25. The system will The angle was adjusted to 25° / s. The fixation / saccade classification results obtained by two subjects on the same set of test materials more accurately reflect their true eye movement states, avoiding systematic misclassification of specific subjects by a fixed threshold.
[0118] Individualization threshold Once determined, it is written into the configuration data structure for this test, replacing the fixed one throughout the entire test. The speed comparison and determination for the S320 are then performed. The calibration process is then repeated for the next subject, and a new result is generated. No system configuration files need to be modified.
[0119] This embodiment only changes step S300. The source method (changed from a preset fixed value to calculation from the individual subject calibration data), the output interface of S300 (a sequence of first state labels or second state labels per frame) and the execution logic of S400 remain unchanged.
[0120] Example 4 The difference between this embodiment and Embodiment 1 is that when multiple subjects are tested sequentially on the same batch of test materials, the vertex simplification processing result of step S200 is calculated by the configuration stage of the first subject, and subsequent subjects can directly reuse the result without repeating S210 to S230 for the same group of static type interest regions.
[0121] Before the first participant's test, the system executed S200 normally, sequentially performing edge detection, vertex simplification, and manual parameter tuning for each static type of interest region, ultimately determining the... Value and Save to batch configuration file. The batch configuration file uses the test material identifier and region of interest identifier as keys, and... Value and The vertex sequence is the value, and the generation timestamp is recorded.
[0122] Starting with the second participant, the system automatically checks for the existence of a valid batch configuration file at the start of the test (the criteria are: the configuration file exists, the test material identifier is the same as the current participant, and the configuration file was generated before the current test). If the conditions are met, steps S210 to S230 are skipped, and each region is directly loaded from the batch configuration file. Proceed to the next steps in S300. If the configuration file does not exist or the test material has been changed, S200 will be executed normally and a new batch configuration file will be generated.
[0123] For example, in a user experience test project, there are 20 participants using the same set of image materials containing 8 static types of regions of interest. Only the first participant needs to execute S200 before the test (requiring approximately 30 seconds of manual parameter tuning). The subsequent 19 participants directly load the batch configuration file (requiring approximately 0.2 seconds). The total configuration time for all 20 participants is from […]. seconds shortened to approximately The efficiency of subject switching in batch testing scenarios is significantly improved in seconds.
[0124] Meanwhile, batch configuration files ensure that all participants within the same project are exposed to the exact same simplified outline of the region of interest, eliminating the impact of different participants' operations on the same simplified outline. The differences in boundary definitions introduced by different value judgments improve the cross-subject comparability of batch test data.
[0125] Those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program controlling the relevant hardware. The program can be stored in a computer-readable storage medium, including a read-only memory, a random access memory, a magnetic disk, or an optical disk.
[0126] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit described above can be implemented in hardware.
[0127] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. An adaptive region of interest processing method based on an eye-tracking test scenario, characterized in that, include: Obtain the type label of each region of interest in the eye-tracking test scenario, wherein the type label indicates whether the region of interest belongs to a static type or a dynamic type; For the first region of interest marked as static, vertex simplification is performed on the contour boundary of the first region of interest before the test run to obtain a simplified contour, which remains unchanged during the test run; During the test run, each frame of eye movement data points was classified into eye movement states to obtain a first state label or a second state label. The first state label corresponds to the fixation state, and the second state label corresponds to the saccade state. Based on the type label of each region of interest and the status label of the current frame data point, a region of interest hit determination is performed on each frame of eye-tracking data points. The region of interest hit determination includes: For the first region of interest, a hit determination is performed based on the simplified outline; For the second region of interest marked as dynamic, when the current frame status label is the first status label, boundary update and hit determination are performed at a first frequency; when the current frame status label is the second status label, boundary update and hit determination are performed at a second frequency, and during the execution frame interval of the second frequency, the line-of-sight coordinates are recorded frame by frame at the first frequency and the hit determination is recorded using the buffered boundary; wherein, the second frequency is lower than the first frequency, and it is restored to the first frequency immediately when the status label is switched from the second status label to the first status label.
2. The method according to claim 1, characterized in that, Performing vertex simplification on the contour boundary includes: Edge detection is performed on the image region corresponding to the region of interest to obtain an original set of edge points, and the contour boundary includes the original set of edge points; Vertex simplification is performed on the original edge point set based on the first distance parameter, and redundant vertices within each arc segment whose vertical offset distance does not exceed the first distance parameter are deleted to obtain the simplified contour.
3. The method according to claim 2, characterized in that, Performing vertex simplification on the original edge point set based on the first distance parameter includes: Select the two farthest vertices on the closed contour as the initial dividing endpoints, and divide the closed contour into two arc segments; For each arc segment, calculate the perpendicular distance from each vertex on the arc segment to the line connecting the two endpoints of the arc segment. If the maximum perpendicular distance is greater than the first distance parameter, then the arc segment is recursively divided and calculated using that vertex as the new dividing endpoint. If the maximum perpendicular distance is not greater than the first distance parameter, then only the two endpoints of the arc segment are retained and the remaining vertices in the arc segment are deleted. The simplified results of each arc segment are combined to obtain the refined outline.
4. The method according to claim 2, characterized in that, The method further includes: The simplified contour generated based on the current first distance parameter is superimposed on the original contour corresponding to the original edge point set for display. In response to the adjustment operation, the first distance parameter is updated, and vertex simplification is re-performed based on the updated first distance parameter to generate an updated simplified outline, and the overlay display is refreshed; In response to the confirmation operation, the current first distance parameter and the corresponding simplified contour are determined as the final boundary configuration of the region of interest.
5. The method according to claim 1, characterized in that, During the test run, the eye movement state classification of each frame of eye movement data points collected included: Calculate the displacement velocity between the current frame eye-tracking data point and the previous frame sampling point; The displacement velocity is compared with a first velocity parameter. If the displacement velocity does not exceed the first velocity parameter and the cumulative duration of the displacement velocity does not exceed the first velocity parameter reaches a first duration threshold, then the current frame data point is assigned a first state label; otherwise, the current frame data point is assigned a second state label.
6. The method according to claim 1, characterized in that, The method further includes: The second frequency is determined based on the first frequency and a preset first coefficient, wherein the first coefficient is a positive integer. During the frame interval when boundary updates and hit determination are performed at the second frequency, the gaze coordinates of each frame's eye-tracking data points are still recorded at the first frequency, and hit determination is performed using the boundary determined during the most recent boundary update.
7. The method according to claim 1, characterized in that, The method further includes: A ray is emitted from the current eye-tracking data point along a preset direction; Count the number of intersections between the ray and each edge segment of the simplified contour; If the number of crossovers is odd, the current eye-tracking data point is determined to fall within the region of interest; if it is even, it is determined to fall outside the region of interest.
8. The method according to claim 1, characterized in that, The type labels for each region of interest in the eye-tracking test scene are obtained as follows: Obtain the test material sequence and type information of each material in the eye-tracking test scenario; Based on the type information, determine the type label of the region of interest contained in each material.
9. An adaptive region of interest processing device based on an eye-tracking test scenario, characterized in that, include: The type tagging module is used to obtain the type tag of each region of interest in the eye-tracking test scenario. The type tag indicates whether the region of interest belongs to a static type or a dynamic type. The contour simplification module is used to perform vertex simplification processing on the contour boundary of the first region of interest marked as static before the test run to obtain a simplified contour, which remains unchanged during the test run. The eye-tracking classification module is used to classify the eye-tracking state of each frame of eye-tracking data points collected during the test run, and obtain a first state label or a second state label. The first state label corresponds to the fixation state, and the second state label corresponds to the saccade state. The adaptive processing module is used to perform differentiated interest region hit determination for each frame of eye-tracking data points based on the type label of each interest region and the status label of the current frame data points: for the first interest region, hit determination is performed based on the simplified contour; For the second region of interest marked as dynamic, boundary updates and hit determination are performed at a first frequency when the first state label is applied, and at a second frequency when the second state label is applied. During the frame interval of the second frequency, the line-of-sight coordinates are recorded frame by frame at the first frequency and the hit determination is performed and recorded with buffered boundaries. The second frequency is lower than the first frequency, and it is restored to the first frequency immediately when the state label is switched from the second state label to the first state label.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 8.