Eye-control adaptive interface design method based on intention recognition

By collecting and analyzing users' eye movements and intention data on the virtual reality eye control interaction test platform, using machine learning algorithms to train intent recognition models, and dynamically adjust the interface layout and functions, the fatigue problems and misoperation risks of users when using the eye control adaptive interface for a long time are solved, and an efficient and accurate interactive experience is achieved.

CN119938204AInactive Publication Date: 2025-05-06ZHEJIANG UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510060324.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

When using the eye-controlled adaptive interface for a long time, users may feel tired due to frequent operating instructions, and traditional methods fail to effectively solve the 'Midas touch' problem, resulting in erroneous operation.

Method used

By building a virtual reality eye control interaction testing platform, users' eye movement feature data and intention data are collected, machine learning algorithms are used to train intent recognition models, identify users' interactive intentions in real time, and dynamically adjust the layout, content and function display of the interface to adapt to the user's usage situation.

Benefits of technology

A more natural interaction method is realized, reducing the complexity of operations and the cognitive burden of users, improving interaction efficiency and accuracy, reducing the risk of misoperation, and improving user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938204A_ABST
    Figure CN119938204A_ABST
Patent Text Reader

Abstract

The invention discloses an eye-controlled adaptive interface design method based on intention recognition, and belongs to the technical field of man-machine interaction interface design, and the method comprises the following steps: constructing a virtual reality eye-controlled interaction experiment platform, and defining various interaction targets; setting various target eye control and intention recognition tasks, performing a data acquisition experiment, and acquiring task response time, recognition precision, eye movement track data and intention data of a user by an experimental platform; identifying an operation intention of a user in the virtual reality eye control interface by using data collected in the various target eye control and intention identification tasks; and integrating the trained intention recognition model to a back-end system of the platform. By integrating various target eye control, the intention recognition model and the adaptive interface optimization strategy, the virtual reality eye control interaction precision and the user experience are improved, and an innovative efficient interaction optimization tool is provided for eye control interface designers and developers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of human-computer interaction interface design, and in particular to an eye-controlled adaptive interface design method based on intention recognition. Background Art

[0002] Gaze-based interaction is a way of interacting with a device or system through the user's eye movements or gaze. This interaction mode relies on information such as the user's eye gaze point, scanning trajectory, and pupil changes. This data is captured and analyzed by eye movement devices to complete interactive operations such as target selection. This interaction method enables users to control interface elements in the virtual environment through gaze, greatly improving the naturalness and intuitiveness of the operation.

[0003] The traditional gaze-based interaction method faces the "Midas touch" problem, that is, the system will mistakenly identify unintentional gaze as a selection behavior, thereby causing misoperation. In order to solve this problem, researchers have proposed a variety of improvement methods, such as introducing a delayed selection mechanism and a reconfirmation step, and enhancing the accuracy of selection by requiring users to perform specific eye movements. For example, the patent CN111214227A with a publication number discloses a method for identifying user operation intentions and cognitive states in human-computer interaction, the method comprising: collecting eye movement feature data and EEG signal data of users during human-computer interaction; pre-processing the collected data and dividing them into a training set and a test set respectively; using the training set to train an SVM classifier, and then using the optimized classifier to classify the test set to obtain the corresponding operation intentions and cognitive states; the recognition method can identify the operation requirements in the interaction process, provide a basis for the adaptive adjustment of the interface, and at the same time judge the cognitive state of the interaction process to ensure the reliability of the completion of the interaction task; however, although these methods have improved the accuracy, they have not been effectively applied to the development of eye-controlled adaptive systems, and often increase the physical burden of users, especially when used for a long time, the user may feel tired due to frequent operation instructions.

[0004] In view of the above problems, it is urgent to carry out innovative designs based on the original eye-controlled adaptive interface design methods. Summary of the invention

[0005] The purpose of the present invention is to provide an eye-controlled adaptive interface design method based on intention recognition to solve the problem raised in the above background technology that users may feel fatigued due to frequent operation instructions during long-term use.

[0006] In order to further optimize the interaction effect, research in recent years has begun to explore the use of eye tracking technology to identify the user's true intention. By analyzing eye movement features such as gaze point, scanning trajectory and pupil changes, and combining machine learning algorithms, the system can more accurately identify and classify the user's behavioral intentions, thereby achieving a more natural interaction; this method, called "implicit interaction", allows the system to make judgments based on the user's natural eye movement patterns without explicit instructions, thereby reducing the complexity of operations and the user's cognitive burden; the eye-controlled adaptive interface is based on this concept. This design method dynamically adjusts the layout, content and function display of the interface by identifying the user's interaction intention in real time to adapt to different usage scenarios. For example, when the system recognizes that the user is focusing on a certain area, it can automatically enlarge or highlight the content of the area; when the user shows a browsing mode, the interface is adjusted to an information browsing mode to reduce meaningless operations. This adaptive design not only improves the interaction efficiency, but also effectively avoids misoperation problems such as "Midas touch", making the user experience smoother and more natural.

[0007] To achieve the above-mentioned purpose, the present invention provides the following technical solutions: a method for designing an eye-control adaptive interface based on intention recognition, comprising the following steps: S1, building an experimental platform, constructing a virtual reality eye-control interaction test platform, configuring an eye-movement interaction interface equipped with a virtual reality screen on the test platform, and defining specific diverse interaction targets on the virtual reality eye-control interaction platform, including icons, letters, paragraphs, and pictures; S2, setting experimental tasks and data collection, setting the layout of eye-control targets, including the number of targets, coordinates, and distribution methods, and designing a diverse target eye-control and intention recognition experiment 1, including five tasks: icon matching, virtual keyboard input, keyword matching, picture selection, and free browsing , recruit experimental participants to complete the task, collect and record the participants' task reaction time, recognition accuracy, eye movement trajectory, and intention data during the experiment; S3, training and integration of intention recognition model, preprocessing and feature extraction of the collected data, applying machine learning methods to train the intention recognition model, including support vector machine, logistic regression, LightGBM, random forest, and deploying the model to the backend of the virtual reality eye control interaction platform; S4, adaptive optimization and system integration, using the trained intention recognition model, set up a variety of target eye control and intention recognition experiments 2, including keyword matching and picture selection, and dynamically adjust the target size and trigger time according to the model prediction results.

[0008] Preferably, S1 specifically includes: configuring a virtual reality eye control interaction interface, defining a variety of interaction targets, including icons, letters, paragraphs, and pictures, and setting two types of tasks: intentional eye control tasks and unintentional eye control tasks; in intentional eye control tasks, the user selects the target area through eye control, and uses a handle to confirm whether the selection is in line with the intention; in unintentional eye control tasks, the user performs random eye control and uses a handle to confirm.

[0009] Preferably, the intentional eye control task includes: A1, obtaining and continuously monitoring the dwell time of the user's gaze point in the target area; A2, judging whether the dwell time reaches a preset trigger time threshold, if not, continuing monitoring, if yes, triggering the target; A3, judging whether the target area is the target that the user intends to select, if yes, triggering the selection and turning the icon green, if not, triggering the selection and turning the icon red, and recording the experimental data; A4, using the handle to confirm whether the selection is in line with the user's intention, and recording the result.

[0010] Preferably, the unintentional eye control task includes: B1, the user selects the target area through eye control and uses the handle to confirm the selection; B2, after a random time of 4s to 6s, the handle input is detected, and if not confirmed, return to step B1; B3, determine that the user has no intention and record other experimental data.

[0011] Preferably, the experiment 1 in S2 is divided into five tasks, and different types of trigger targets are set, specifically including: Task 1, icon matching, the system displays a prompt icon, and the participant selects the same icon as the prompt icon from the menu by gazing; Task 2, virtual keyboard input, the system presents a simplified QWERTY keyboard, and the participant enters text by gazing on the corresponding letter key to respond to the prompt information; Task 3, keyword matching, a keyword is displayed at the top of the interface, and six descriptive phrases are listed below, and the participant needs to select the phrase that best matches the keyword; Task 4, picture selection, the system displays a prompt image, and the participant selects the picture that matches the prompt image by gazing; Task 5, the participant freely browses the pictures in the virtual reality interface, and uses the handheld controller to trigger the target operation to enter the next round.

[0012] Preferably, Tasks 1 to 4 are intentional eye control tasks. After finding the target button, the subject triggers the target button by gazing at it. After continuously gazing at the button for a preset time, the target button will be displayed in different colors depending on whether the triggering operation is successful or not, a green button for success and a red button for failure. Subsequently, the participant uses the handle to select whether the picture meets his or her intention.

[0013] Preferably, the task five is an unintentional eye control task, and the subject uses the handle to assist in selecting to enter the next selection interface, and the user intention data is recorded as unintentional.

[0014] Preferably, S3 specifically includes: cleaning and standardizing the eye tracking data collected in Experiment 1, extracting pupil diameter, line of sight coordinates and basic eye movement features; using a variety of machine learning algorithms to train the intention recognition model, and integrating the model into the back end of the virtual reality eye control interaction platform to achieve real-time intention recognition and interface optimization.

[0015] Preferably, S4 specifically includes: setting a second experiment of multi-target eye control and intention recognition, setting a gaze control interface for an online shopping platform, wherein users trigger interactive elements by gazing, and the system determines the trigger result by detecting the gaze duration threshold and the activation position. When users want to activate a button, they need to keep gazing until the action is triggered; the intention recognition model determines whether the user intends to trigger the interactive target by analyzing the user's current eye movement behavior, and the judgment result is represented by the Iswish indicator. The difficulty of the task is evaluated by combining the accuracy and intention value of the task, and the trigger target size and trigger time are adjusted according to the difficulty of different tasks to improve the user experience and interaction accuracy.

[0016] Preferably, the interface partitions include: area A, product category area; area B, text input area; area C, product recommendation area; area D, product display area.

[0017] Compared with the prior art, the beneficial effect of the present invention is that the eye control adaptive interface design method based on intention recognition combines intentional eye control and unintentional eye control, collects experimental data and establishes a user intention recognition model, and the system can accurately predict the user's intention.

[0018] Furthermore, on this basis, an eye-controlled adaptive interface system that can adapt to various target applications was developed. The system can adjust the target size and interface layout according to user intention prediction and target triggering accuracy.

[0019] Different from traditional static interfaces, the system of the present invention can respond to the user's eye movement behavior in real time, provide an interaction method that better meets the user's needs, reduce false triggers and operation delays, and improve the accuracy and efficiency of eye-controlled interaction.

[0020] Furthermore, this method has strong adaptability and can be widely used in virtual reality, augmented reality and other eye-controlled interaction systems, with good practicality and broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 Schematic diagram of the eye-controlled adaptive interface design method based on intention recognition of the present invention.

[0022] Figure 2 This is a flow chart of the intentional eye control task of the present invention.

[0023] Figure 3 This is a flowchart of the unintentional eye control task of the present invention.

[0024] Figure 4 This is a flow chart of Experiment 2 of the present invention.

[0025] Figure 5 This is a bar chart of model indicators of experiment 1 of the present invention.

[0026] Figure 6 This is a characteristic bar graph of Experiment 1 of the present invention.

[0027] Figure 7 This is a schematic diagram of the functional zoning of the interface of the present invention.

[0028] Figure 8 This is a flow chart of Experiment 1 of the present invention.

[0029] Fig. 9 This is a schematic diagram of the layout of diverse interactive targets in Experiment 2 of the present invention.

[0030] Fig.10 This is a flow chart for optimizing and adjusting the adaptive interface of the present invention.

[0031] Fig.11 This is a graph of the trigger accuracy of the second and fourth stages of the experiment of the present invention.

[0032] Fig.12 This is a trigger time diagram for the second and fourth stages of the experiment of the present invention. DETAILED DESCRIPTION

[0033] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0034] Example 1: Please refer to Figure 1-Figure 12 The present invention provides the following technical solution: a method for designing an eye-controlled adaptive interface based on intention recognition, comprising the following steps:

[0035] S1. Build an experimental platform, construct a virtual reality eye-controlled interaction test platform, configure an eye-movement interaction interface equipped with a virtual reality screen on the test platform, and define specific and diverse interaction targets on the virtual reality eye-controlled interaction platform, including icons, letters, paragraphs, and pictures.

[0036] S2. Experimental task setting and data collection: set the layout of eye control targets, including the number of targets, coordinates, and distribution methods; design a diverse target eye control and intention recognition experiment, including five tasks: icon matching, virtual keyboard input, keyword matching, picture selection, and free browsing; recruit experimental participants to complete the tasks; collect and record the participants' task reaction time, recognition accuracy, eye movement trajectory, and intention data during the experiment.

[0037] S3. Training and integration of intent recognition model. Preprocessing and feature extraction of collected data. Application of machine learning methods to train intent recognition model, including support vector machine, logistic regression, LightGBM, random forest, and deployment of the model to the backend of the virtual reality eye-controlled interaction platform.

[0038] S4, adaptive optimization and system integration, using the trained intention recognition model, set up the second experiment of eye control and intention recognition of various targets, including keyword matching and picture selection, and dynamically adjust the size and trigger time of the target according to the model prediction results.

[0039] S1 specifically includes: configuring the virtual reality eye control interaction interface, defining a variety of interaction targets, including icons, letters, paragraphs, and pictures, and setting two types of tasks: intentional eye control tasks and unintentional eye control tasks; in intentional eye control tasks, users select the target area through eye control and use the handle to confirm whether the selection is in line with the intention; in unintentional eye control tasks, users perform random eye control and use the handle to confirm.

[0040] The intentional eye control task includes: A1. Acquire and continuously monitor the dwell time of the user's gaze point in the target area; A2. Determine whether the dwell time reaches the preset trigger time threshold. If not, continue monitoring. If yes, trigger the target; A3. Determine whether the target area is the target that the user intends to select. If yes, trigger the selection and turn the icon green. If not, trigger the selection and turn the icon red, and record the experimental data; A4. Use the handle to confirm whether the selection is in line with the user's intention and record the result.

[0041] The unintentional eye control task includes: B1. The user selects the target area through eye control and confirms the selection with the handle; B2. After a random time of 4s to 6s, the handle input is detected. If not confirmed, return to step B1; B3. It is determined that the user has no intention and other experimental data are recorded.

[0042] Experiment 1 in S2 is divided into five tasks, with different types of trigger targets, including: Task 1, Icon matching, the system displays a prompt icon, and the participant selects the same icon as the prompt icon from the menu by gazing; Task 2, Virtual keyboard input, the system presents a simplified QWERTY keyboard, and the participant enters text by gazing on the corresponding letter keys to respond to prompt information; Task 3, Keyword matching, a keyword is displayed at the top of the interface, and six descriptive phrases are listed below. The participant needs to select the phrase that best matches the keyword; Task 4, Picture selection, the system displays a prompt image, and the participant selects the picture that matches the prompt image by gazing; Task 5, the participant freely browses the pictures in the virtual reality interface and uses the handheld controller to trigger the target operation to enter the next round.

[0043] Tasks one to four are intentional eye control tasks. After finding the target button, the subjects trigger the target button by gazing at it. After staring at the button for a preset time, the target button will be displayed in different colors depending on whether the triggering operation is successful or not. A green button will be displayed for success and a red button will be displayed for failure. Then, the participants use the handle to choose whether the picture meets their intention.

[0044] Task five was an unintentional eye control task, in which the subjects used the handle to assist in selecting the next selection interface, and the user intention data were all recorded as unintentional.

[0045] S3 specifically includes: cleaning and standardizing the eye tracking data collected in Experiment 1, extracting pupil diameter, line of sight coordinates and basic eye movement features; using a variety of machine learning algorithms to train the intention recognition model, and integrating the model into the back-end of the virtual reality eye-controlled interaction platform to achieve real-time intention recognition and interface optimization.

[0046] S4 specifically includes: setting up a second experiment of eye control and intention recognition with multiple targets, setting up a gaze control interface for an online shopping platform, where users trigger interactive elements by gazing, and the system determines the trigger result by detecting the gaze duration threshold and activation position. When users want to activate a button, they need to keep gazing until the action is triggered; the intention recognition model determines whether the user intends to trigger the interactive target by analyzing the user's current eye movement behavior. The judgment result is represented by the Iswish indicator, which combines the accuracy of the task with the intention value to evaluate the difficulty of the task, and adjusts the trigger target size and trigger time according to the difficulty of different tasks to improve user experience and interaction accuracy.

[0047] The interface divisions include: Area A, product category area; Area B, text input area; Area C, product recommendation area; Area D, product display area.

[0048] Embodiment 2

[0049] (1) Building an experimental platform

[0050] ① The experimental program was developed using Unity and C# and ran on an HTC VIVE Pro Eye (HTC, Taiwan, China) virtual reality (VR) head-mounted display device equipped with a Tobii eye tracking module; the VR device captured gaze data at a sampling rate of 90 Hz and provided a 110° field of view; the device was connected to a desktop computer equipped with an Intel Core i7-10800F processor and an NVIDIA GeForce RTX 3090 graphics card; the background color in the VR environment was light gray (RGB164, 164, 164).

[0051] ② Participants were instructed to complete the tasks in each scenario using the gaze selection mechanism established in our previous study, which activates the interactive object by detecting the gaze area and maintaining a dwell time of 800 milliseconds. For example, in the text input task, in order to input the letter "A", participants needed to find the "A" button and maintain gaze on it for 800 milliseconds to complete the input.

[0052] ③ Collect icons, texts, and pictures as search targets for different tasks; In order to collect search targets for different tasks, the present invention collects icons, letters, texts, and pictures as interactive targets; In the icon selection task (task one), 50 icons were collected, and the system displayed a prompt icon. The participant selected the same icon as the prompt icon from the menu by gazing; The text input task (task two) used a simplified QWERTY keyboard, and the participant entered text by gazing on the corresponding letter key to respond to the prompt information; In the prompt selection task (task three), a word was displayed at the top of the interface, and six descriptive phrases were listed below. The participant was required to select the phrase that best matched the word. In the image selection task (task four), the participant selected the picture that matched the prompt image by gazing; and in the browsing task (task five), the participant browsed the pictures freely and used the handheld controller to enter the next round.

[0053] ④ Set the layout of interactive targets. According to different experimental tasks, each layout randomly selects the same number of icons or texts from all task search targets.

[0054] (2) Experimental task setting and data collection

[0055] ① Design of Experiment 1 with diverse target eye control and intention recognition. Participants were required to complete five different types of tasks. Each participant completed 40 intentional eye control tasks in Tasks 1 to 4, and 70 unintentional eye control tasks in Task 5. In the intentional eye control task, a prompt image (Tasks 1, 2, and 4) or a prompt phrase (Task 3) was displayed on the screen, and participants were required to find and trigger the corresponding target according to the prompt. If the correct target was triggered, the trigger object would turn green, and if the trigger was wrong, the trigger object would turn red. Subsequently, participants would use the handle to confirm whether the selection was in line with the user's intention. In the unintentional eye control task (Task 5), no target would be displayed on the screen, and the user would browse the images randomly and use the handle to confirm the end.

[0056] Figure 2 It is an intentional eye control task, specifically including: A1. Obtain and continuously monitor the dwell time of the user's gaze point in the target area; A2. Determine whether the dwell time reaches the preset trigger time threshold. If not, continue monitoring. If yes, trigger the target; A3. Determine whether the target area is the target that the user intends to select. If yes, trigger the selection and turn the object green. If not, trigger the selection and turn the object red, and record the experimental data. A4. Use the handle to confirm whether the selection meets the user's intention and record the result.

[0057] Figure 3 It is unintentional eye control, which specifically includes the following steps: B1, the user selects the target area through eye movement control, and uses the handle to confirm whether the target selection is in line with the intention; B2, obtains the handle input signal, and determines whether the user has performed the selection confirmation operation. If the selection confirmation is not performed, returns to step B1; B3, triggers the target selection and records the eye movement data and trigger time.

[0058] ②Forty college students were recruited through the campus website to participate in the experiment (20 males and 20 females, with an average age of 23.07 years and an age standard deviation of 3.15). All participants had normal or corrected-to-normal vision. To ensure the accuracy of eye tracking data, myopia correction was limited to less than 600 degrees. In addition, participants were required to have no known symptoms of 3D vertigo and were advised to get enough rest before the experiment to avoid eye fatigue. The research protocol was approved by the university ethics committee, and all participants signed an informed consent form before the experiment. After completing the experiment, participants were paid $8 per hour.

[0059] ③ The subjects followed the following steps to perform the multi-target eye control and intention recognition experiment 1. The virtual reality screen first displayed the task introduction, followed by the interactive target for 1000 milliseconds. Then the multi-target eye control interface appeared and the timer started. After the trigger was successful or failed, the intention recognition interface was entered. The system experimental program recorded the trigger time, and the next trial started after the end. The experimental process is as follows Figure 8 shown.

[0060] (3) Training and integration of intent recognition models

[0061] ① The extended feature dataset collected in Experiment 1 was preprocessed by deleting columns irrelevant to intent recognition (such as Task, Trail, TriTime, and GazeCnt), and deleting rows containing missing values ​​to ensure data integrity and accuracy. Subsequently, the features in the dataset were separated from the labels, with the label being isWish, and divided into a training set of 80% and an independent test set of 20%. A stratified sampling method was used to maintain the consistency of label distribution.

[0062] ②In the feature engineering stage, detailed feature extraction is performed on the cleaned data, including 38 pupil diameter features (covering 20 statistical features, 9 trend features and 9 change rate features), 324 gaze coordinate features (including 40 statistical features, 142 trend features and 142 change rate features) and 2 basic eye movement features (GazeTime and GazeDensity); all extracted features are standardized to eliminate the impact of dimensional differences on model training.

[0063] Next, four algorithms, LightGBM, random forest, logistic regression and support vector machine, were used for model training. For each model, StratifiedKFold was used for evaluation, and multiple evaluation indicators such as accuracy, precision, recall, F1 score and AUC were recorded.

[0064] ④ In order to intuitively compare the performance of each model, a visual analysis of the performance of different algorithms on each evaluation indicator was performed; Figure 5 The average scores and standard deviations of the four models on five indicators, namely AUC, accuracy, precision, recall and F1 score, are shown; LightGBM performs best in all evaluation indicators, with an AUC value of 0.9698, accuracy of 0.9443, precision of 0.9352, recall of 0.9708 and F1 score of 0.9527, which are significantly higher than other models, indicating that LightGBM has stronger predictive and generalization capabilities in intent recognition tasks.

[0065] ⑤After training the model using the LightGBM algorithm, the importance of the features was analyzed. According to the trained LightGBM model, the relative importance of each feature was extracted, and the top 20 most important features were displayed. From the results, the GazeTime feature has the highest importance, followed by y_all_std, x_abs_std, and y_abs_mean, showing the key role of these eye movement-related features in predicting the target variable (isWish).

[0066] ⑥ Model preservation and deployment. After determining the best intent recognition model, it is serialized and saved so that it can be called in subsequent system integration. Pickle is used as a serialization tool to serialize and save the trained LightGBM model to the server storage system to ensure the integrity and accessibility of the model data. Through Pickle serialization, the model can be efficiently loaded and deployed to support subsequent real-time application requirements. The saved intent recognition model is integrated into the back-end system of the virtual reality eye-controlled interaction experimental platform. The back-end system is developed using the Flask framework. The communication interface between the front-end and the back-end is configured through Flask to achieve real-time data transmission and seamless connection of intent recognition functions. The Flask back-end is responsible for receiving eye movement data requests from the front-end, calling the serialized intent recognition model for analysis, and returning the prediction results to the front-end for interface optimization.

[0067] (4) Build and evaluate adaptive interfaces.

[0068] ① Set up the second experiment of multi-target eye control and intention recognition, and set up the gaze control interface for the online shopping platform: the interface divides different functions into different areas by color blocks, including product category area (area A), text input area (area B), product recommendation area (area C) and product display area (area D). Users trigger interactive elements by gazing, and the system determines the trigger result by detecting the gaze duration threshold and activation position; when users want to activate a button, they need to keep gazing until the action is triggered; the process of the second experiment of multi-target eye control and intention recognition is as follows Figure 4 As shown, the interface layout is as follows Fig. 9 shown.

[0069] ②The task simulated a shopping scenario. Participants needed to complete a product purchase task such as “buy a double-door refrigerator”. Two methods of completing the task were provided: category-based selection or direct text input.

[0070] In the category-based method, participants were required to select “refrigerator” from the product category area, then select “two-door” from the product recommendation area, and finally select a two-door refrigerator from the product display area.

[0071] In the text input method, participants directly input “two-door refrigerator” in the text input area and then select the desired product from the product display area.

[0072] ③ The intention recognition model determines whether the user intends to trigger the interaction target by analyzing the user's current eye movement behavior, which is represented by the Iswish indicator; by combining the average accuracy and average intention value of the last three tasks, the difficulty of the task can be evaluated, and the gaze control interface can be adjusted accordingly according to different task difficulties. The process is as follows: Fig.10 shown.

[0073] The model considers two specific cases:

[0074] When Iswish = yes and ACC > 60%: This indicates that the user can complete the trigger task easily and accurately, which means that the current interface is relatively simple for the user and there is no need to modify the interface.

[0075] When Iswish = yes and ACC < 60%: This indicates that the user has difficulty completing the task, indicating that the current interface is quite challenging. To improve accuracy and reduce false triggers, increase the target size by 20% and the trigger time by 20%.

[0076] ④ To further explore the changing rules of trigger performance over time in the adaptive interface, this study divided the experimental process into four stages in chronological order, and systematically analyzed the changing trends and significant differences of trigger accuracy and trigger time in each stage. At the same time, these differences were further explained in combination with the dynamic adjustment of button size in the adaptive interface.

[0077] The analysis results of the trigger accuracy showed that there were significant differences in the trigger accuracy in different experimental stages (F(3,03) = 17.232, p < 0.001,η² = 0.334). Fig.11 As shown in the figure, the trigger accuracy rate significantly increased from stage 1 to stage 2 and stabilized after stage 2. Pairwise comparison results showed that the difference between stage 1 and stages 2, 3, and 4 was significant (p < 0.05), while the difference between stages 2, 3, and 4 was not significant. This result shows that in the adaptive interface, when the button size is small, the trigger accuracy rate is relatively low. However, as the button size increases, the trigger accuracy rate increases significantly. When the button size reaches a certain threshold, the trigger accuracy rate tends to stabilize, indicating that further increasing the button size has limited effect on improving the accuracy rate. This result verifies the important role of button size in improving interaction accuracy.

[0078] There were also significant differences in triggering time between different experimental phases (F(3,103) = 34.509, p <0.001,η² = 0.501). Fig.12 As shown in the figure, the trigger time decreased significantly from stage 1 to stage 2 and gradually stabilized after stage 2. The results of pairwise comparisons showed that the difference between stage 1 and stages 2, 3, and 4 was significant (p < 0.05), while the difference between stages 2, 3, and 4 was not significant. This result shows that increasing the button size in the adaptive interface can significantly shorten the trigger time, among which the decrease from stage 1 to stage 2 is the most obvious. When the button size reaches a certain threshold, the trigger time gradually stabilizes, which means that the effect of further increasing the button size on improving the trigger efficiency tends to saturate.

[0079] The present invention can be verified by objective results:

[0080] First, in the process of virtual reality eye-controlled interaction, the trigger time and target size of the interactive target have a significant impact on the accuracy of various target eye-controlled tasks. Through multi-dimensional adaptive optimization, the accuracy and response time of gaze-triggered interaction can be significantly improved, thereby improving the overall interaction efficiency and user experience.

[0081] Second, eye tracking data has high accuracy and reliability in predicting user operation intentions through model training.

[0082] Third, the training method of 5-fold cross-validation effectively utilizes data resources and significantly reduces the risk of overfitting of the model, ensuring the good generalization ability of the model on different data sets. In addition, the performance of four machine learning models, random forest, logistic regression, support vector machine (SVM) and LightGBM, was compared. The results showed that the LightGBM model outperformed other models in terms of AUC value and accuracy, verifying its superiority and applicability in the present invention.

[0083] The present invention verifies the prediction accuracy of the intention recognition model through objective experimental data and model performance evaluation, develops an adaptive interface optimization system, and provides a scientific and effective eye-controlled interface design optimization method.

[0084] In the description of the present invention, it should be noted that, unless otherwise clearly specified and limited, the terms "connected" and "connection" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0085] Although the present invention has been described in detail with reference to the aforementioned embodiments, it is still possible for those skilled in the art to modify the technical solutions described in the aforementioned embodiments, or to make equivalent substitutions for some of the technical features therein. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for designing an eye-controlled adaptive interface based on intention recognition, characterized in that: The steps include: S1. Build an experimental platform, construct a virtual reality eye-controlled interaction test platform, configure an eye-movement interaction interface equipped with a virtual reality screen on the test platform, and define specific and diverse interaction targets on the virtual reality eye-controlled interaction platform, including icons, letters, paragraphs, and pictures; S2. Experimental task setting and data collection: set the layout of eye control targets, including the number, coordinates, and distribution of targets; design a multi-target eye control and intention recognition experiment, including five tasks: icon matching, virtual keyboard input, keyword matching, picture selection, and free browsing; recruit experimental participants to complete the tasks; collect and record the task reaction time, recognition accuracy, eye movement trajectory, and intention data of the participants during the experiment; S3, training and integration of intent recognition model, preprocessing and feature extraction of collected data, applying machine learning methods to train intent recognition model, including support vector machine, logistic regression, LightGBM, random forest, and deploying the model to the backend of the virtual reality eye-controlled interaction platform; S4, adaptive optimization and system integration, using the trained intention recognition model, set up the second experiment of eye control and intention recognition of various targets, including keyword matching and picture selection, and dynamically adjust the size and trigger time of the target according to the model prediction results.

2. The method for designing an eye-controlled adaptive interface based on intention recognition according to claim 1, characterized in that: The S1 specifically includes: Configure the VR eye-control interaction interface, define various interaction targets, including icons, letters, paragraphs, and pictures, and set two types of tasks: intentional eye-control tasks and unintentional eye-control tasks; In the intentional eye control task, the user selects the target area through eye control and uses the handle to confirm whether the selection is in line with the intention; In the unintentional eye control task, the user made random eye controls and confirmed them using the handle.

3. The method for designing an eye-controlled adaptive interface based on intention recognition according to claim 2, characterized in that: The intentional eye control task includes: A1. Obtain and continuously monitor the time the user's gaze stays in the target area; A2, determine whether the stay time reaches the preset trigger time threshold, if not, continue monitoring, if yes, trigger the target; A3. Determine whether the target area is the target that the user intends to select. If so, trigger the selection and turn the icon green. If not, trigger the selection and turn the icon red, and record the experimental data. A4. Use the handle to confirm whether the selection is consistent with the user's intention and record the result.

4. The method for designing an eye-controlled adaptive interface based on intention recognition according to claim 2, characterized in that: The unintentional eye control task includes: B1. The user selects the target area through eye control and confirms the selection using the handle; B2. After a random time of 4 to 6 seconds, the controller input is detected. If not confirmed, the process returns to step B1. B3. Determine that the user has no intention and record other experimental data.

5. The method for designing an eye-controlled adaptive interface based on intention recognition according to claim 1, characterized in that: The experiment in S2 is divided into five tasks, with different types of trigger targets, including: Task 1: Icon matching: the system displays a prompt icon, and the participants select the same icon as the prompt icon from the menu by gazing; Task 2: virtual keyboard input. The system presents a simplified QWERTY keyboard. Participants input text by gazing on the corresponding letter keys and responding to prompts. Task 3: Keyword matching: a keyword is displayed at the top of the interface, and six descriptive phrases are listed below. Participants need to choose the phrase that best matches the keyword; Task 4: Picture selection: the system displays a prompt image and participants select the picture that matches the prompt image by gazing; Task 5: Participants browsed pictures freely in the VR interface and used handheld controllers to trigger target actions to enter the next round.

6. The method for designing an eye-controlled adaptive interface based on intention recognition according to claim 5, characterized in that: Tasks one to four are intentional eye control tasks. After finding the target button, the subjects trigger the target button by gazing at it. After staring at the button for a preset time, the target button will be displayed in different colors depending on whether the triggering operation is successful or not. A green button is displayed for success and a red button is displayed for failure. Then, the participants use the handle to choose whether the picture meets their intentions.

7. The method for designing an eye-controlled adaptive interface based on intention recognition according to claim 5, characterized in that: Task five is an unintentional eye control task, in which the subjects use the handle to assist in selecting the next selection interface, and the user intention data are all recorded as unintentional.

8. The method for designing an eye-controlled adaptive interface based on intention recognition according to claim 1, characterized in that: The S3 specifically includes: The eye tracking data collected in Experiment 1 were cleaned and standardized to extract pupil diameter, gaze coordinates, and basic eye movement features. A variety of machine learning algorithms are used to train the intention recognition model, and the model is integrated into the backend of the virtual reality eye-controlled interaction platform to achieve real-time intention recognition and interface optimization.

9. The method for designing an eye-controlled adaptive interface based on intention recognition according to claim 1, characterized in that: The S4 specifically includes: Experiment 2: Setting up a gaze control interface for an online shopping platform. Users trigger interactive elements by gazing. The system determines the trigger result by detecting the gaze duration threshold and activation position. When users want to activate a button, they need to keep gazing until the action is triggered. The intention recognition model analyzes the user's current eye movement behavior to determine whether the user intends to trigger the interaction target. The judgment result is represented by the Iswish indicator. The task difficulty is evaluated by combining the accuracy and intention value of the task. The trigger target size and trigger time are adjusted according to the difficulty of different tasks to improve user experience and interaction accuracy.

10. The method for designing an eye-controlled adaptive interface based on intention recognition according to claim 9, characterized in that: The interface partitioning includes: Area A, product category area; Area B, text input area; Area C, product recommendation area; Area D, product display area.

Citation Information

Patent Citations

  • Eye movement interaction method, system and device based on eye movement tracking technology

    CN111949131A

  • Interactive target layout method and system for eye control interface in virtual environment

    CN117632330A

  • Eye movement selection interaction intention recognition method

    CN117742490A