Method for eye movement feature fusion for user identity authentication under VR
By collecting and fusing eye-tracking features in a virtual reality environment, and utilizing eye-tracking behavior and gaze point distribution features, the problem of traditional biometric identification technology being easily deceived in virtual reality has been solved, achieving high-precision user identity verification.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHONGQING UNIV
- Filing Date
- 2023-11-27
- Publication Date
- 2026-07-21
AI Technical Summary
Traditional biometric authentication technologies are susceptible to being deceived by replicas or artificially created models in virtual reality environments, resulting in insufficient reliability of identity verification.
An eye-tracking feature fusion method is adopted. By collecting users' eye-tracking data in a virtual reality environment, user features are extracted and preprocessed. The user's eye-tracking behavior features and gaze point distribution map features are combined, and a machine learning classifier is used for identity identification. The final user classification probability is calculated through weighted fusion.
It improves the success rate and accuracy of user identification, and achieves high-precision user identification by taking advantage of the complexity of eye-tracking behavior patterns that are difficult to be mechanically replicated.
Smart Images

Figure CN117574267B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of machine learning technology and relates to a method for eye-tracking feature fusion for user identification in VR. Background Technology
[0002] In recent years, advancements in virtual reality (VR) technology and the widespread adoption of VR devices have garnered significant attention, particularly the combination of eye tracking and head-mounted displays. These VR headsets can be used in various applications, such as education, training, business, and collaboration. Past research has shown that gaze data is unique to each individual and can reveal personal identity information. Therefore, the ability to identify users based on gaze data opens up new possibilities for interaction in VR applications. For example, personalized experiences can be tailored to users based on their identified identities, and frequent, implicit authentication processes can enhance system security without distracting users when using VR applications.
[0003] Traditional biometric identification technologies mostly use human physiological characteristics, such as fingerprints and irises. These systems based on human physiological characteristics are susceptible to being deceived by replicas or artificially created models. Summary of the Invention
[0004] In view of this, the purpose of this invention is to provide a method for fusing user identification features in virtual reality to improve the success rate of user identification based on eye-tracking data.
[0005] To achieve the above objectives, the present invention provides the following technical solution:
[0006] A method for eye-tracking feature fusion for user authentication in VR includes the following steps:
[0007] S1: Provide a verification code input system for VR, in which eye-tracking data is collected when the user enters the verification code;
[0008] S2: Preprocess eye-tracking data to extract user features;
[0009] S3: Fuse user identity features and calculate the probability of the identification result;
[0010] S4: Output the user authentication result.
[0011] Furthermore, in the VR verification code input system described in step S1, the user is required to input a verification code. One digit of the verification code appears at a random location in the virtual scene, and the user inputs the corresponding verification code by pressing a button on the controller. The user stares at one of the verification codes for a certain period of time, triggering the next digit of the verification code to appear at a random location, until all verification codes are input. During this process, the user's eye movement data is collected.
[0012] Furthermore, the eye-tracking data consists of the following: timestamp T, and spatial coordinates of the pupils of both eyes (x, y, y). left y left , z left ), (x right y right , z right Pupil diameter p d Pupil gaze direction vector V left (x vleft y vleft , z vleft V right (x vright y vright , z vright ), gaze point spatial coordinate data (x p y p , z p ), the object name tag of the fixation point, and the degree of eye opening.
[0013] Furthermore, the user features mentioned in step S2 include user eye movement behavior features and user gaze point distribution map features;
[0014] The user eye movement behavior features include a series of user-related eye movement behavior features, which together form a user feature vector, used as input to a machine learning classifier to output the probability of user classification.
[0015] The user gaze point distribution map features are used to directly calculate the probability of user classification;
[0016] The two types of features are fused using a weighted method to generate the final user classification probability for identification.
[0017] Furthermore, the user's eye-tracking behavior features include:
[0018] Pupil diameter range: The range of pupil diameter when the user is using the VR verification code input system;
[0019] Scanning speed: The speed at which a user scans a CAPTCHA from one CAPTCHA to another, extracting features v in the x, y, and z directions. sx V sy V sz ;in i and j are the sequence indices at different times, x pi Let V be the x-coordinate of the gaze point at that moment. sy and V sz Calculation formula and v sx similar;
[0020] Gazing velocity: The speed of eye tremors when a user gazes at a number block; features v in the x, y, and z directions are extracted. fx V fy V fz ;in i and j are the sequence indices at different times, x pi V represents the x-coordinate of the gaze point at that moment. fy and V fz The calculation formula and v fx similar;
[0021] Single-digit fixation duration: The time from when the user's gaze falls on a particular digit square to when the user leaves that digit square: T d =T i -T j T i T is the moment when the user begins to look at the digital cube. j For the user to watch the final moment of the digital cube;
[0022] Binocular line-of-sight difference: The angular difference between the direction vectors of the binocular lines of sight. Where V lefti V righti Let i be the direction vector of the left and right eye gazes for sequence number i.
[0023] Furthermore, the user gaze point distribution map feature is used to measure the similarity of the gaze point coordinate distribution when different users gaze at the same digital cube. The calculation steps are as follows:
[0024] Step 1: Consider two samples F from the same number block. y and F y F x For samples with known identity labels, F y It is a test sample;
[0025] Step 2: Initialize test sample F y Length n, initial distance and distance = 0;
[0026] Step 3: Select from F y Data point x i Calculation from F y The point with the minimum Euclidean distance y i ;
[0027] Step 4: Calculate data point x i and y i Euclidean distance d i ;
[0028] Step 5: Calculate the distance and sum
[0029] Step 6: Calculate and output the similarity of the coordinate sequences. That is, several test samples can be divided into different user categories based on the similarity p of their coordinate sequences, thereby calculating the probability of user classification.
[0030] Furthermore, in step S3, the probability of a sample being classified into different user categories is calculated based on the user's eye-tracking behavior features and the user's gaze point distribution map features. The two types of features are then fused using a weighted average to obtain the final user classification probability, as shown in the following formula:
[0031] score i =μp ai +(1-μ)p bi
[0032] Where μ is the weight of feature fusion; p ai Let p be the probability of classifying user i based on user eye-tracking behavior features. bi To calculate the probability of user i based on the characteristics of the user gaze point distribution map.
[0033] Furthermore, in step S4, the user is identified as the category with the highest probability. The user identification process is as follows:
[0034] label = i, score i = max(score).
[0035] The beneficial effects of this invention are as follows: This invention provides a method for user identification feature fusion based on eye movement features in virtual reality. Eye movement user identification technology utilizes the behavioral pattern features of eye movement to provide a large amount of information about brain cognitive functions and neural signals controlling eye movement. It is difficult to mechanically replicate such complex eye movement behavioral pattern features. Therefore, based on the collected user eye movement dataset and machine learning classification method, this invention achieves high accuracy and high user identification success rate when identifying users.
[0036] Other advantages, objectives, and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination, or may be learned from practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description
[0037] To make the objectives, technical solutions, and advantages of the present invention clearer, the preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings, wherein:
[0038] Figure 1 A flowchart of an eye-tracking feature fusion method for user identification in VR;
[0039] Figure 2 For the user feature fusion and identification process. Detailed Implementation
[0040] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0041] The accompanying drawings are for illustrative purposes only and are schematic diagrams, not actual pictures. They should not be construed as limiting the invention. To better illustrate the embodiments of the invention, some parts in the drawings may be omitted, enlarged, or reduced, and do not represent the actual product dimensions. It is understandable to those skilled in the art that some well-known structures and their descriptions may be omitted in the drawings.
[0042] In the accompanying drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components. In the description of the present invention, it should be understood that if terms such as "upper," "lower," "left," "right," "front," and "rear" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, they are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, the terms used to describe positional relationships in the drawings are only for illustrative purposes and should not be construed as limiting the present invention. For those skilled in the art, the specific meaning of the above terms can be understood according to the specific circumstances.
[0043] Please see Figure 1 This invention provides a method for user identification feature fusion in virtual reality, the method comprising the following steps: S1, providing a verification code input system in VR, where eye-tracking data is collected when the user inputs the verification code; S2, preprocessing the eye-tracking data to extract user features; S3, calculating the probability of the identification result by fusing user identity features; S4, outputting the user identification result.
[0044] In step S1, the detailed process includes:
[0045] VR CAPTCHA Input System: Users are required to enter a 4-digit CAPTCHA. When the virtual scene starts, a number square (0-9) will randomly appear. Users can enter the corresponding CAPTCHA by pressing a button on the controller. The user must stare at the number square for 1.5 seconds before the next number square appears. During the user's CAPTCHA input process, eye-tracking data is collected for identity verification.
[0046] Eye-tracking data consists of the following: timestamp T, and spatial coordinates of the pupils in both eyes (x, y, y). left y left , z left ), (x right y right , z right Pupil diameter p d Pupil gaze direction vector V left (x vleft y vleft , z vleft V right (x vright y vright , z vright ), gaze point spatial coordinate data (x p y p , z p ), the object name tag of the fixation point, and the degree of eye opening.
[0047] In step S2, user features for identification are generated based on the eye-tracking data extracted in step S1. These user features are divided into two categories: user eye-tracking behavior features and user gaze point distribution map features. The user eye-tracking behavior features contain a series of user-related eye-tracking behavior features, which together form a user feature vector, used as input to a machine learning classifier to output the probability of user classification. The user gaze point distribution map features can directly calculate the probability of user classification. The two types of features are fused using a weighted method to generate the final user classification probability for identification. Detailed feature descriptions are as follows:
[0048] User eye-tracking behavior characteristics include:
[0049] Pupil diameter range: The range of pupil diameter when the user uses the VR verification code input system.
[0050] Scanning speed: The speed at which a user scans their eyes while moving from one number tile to another, extracting features in the x, y, and z directions; where... i and j are the sequence indices at different times, x pi V represents the x-coordinate of the gaze point at that moment. sy V sz The calculations are similar.
[0051] Gazing velocity: The speed of eye tremors when a user gazes at a number block; features are extracted in the x, y, and z directions. i and j are the sequence indices at different times, x pi V represents the x-coordinate of the gaze point at that moment. fy V fz The calculations are similar.
[0052] Single-digit fixation duration: The time from when the user's gaze falls on a particular digit square to when the user leaves that digit square: T d =T i -T j T i T is the moment when the user begins to look at the digital cube. j This allows users to observe the final moments of the digital cube.
[0053] Binocular line-of-sight difference: The angular difference between the direction vectors of the binocular lines of sight. Where V lefti V righti Let i be the direction vector of the left and right eye gazes for sequence number i.
[0054] User gaze point distribution map features: These features measure the similarity of the gaze point coordinate distribution when different users gaze at the same number square. The following algorithm is used to calculate this: Consider two samples F from the same number square. x and F y F x For samples with known identity labels, F y These are test samples. The pseudocode for the algorithm to calculate the matching degree between two coordinate samples is as follows:
[0055]
[0056] The test sample of the user to be identified is categorized into the known user category with the highest matching degree (i.e., the smallest return value p). The probability of identifying the user as a different user can be calculated from several samples of the user to be identified.
[0057] In step S3, as Figure 2 As shown, feature fusion is performed based on the user classification probabilities calculated from the user features extracted in step S2. The detailed process is as follows:
[0058] Feature fusion process: Based on user eye-tracking behavior features and user gaze point distribution map features, the probability of a sample being classified into different user categories can be calculated. The two types of features need to be fused using weighted fusion to obtain the final user classification probability. This is described below:
[0059] score i =μp ai+(1-μ)p bi
[0060] Where μ is the weight of feature fusion, which is 0.7 in this embodiment; p ai Let p be the probability of classifying user i based on user eye-tracking behavior features. bi To calculate the probability of user i based on the characteristics of the user gaze point distribution map.
[0061] In step S4, the end user is identified as the category with the highest probability. The user identification process is as follows:
[0062] label = i, score i =max(score)
[0063] In this embodiment, the dataset used comes from the eye movement data of 20 volunteers participating in a virtual reality user eye movement recognition experiment. The experimental procedure is the VR verification code input system and eye movement data acquisition program provided by this invention. Each volunteer repeats the experiment to collect eye movement data four times, and eye movement calibration is performed before the experiment begins. The experimental procedure is as follows: Figure 1 As shown in the figure. During the experiment, since users may blink, abnormal data such as left and right eye closure will be removed from the collected eye movement data in the data preprocessing stage.
[0064] The proposed identity feature fusion method was compared with user eye-tracking behavior features and user gaze point distribution map features. User identification accuracy was used to evaluate the performance of this method on the dataset. Machine learning classification methods were employed in the user identification process. Table 1 shows the performance of different features in the user identification process.
[0065] Table 1
[0066] Eye movement behavior characteristics 70% Features of gaze point distribution map 30% Fusion of eye movement behavior features and gaze point distribution features 85%
[0067] It can be seen that the accuracy of the gaze point distribution map feature alone is not high, only 30%, but the accuracy is improved after fusion with the eye movement behavior feature. This shows the effectiveness of the gaze point distribution map feature alone in the identity feature fusion method, and also shows the effectiveness of the identity feature fusion method.
[0068] Those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium. When the program is executed, it can implement the steps of the method. The storage medium may be, for example, ROM / RAM, magnetic disk, optical disk, etc.
[0069] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A method for eye-tracking feature fusion for user identification in VR, characterized in that: Includes the following steps: S1: Provide a verification code input system for VR, in which eye-tracking data is collected when the user enters the verification code; S2: Preprocess the eye-tracking data to extract user features; the eye-tracking data consists of the following: timestamp T, spatial coordinates of both pupils (… , , ), ( , , ), pupil diameter Pupil gaze direction vector ( , , ), ( , , ), gaze point spatial coordinate data ( , , ), object name tag at the point of fixation, and degree of eye opening; The user features mentioned in step S2 include user eye movement behavior features and user gaze point distribution map features; The user eye movement behavior features include a series of user-related eye movement behavior features, which together form a user feature vector, used as input to a machine learning classifier to output the probability of user classification. The user gaze point distribution map features are used to directly calculate the probability of user classification; The two types of features are fused using a weighted method to generate the final user classification probability for identification. The user's eye-tracking behavior features include: Pupil diameter range: The range of pupil diameter when the user is using the VR verification code input system; Scanning speed: The speed at which a user scans a CAPTCHA from one CAPTCHA to another, extracting features in the x, y, and z directions. , , ;in , where i and j are the sequence indices at different times. Let x be the x-coordinate of the gaze point at that moment. and Calculation formula and similar; Gazing velocity: The speed of eye tremors when a user gazes at a number block; features are extracted in the x, y, and z directions. , , ;in , where i and j are the sequence indices at different times. Let x be the x-coordinate of the gaze point at that moment; and The calculation formula and similar; Single-digit gaze duration: The time from when the user's gaze falls on a particular digit square to when it leaves that digit square. ,in The moment when the user begins to look at the digital cube, For the user to watch the final moment of the digital cube; Binocular line-of-sight difference: The angular difference between the direction vectors of the binocular lines of sight. ,in , Let i be the direction vector of the left and right eye gazes. S3: Fuse user identity features and calculate the probability of the identification result; S4: Output the user authentication result.
2. The eye-tracking feature fusion method for user identification in VR according to claim 1, characterized in that: In the VR verification code input system described in step S1, the user is required to input a verification code. One digit of the verification code appears at a random location in the virtual scene, and the user inputs the corresponding verification code by pressing a button on the controller. The user stares at one of the verification codes for a certain period of time, triggering the next digit of the verification code to appear at a random location, until all verification codes are input. During this process, the user's eye movement data is collected.
3. The eye-tracking feature fusion method for user identification in VR according to claim 1, characterized in that: The user gaze point distribution map feature is used to measure the similarity of the gaze point coordinate distribution when different users gaze at the same digital square. The calculation steps are as follows: Step 1: Consider two samples from the same number block and ,in These are samples with known identity tags. It is a test sample; Step 2: Initialize the test sample Length n, initial distance and distance = 0; Step 3: Select from data points Calculation from The point with the minimum Euclidean distance ; Step 4: Calculate data points and European distance d i ; Step 5: Calculate the distance and sum Step 6: Calculate and output the similarity of the coordinate sequences. That is, several test samples can be divided into different user categories based on the similarity p of their coordinate sequences, thereby calculating the probability of user classification.
4. The eye-tracking feature fusion method for user identification in VR according to claim 1, characterized in that: In step S3, the probability of a sample being classified into different user categories is calculated based on user eye-tracking behavior features and user gaze point distribution map features. The two types of features are then fused using a weighted average to obtain the final user classification probability, as shown in the following formula: in These are the weights for feature fusion; Let i be the probability of classifying user i based on user eye-tracking behavior features. To calculate the probability of user i based on the characteristics of the user gaze point distribution map.
5. The eye-tracking feature fusion method for user identification in VR according to claim 1, characterized in that: In step S4, the user is identified as the category with the highest probability. The user identification process is as follows: 。