A deep learning-based VRAR binocular 3D target positioning method

By combining the pupil-corneal reflection vector method and deep learning in virtual reality, a binocular 3D target localization model is constructed, which solves the shortcomings of existing 3D localization technologies, achieves accurate localization of personalized users, and improves robustness and adaptability.

CN116503475BActive Publication Date: 2025-11-18NANJING BOTUO VISION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310357710.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-06
Publication Date
2025-11-18
Estimated Expiration
2043-04-06

AI Technical Summary

Technical Problem

Existing eye-tracking technology cannot achieve 3D positional positioning, and its robustness is poor, especially for users with abnormal eyes or abnormal eye habits. Furthermore, the need for short-range 3D target positioning in virtual reality technology remains unmet.

Method used

Based on the physical structure of the pupil-corneal reflection vector method, combined with binocular eye image processing and deep learning, a binocular 3D target localization model is constructed. By deploying points of interest in virtual space, collecting and analyzing eye map video data, and using a feature fusion module to achieve 3D localization, the robustness is improved.

Benefits of technology

It achieves accurate positioning of 3D targets in virtual reality environments, enhances adaptability to individual users, and does not increase additional hardware costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116503475B_ABST
    Figure CN116503475B_ABST
Patent Text Reader

Abstract

The application discloses a VRAR binocular 3D target positioning method based on deep learning, which comprises the following steps: collecting an eye chart video of a human eye tracking interest point Pc variation, analyzing the state when the eyes are stable, and analyzing the time required for a position change of the eye tracking interest point; a binocular 3D positioning model is constructed, including a feature extraction model based on a pupil-corneal reflection vector method, a 3D positioning model and a feature fusion module; each small eye chart change video and various features are taken as input parameters, and the three-dimensional coordinates of the interest point are taken as output parameters, which are input into a position recognition model for training and learning to obtain a trained position recognition model which is used in the practical stage; then, the position recognition model is saved and updated to a data set and used as a personal data set of a user to improve the adaptability of the personal model to the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a localization technology for the eye's ability to track 3D targets in the field of virtual reality. Specifically, it is a deep learning-based scheme that obtains the 3D position of the target being focused on by taking pictures of both eyes. Technical Background

[0002] In the current field of eye-tracking technology, the main research is based on monocular localization studies. The main methods include eye movement measurement methods, which have gradually developed from early direct observation and subjective perception methods to pupil-corneal reflection vector method, electrooculography (EOG) method, iris-scleral edge method, corneal reflection method, double Pukin elephant method, contact lens method, etc.

[0003] The main approach of these methods is based on a precise modeling architecture. These approaches achieve pixel-level accuracy primarily through precise measurement and calculation. However, such approaches have two problems:

[0004] [1]. Existing solutions do not conduct research on 3D position localization. This is because existing eye-tracking technology uses a precise measurement scheme, which can accurately measure on a 2D plane. However, it cannot obtain motion parameters such as eye convergence to perform depth localization. The details of eye convergence are related to individual eye size, muscle changes, and movement habits, making it a personalized field of motion recognition.

[0005] [2]. If the user's eyes are not normal, or the user does not have normal eye habits, such as a single artificial eye or strabismus, the measurement cannot be accurate, that is, it does not have good robustness.

[0006] Currently, the rapid development of virtual reality technology has placed demands on short-range 3D target positioning technology, especially on the short-range, light-load product performance requirements brought by VRAR structures based on the existing PANCAKE solution. Summary of the Invention

[0007] Based on existing image processing and machine learning theories, this invention proposes an algorithm that can achieve 3D positioning, improve robustness, and enable personalized customization using binocular vision, building upon the physical and algorithmic structure of the original pupil-corneal reflection vector method, without increasing additional solution costs.

[0008] The main components of this plan consist of two phases: a learning phase and an application phase. The learning phase includes steps such as learning data collection, learning data analysis and segmentation, dataset updates, and model training. The application phase includes steps such as practical data collection, practical model analysis, and feedback updates.

[0009] Specifically, the present invention provides a VR / AR binocular 3D target localization method based on deep learning, comprising the following steps:

[0010] Step 1, Build the dataset

[0011] Points of interest (Pcs) with constantly changing positions are deployed in virtual space. The user's eyes track and gaze at the Pcs with constantly changing positions. An eye diagram camera records eye diagram video data during this process.

[0012] The time interval for each change in the position of the point of interest Pc is TFreq1, and the corresponding number of video frames is sf_TFreq1;

[0013] The eye diagram videos of the left and right eyes changing with the point of interest Pc within the time period TFreq1 are denoted as Study_Lefteye_V(i,userid) and Study_Righteye_V(i,userid), respectively; where i represents the i-th position of the point of interest Pc, and userid is the user ID; the position of the point of interest Pc(i) is represented as: Pc(i)=(xi,yi,zi);

[0014] Step 2: Analyze the images in the eye diagram videos Study_Lefteye_V(i,userid) and Study_Righteye_V(i,userid) that tend to stabilize after data changes, and obtain the frame number isteady(framei,i) of the eye diagram in the i-th tracking video of the user id when the human eye begins to stabilize its gaze.

[0015] Find the corresponding stable frame images Study_Lefteye_V(isteady(framei,i),userid) and Study_Righteye_V(isteady(framei,i),userid) in the left and right eye diagram videos;

[0016] Step 3: Analyze the image frames with the greatest intensity of human eye movement change in the eye map videos Study_Lefteye_V(i,userid) and Study_Righteye_V(i,userid), i.e., isummax_left(framei,i,userid) and i.e. isummax_right(framei,i,userid). The image corresponding to this frame number represents the image with the greatest intensity of human eye movement change when the userid is tracking the i-th position.

[0017] Step 4: Only retain the eye map videos between the frames with the greatest intensity of human eye movement changes and the stable frame images in the eye map videos Study_Lefteye_V(i,userid) and Study_Righteye_V(i,userid) for model training;

[0018] Step 5: Construct a binocular 3D target localization model

[0019] The model includes a feature extraction model based on the pupil-corneal reflection vector method, a 3D localization model, and a feature fusion module;

[0020] The feature extraction model based on the pupil-corneal reflection vector method is used to extract the coordinates of the pupil center and the corneal reflection center in eye diagram videos.

[0021] The 3D localization model is used to predict the z-coordinate in the location of the point of interest P1_3D, and to output a high-order feature map to the feature fusion module.

[0022] The feature fusion module is based on time-series feature data. It fuses and analyzes the high-order features obtained from the 3D positioning model with the features of the pupil center and corneal reflection center extracted based on the pupil-corneal reflection vector method to predict the x and y coordinates of the interest point position P1_3D.

[0023] Step 6: Train the binocular 3D target localization model.

[0024] The eye diagram video used for model training in step 4 is input into the feature extraction model based on the pupil-corneal reflection vector method to extract the pupil center and corneal reflection center data;

[0025] The eye map video used for model training in step 4 is input into the 3D localization model to predict the z-coordinate of the interest points;

[0026] Simultaneously, high-order feature maps from the left and right eye images are extracted, concatenated into FF(2*m,(framei,i,userid)), and input together with the pupil center and corneal reflection center data into the feature fusion module for predicting the x and y coordinates of interest points; where m is the number of features in a high-order feature map of an image;

[0027] Finally, a well-trained binocular 3D target localization model is obtained.

[0028] Furthermore, step 7 includes acquiring user eye map videos, finding stable frame images and frames with the greatest intensity of changes in human eye movements, inputting the eye map videos between the stable frame images and frames with the greatest intensity of changes in human eye movements into the trained binocular 3D target localization model, and outputting the localization.

[0029] Furthermore, the 3D positioning model is the VGG+TLE model.

[0030] Furthermore, in steps 2 and 3, optical flow analysis is specifically used to analyze images in eye diagram videos that tend to stabilize after data changes, as well as images with the greatest intensity of changes in human eye movements.

[0031] Furthermore, the feature fusion module in the binocular 3D target localization model includes an input layer, a bidirectional LSTM network layer, a DropOut layer, a fully connected layer, an x / y connection layer, and a softmax regression layer, connected in sequence.

[0032] Furthermore, the eye map video data mentioned in step 1 includes data on changes in the eyeball and data on changes in the muscles around the eyeball, including data on changes in the upper eyelid and eye bags, thereby collecting information reflecting changes in the human eye's gaze at depth information.

[0033] Furthermore, step 2 uses optical flow analysis to identify images in the eye diagram video that tend to stabilize after data changes, as well as images showing the greatest intensity of changes in human eye movement. This specifically includes the following steps:

[0034] Step 2.1: Calculate the optical flow map for each frame in the left and right eye video starting from the second frame;

[0035] Step 2.2, then calculate the sum of the two components (u,v) of all points in a single optical flow diagram, where u and v are the changes on the X and Y axes of the optical flow diagram;

[0036] Step 2.3: Find the maximum value of the sum of components (u,v) in each eye diagram video segment, and the corresponding frame numbers of the left and right eyes, isummax_left(framei,i,userid) and isummax_right(framei,i,userid). The images corresponding to these two frame numbers represent the maximum change in human eye movement when the userid is tracking the i-th position. Here, framei represents the frame number, and framei = 2 to sf_TFreq1.

[0037] Step 2.4: Starting from the frame numbers of the maximum sum of components (u,v) isummax_left(framei,i,userid) and isummax_right(framei,i,userid), and moving forward to the last sf_TFreq1 frame, find stable frames.

[0038] Set thresholds T1 and T2. If the sum of the components (u,v) of the left eye video frame within this time range is less than or equal to T1 * the maximum value of the sum of the components (u,v) of the left eye video frame, and the sum of the components (u,v) of the right eye video frame is less than or equal to T1 * the maximum value of the sum of the components (u,v) of the right eye video frame, and this condition is maintained for T2 frames, then the frame starting from frame T2 is the number of the stable frame.

[0039] Beneficial Effects: This invention employs a physical structure based on the pupil-corneal reflection vector method. This structure involves using infrared illumination, capturing an image of the eye with a miniature camera, and then analyzing the image. The principle is that under infrared illumination, the human retina is insensitive to infrared light and will not interfere with the eye. Because different parts of the eye have different reflectivities and absorptivities for infrared light, the cornea has a high reflectivity, and the pupil and iris regions have significantly different reflectivities and absorptivities. Based on this characteristic, a reflected light spot (Pulchin spot) and a clear pupil will appear in the eye image acquired under an infrared light source. Image processing of the acquired eye image yields the pupil center and the corneal reflected light spot center. As the eyeball rotates, the pupil and corneal reflected light spot change position. Based on the relative offset of the pupil center and the corneal reflected light spot, a relatively accurate fixation point coordinate can be estimated using a specific mapping model. This method causes minimal interference to the user during measurement and is highly accurate, making it a relatively ideal eye movement measurement method.

[0040] This invention improves upon the physical and algorithmic structure of the pupil-corneal reflection vector method by using binocular vision to build a shallow, computationally inefficient deep learning solution. This solution is deployed on existing external VR / AR computing resources. Video of eye changes is captured by a camera on the VR / AR device and sent to the external VR / AR computing resources for computation. This eliminates the need for hardware upgrades, achieving 3D positioning and improved robustness without incurring additional costs. Attached Figure Description

[0041] Figure 1 This is a hardware device diagram for VRAR display and human eye image acquisition according to a specific embodiment of the present invention.

[0042] Figure 2 This is an eye image of a human eye collected by VRAR in a specific embodiment of the present invention.

[0043] Figure 3 This is a diagram showing the positional relationship of several key frames of the left eye in step 1.2 of a specific embodiment of the present invention.

[0044] Figure 4 This is a flowchart of the learning phase in a specific embodiment of the present invention.

[0045] Figure 5 This is a flowchart of the practical stages of a specific embodiment of the present invention.

[0046] Figure 6 This is a structural diagram of a binocular 3D target localization model according to a specific embodiment of the present invention.

[0047] Figure 7 This is a structural diagram of the feature fusion module in a specific embodiment of the present invention. Detailed Implementation

[0048] The following is a detailed explanation. Figures 1-7 The specific embodiments of the present invention will be further explained below.

[0049] In summary, the VRAR binocular 3D target localization method based on deep learning proposed in this invention is based on existing image processing and machine learning theories. On the basis of the original pupil-corneal reflection vector method's physical and algorithmic structure, it uses binocular vision to propose an algorithm that can achieve 3D localization, improve robustness, and realize personalized customization, without increasing the additional solution cost.

[0050] This invention comprises two main phases: a learning phase and a practical phase. The learning phase includes steps such as collecting learning data, analyzing and segmenting the learning data, updating the dataset, and training the model. The practical phase includes steps such as collecting practical data, analyzing practical models, and updating based on feedback.

[0051] Specifically, the present invention provides a VR / AR binocular 3D target localization method based on deep learning, comprising the following steps: Step 1, learning phase

[0052] like Figure 3 As shown, the learning phase involves deploying a point of interest (Pc) with a constantly changing position in VR / AR virtual space; the user inputs accurate refractive error data for their left and right eyes obtained through medical testing; the frequency of the Pc's position change can be set by the user, such as... Figure 1 As shown, an eye imager is set up on the VR / AR device to capture eye map videos, recording the eye changes as the eye tracks a point of interest (Pc) whose position is constantly changing. The user tracks and gazes at the Pc, achieving a stable eye image (eye map). Each time the Pc changes, the data is re-recorded. After acquiring this video data, the starting point of each video is the time of change at each Pc, and the ending point is the image where stable eye tracking is achieved. The entire eye map tracking video data is then segmented into smaller videos corresponding to different Pc changes, each showing the transition from the start of the change to a stable state.

[0053] Then, we analyzed the state of the eye when it was stable and the time required for the eye to track a single change in the position of a point of interest.

[0054] Then, update the data and save it;

[0055] Then, the binocular 3D localization model is trained, including a feature extraction model based on the pupil-corneal reflection vector method, a 3D localization model, and a feature fusion module.

[0056] Then, each segment of eye diagram change video and various features are used as input parameters, and the 3D coordinates of the interest point are used as output parameters. These are input into the location recognition model for training and learning, resulting in a trained location recognition model that can be used in the practical stage.

[0057] Then, the location recognition model is saved and updated into a dataset, which is used as a user's personal dataset to improve the personal model's adaptability to the individual.

[0058] The learning phase includes the following steps:

[0059] Step 1.1, learn data collection.

[0060] The main task of this step is to collect the user's eye usage data. The process involves setting up a VRAR (Virtual Reality Array) to deploy a point of interest (POC) whose position is constantly changing in virtual space; then, the user inputs their eye refractive error data; and then, with the POC's position change frequency (TFreq1) set by the user, the eye-tracking camera on the VRAR continuously captures and records video of the eyes tracking the changing POC, collecting data. The user is required to track and fixate on the POC in a comfortable position to achieve stable tracking.

[0061] Definition of space and coordinate system.

[0062] Define the virtual space (VS) displayed by VR / AR and the real space (RS) in the actual physical space of the human eye.

[0063] Assuming the size and spacing of the VRAR are suitable for the user, and the center of the VRAR's two lenses coincides with the center of the human eye, the origin O of the real space RS coordinate system is defined as the center point of the line connecting the centers of the two eyes. Parallel to the face plane, starting from point O and extending to the right eye, is the X-axis of the RS coordinate system. Parallel to the face plane, starting from point O and perpendicularly downwards, is the Y-axis. Perpendicular to the face plane, starting from point O and moving away from the human eye, is the Z-axis.

[0064] The physical range of the virtual space (VS) in the RS coordinate system is generally: -5mm to 5mm on the X-axis, -4mm to 4mm on the Y-axis, and 0mm to 12mm on the Z-axis. Its visual psychological size is related to the display screen. Currently, the pixel size of the virtual space VS is around 1000(X)*800(Y), with no pixels on the Z-axis. In terms of measurement accuracy, the pupil-corneal reflection vector method generally requires an accuracy of 1 pixel on the X and Y axes, but there is no measurement scheme for the Z-axis, so there is no requirement for it.

[0065] The i-th position of the point of interest Pc is represented as Pc(i) = (xi, yi, zi), and the time interval for position changes is TFreq1 (assuming the corresponding frame number is sf_TFreq1). The point of interest Pc is randomly positioned in X: (-5mm~5mm), Y: (-4mm~4mm), and Z: (0mm~12mm). The number of position changes reaches 5000, resulting in the video data group Study_Pc(i) containing the specific positions of the point of interest in the three coordinates X, Y, and Z, where i = 1~5000.

[0066] At the same time, such as Figure 2 As shown, the eye map camera records images of eye changes as the eye tracks the position of a point of interest (Pc), thus acquiring video data. The physical size range of the data includes the range of changes in the eyeball and the muscles surrounding the eyeball. Figure 2 Image (b) in the image is a photograph of the eyeball. Figure 2 (a) in the image shows the muscle diagram around the eye. Changes in the muscles around the eyeball include changes in the upper eyelid, eye bags, etc., to collect information reflecting changes in the human eye's gaze at depth. The eye diagram changes for both eyes within each time period TFreq1 (total frame count sf_TFreq1) are obtained and saved as video Study_Lefteye_V(i,userid) and video Study_Righteye_V(i,userid), respectively. Video Study_Lefteye_V(i,userid) refers to the learning video of the left eye of user id tracking the i-th change in the position of the point of interest Pc, and video Study_Righteye_V(i,userid) refers to the learning video of the right eye of user id tracking the i-th change in the position of the point of interest Pc.

[0067] It should be noted that the images acquired in this invention patent differ from traditional machine learning methods. This invention patent not only collects changes in the eyeball, but also includes changes in the muscles around the eyeball, such as changes in the upper eyelid and eye bags, thereby collecting information reflecting changes in the human eye's gaze regarding depth. The images acquired in this invention, such as... Figure 2Image (a) shows an eye diagram of a person captured by a camera, as shown below. Figure 2 Image (b) shows the pupil portion of the extracted human eye image.

[0068] Step 1.2 Learning Data Analysis and Segmentation

[0069] The user's changing video learning data (Study_Lefteye_V(i,userid)) within the TFreq1 time period in step 1.1 is analyzed to determine whether the eye state has reached stability and the time of change when stability is reached. The changed video learning data is then removed after stabilization, thereby achieving fine segmentation and ensuring that the 3D positioning model is not interfered with by unnecessary data during training.

[0070] To analyze whether the eye's state stabilizes with each change of the interest point in the constantly changing position during observation step 1.1, a method for extracting key stable frames from a video can be referenced. Such methods include those based on camera boundaries, motion analysis, image information, camera movement, and video clustering. This invention employs a motion analysis-based approach:

[0071] Step 1.2.1 Downsampling.

[0072] Downsampling reduces the resolution of the image, thereby reducing the computational load of subsequent steps 1.2.2. The downsampling scheme involves reserving one pixel every fixed number of pixels along the X and Y axes from both the video Study_Lefteye_V(i,userid) and the video Study_Righteye_V(i,userid), and deleting the rest, thus retaining a low-resolution image Downs_Study_Lefteye_V(i,userid) and Downs_Study_Righteye_V(i,userid). In this embodiment, downsampling is performed by a factor of 4, meaning that one pixel is retained every three pixels along the X and Y axes (after discarding these three pixels).

[0073] Step 1.2.2 Find stable frames using optical flow method.

[0074] Optical flow analysis is used to analyze the stable images of the videos (Downs_Study_Lefteye_V(i,userid), Downs_Study_Righteye_V(i,userid)) after changes, thus determining the stable frame number isteady(framei) of the eye diagram in the i-th tracking video where the human eye begins to fixate. Framei ranges from 2 to sf_TFreq1. Then, the original images (Study_Lefteye_V(isteady(framei),i,userid), Study_Righteye_V(isteady(framei),i,userid)) that have not experienced resolution degradation are retrieved for further analysis. In this embodiment, the optical flow method is the classic HS method.

[0075] The process is as follows:

[0076] [1]. Using the well-known optical flow HS method, the optical flow graphs O_Downs_Study_Lefteye_V(i,userid) and O_Downs_Study_Righteye_V(i,userid) for each frame starting from the second frame of the two videos Downs_Study_Lefteye_V(i,userid) and Downs_Study_Righteye_V(i,userid) are obtained. At this moment, it represents the optical flow graph image of the i-th frame when user userid is tracking the i-th point of interest, where framei = 2 ~ sf_TFreq1.

[0077] [2]. Then calculate the sum of the two components (u, v) of all points in a single optical flow graph (u and v are the changes on the X and Y axes in the optical flow graph): sum_O_Downs_Study_Lefteye_V(framei,i,userid) and sum_O_Downs_Study_Righteye_V(framei,i,userid). These values ​​represent the intensity of the change in human eye motion between framei and the previous framei-1.

[0078] [3]. Then, among all (sf_TFreq1–1) sum_O_Downs_Study_Lefteye_V(framei,i,userid) and sum_O_Downs_Study_Righteye_V(framei,i,userid) in this framei=2~sf_TFreq1, find the maximum values ​​summax_O_Downs_Study_Lefteye_V(i,userid) and summax_O_Downs_Study_Righteye_V(i,userid), and their corresponding left and right eye frame numbers isummax_left(i,userid) and isummax_right(i,userid), as follows: Figure 3 As shown, at this moment, the image corresponding to this frame number represents the maximum change in eye movement (i.e., optical flow value) of the user-id when tracking the i-th position. This is generally the moment when the eye and surrounding muscles undergo the greatest changes while focusing on the target. This value is also used later in the dataset as a feature for the model to learn.

[0079] [4]. Then, as Figure 3 As shown, from the frame number starting from this maximum value to the last frame sf_TFreq1, a stable frame is searched. In principle, the intensity of changes in human eye when focusing is stable is related to the user's own eye habits and physical condition, and it basically remains unchanged for a period of time. However, there may be occasional unconscious blinking (the information of unconscious blinking after stabilization is what this invention aims to eliminate, thereby avoiding the model learning unnecessary information and improving accuracy). Therefore, the solution in this invention is to find the starting point of a relatively stable state within isummax_left(i,userid)~sf_TFreq1 and isummax_right(i,userid)~sf_TFreq1. The solution adopted in this invention is to set two thresholds T1 and T2. When the sum of the components (u,v) of the left eye video frame within this time range is <= T1 * the maximum value of the sum of the components (u,v) of the left eye video frame, and the sum of the components (u,v) of the right eye video frame is <= T1 * the maximum value of the sum of the components (u,v) of the right eye video frame, and this is maintained for T2 frames, then the frame starting from this T2 frame is the stable frame number isteady(i,userid). In this embodiment, T1 is 10%, and T2 is 5 frames.

[0080] Steps 1, 2, and 3. Calculate the time it takes for the human eye to change.

[0081] like Figure 3As shown, the time it takes for the human eye to change from the beginning to the stable frame is the time of isteady(i,userid). The frame number is changed to the time and output as T_isteady(i,userid).

[0082] Then, all i values ​​of T_isteady(i, userid) are statistically analyzed to calculate their mean Tavg_isteady and variance Tdev_isteady. The calculation of the mean and variance is well-known and will not be elaborated here. This value is also used later in the dataset as a feature for the model to learn.

[0083] Step 1.2.4 Refined segmentation.

[0084] like Figure 4 As shown, the purpose of fine segmentation is to delete the stable video frames in the original left and right eye diagram videos Study_Lefteye_V(i,userid) and Study_Righteye_V(i,userid), that is, to delete the images from isteady(i,userid) to sf_TFreq1. The purpose of doing this is to remove the information of the stable subconscious blinking action during the later deep learning recognition and localization, so as to avoid the model learning unnecessary information and improve accuracy, thus obtaining Study2_Lefteye_V(i,userid) and Study2_Righteye_V(i,userid).

[0085] Step 1.3 Dataset Update

[0086] Add the video data from the maximum frames isummax_left(i,userid) and isummax_right(i,userid) in Study2_Lefteye_V(i,userid) to the stable frame isteady(i,userid) in Study2_Righteye_V(i,userid) to the dataset, where i represents the location tracking process of the i-th point of interest, and userid is the user ID. Also save Study2_Lefteye_V(isteady(i,userid),i,userid) as the stable frame image of the left eye diagram, Study2_Righteye_V(isteady(i,userid),i,userid) as the stable frame image of the right eye diagram, and Pc(i) as the position of the i-th point of interest.

[0087] Additionally, for the same userid, among the maximum values ​​of change in all i tracking records, summax_O_Downs_Study_Lefteye_V(i,userid) and summax_O_Downs_Study_Righteye_V(i,userid), find all the maximum values, minimum values, and variances, maxofsummax, minofsummax, and devofsummax. These, along with the maximum values ​​of change in each of the i tracking records, summax_O_Downs_Study_Lefteye_V(i,userid) and summax_O_Downs_Study_Righteye_V(i,userid), are also saved to the corresponding dataset.

[0088] In addition, for the same userid, record the time T_isteady(i,userid) of each change in human eye in step 1.2.3, the average value Tavg_isteady, and the variance Tdev_isteady.

[0089] Step 1.4, Training the binocular 3D target localization model

[0090] like Figure 5 As shown, the binocular 3D target localization model includes a feature extraction model based on the pupil-corneal reflection vector method, a 3D localization model, and a feature fusion module.

[0091] Step 1.4.1 Feature extraction model based on pupil-corneal reflection vector method

[0092] When improving upon the existing physical and algorithmic structure of the pupil-corneal reflection vector method, a feature enhancement approach was adopted to incorporate the superior capabilities of the original algorithm, given the superior performance of the original scheme in calculating precise 2D coordinates. Therefore, features need to be extracted according to the original scheme, and these feature values ​​are used as input and fused with other extracted information to improve robustness.

[0093] Its feature extraction scheme adopts the scheme in the master's thesis "Research on Automated Micro-operation Method Based on Operator's Visual Positioning" by Ma Hui of Wuhan University of Technology, such as... Figure 2 As shown in the right figure, the coordinates of the pupil center and the corneal reflection center are extracted.

[0094] Step 1.4.2, 3D localization model training

[0095] Considering that the area of ​​muscles involved in focusing depth information near each person's eye varies, and that the degree of focusing action differs due to factors such as refractive error and personal habits, we use eye-map videos of individual attention movements as input and employ a temporal network for recognition. The recognition result is the predicted value Pcp(i) in three-dimensional coordinates. The goal is to minimize the difference between the predicted value Pcp(i) and the original true value Pc(i) in three-dimensional coordinates.

[0096] The input image for this scheme is a video of the eye diagrams of the left and right eyes, input frame by frame. The VR / AR glasses emit illumination light from infrared LEDs (which are imperceptible to the human eye) to illuminate the eyes, which are then captured by an eye diagram camera. The captured images are grayscale images, lacking the RGB channels. The content is the same as Study2_Lefteye_V(i,userid) and Study2_Righteye_V(i,userid) in step 1.2.4.

[0097] The output of the model learning in this scheme is the three-dimensional predicted coordinates Pcp(i) of the point of interest. The true values ​​of these data are the coordinates of the point of interest in step 1.1 of Pc(i). The goal of learning is to make the difference between Pcp(i) and Pc(i) as small as possible.

[0098] In terms of the model structure of this scheme, considering that the input of this scheme is three angles, namely the eye diagrams of the left and right eyes, the coordinates of the pupil center and corneal reflection center extracted in step 1.4.1, and some personalized data in step 1.3, the module for position recognition of this invention mainly includes a 3D positioning module and a feature fusion module based on the pupil-corneal reflection vector method. The 3D positioning module is used to analyze the z coordinate of the 3D positioning module from the eye diagrams of the left and right eyes, and the feature fusion module is used to output the x and y coordinates.

[0099] The 3D positioning module is used to analyze the z-coordinate from the left and right eye diagrams.

[0100] like Figure 6 As shown, the 3D positioning module learns the correspondence between input and output from the input eye change video data and its corresponding output original true value Pc(i) three-dimensional coordinates; the video data refers to the video data in Study2_Lefteye_V(i,userid) and Study2_Righteye_V(i,userid) from the maximum frames isummax_left(i,userid) and isummax_right(i,userid) to the stable frame isteady(i,userid).

[0101] The 3D localization module uses the VGG9+TLE model (Deep Temporal Linear Encoding Networks, CVPR 2017) which combines VGG-9 with Temporal Linear Encoding (TLE). The intermediate layer extracts features for the fusion module. This model has relatively low computational requirements. The VGG9+TLE model used in this invention is relatively easy to deploy on external VR / AR computing resources, thereby reducing the need for upgrading existing hardware.

[0102] By training a VGG9+TLE model, a predicted coordinate Pcp(i) with x, y, and z directions can be obtained. Its error compared to the ideal value Pc(i) is small, and it can already be used for prediction. However, considering that its x and y coordinates do not fully utilize the features extracted based on the pupil-corneal reflection vector method (i.e., it does not utilize the results of feature enhancement), and these features are relatively accurate in existing 2D detection schemes, it is necessary to fully utilize these features in the x and y coordinate output to further improve accuracy. This model is publicly available, and its learning process is common knowledge, so it will not be described in detail here.

[0103] The present invention uses a method that fuses the features extracted from the intermediate layer with features based on the pupil-corneal reflection vector method to obtain a coordinate positioning scheme for x and y that has learned relevant information on the time axis for prediction.

[0104] The function of the feature fusion module is to perform fusion analysis on the video data from the maximum frames isummax_left(i,userid) and isummax_right(i,userid) to the stable frame isteady(i,userid) in Study2_Lefteye_V(i,userid) and Study2_Righteye_V(i,userid) based on time-series feature data. It combines the features obtained by the VGG-9 model with two features extracted based on the pupil-corneal reflection vector method: pupil center and corneal reflection center. The fusion results are obtained by fusing the features at the feature layer to obtain a set of appropriate x and y coordinate positioning results.

[0105] like Figure 6 As shown, the input to this module consists of 3 sets:

[0106] Group 1 uses personal data from the dataset in step 1.3. This includes...

[0107] The video data from the maximum frames isummax_left(i,userid) and isummax_right(i,userid) in Study2_Lefteye_V(i,userid) to the stable frame isteady(i,userid) in Study2_Righteye_V(i,userid) are added to the dataset, where i represents the location tracking process of the i-th point of interest, and userid is the user ID. Study2_Lefteye_V(isteady(framei),i,userid) is the stable frame image of the left eye diagram, and Study2_Righteye_V(isteady(framei),i,userid) is the stable frame image of the right eye diagram.

[0108] It also includes the following from step 1.3: for the same userid, the maximum value maxofsummax, minimum value minofsummax, variance devofsummax in all i tracking records, the maximum change value summax_O_Downs_Study_Lefteye_V(i,userid), summax_O_Downs_Study_Righteye_V(i,userid) in each i tracking record, the time of change of human eye T_isteady(i,userid) in each tracking record, the average value Tavg_isteady, and the variance Tdev_isteady, which are also input here.

[0109] Furthermore, considering that the first group consists entirely of statistical data, and statistical data is initially not very reliable, therefore, if... Figure 6 As shown in (a), during the first 100 position learning iterations, the data is not input into the model; instead, a statistical update is performed after each learning iteration. The data is then input into the model after the 100th position learning iteration. Figure 6 As shown in (b).

[0110] The second group adopts the scheme in step 1.4.1. For the input video data Study2_Lefteye_V(i,userid) and Study2_Righteye_V(i,userid), from the maximum frames isummax_left(i,userid) and isummax_right(i,userid) to the stable frame isteady(i,userid), the pupil center FF(11,(framei,i,userid)) and corneal reflection center FF(12,(framei,i,userid)) are extracted from each frame i in the video data. FF represents the feature sequence, and 11 and 12 represent the first and second features of the i-th user in the video tracking the position of the i-th point of interest.

[0111] The third group uses the VGG9 algorithm from step 1.4.2 (VGG9+TLE) to extract the high-order features from the second layer for each frame i of the input video data Study2_Lefteye_V(i,userid) and Study2_Righteye_V(i,userid). Then, the high-order feature maps of the left and right eye images are concatenated into FF(2*m,(framei,i,userid)), where m is the number of features in a high-order feature map, which depends on the VGG-9 settings.

[0112] Then, concatenate the features of groups 1, 2, and 3 to obtain a feature sequence FF(mm,(framei,i,userid)) with 2*m+2 numbers, where mm = 1 to 2*m+2.

[0113] like Figure 5 Figure 6 As shown, its output is the position Pc(i) of the i-th point of interest in step 1.3, including the planar coordinates of x and y, and a depth z coordinate.

[0114] like Figure 7 As shown, the module's structure is a time-series network based on action recognition. The network structure is a simple one: an input layer, a bidirectional LSTM network layer, a dropout layer, a fully connected layer, an x / y connection layer, and a softmax regression layer. The learning method uses the Adam optimizer.

[0115] The number of input layers is the maximum value of the stable frames isteady(framei) during all interest point tracking in step 1.2.2, that is, the maximum number of stable frames isteady(framei) during the 5000 tracking actions in step 1.1.

[0116] A bidirectional LSTM (Bi-LSTM) network has 128 layers.

[0117] The ratio of a dropout layer is 0.5.

[0118] A fully connected layer has 128 neurons.

[0119] A connection layer of x\y consists of 2 neurons.

[0120] A softmax regression layer is a standard softmax regression layer.

[0121] Through this module 2, the present invention can perform fusion at the feature layer to obtain a set of suitable x and y coordinate positioning results in a time series manner.

[0122] Finally, the x and y coordinate positioning results predicted by module 2 and the z coordinate positioning results calculated by module 1 are merged to output Pcp. Since it is a classic bidirectional LSTM model, its learning objective is to minimize the difference between the X and Y values ​​of Pcp(i) and Pc(i). The learning process is well known and will not be described in detail here.

[0123] Save the model parameter for userid to the dataset.

[0124] Step 2, Practical Application Stage

[0125] The practical application phase includes steps such as practical data collection, practical model analysis, and feedback updates. Overall, as follows... Figure 5 As shown.

[0126] Step 2.1 Practical Data Collection

[0127] The practical data acquisition phase involves performing motion localization on existing, continuously acquired images to identify a complete and changing eye map-based motion video.

[0128] [1]. Retrieve all previous data of the current user ID recorded in step 1.3 from the dataset.

[0129] [2]. i = 0, for the currently acquired image (App_Lefteye_V(i,userid), App_Righteye_V(i,userid)), perform downsampling as in step 1.2.1 to obtain (Downs_App_Lefteye_V(i,userid), Downs_App_Righteye_V(i,userid));

[0130] [3]. For the downsampled image, the optical flow method is used in step 1.2.2 to analyze the image of the current application video that tends to stabilize after the change. The statistics are performed according to the method of calculating the sum of u and v in step [2] of step 1.2.2.

[0131] If there exists a value less than minofsummax-3*devofsummax stored in step 1.3, and this value is maintained for T2 frames, then it marks the beginning of a stable frame appsteadyi.

[0132] Starting from this stable frame number, trace back Tavg_isteady+3*Tdev_isteady frames (this value comes from the latest version of step 1.3), and obtain the corresponding frame images from the original image (App_Lefteye_V(framei,i,userid), App_Righteye_V(framei,i,userid)) (framei ranges from the appsteadyi-Tavg_isteady-3*Tdev_isteady frame to the appsteadyi frame). These images represent the changing motion videos. At this point, i = i+1, and proceed to the next motion detection process until all acquired video data has been detected.

[0133] Step 2.2 Practical Model Analysis

[0134] The corresponding frame images (framei from the appsteadyi-Tavg_isteady-3*Tdev_isteady frame to the appsteadyi frame) obtained from the action extraction (App_Lefteye_V(framei,i,userid)) are input into the model in step 1.4 to obtain Pcp(i).

[0135] Step 2.3 Feedback Update

[0136] Once the data from step 2.2 is obtained, the user can input the video data from step 2.1 and the Pcp(i) from step 2.2 into the dataset, update the dataset, update the user's model parameters according to the scheme in step 1.3, and save them into the dataset.

[0137] By implementing the solution of this invention, users can build an algorithm that can achieve 3D positioning, improve robustness, and enable personalized customization based on the existing physical and algorithmic structure of the pupil-corneal reflection vector method, using binocular vision, without increasing the additional solution cost.

Claims

1. A VR / AR binocular 3D target localization method based on deep learning, characterized in that, Includes the following steps: Step 1, Build the dataset Points of interest (Pcs) with constantly changing positions are deployed in virtual space. The user's eyes track and gaze at the Pcs with constantly changing positions. An eye diagram camera records eye diagram video data during this process. The time interval for each change in the position of the point of interest Pc is TFreq1, and the corresponding number of video frames is sf_TFreq1; The eye diagram videos of the left and right eyes changing with the point of interest Pc within the time period TFreq1 are denoted as Study_Lefteye_V(i,userid) and Study_Righteye_V(i,userid), respectively; where i represents the i-th position of the point of interest Pc, and userid is the user ID; the position of the point of interest Pc(i) is represented as: Pc(i)=(xi,yi,zi); Step 2: Analyze the images in the eye diagram videos Study_Lefteye_V(i,userid) and Study_Righteye_V(i,userid) that tend to stabilize after data changes, and obtain the frame number isteady(framei,i) of the eye diagram in the i-th tracking video of user id when the human eye begins to stabilize its gaze. Find the corresponding stable frame images Study_Lefteye_V(isteady(framei,i),userid) and Study_Righteye_V(isteady(framei,i),userid) in the left and right eye diagram videos; Step 3: Analyze the image frames with the greatest intensity of human eye movement change in the eye map videos Study_Lefteye_V(i,userid) and Study_Righteye_V(i,userid), i.e., isummax_left(framei,i,userid) and i.e. isummax_right(framei,i,userid). The image corresponding to this frame number represents the image with the greatest intensity of human eye movement change when the userid is tracking the i-th position. Step 4: Only retain the eye map videos between the frames with the greatest intensity of human eye movement changes and the stable frame images in the eye map videos Study_Lefteye_V(i,userid) and Study_Righteye_V(i,userid) for model training; Step 5: Construct a binocular 3D target localization model The model includes a feature extraction model based on the pupil-corneal reflection vector method, a 3D localization model, and a feature fusion module; The feature extraction model based on the pupil-corneal reflection vector method is used to extract the coordinates of the pupil center and the corneal reflection center in the eye diagram video. The 3D localization model is used to predict the z-coordinate in the location of the point of interest P1_3D, and to output a high-order feature map to the feature fusion module. The feature fusion module is based on time-series feature data. It fuses and analyzes the high-order features obtained from the 3D positioning model with the features of the pupil center and corneal reflection center extracted based on the pupil-corneal reflection vector method to predict the x and y coordinates of the interest point position P1_3D. The feature fusion module includes an input layer, a bidirectional LSTM network layer, a DropOut layer, a fully connected layer, an x\y connection layer, and a softmax regression layer connected in sequence. Step 6: Train the binocular 3D target localization model. The eye diagram video used for model training in step 4 is input into the feature extraction model based on the pupil-corneal reflection vector method to extract the pupil center and corneal reflection center data; The eye map video used for model training in step 4 is input into the 3D localization model to predict the z-coordinate of the interest point. At the same time, the high-order feature maps of the left and right eye images are extracted, concatenated into FF(2*m,(framei,i,userid)), and input together with the pupil center and corneal reflection center data into the feature fusion module to predict the x and y coordinates of the interest point. Here, m is the number of features in the high-order feature map of an image. Finally, a well-trained binocular 3D target localization model is obtained.

2. The VR / AR binocular 3D target localization method based on deep learning according to claim 1, characterized in that, Also includes: Step 7: Collect user eye map video, find stable frame images and frames with the greatest intensity of changes in human eye movements, input the eye map video between the stable frame images and frames with the greatest intensity of changes in human eye movements into the trained binocular 3D target localization model, and output the localization.

3. The VR / AR binocular 3D target localization method based on deep learning according to claim 1, characterized in that, The 3D positioning model is the VGG+TLE model.

4. The VR / AR binocular 3D target localization method based on deep learning according to claim 1, characterized in that, In steps 2 and 3, optical flow analysis is used to analyze images in eye diagram videos that tend to stabilize after data changes, as well as images with the greatest intensity of changes in human eye movements.

5. The VR / AR binocular 3D target localization method based on deep learning according to claim 1, characterized in that, The eye image video data mentioned in step 1 includes data on changes in the eyeball and data on changes in the muscles around the eyeball, wherein the changes in the muscles around the eyeball include data on changes in the upper eyelid and eye bags.

6. The VR / AR binocular 3D target localization method based on deep learning according to claim 4, characterized in that: Step 2 uses optical flow analysis to identify images in the eye diagram video that tend to stabilize after data changes, as well as images with the greatest intensity of changes in human eye movements. This includes the following steps: Step 2.1: Calculate the optical flow map for each frame in the left and right eye video starting from the second frame; Step 2.2, then calculate the sum of the two components (u,v) of all points in a single optical flow diagram, where u and v are the changes on the X and Y axes of the optical flow diagram; Step 2.3: Find the maximum value of the sum of components (u,v) in each eye diagram video segment, and the corresponding frame numbers of the left and right eyes, isummax_left(framei,i,userid) and isummax_right(framei,i,userid). The images corresponding to these two frame numbers represent the maximum change in human eye movement when the userid is tracking the i-th position. Here, framei represents the frame number, and framei = 2 to sf_TFreq1. Step 2.4: Starting from the frame number of the maximum sum of components (u,v) in the frame numbers isummax_left(framei,i,userid) and isummax_right(framei,i,userid), and moving forward to the last frame sf_TFreq1, find stable frames. Set thresholds T1 and T2. If the sum of components (u,v) of the left-eye video frame within this time range is less than or equal to T1 * the maximum sum of components (u,v) of the left-eye video frame, and the sum of components (u,v) of the right-eye video frame is less than or equal to T1 * the maximum sum of components (u,v) of the right-eye video frame, and this condition is maintained for T2 frames, then the frame starting from frame T2 is the number of the stable frame.

Citation Information

Patent Citations

  • Binocular eye tracking from video frame sequences

    US9775512B1

  • Eye tracking with prediction and late update to GPU for fast foveated rendering in an HMD environment

    WO2019221979A1