Three-dimensional eyeball model reconstruction method and system based on multi-modal fusion

Through a multimodal fusion three-dimensional eye model reconstruction method, combined with feature extraction of infrared eye diagrams and event diagrams, the Transformer network and three-dimensional convolutional network are used to solve the challenge of line of sight estimation in the prior art and achieve high-precision and robust line of sight prediction.

CN120259536APending Publication Date: 2025-07-04NAT INNOVATION INST OF DEFENSE TECH PLA ACAD OF MILITARY SCI
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510302587.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In the prior art, model-based eye tracking methods have challenges in viewpoint occlusion, lighting changes, head movements, and cross-device adaptability. Single-modal information acquisition is limited, making it difficult to achieve high-precision and robust line-of-sight estimation.

Method used

The three-dimensional eye model reconstruction method with multimodal fusion is adopted, combined with feature extraction of infrared eye diagrams and event diagrams, and time-series feature fusion is performed through the Transformer network, and a three-dimensional convolutional network and a small sample weak supervision technology are used to generate a differentiable 3D eye model to predict the line of sight direction.

Benefits of technology

High-precision and robust line of sight estimation are achieved in a small number of samples, improving the generalization ability of the model and cross-device adaptability, and enhancing the modeling ability of the eyeball model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259536A_ABST
    Figure CN120259536A_ABST
Patent Text Reader

Abstract

The invention discloses a three-dimensional eyeball model reconstruction method and system based on multi-modal fusion, and the method comprises the following steps: S1, feature extraction: respectively extracting features of an original infrared eye diagram and a generated event diagram, carrying out the calculation of channel attention, and fusing the features of two modals in a spatial dimension; s2, time sequence feature fusion: compressing spatial dimensions through global average pooling, taking an ordered sequence of eye features and an ordered sequence of event graph features after spatial multi-modal fusion, and outputting corresponding eye parameters; s3, three-dimensional eyeball model reconstruction: defining an eyeball as a sphere with an eyeball center Oe and an eyeball radius re; generating discrete point clouds from pupils and irises of the 3D eye model; and S4, carrying out weak supervision on few samples. According to the method, the model can better learn a deformable eyeball model, and a robust and high-precision sight line estimation task can be carried out under the condition of a small number of samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of human-computer interaction and can be applied in the fields of gaze estimation, eye movement interaction, computer vision, etc., and particularly relates to a three-dimensional eyeball model reconstruction method and system based on multimodal fusion. Background Art

[0002] With the rapid development of virtual reality and augmented reality technologies, head-mounted devices, as an important human-computer interaction tool, have gradually become the focus of research and commercial fields. Knowledge about human eye fixation plays an important role in behavior analysis, emotion computing, human-computer interaction, and extended reality. Gaze tracking is crucial in many fields, and due to its importance, there have been many studies on improving tracking accuracy, including the use of traditional infrared / RGB cameras and emerging event cameras.

[0003] Eye movement tracking methods are mainly divided into two categories: model-based methods and appearance-based methods. Appearance-based methods directly learn the mapping from eye images to gaze directions. The present invention focuses on model-based methods, which use physiologically inspired eye models to predict fixation. Model-based methods are generally considered to provide better accuracy than appearance-based methods. Model-based methods are about the estimation of the eyeball and the optical axis.

[0004] Most previous eye movement tracking algorithms are frame-by-frame. However, eye movement tracking is a continuous sequence of frames in almost all use cases, and the temporal information is often ignored.

[0005] Therefore, the technical problems to be solved by the present invention are as follows:

[0006] (1) The difficulty in accurately fitting the model is mainly due to the fact that viewpoint changes often lead to occlusion, the ill-posedness of reconstructing a 3D model from an image with only eye semantics, and the change of lighting conditions.

[0007] (2) Appearance-based gaze estimation faces many challenges, such as head movement, subject differences, etc., especially in an unconstrained environment. These factors have a greater impact on the eye appearance and complicate the eye appearance.

[0008] (3) Model-based methods usually have good generalization ability, but they are not differentiable, so effective gaze prediction across systems and devices cannot be performed, that is, knowledge cannot be transferred to new devices.

[0009] (4) The characterization ability of a single-modal image for eye movement features is limited, resulting in limitations in information acquisition. Different-modal data can provide complementary information, and single-modal data cannot fully utilize these potential complementary information. Summary of the Invention

[0010] Aiming at the problems existing in the prior art, the purpose of the present invention is to provide a three-dimensional eyeball model reconstruction method and system based on multi-modal fusion, applying the information of the event graph modality to gaze estimation, combining the advantages of model-based and appearance-based gaze estimation methods, and predicting the gaze direction by predicting a fully differentiable 3D eyeball model, and using event information to enhance the modeling ability of the eyeball model.

[0011] To achieve the above object, the present invention provides a three-dimensional eyeball model reconstruction method based on multi-modal fusion, and the method includes the following steps:

[0012] S1. Feature extraction: Extract features from the original infrared eye image and the generated event graph, respectively extract the features of the original infrared eye image and the generated event graph, calculate the channel attention of the two different modalities of features extracted, and fuse the features of the two modalities in the spatial dimension;

[0013] S2. Temporal feature fusion: Compress in the spatial dimension through global average pooling, take the ordered sequences of the eye features after spatial multi-modal fusion and the ordered sequence of the event graph features, and output the corresponding eye parameters;

[0014] S3. Three-dimensional eyeball model reconstruction: Define the eyeball as a sphere with its eyeball center O e and the eyeball radius r e ; Initialize the eyeball center as Initialize the gaze vector as g c =(0, 0, -1), the eyeball is located at the center of the regular coordinate system, the optical axis is collinear with the regular z-axis and points in the negative direction, and generate discrete point clouds for the pupil and iris of the 3D eye model;

[0015] S4. Few-shot weak supervision: After the eyeball model output by the network is rendered to the 2D plane, in order to supervise the generated point cloud with semantic labels, set an edge loss function to minimize the Euclidean distance between the edge pixels of the projected semantic point cloud and the edge pixels of the true semantic mask.

[0016] Furthermore, in step S1, use ResNet18, EfficientNet, DenseNet, MobileNet or a feature pyramid network as the backbone of the feature extraction network.

[0017] Furthermore, in step S1, the feature extraction and channel attention fusion methods are as follows:

[0018] S1.1 Input the human eye bimodal image I∈R HxWx1, where H and W are the height and width of the image respectively, and each picture is a grayscale image composed of one channel; initialize two backbone feature extractors, and use the infrared image I and the event image E as inputs respectively;

[0019] S1.2 Generate two mappings F1 = B(I) and F2 = B(E); where, contains D channels, and the original spatial dimension is downsampled by a factor of s;

[0020] S1.3 The features obtained from the original infrared eye diagram and the event diagram are fused in the spatial dimension through a lightweight channel attention mechanism ECA to obtain a feature map.

[0021] Furthermore, in step S2, the spatial dimension is compressed through global average pooling to obtain F1 ∈ R 1x1xD , F2 ∈ R 1x1xD represents the global features of the fused event and infrared two-modal eye images and the individual features of the event diagram;

[0022] Use two Transformer-based joint processing networks, convolutional neural networks, long short-term memory networks or gated recurrent units, take the ordered sequence F of the eye features after spatial multi-modal fusion, take the ordered sequence D of the event diagram features, and output the corresponding eye parameters E = T(F1) + T(F2).

[0023] Furthermore, in step S3, it specifically includes:

[0024] S3.1 First, perform mathematical modeling of the three-dimensional eyeball model, and define the eyeball as a sphere with its center O e and eyeball radius r e ; at the same time, learn the focal length f of the camera, and set the camera center to (c x , c y ) = (w / 2, H / 2), so that the center of the camera coordinate system and the center of the image coordinate system coincide on the two-dimensional plane;

[0025] S3.2 Initialize the eyeball center as Initialize the gaze vector as g c = (0, 0, -1), the eyeball is located at the center of the regular coordinate system, the optical axis is collinear with the regular z-axis and points in the negative direction, and the pupil and iris of the 3D eye model are generated into a discrete point cloud;

[0026] S3.3 The three-dimensional eyeball model projects the deformed iris and pupil point cloud regions onto the image plane and obtains the corresponding 2D segmentation mask.

[0027] Furthermore, in step S3.2, it specifically includes:

[0028] S3.2.1 First, generate the discrete point cloud of the pupil disk:

[0029]

[0030] In the same way, the present invention generates the discrete point cloud of the iris disk:

[0031]

[0032] where r p is the radius of the pupil, ρ is the radial distance in polar coordinates, and θ is the polar angle; L p is the distance from the center of the eye to the center of the iris or the distance from the center of the eye to the center of the pupil;

[0033] S3.2.2 Secondly, transform the pupil disk and iris disk generated in the previous step into the camera coordinate system to obtain the 3D deformed point clouds of the pupil and iris:

[0034]

[0035] where R is the rotation matrix of the camera, that is, the matrix obtained by rotating the predicted gaze matrix to be in the same direction as the negative z-axis direction of the camera coordinate system; T is the translation distance from the center of the eye to the center of the camera;

[0036] S3.2.3 Finally, project the deformable pupil point cloud in the camera coordinate system onto the 2D plane:

[0037]

[0038] In the same way, project the deformable iris point cloud in the camera coordinate system onto the 2D plane:

[0039]

[0040] where K is the internal parameter matrix of the camera, which is composed of the focal length f of the camera and the corresponding relationship between the center of the camera coordinate system and the center of the image coordinate system (c x , c y ) = (W / 2, H / 2), where W and H are the width and length of the image respectively.

[0041] Furthermore, in step S3, an eye movement estimation based on optical flow is adopted: the optical flow method is used to estimate the eye movement, thereby deriving the parameters of the three-dimensional eye model;

[0042] and directly predicting the 3D model parameters using a three-dimensional convolutional network: 3D-CNN can directly process the information in the spatial and temporal dimensions and directly predict the 3D model parameters of the eye.

[0043] Further, in step S4, the few-shot weak supervision module decodes the features after multi-modal fusion encoding to generate a mask for supervision, which is used to segment the eye image into the pupil, iris ellipse, and background.

[0044] Further, in step S4, it specifically includes: using a combination of loss functions in RITnet as the supervision strategy; this strategy is implemented using a weighted combination of four loss functions, including:

[0045] Standard cross-entropy loss: the default choice for applications with balanced class distributions;

[0046] Generalized Dice Loss: The Dice score coefficient measures the degree of overlap between true pixels and their predicted values;

[0047] BoundaryAware Loss: The semantic boundary divides regions based on class labels, and edge awareness is introduced by weighting the loss of each pixel according to its distance to the two nearest segments;

[0048] Surface Loss: Based on a distance metric in the image contour space.

[0049] Further, after the eyeball model output by the network is rendered onto a 2D plane, in order to supervise the generated point cloud with semantic labels, an edge loss function is designed as follows:

[0050]

[0051] where x ji is the coordinate of the predicted edge point, and y ji is the true coordinate value;

[0052] The total loss function is:

[0053] L total = L seg + L gaze + L edge (10)

[0054] where L seg is the segmentation loss, L gaze is the line-of-sight vector loss, and L edge is the edge loss.

[0055] On the other hand, the present invention provides a three-dimensional eyeball model reconstruction system based on multi-modal fusion, which is used to implement the three-dimensional eyeball model reconstruction method based on multi-modal fusion according to the present invention.

[0056] The system includes a spatial feature extraction module, a temporal feature fusion module, a three-dimensional eyeball model reconstruction module, and a few-shot weak supervision module.

[0057] Therefore, the beneficial effects of the present invention are as follows:

[0058] By introducing the preamble information provided by the event, the method and system of the present invention fuse the event graph and the original eye diagram in the spatial dimension at the feature fusion stage. Then, through the temporal transformer feature encoding network, the event information and the image information are fused in the temporal dimension. This enables the model to better learn the deformable eyeball model.

[0059] Meanwhile, the present invention uses an encoding and decoding architecture to perform self-supervision on the masks of the pupil and iris formed in the intermediate stage and the final stage, and uses a small number of 3D gaze labels for weak supervision. The method of the present invention has the potential and creativity for performing robust and high-precision gaze estimation tasks with a small number of samples. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 is a system network architecture diagram according to the present invention;

[0061] Figure 2 is a network structure diagram of the feature extraction module according to the present invention;

[0062] Figure 3 is a three-dimensional eyeball model structure diagram according to the present invention;

[0063] Figure 4 is a 2D rendering diagram of the three-dimensional eyeball according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0064] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0065] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, and does not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention. In addition, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.

[0066] In the description of the present invention, it should be noted that unless otherwise clearly specified and defined, the terms "installation", "connection", and "linkage" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0067] The following will Figures 1 - 4 describe in detail the specific embodiments of the present invention. It should be understood that the specific embodiments described herein are only used to illustrate and explain the present invention and are not used to limit the present invention.

[0068] The inventive concept of the present invention lies in applying the information of the modality of event graphs to gaze estimation and designing a system for reconstructing a three-dimensional eyeball model based on multimodal fusion. Event graphs refer to the technology of representing and analyzing events and their relationships through graph structures. The research prospects of event graphs are broad, covering multiple fields and application scenarios such as knowledge graph construction, natural language processing, event graph analysis and processing, and cross-modal action capture.

[0069] The system of the present invention uses temporal correlation for real-time tracking. An event camera only reports the changes in pixel brightness levels, that is, events. There are two reasons for the present invention to use events as features. First, events inherently encode temporal changes, thus providing natural clues for predicting moving regions. Second, events have natural sparsity and can be efficiently encoded, enabling the neural network to have a lower overhead in predicting moving regions.

[0070] The present invention takes semantics only as a weak supervision signal for learning 3D eye gaze estimation, and then solves the problem of the small amount of labels in three-dimensional gaze estimation through weak supervision with a small number of 3D gaze labels.

[0071] Combining the advantages of model-based and appearance-based gaze estimation methods, and predicting the gaze direction by predicting a fully differentiable 3D eyeball model, using event information to enhance the modeling ability of the eyeball model. The model can also be weakly supervised through semantics. It is mainly divided into four major modules: a spatial feature extraction module, a temporal feature fusion module, a three-dimensional eyeball model reconstruction module, and weak supervision.

[0072] As Figure 1 shown, the three-dimensional eyeball model reconstruction system based on multimodal fusion according to the present invention is composed of four major modules: a spatial feature extraction module, a temporal feature fusion module, a three-dimensional eyeball model reconstruction module, and a few-shot weak supervision module. The following will describe each module in detail:

[0073] 1. Spatial Feature Extraction Module:

[0074] This module is responsible for feature extraction from the original infrared eye diagram and the generated event diagram. In the fields of computer vision and deep learning, the feature extraction module is a key part of the model for extracting features from input data (such as images or videos). Usually, the feature extraction module will use pre-trained backbone networks, such as ResNet, VGG, Inception, etc. These backbone networks are trained on large-scale datasets (such as ImageNet) and have good feature representation capabilities. One of the most commonly used feature extraction networks in the field of eye movement interaction is ResNet18. As a relatively lightweight network in the ResNet network series, ResNet18 already has strong feature extraction capabilities. In order to reduce the number of parameters of the method model of the present invention, the present invention selects ResNet18 as the backbone of the feature extraction network. After respectively extracting the features of the original infrared eye diagram and the generated event diagram, the present invention will calculate the channel attention of the two different modalities of features extracted, and fuse the features of the two modalities in the spatial dimension. Enhance the complementary ability of the two modalities of information, so that the extracted features are more significant in terms of motion features.

[0075] In addition to ResNet18, other deep network architectures such as EfficientNet, DenseNet, or MobileNet can also be used. These networks have different advantages in terms of feature extraction capabilities and computational efficiency. Using a Feature Pyramid Network (FPN) or other multi-scale feature fusion methods can enhance the detection ability for targets of different sizes.

[0076] Figure 1 The encoder module in Figure 2 is composed of a feature extraction and channel attention fusion network, and the network structure is as HxWx1 shown. The input human eye bimodal image is I ∈ R Wherein B represents the feature extractor, I is the original infrared image, E is the generated event image, F1 and F2 represent the extracted infrared eye image features and the extracted event image features, respectively, and R represents the feature dimension, which usually contains D channels, and the original spatial dimension is downsampled according to the scale factor of s. The features obtained from the original infrared eye image and the event image are fused in the spatial dimension through a lightweight channel attention mechanism ECA. The obtained feature map F can be directly used for various computer vision tasks (for example: classification, regression, dense prediction, etc.). In the method proposed in the present invention, the last layer of features obtained is decoded for the task of classification, that is, the pixels are divided into three categories: pupil area, iris area and background, and the obtained mask is used for subsequent weak supervision.

[0077] 2. Temporal feature fusion module:

[0078] In the method of the present invention, the spatial dimension is compressed by global average pooling to obtain F1∈R 1x1xD , F2∈R 1x1xD Represents the global features of the fused event and infrared modality eye images as well as the individual features of the event map.

[0079] Studies have shown that applying a transformer to features extracted independently from consecutive frames is very effective for predicting each frame of video. The present invention uses two joint processing networks based on Transformer, namely Figure 1 The temporal fusion processing dual-stream transformer in the image is used. It takes the ordered sequence F of eye features after spatial multimodal fusion and the ordered sequence D of event graph features. It outputs the corresponding eye parameters E = t(F1) + t(F2). Where t represents the Transformer joint processing network, and E is the final network output, i.e., the three-dimensional eyeball parameters, including: eyeball radius r p , iris radius r i , rotation matrix (gaze direction) R, camera parameters f, eyeball radius r e , translation matrix T. For the same individual (such as eyeball radius, iris radius, etc.), some eye parameters do not change between different frames, while some eye parameters are different for each frame (such as gaze direction, pupil radius, etc.). However, in consecutive frames, different eye parameters are very correlated. Through joint processing, valuable information is shared between consecutive frames. At the same time, in the two timing modeling stages, the information of the time flow represented by the event graph is passed into the information of the spatial flow represented by the original eye graph to supplement the learning scope of the entire eye model.

[0080] By jointly processing eye features and event map features in consecutive frames, valuable information can be shared between frames, enhancing the model's prediction ability. The Transformer network is utilized to capture temporal relationships, and information sharing and fusion are achieved through the cross-attention mechanism, ultimately enabling accurate prediction of eye parameters. This method effectively combines temporal and spatial information, providing a powerful solution for complex computer vision tasks.

[0081] In addition, a convolutional neural network (CNN) can also be used to process temporal information: a three-dimensional convolutional neural network (3D-CNN) is adopted to capture temporal information, replacing the Transformer network.

[0082] Based on long short-term memory networks (LSTM) or gated recurrent units (GRU): These recurrent neural network (RNN) models also perform well in processing temporal data and can replace the Transformer for temporal feature processing.

[0083] 2. Three-dimensional eyeball model reconstruction module:

[0084] First, a mathematical model of the three-dimensional eyeball is established. In the present invention, the eyeball is defined as a sphere with its center O e and an eyeball radius of r e . It should be noted that in eye movement interaction, for the same individual, for example: the eyeball radius r e , the iris radius r i , the eyeball translation amount T (the translation vector from the eyeball center to the camera center), etc., these eye parameters do not change between different frames; while for each frame, for example, the gaze direction gaze / R n (the eyeball rotation matrix, making the gaze vector collinear with the negative direction of the z-axis of the camera coordinate system), the pupil radius r pn , these eye parameters are different. However, in consecutive frames, different eye parameters are highly correlated. Through joint processing, valuable information is shared between consecutive frames. At the same time, the internal parameters of the camera, namely the focal length f, can also be learned. In the present invention, it is assumed that the camera center is (c x , c y ) = (w / 2, H / 2), that is, it is assumed that the center of the camera coordinate system and the center of the image coordinate system coincide on the two-dimensional plane. In fact, the non-coincident situation can be made coincident through a learnable translation matrix T.

[0085] The remaining three-dimensional eyeball information can be represented by these parameters that are invariant and variable for each frame. Due to the constraint of the rotation angle in human eye movement, the three-dimensional eyeball parameters predicted by the model in the present invention: the eyeball radius r p , the iris radius r i , the rotation matrix (gaze direction) R, the camera parameter f, the eyeball radius re , the translation matrix T. Scale and constrain them within reasonable physical ranges.

[0086] The present invention first initializes the eye center as Initialize the fixation vector as g c =(0, 0, -1). The eye is located at the center of the regular coordinate system, and the optical axis is collinear with the negative direction of the regular z-axis. To weakly supervise 3D eye parameters with 2D semantic labels in a fully differentiable manner, discrete point clouds are generated for the pupil and iris of the 3D eye model.

[0087] First, generate the discrete point cloud of the pupil disk:

[0088]

[0089] In the same way, the present invention generates the discrete point cloud of the iris disk:

[0090]

[0091] where r p is the radius of the pupil, ρ is the radial distance in polar coordinates, and θ is the polar angle. L p is the distance from the eye center to the iris center or from the eye center to the pupil center.

[0092] Secondly, the present invention transforms the pupil disk and iris disk generated in the previous step into the camera coordinate system to obtain the 3D deformed point clouds of the pupil and iris:

[0093]

[0094] where R is the rotation matrix of the camera, that is, the matrix obtained by rotating the predicted fixation matrix to be in the same direction as the negative direction of the z-axis of the camera coordinate system; T is the translation matrix of the camera, that is, the translation distance from the eye center to the camera center. Finally, the present invention projects the deformable pupil point cloud in the camera coordinate system onto the 2D plane:

[0095]

[0096] In the same way, the present invention projects the deformable iris point cloud in the camera coordinate system onto the 2D plane:

[0097]

[0098] where K is the internal parameter matrix of the camera, which is composed of the focal length f of the camera and the corresponding relationship between the center of the camera coordinate system and the center of the image coordinate system (c x , c y )=(W / 2, H / 2), where W and H are the width and length of the image respectively.

[0099] The images of the three-dimensional eyeball model reconstruction are as follows Figure 3 shown, where

[0100] Standard eyeball model: The eyeball is modeled as a sphere with its center at O e and radius r e . The second sphere intersecting the eyeball represents the iris with its center at O i and radius r i . There is also a concentric pupil circle inside the iris with its center at O p = O i and radius r p .

[0101] Deformable eyeball model: Based on the standard eyeball model, a deformable eyeball model with real-time updated optimization is reconstructed through learnable eyeball parameters

[0102] Optical axis calculation: The optical axis g is defined as the unit vector passing through the two centers O i and O e and is regarded as an approximate fixation vector .

[0103] Parameter estimation: In each frame of image I n , five eyeball parameters E n = {r e , r i , r pn , T, R n} are estimated. Among them, the eyeball radius r e , the iris radius, r i and the translation vector T of the eyeball remain unchanged throughout the sequence, while the pupil radius, r pn and the rotation matrix R of the eyeball n are estimated separately in each frame

[0104] Rotation matrix R n : R n rotates the camera coordinate system so that the negative z-axis is collinear and in the same direction as the fixation vector g n . In the camera coordinate system, the x-axis is positive to the right, the y-axis is positive downward, and the z-axis is positive forward. Therefore, the negative z-axis faces the camera, and the optical axis can be represented as g n = R n [0, 0, -1] T .

[0105] Point cloud generation: The distance from the eyeball center O e to the pupil center Next, the 3D pupil dot cloud P C p and the 3D iris circle dot cloud P are generated according to the estimated parametersC i In the standard coordinate system, the eyeball is located at the origin O e C =(0, 0, 0), and the optical axis g c =(0, 0, -1) corresponds to the negative z-axis.

[0106] Point cloud transformation: The point cloud is transformed in the camera coordinate system through P p 3D =[R][T]P C p and P i 3D =[R][T]P i c transformed into the camera coordinate system. Where R is the rotation matrix, T is the translation matrix, and P C p is the 3D pupil dot cloud, and P C i is the 3D iris dot cloud.

[0107] The schematic diagram of the three-dimensional eyeball model projecting the deformed iris and pupil point cloud regions onto the image plane and obtaining the corresponding 2D segmentation mask is as shown Figure 4 shown.

[0108] Projection process: Project the transformed 3D point cloud onto the camera screen, and the projection matrix used is where f is the focal length, and W and H are the width and height of the image respectively.

[0109] 2D projection: For each point cloud, apply the camera intrinsic matrix K to project the three-dimensional point cloud P p 3D and P i 3D onto the two-dimensional screen to obtain the 2D pupil point cloud P p 2D and the 2D iris point cloud P i 2D .

[0110] In addition, the present invention can also directly predict 3D model parameters using a three-dimensional convolutional network (3D-CNN): The 3D-CNN can directly process information in the spatial and temporal dimensions and directly predict the 3D model parameters of the eyeball.

[0111] The specific implementation process of the three-dimensional convolutional network directly predicting 3D model parameters is as follows:

[0112] 1. 3D-CNN network design

[0113] Network architecture: Design a 3D-CNN architecture suitable for processing spatio-temporal data. A typical 3D-CNN consists of 3D convolutional layers, 3D pooling layers, batch normalization layers, activation layers, etc. The final output of the network is a regression layer corresponding to the 3D model parameters.

[0114] 3D convolutional layer: Used to extract spatio-temporal features of the input data, and the convolutional kernel slides in three dimensions of time and space.

[0115] 3D pooling layer: Used to downsample spatio-temporal data, reduce the computational amount and extract key features.

[0116] Fully connected layer: After the convolutional layer, use the fully connected layer to map the features to specific 3D model parameters.

[0117] 2. Network training

[0118] 3. Model inference

[0119] Predict 3D parameters: Input a new sequence of eyeball images into the trained 3D-CNN model to directly predict the corresponding 3D model parameters.

[0120] Parameter decoding: Decode the parameters output by the network into a specific three-dimensional eyeball model (such as rotation matrix, displacement, etc.).

[0121] The advantage of using 3D-CNN is that for scenarios with short time series or limited computing resources, 3D-CNN can perform calculations and training more efficiently while maintaining high accuracy, and performs well in tasks with dense spatial and temporal information. The disadvantages are: The receptive field is limited by the convolutional kernel and the number of network layers, and may not be able to capture long-term dependencies. It relies on local convolutional operations. Although larger contexts can be captured through multiple layers of convolution, it still tends to extract local features.

[0122] 4. Few-shot weak supervision module:

[0123] The few-shot number refers to being between 10% and 20% of the normal sample number, and is defined according to the size of the training dataset.

[0124] Weak supervision learning: Different from traditional supervised learning, the weak supervision learning adopted in the present invention does not rely on comprehensive and accurate labeled data, but improves the performance of the model on weakly labeled data by introducing additional constraints, regularization or prior knowledge.

[0125] The purpose of setting up the weak supervision module is that due to the required hardware settings and computational requirements, it is very cumbersome to obtain large-scale 3D gaze data, and there are few publicly available datasets that contain both eye semantic labels and 3D gaze vector labels at the same time. Setting up this module can fine-tune a semantic segmentation network through pre-training and then use a small number of 3D gaze labels, or supervise the losses of segmentation and gaze simultaneously. The former only requires a small number of 3D sample ground truths for fine-tuning, and the latter introduces the gaze label ground truth at the beginning of the model. After training for several rounds, the gaze loss is no longer supervised, and the effect of few-shot weak supervision can also be achieved.

[0126] To supervise the decoding of the features encoded by multi-modal fusion to generate a mask, which is mainly used to segment the eye image into the pupil and iris ellipses, as well as the background. To train such an architecture, the present invention uses a combination of loss functions proposed in RITnet. This strategy involves using a weighted combination of four loss functions;

[0127] Standard Cross Entropy Loss (CEL): It is the default choice for applications with balanced class distributions.

[0128] Generalized Dice Loss (GDL): The Dice score coefficient measures the degree of overlap between the true pixels and their predicted values.

[0129] BoundaryAware Loss (BAL): The semantic boundary divides the region based on the class label. Introducing edge awareness by weighting the loss of each pixel according to its distance to the two nearest segments.

[0130] Surface Loss (SL): It is a distance metric based on the image contour space, which preserves small and infrequent structures with higher semantic values.

[0131] The loss L of this segmentation head seg is composed of a weighted combination of the above four losses:

[0132] L seg = L CEL (w CEL + w BAL L BAL ) + w GDL L GDL + w SL L SL (7)

[0133] where w CEL , w BAL , w GDL , w SL are the initial weights of the four different losses respectively.

[0134] When supervising the estimated gaze vector, the present invention uses the mean squared error loss with weight wgaze:

[0135]

[0136] where w gaze is the weight of this loss function, g n is the predicted gaze vector, and is the true value of the gaze vector.

[0137] Finally, in this architecture, after the eyeball model output by the network is rendered onto a 2D plane, in order to supervise the generated point cloud with semantic labels, the present invention designs an edge loss to minimize the Euclidean distance between the edge pixels of the projected semantic point cloud and the edge pixels of the true semantic mask.

[0138]

[0139] where x ji is the coordinate of the predicted edge point, and y ji is the true value of the coordinate.

[0140] Therefore, the total loss function of this method is:

[0141] L total = L seg + L gaze + L edge (10)

[0142] where L seg is the segmentation loss, L gaze is the gaze vector loss, and L edge is the edge loss.

[0143] Widely available semantic labels can be used to train a model from scratch as a good starting point. Then, the network weights of this kind of model can be fine-tuned on a small number of available gaze labels in order to impose more 3D supervision and constraints.

[0144] The overall process of the three-dimensional eyeball model reconstruction method based on multi-modal fusion according to the present invention is as follows:

[0145] First, an event camera emulator is constructed to obtain event information between every two frames. Then, a backbone feature extraction network is used to extract features from the original eye image and the event image. The extracted features are decoded by a decoder to obtain a mask including the pupil and iris. The features of the extracted eye image and event image are jointly processed to estimate 3D eye parameters, which are then used for the learning of a differentiable eyeball model and transformed into camera coordinates. Then, the eyeball model is rendered onto the image plane through a rendering method, using the 2D semantic masks of the projected pupil and iris. The rendered 2D semantic masks are self-supervised with the pupil and iris masks decoded by the previous decoder. Then, weak supervision is performed using a small number of 3D gaze labels. Specifically, the method includes the following steps:

[0146] S1. Feature extraction: Feature extraction is performed on the original infrared eye image and the generated event image, and the features of the original infrared eye image and the generated event image are respectively extracted. Channel attention is calculated for the two different modalities of features, and the features of the two modalities are fused in the spatial dimension. In the present invention, ResNet18 is selected as the backbone of the feature extraction network. Fusing the features of the two modalities can enhance the complementary ability of the two modalities of information, making the extracted features more significant in terms of motion features.

[0147] The feature extraction and channel attention fusion method is as follows:

[0148] S1.1 Input the binocular human eye image I∈R HxWx1 , where H and W are the height and width of the image respectively, and each picture is a grayscale image composed of one channel. Initialize two backbone feature extractors, and take the infrared image I and the event image E as inputs respectively.

[0149] S1.2 Generate two mappings F1 = B(I) and F2 = B(E). Where, contains D channels, and the original spatial dimension is downsampled by a factor of s.

[0150] S1.3 The features obtained from the original infrared eye image and the event image are fused in the spatial dimension through a lightweight channel attention mechanism ECA to obtain the feature map F. The feature map F can be directly used for various computer vision tasks (such as: classification, regression, dense prediction, etc.). In the method proposed in the present invention, the last layer of features obtained is decoded for the classification task, that is, the pixels are divided into three categories: pupil area, iris area, and background, and the obtained mask is used for subsequent weak supervision.

[0151] S2. Temporal Feature Fusion: The spatial dimension is compressed by global average pooling. Take the ordered sequences of the eye features after spatial multi-modal fusion and the ordered sequence of the event graph features, and output the corresponding eye parameters.

[0152] In the method of the present invention, the spatial dimension is compressed by global average pooling to obtain F1 ∈ R 1x1xD , F2 ∈ R 1x1xD represents the global features of the fused event and infrared two-modal eye images and the individual features of the event graph, where R is the representation form of the feature vector, F1 represents the features extracted from the original eye image by the feature extractor, F2 represents the features extracted from the event graph by the feature extractor, and D represents the feature dimension.

[0153] Applying a transformer to the features independently extracted from consecutive frames is very effective for predicting each frame of video. The present invention uses two joint processing networks based on Transformer. Take the ordered sequence F of the eye features after spatial multi-modal fusion and the ordered sequence D of the event graph features, and output the corresponding eye parameter E = t(F1) + t(F2). Where t represents the Transformer joint processing network, and E is the final network output, that is, the three-dimensional eyeball parameters, including: eyeball radius r p , iris radius r i , rotation matrix (gaze direction) R, camera parameter f, eyeball radius r e , translation matrix T.

[0154] S3. Three-dimensional Eyeball Model Reconstruction: Define the eyeball as a sphere with its eyeball center O e and an eyeball radius of r e ; Initialize the eyeball center as Initialize the gaze vector as g c =(0, 0, -1). The eyeball is located at the center of the regular coordinate system, and the optical axis is collinear with the regular z-axis and points in the negative direction. Generate discrete point clouds for the pupil and iris of the 3D eye model. Specifically include:

[0155] S3.1 First, perform mathematical modeling of the three-dimensional eyeball model. Define the eyeball as a sphere with its eyeball center O e and an eyeball radius of r e . It should be noted that in eye movement interaction, for the same individual, for example: eyeball radius r e , iris radius r i , eyeball translation amount T (translation vector from the eyeball center to the camera center), etc., these eye parameters do not change between different frames; while for each frame, for example, gaze direction gaze / R n (eyeball rotation matrix, making the gaze vector collinear with the negative direction of the z-axis of the camera coordinate system) and pupil radius r pn, these eye parameters are different. However, in consecutive frames, different eye parameters are highly correlated. Through joint processing, valuable information is shared between consecutive frames. At the same time, the internal parameters of the camera, namely the focal length f, are learned, and the camera center is set as (c x , c y ) = (w / 2, H / 2), so that the center of the camera coordinate system and the center of the image coordinate system coincide on the two-dimensional plane. In addition, for the non-coincident case, a learnable translation matrix T is used to achieve coincidence.

[0156] The remaining 3D eyeball information can be represented by these per-frame invariant and per-frame variable parameters. Due to the rotational angle constraint in human eye movement, the present invention scales and constrains the 3D eyeball parameters predicted by the model to keep them within a reasonable physical range.

[0157] S3.2 Initialize the eyeball center as Initialize the fixation vector as g c = (0, 0, -1), with the eyeball located at the center of the regular coordinate system, and the optical axis collinear with the negative direction of the regular z-axis. To weakly supervise the 3D eye parameters with 2D semantic labels in a fully differentiable manner, discrete point clouds are generated for the pupil and iris of the 3D eye model. Specifically as follows:

[0158] S3.2.1 First, generate the discrete point cloud of the pupil disk:

[0159]

[0160] In the same way, the present invention generates the discrete point cloud of the iris disk:

[0161]

[0162] Among them, r p is the radius of the pupil, ρ is the radial distance in polar coordinates, and θ is the polar angle. L p is the distance from the eyeball center to the iris center or the distance from the eyeball center to the pupil center.

[0163] S3.2.2 Secondly, transform the pupil disk and iris disk generated in the previous step into the camera coordinate system to obtain the 3D deformed point clouds of the pupil and iris:

[0164]

[0165] Among them, R is the rotation matrix of the camera, that is, the matrix obtained by rotating the predicted fixation matrix to the same direction as the negative direction of the z-axis of the camera coordinate system; T is the translation matrix of the camera, that is, the translation distance from the eyeball center to the camera center.

[0166] S3.2.3 Finally, the present invention projects the deformable pupil point cloud in the camera coordinate system onto a 2D plane:

[0167]

[0168] Using the same method, the present invention projects the deformable iris point cloud in the camera coordinate system onto a 2D plane:

[0169]

[0170] where K is the internal parameter matrix of the camera, which is composed of the focal length f of the camera and the correspondence between the center of the camera coordinate system and the center of the image coordinate system (c x ,c y ) = (W / 2, H / 2), where W and H are the width and length of the image respectively.

[0171] S3.3 The three-dimensional eyeball model projects the deformed iris and pupil point cloud regions onto the image plane and obtains the corresponding 2D segmentation mask.

[0172] S4. Few-shot weak supervision: After the eyeball model output by the network is rendered onto the 2D plane, in order to supervise the generated point cloud with semantic labels, an edge loss function is set to minimize the Euclidean distance between the edge pixels of the projected semantic point cloud and the edge pixels of the ground truth semantic mask.

[0173] In order to supervise the decoding of the features encoded by multi-modal fusion to generate a mask, which is mainly used to segment the eye image into pupil and iris ellipses, and the background. To train such an architecture, the present invention uses a combination of loss functions proposed in RITnet. This strategy involves using a weighted combination of four loss functions:

[0174] Standard Cross Entropy Loss (CEL): The default choice for applications with balanced class distributions.

[0175] Generalized Dice Loss (GDL): The Dice score coefficient measures the degree of overlap between the true pixels and their predicted values.

[0176] BoundaryAware Loss (BAL): The semantic boundary divides the region based on the class label. Edge awareness is introduced by weighting the loss of each pixel by its distance to the two nearest segments.

[0177] Surface Loss (SL): It is a distance metric based on the image contour space, which preserves small and infrequent structures with higher semantic values.

[0178] The loss L of this segmentation head seg is composed of a weighted combination of the above four losses:

[0179] L seg = L CEL (w CEL + w BAL L BAL ) + w GDL L GDL + w SL L SL (7)

[0180] where w CEL , w BAL , w GDL , w SL are the initial weights of four different losses respectively.

[0181] When supervising and estimating the gaze vector, the present invention uses the mean square error loss with weight wgaze:

[0182]

[0183] where w gaze is the weight of this loss function, g n is the predicted gaze vector, is the true value of the gaze vector.

[0184] Finally, after the eyeball model output by the network in this architecture is rendered onto a 2D plane, in order to supervise the generated point cloud with semantic labels, the present invention designs an edge loss to minimize the Euclidean distance between the edge pixels of the projected semantic point cloud and the edge pixels of the true value semantic mask.

[0185]

[0186] Therefore, the total loss function of this method is:

[0187] L total = L seg + L gaze + L edge (10)

[0188] Widely available semantic labels can be used to train a model from scratch as a good starting point. Then, the network weights of such a model can be fine-tuned on a small number of available gaze labels to impose more 3D supervision and constraints.

[0189] The technical advantages of the present invention are as follows:

[0190] 1. Multi-modal Fusion: Combine the infrared eye diagram and the event diagram. Through feature extraction and fusion, make full use of the complementary information of the two modalities for fusion in space. Use a Transformer-based network to process consecutive frames, capture temporal relationships, enhance the utilization of relevant information between different frames, and perform fusion in time series.

[0191] 2. 3D Eyeball Model Reconstruction: Predict the 3D eyeball model parameters and use event information to enhance the accuracy and robustness of the model. Render the eyeball model onto the image plane to achieve high-precision eyeball parameter estimation.

[0192] 3. Weak Supervision and Self-Supervision Mechanisms: Design a combination of multiple loss functions (such as cross-entropy loss, Dice loss, boundary-aware loss, etc.) to supervise the masks generated by feature decoding. At the same time, perform edge loss supervision on the generated point clouds with semantic labels.

[0193] 4. Application of Event Information: Generate event information through an event camera simulator. Utilize the temporal change clues and sparsity of event information to enhance the modeling ability of the eyeball model in feature extraction and temporal modeling, and improve the efficiency and accuracy of motion area prediction.

[0194] Any process or method description shown in the flowchart of the present invention or described in other ways herein can be understood as representing a module, segment, or part of code including one or more executable instructions for implementing a specific logical function or process, which can be implemented in any computer-readable medium for an instruction execution system, apparatus, or device. The computer-readable medium can be any medium including storage, communication, propagation, or transmission of a program for use by an instruction execution system, apparatus, or device, including read-only memory, magnetic disks, or optical discs, etc.

[0195] In the description of this specification, the description referring to terms such as "embodiment", "example", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. In addition, those skilled in the art can combine or combine different embodiments or examples described in this specification and the features therein without contradiction.

[0196] Although the above content has shown and described embodiments of the present invention, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can perform operations such as changes, modifications, substitutions, and variations on the above embodiments within the scope of the present invention.

Claims

1. A three-dimensional eyeball model reconstruction method based on multimodal fusion, characterized in that The method includes the following steps: S1. Feature extraction: Feature extraction is performed on the original infrared eye diagram and the generated event diagram, respectively extracting the features of the original infrared eye diagram and the generated event diagram. Channel attention calculation is performed on the extracted features of the two different modalities, and the features of the two modalities are fused in the spatial dimension; S2. Temporal feature fusion: Compression is performed through global average pooling in the spatial dimension. The ordered sequences of the eye features and the event diagram features after spatial multi-modal fusion are taken, and the corresponding eye parameters are output; S3. Three-dimensional eyeball model reconstruction: Define the eyeball as a sphere with its center O e and an eyeball radius of r e . Initialize the eyeball center as Initialize the fixation vector as g c = (0, 0, -1). The eyeball is located at the center of the regular coordinate system, the optical axis is collinear with the regular z-axis and points in the negative direction. Generate discrete point clouds for the pupil and iris of the 3D eye model; S4. Few-shot weak supervision: After the eyeball model output by the network is rendered onto a 2D plane, in order to supervise the generated point cloud with semantic labels, an edge loss function is set to minimize the Euclidean distance between the edge pixels of the projected semantic point cloud and the edge pixels of the ground-truth semantic mask.

2. The three-dimensional eyeball model reconstruction method based on multi-modal fusion according to claim 1, characterized in that In step S1, ResNet18, EfficientNet, DenseNet, MobileNet, or a feature pyramid network is used as the backbone of the feature extraction network.

3. The three-dimensional eyeball model reconstruction method based on multimodal fusion according to claim 2, wherein In step S1, the feature extraction and channel attention fusion method is as follows: S1.1 Input the binocular human eye image I ∈ R HxWx1 , where H and W are the height and width of the image respectively, and each picture is a grayscale image composed of one channel; Initialize two backbone feature extractors, taking the infrared image I and the event image E as inputs respectively; S1.2 generates two mappings F1 = B(I) and F2 = B(E); where, contains D channels, and the original spatial dimension is downsampled by a factor of s; S1.3 The features obtained from the original infrared eye diagram and the event diagram are fused in the spatial dimension through a lightweight channel attention mechanism ECA to obtain a feature map.

4. The three-dimensional eyeball model reconstruction method based on multimodal fusion according to claim 1, characterized in that In step S2, the spatial dimension is compressed by global average pooling to obtain F1 ∈ R 1x1xD , F2 ∈ R 1x1xD which represents the global features of the fused event and infrared eye images and the individual features of the event map; Two joint processing networks based on Transformer, convolutional neural network, long short-term memory network, or gated recurrent unit are used. The ordered sequence F of the eye features after spatial multi-modal fusion is taken, and the ordered sequence D of the event diagram features is taken, and the corresponding eye parameter E = T(F1) + T(F2) is output.

5. The three-dimensional eyeball model reconstruction method based on multi-modal fusion according to claim 1, characterized in that In step S3, it specifically includes: S3.1 First, perform mathematical modeling of the three-dimensional eyeball model, defining the eyeball as a sphere with its center O e and an eyeball radius of r e . At the same time, learn the focal length f of the camera, and set the camera center as (c x , c y ) = (w / 2, H / 2), so that the center of the camera coordinate system and the center of the image coordinate system coincide on the two-dimensional plane; S3.2 Initialize the center of the eyeball as Initialize the fixation vector as g c =(0, 0, -1). The eyeball is located at the center of the regular coordinate system, the optical axis is collinear with the regular z-axis and points in the negative direction, and discrete point clouds are generated for the pupil and iris of the 3D eye model; S3.3 The three-dimensional eyeball model projects the deformed iris and pupil point cloud regions onto the image plane and obtains the corresponding 2D segmentation mask.

6. The three-dimensional eyeball model reconstruction method based on multimodal fusion according to claim 1, characterized in that In step S3, optical flow-based eyeball movement estimation is adopted: The optical flow method is used to estimate the eyeball movement, thereby deriving the parameters of the three-dimensional eyeball model; And a three-dimensional convolutional network is used to directly predict the 3D model parameters: The 3D-CNN directly processes the information in the spatial and temporal dimensions and directly predicts the 3D model parameters of the eyeball.

7. The three-dimensional eyeball model reconstruction method based on multimodal fusion according to claim 1, wherein, In step S4, the few-shot weak supervision module supervises the mask generated by decoding the features after multi-modal fusion encoding, which is used to segment the eye image into the pupil, iris ellipse, and background.

8. The three-dimensional eyeball model reconstruction method based on multi-modal fusion according to claim 1, wherein In step S4, it specifically includes: Using the combination of the loss functions in RITnet as the supervision strategy; This strategy is implemented using the weighted combination of four loss functions; including: Standard cross-entropy loss: The default choice for applications with balanced class distributions; GeneralizedDice Loss: The Dice score coefficient measures the degree of overlap between the true pixels and their predicted values; BoundaryAware Loss: The semantic boundary divides the region based on the class label, and edge awareness is introduced by weighting the loss of each pixel according to its distance to the two nearest segments; Surface Loss: Based on a distance metric in the image contour space.

9. The three-dimensional eyeball model reconstruction method based on multimodal fusion according to claim 8, characterized in that After the eye model output by the network is rendered onto a 2D plane, in order to supervise the generated point cloud with semantic labels, the edge loss function is designed as follows: where x ji are the coordinates of the predicted edge points, and y ji are the true coordinates; The total loss function is: L total = L seg + L gaze + L edge (10) where L seg is the segmentation loss, L gaze is the gaze vector loss, and L edge is the edge loss.

10. A three-dimensional eyeball model reconstruction system based on multimodal fusion, characterized in that, The system is used to implement the three-dimensional eye model reconstruction method based on multimodal fusion according to any one of claims 1-9; the system includes a spatial feature extraction module, a temporal feature fusion module, a three-dimensional eye model reconstruction module, and a few-shot weak supervision module.

Citation Information

Cited By

  • Sight line point coordinate calculation method and system based on multi-modal fusion

    CN121541788A

  • A method and system for calculating line-of-sight coordinates based on multimodal fusion

    CN121541788B