Sight line estimation method and device and computer readable storage medium
By extracting the eye area image and performing image segmentation in gaze estimation, and using a lightweight gaze estimation model, the problem of high computational complexity of traditional methods is solved, and real-time gaze estimation on embedded devices is achieved.
Patent Information
- Application Number
- CN202510838773.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-10-17
AI Technical Summary
Traditional line of sight estimation methods have high computational complexity, resulting in huge computing resource requirements and making them difficult to apply in real time on embedded devices and low-power platforms.
By obtaining the image and status of the eye area in the original image, image segmentation processing is performed to obtain the feature maps of the pupil, iris and background. A lightweight gaze estimation model is used to perform gaze estimation, replacing the complex feature extraction process of the traditional three-dimensional eyeball model.
The computational complexity is significantly reduced, enabling line of sight estimation to run in real time under the limited computing resources of embedded devices, thus achieving lightweight line of sight estimation.
Smart Images

Figure CN120808427A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of eye tracking technology, and in particular to a gaze estimation method, device, and computer-readable storage medium. Background Art
[0002] Gaze tracking technology, also known as eye tracking technology, is a technology that uses software algorithms, mechanics, electronics, optics and other detection methods to obtain people's eye movement process. The movement of the eyeballs not only changes people's field of vision but also reflects the individual's cognitive process. Therefore, gaze tracking technology has very broad application prospects in the field of human-computer interaction.
[0003] Gaze estimation, a key component of gaze tracking technology, is used to estimate information such as the direction and position of a person's gaze. Traditional gaze estimation methods calculate gaze direction through three-dimensional eye modeling. Specifically, by constructing a digital model of the eye's geometric structure, physical properties, and motion patterns, the three-dimensional eye posture is directly fitted or inferred from image data to achieve gaze estimation.
[0004] However, gaze estimation methods based on 3D eye models require the creation and solution of a detailed 3D eye model (including geometry, physics, and kinematics), processing a large number of parameters, and employing complex optimization algorithms. This results in high model complexity and significant computational resource requirements, leading to a heavy computational burden on the system. This often requires high-performance computing devices, significantly limiting their application in embedded devices (such as virtual reality headsets, extended reality headsets, mobile devices, and wearables), low-power platforms, or real-time interactive scenarios requiring fast responses.
[0005] Therefore, how to reduce the computational complexity of line of sight estimation to achieve lightweight line of sight estimation is a technical problem that needs to be solved urgently. Summary of the Invention
[0006] The main purpose of this application is to provide a line of sight estimation method, device and computer-readable storage medium, aiming to solve the technical problem of how to reduce the computational complexity of line of sight estimation to achieve lightweight line of sight estimation.
[0007] To achieve the above objectives, the present application provides a line of sight estimation method, which includes:
[0008] Acquire an original image for which sight line estimation is to be performed, and extract an eye region image and an eye state corresponding to the eye region image from the original image;
[0009] if the eye state is an open-eye state, performing image segmentation processing on the eye region image to obtain an eye feature map, wherein the eye feature map is a classification representation feature map of the pupil, the iris, and the background in the eye region image;
[0010] inputting the eye feature map into a pre-trained gaze estimation model to output a gaze estimation result.
[0011] In an embodiment, after the step of obtaining the original image to be used for gaze estimation, the method further comprises:
[0012] obtaining an ambient light intensity;
[0013] if the ambient light intensity is less than or equal to a first preset threshold, performing image enhancement processing on the original image, and performing the steps of extracting the eye region image in the original image and the eye state corresponding to the eye region image based on the original image after image enhancement processing;
[0014] if the ambient light intensity is greater than a second preset threshold, performing desaturation processing on the original image, and performing the steps of extracting the eye region image in the original image and the eye state corresponding to the eye region image based on the original image after desaturation processing;
[0015] wherein the second preset threshold is greater than or equal to the first preset threshold.
[0016] In an embodiment, after the step of inputting the eye feature map into a pre-set gaze estimation model to output a gaze estimation result, the method further comprises:
[0017] obtaining a pre-set calibration coefficient, wherein the calibration coefficient is used to calibrate the deviation between the gaze direction estimated by the model and the actual gaze direction of the user;
[0018] calibrating the gaze estimation result based on the pre-set calibration coefficient to obtain a calibrated gaze direction.
[0019] In an embodiment, the step of obtaining a pre-set calibration coefficient comprises:
[0020] obtaining a device identifier of a gaze estimation device, obtaining a calibration coefficient corresponding to the device identifier based on a first preset mapping relationship, and determining the calibration coefficient corresponding to the device identifier as the pre-set calibration coefficient, wherein the first preset mapping relationship is a mapping relationship between different device identifiers and calibration coefficients; or
[0021] Obtain the account identifier of the user login account in the line of sight estimation device, obtain the calibration coefficient corresponding to the account identifier based on a second preset mapping relationship, and determine the calibration coefficient corresponding to the account identifier and the preset calibration coefficient, wherein the second preset mapping relationship is a mapping relationship between different account identifiers and calibration coefficients.
[0022] In one embodiment, before the step of obtaining a preset calibration coefficient, the method further includes:
[0023] Acquire the actual sight line direction of the user when gazing at a preset reference point and the reference image when gazing at the preset reference point;
[0024] performing line of sight estimation on the reference image based on the line of sight estimation model to obtain an estimated line of sight direction;
[0025] The deviation between the estimated sight line direction and the actual sight line direction is fitted to obtain a preset calibration coefficient.
[0026] In one embodiment, the step of extracting the eye region image from the original image and the eye state corresponding to the eye region image includes:
[0027] Input the original image into a preset object detection model, and output the eye area position and eye prediction state;
[0028] The original image is cropped based on the eye region position to obtain an eye region image, and the predicted eye state is determined to be the eye state corresponding to the eye region image.
[0029] In one embodiment, the step of performing image segmentation processing on the eye region image to obtain an eye feature map includes:
[0030] Inputting the eye region image into an improved U-Net network, performing image segmentation processing on the eye region image through the improved U-Net network, and outputting an eye feature map;
[0031] The improved U-Net network uses 10×10 convolution layers for convolution processing in the upsampling path and the downsampling path.
[0032] In one embodiment, the gaze estimation model is a lightweight DenseNet network, and the step of inputting the eye feature map into a pre-trained gaze estimation model and outputting a gaze estimation result includes:
[0033] Input the eye feature map into the pre-trained lightweight Densenet network, and output a gaze estimation result;
[0034] The lightweight Densenet network includes five dense blocks, and the dense blocks are connected through transition layers.
[0035] In an embodiment, before the step of inputting the eye feature map into the pre-trained gaze estimation model, the method further includes:
[0036] obtaining a training data set and a gaze estimation model to be trained, wherein the training data set includes a training eye feature map obtained through image segmentation processing and a gaze direction corresponding to the training eye feature map;
[0037] training the gaze estimation model by taking the training eye feature map as input and the gaze direction as label to obtain the trained gaze estimation model.
[0038] In addition, to achieve the above-mentioned purpose, the present application also provides a gaze estimation device, which comprises a memory, a processor and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the gaze estimation method as described above.
[0039] In addition, to achieve the above-mentioned purpose, the present application also provides a readable storage medium, which is a computer readable storage medium, and a computer program is stored on the computer readable storage medium, and the computer program is executed by a processor to implement the steps of the gaze estimation method as described above.
[0040] The present application also provides a computer program product, which comprises a computer program, and the computer program is executed by a processor to implement the steps of the gaze estimation method as described above.
[0041] The one or more technical solutions provided by the present application have at least the following technical effects:
[0042] The present application obtains an original image for gaze estimation, extracts an eye region image and the eye state corresponding to the eye region image from the original image; if the eye state is open, performs image segmentation processing on the eye region image to obtain an eye feature map, wherein the eye feature map is a classification representation feature map of the pupil, iris, and background in the eye region image; the eye feature map is input into a pre-trained gaze estimation model, and the gaze estimation result is output. In this way, by extracting the eye region image and the eye state from the image, the embodiment of the present application can filter the open-eye image by eye state, first avoiding redundant calculations in invalid states such as closed eyes, and preliminarily optimizing the allocation of computing resources. Subsequently, the eye area image is segmented to obtain a semantic classification representation feature map of the pupil, iris and background area in the eye area image. This step essentially replaces the feature extraction process in the traditional three-dimensional eyeball model that relies on solving complex geometric / physical parameters. The eye feature map generated by segmentation extracts anatomical structures (pupil, iris) related to the gaze direction, while stripping away background noise and detail information that is not related to gaze estimation. This data-driven structured feature representation significantly reduces the dimensionality and complexity of the feature space, providing highly compressed and task-related input for subsequent processing. Based on this, the solution further uses a gaze estimation model to act on the eye feature map. Since the input is a low-dimensional, high-semantic feature representation that has been segmented and refined (without the need for complex modeling and parsing of original images or three-dimensional parameters), the model no longer needs to have built-in heavy physical constraints or iterative optimization algorithms to simulate eye movements. The model only needs to learn the mapping relationship from the streamlined feature space to the gaze direction, which greatly compresses the computational complexity and memory usage required for inference. This allows the use of a lightweight gaze estimation model for gaze estimation, thereby greatly reducing the computational complexity of gaze estimation and enabling real-time operation under the limited computing power resources of embedded devices, thereby realizing lightweight gaze estimation. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0044] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0045] Figure 1 This is a flow chart of the first embodiment of the line of sight estimation method of the present application;
[0046] Figure 2A schematic diagram of a line-of-sight estimation process related to an embodiment of the line-of-sight estimation method of the present application;
[0047] Figure 3 A schematic diagram of an improved U-Net network structure related to an embodiment of the line-of-sight estimation method of the present application;
[0048] Figure 4 A schematic diagram of the system structure of the line-of-sight estimation system of the present application;
[0049] Figure 5 A schematic diagram of the device structure of the hardware running environment related to the line-of-sight estimation method and device in an embodiment of the present application.
[0050] The purposes, functional features and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION
[0051] In order to make the above-mentioned purposes, features and advantages of the present application more apparent and easy to understand, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.
[0052] Eye tracking technology can accurately measure the fixation point, saccade trajectory and other information of individuals when observing specific targets by capturing and analyzing eye movements. This technology has shown important value in many fields, such as human-computer interaction interface optimization, psychology and cognitive science research, consumer behavior analysis, advertising effectiveness evaluation, reading pattern exploration, and assisting disabled people in communication and controlling devices. With the continuous evolution of technology, the application potential of eye tracking is increasingly expanding, which has far-reaching significance for improving the depth of scientific research and improving the quality of human life, and line-of-sight estimation (i.e., determining the actual direction of gaze or the specific position in the screen / space) is the most core function and application basis of eye tracking technology.
[0053] Currently, the techniques for realizing gaze estimation mainly develop in two directions: feature point detection-based, such as the Pupil-Corneal Reflection method (PCCR), and three-dimensional eye modeling-based. The Pupil-Corneal Reflection method is one of the most commonly used gaze estimation methods, which calculates the gaze direction of the eye by detecting the positions of the pupil and corneal reflection points. However, this method is highly dependent on stable lighting conditions. Strong light or changing lighting environments can cause the detection of reflection points to fail or be misdetected, thereby affecting the accuracy of tracking. For users wearing glasses, the lenses can reflect light, interfering with the detection of the pupil and corneal reflection points. The three-dimensional eye modeling-based method aims to construct a digital model of the fine eye geometry, physical properties, and motion laws, and directly fit or infer the three-dimensional pose of the eye from image data.
[0054] However, the high-precision gaze estimation method based on three-dimensional eye modeling faces a significant technical bottleneck: the model complexity is too high, the computational resource demand is huge, and it is difficult to realize lightweight deployment and real-time efficient operation. Establishing and solving a detailed three-dimensional eye model requires processing a large number of parameters and complex optimization algorithms, resulting in a heavy computational burden for the system. This makes such methods usually need to rely on high-performance computing devices, greatly limiting their application in embedded devices, low-power platforms, or real-time interactive scenarios that require fast response.
[0055] Based on this, the main solution of the present application is: obtaining an original image to be estimated for gaze, extracting an eye region image in the original image and an eye state corresponding to the eye region image; if the eye state is an open-eye state, performing image segmentation processing on the eye region image to obtain an eye feature map, wherein the eye feature map is a classification representation feature map of the pupil, iris and background in the eye region image; inputting the eye feature map into a pre-trained gaze estimation model to output a gaze estimation result.
[0056] The application extracts an eye region image and an eye state in an image, thereby screening open-eye images through the eye state, firstly avoiding redundant calculation in invalid states such as closed eyes, and preliminarily optimizing the allocation of computing resources. Subsequently, the eye region image is subjected to image segmentation processing, thereby obtaining a semantic classification feature map of the pupil, iris and background region in the eye region image. This step substantially replaces the feature extraction process in the traditional three-dimensional eyeball model that relies on complex geometric / physical parameters for solving. The segmented eye feature map extracts the anatomical structure (pupil, iris) related to the gaze direction, while stripping the background noise and detail information irrelevant to the gaze estimation. This data-driven structured feature representation significantly reduces the dimension and complexity of the feature space, providing highly compressed and task-related input for subsequent processing. Based on this, the scheme further uses a gaze estimation model to act on the eye feature map. Since the input is a low-dimensional, high-semantic feature representation (without the need for complex modeling and analysis of the original image or three-dimensional parameters), the model does not need to build heavy physical constraints or iterative optimization algorithms to simulate eyeball movement. The model only needs to learn the mapping relationship from the simplified feature space to the gaze direction, greatly reducing the computational complexity and memory occupation required for inference, so that a lightweight gaze estimation model can be used for gaze estimation, thereby greatly reducing the computational complexity of gaze estimation and enabling real-time operation under the limited computing resources of embedded devices, thereby realizing lightweight gaze estimation.
[0057] It should be noted that the execution subject of each embodiment of the gaze estimation method of the application can be a computing service device with data processing, network communication and program running functions, such as a server, a tablet computer, a personal computer, a mobile phone, etc., or a gaze estimation device capable of realizing the above functions, such as an AR (Augmented Reality) helmet, a VR (Virtual Reality) helmet, etc. The embodiments of the gaze estimation method of the application do not make specific limitations on this.
[0058] Based on this, the application proposes a gaze estimation method of the first embodiment, as shown in Figure 1 , please refer to Figure 1 , the gaze estimation method has the following steps S10-S30:
[0059] Step S10, acquiring an original image to be estimated, extracting an eye region image in the original image and an eye state corresponding to the eye region image;
[0060] The original image can be specifically an image frame originally collected by an image collection component. The image collection component refers to a hardware module with an image collection function, for example, including but not limited to a near-infrared sensor, a visible light camera, etc. Preferably, the field of view direction of the image collection component is configured to face the user's face region, so as to ensure that the original image collected contains the user's eye region.
[0061] The eye region image can be specifically an image region containing at least one eye of the user extracted by image processing on the original image. The eye state includes an open eye state and a closed eye state. The open eye state refers to an open eye state, which is characterized in that the iris and pupil are not covered by the eyelid and are clearly visible (or can be recognized by an algorithm), thereby providing an effective input for subsequent gaze estimation. The closed eye state refers to a closed eye state, which is characterized in that the upper eyelid and the lower eyelid are closed, and the iris and pupil region are completely covered by the eyelid and cannot be seen (or cannot be effectively recognized by an algorithm).
[0062] If the eye region is not successfully recognized in the original image after processing, or the eye state corresponding to the recognized eye region is determined to be a closed eye state, the gaze estimation process for this time (i.e., for the current frame image) can be ended. After obtaining the next frame of original image to be estimated, the step S10 and the subsequent gaze estimation steps (if applicable) can be repeatedly executed, so as to realize continuous tracking of the user's eye movement gaze trajectory by continuously processing and judging each frame of original image, thereby realizing eye movement tracking.
[0063] In step S20, if the eye state is an open eye state, an image segmentation process is performed on the eye region image to obtain an eye feature map, wherein the eye feature map is a classification representation feature map of the pupil, iris and background in the eye region image.
[0064] When the eye state is an open eye state, the eye region image is subjected to an image segmentation process to segment the pupil, iris and background region in the eye region image, thereby obtaining a segmented eye feature map. The eye feature map is a classification representation feature map of the pupil, iris and background in the eye region image. Specifically, the eye feature map can be a classification representation multi-channel probability matrix of the pupil, iris and background, wherein each channel carries a single-class pixel-level confidence, and the sum of the probability values between channels after normalization processing can be 1.
[0065] Exemplarily, the image segmentation processing can be implemented by a pre-trained deep learning segmentation network, for example, a U-Net network (Convolutional Networks for Biomedical Image Segmentation) based on an encoder-decoder structure, a DeepLabv3+ network (Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation) and the like, the segmentation network takes the eye region image as input and outputs a multi-channel probability feature map with the same resolution as the input image, wherein each channel corresponds to a pixel-level probability distribution of a preset class; the preset classes include three classes of pupils, irises and backgrounds, so that in the output feature map: the first channel represents the probability that each pixel belongs to the pupil region, the second channel represents the probability that each pixel belongs to the iris region, and the third channel represents the probability that each pixel belongs to the background region; the eye feature map accurately represents the pixel-level classification probability of the pupils, irises and backgrounds in the original eye image in a quantitative manner, providing a robust feature input for subsequent gaze direction calculation.
[0066] In step S30, the eye feature map is input into a pre-trained gaze estimation model to output a gaze estimation result.
[0067] After obtaining the eye feature map, the eye feature map is input into a pre-trained gaze estimation model to output a gaze estimation result.
[0068] The gaze estimation result can specifically be a three-dimensional direction vector (θ, φ, r), wherein θ is the pitch angle, φ is the yaw angle, and r is the gaze depth. Further, the pitch angle θ ∈ [-90 degrees, 90 degrees], representing the vertical direction line of sight offset (positive value for upward tilt); the yaw angle φ ∈ [-180 degrees, 180 degrees], representing the horizontal direction line of sight offset (positive value for right turn); the gaze depth r ∈ [0, 1], representing the normalized gaze point distance (0 represents infinite distance / minimum distance, 1 represents the nearest / maximum distance).
[0069] The gaze estimation model refers to a model for estimating the gaze direction. Exemplarily, the type of the gaze estimation model can be a Densenet network (Densely Connected Convolutional Networks), which is not specifically limited in this embodiment.
[0070] Further, the line-of-sight estimation model can be a lightweight model, which means that the model meets certain lightweight constraints, such as a parameter amount ≤ 5M, a floating-point operation amount (FLOPs) ≤ 1GMac, a model volume (after INT8 quantization) ≤ 2MB, and the like.
[0071] It should be noted that the line-of-sight estimation model can be supervised learning using an eye feature map data set containing real line-of-sight direction labels in the training stage. Specifically, in one possible implementation, before the step of inputting the eye feature map into the pre-trained line-of-sight estimation model, the method further comprises:
[0072] Step S100, obtaining a training data set and a line-of-sight estimation model to be trained, wherein the training data set includes a training eye feature map obtained by image segmentation processing and a line-of-sight direction corresponding to the training eye feature map;
[0073] The training data set can be a pre-prepared data set, which can specifically include a plurality of training eye feature maps and line-of-sight directions corresponding to each training eye feature map. Wherein, the training eye feature map can be an eye feature map obtained by image segmentation processing on an image (hereinafter referred to as a training image), and the corresponding training eye feature map is also a classification representation feature map of the pupil, iris and background. The image segmentation processing includes but is not limited to using a segmentation network (such as U-Net, DeepLab, etc.) for image segmentation. For each training eye feature map, its corresponding line-of-sight direction is labeled. The line-of-sight direction labeling can use a manual labeling method, in which professional personnel label the line-of-sight direction of each training eye feature map; or use a device with a line-of-sight direction sensor (such as an eye tracker) to collect the training image and its corresponding line-of-sight direction; or use a computer vision algorithm (such as a line-of-sight estimation algorithm based on feature points) to estimate the line-of-sight direction of the training image, and use the estimation result as the labeled data, etc., which is not limited here.
[0074] Step S200, training the line-of-sight estimation model with the training eye feature map as input and the line-of-sight direction as label, to obtain a trained line-of-sight estimation model.
[0075] Further, before training the line-of-sight estimation model, the training data set can also be pre-processed, such as dividing the training data set into a training set and a validation set, which are respectively used for model training and performance evaluation.
[0076] After obtaining the training data set, the gaze estimation model is trained by taking the trained eye feature map as input and the gaze direction as label. The mature supervised learning method can be used to train the gaze estimation model. For example, the specific training process of the gaze estimation model can be as follows: the training eye feature map in the training data set is input into the gaze estimation model to be trained, the output of the model is calculated through forward propagation, then the loss function value between the model output and the corresponding gaze direction label is calculated, the parameters of the model are updated according to the loss function value using the back propagation algorithm, and the performance of the model is optimized. During the training process, the performance of the model is evaluated regularly using the validation set. According to the performance index on the validation set, the above training and validation process is repeated until the model meets the preset training end condition, and the trained gaze estimation model is obtained. The loss function can be mean square error loss, cross entropy loss, etc., which is not limited here. The preset training end condition can be a condition set in advance, such as reaching a predetermined number of iterations, the loss function value being lower than a predetermined threshold, running out of computing resources, reaching a time limit, the accuracy reaching a predetermined threshold, etc., which is not limited here.
[0077] In this embodiment, the eye region image in the image and the eye state are extracted, so that the open-eye image can be screened through the eye state. Firstly, redundant calculation in invalid states such as closed eyes is avoided, and the allocation of computing resources is preliminarily optimized. Then, the image segmentation processing is performed on the eye region image, so as to obtain a semantic classification feature map of the pupil, iris and background region in the eye region image. This step substantially replaces the feature extraction process in the traditional three-dimensional eyeball model which depends on complex geometric / physical parameters for solving. The segmented eye feature map extracts the anatomical structure (pupil, iris) related to the gaze direction, while stripping the background noise and detail information irrelevant to the gaze estimation. This data-driven structured feature representation significantly reduces the dimension and complexity of the feature space, providing a highly compressed and task-related input for subsequent processing. Based on this, the gaze estimation model is further used on the eye feature map. Since the input is a low-dimensional and high-semantic feature representation extracted by segmentation (without complex modeling and analysis of the original image or three-dimensional parameters), the model does not need to build heavy physical constraints or iterative optimization algorithms to simulate eyeball movement. The model only needs to learn the mapping relationship from the simplified feature space to the gaze direction, greatly reducing the computational amount and memory occupation required for inference. This enables the use of a lightweight gaze estimation model for gaze estimation, thereby greatly reducing the computational complexity of gaze estimation and enabling real-time operation on embedded devices with limited computing resources, thereby realizing lightweight gaze estimation.
[0078] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as the above embodiment one can be referred to the above introduction, and the subsequent will not be described. On this basis, the step of extracting the eye region image in the original image and the eye state corresponding to the eye region image comprises:
[0079] Step A10, inputting the original image into a preset target detection model, and outputting an eye region position and an eye prediction state;
[0080] The target detection model refers to a model for identifying the eye region in the image and predicting the eye state, which can be specifically MobileNet-SSD (MobileNet-based Single Shot MultiBox Detector, MobileNet-based single shot multi-box detector) model, YOLO model, etc. In a preferred embodiment, considering that MobileNet-SSD is a target detection network model combining MobileNet and SSD, by using deep separable convolution and feature pyramid network, MobileNet-SSD has lower calculation and storage cost while maintaining high precision. Therefore, the MobileNet-SSD model is selected as the preset target detection model.
[0081] The eye region position can be represented in the form of coordinates, such as including the upper left corner coordinates and the lower right corner coordinates of the eye region, or represented by center point coordinates and width, height and other parameters, which are used for accurate cutting of the eye region image in the subsequent process. The eye prediction state is the open-close state of the eye (such as open eyes, closed eyes) and the like.
[0082] It should be noted that in the training process of the target detection model, images labeled with eye region positions and corresponding eye states are used as training data. By learning the features of the eye region and the features of the eye state, the model can accurately locate the eye region in the input original image and predict its state.
[0083] Step A20, based on the eye region position, cutting the original image to obtain an eye region image, and determining the eye prediction state as the eye state corresponding to the eye region image.
[0084] In the image cutting process, the eye region image can be extracted from the original image according to the coordinate information of the eye region position, ensuring the integrity of the eye region image.
[0085] In a possible implementation, the step of performing image segmentation processing on the eye region image to obtain an eye feature map comprises:
[0086] Step B10, inputting the eye region image into the improved U-Net network, performing image segmentation processing on the eye region image through the improved U-Net network, and outputting an eye feature map;
[0087] In the improved U-Net network, 10*10 convolution layers are used for convolution processing in the up-sampling path and the down-sampling path.
[0088] It should be noted that before the eye region image is input into the improved U-Net network, if the eye region image is a non-gray image, the eye region image can also be processed to be a gray image, so as to convert the eye region image into a gray image. The gray processing can reduce the number of channels of the image, reduce the calculation amount, and at the same time retain the main structure information of the image, which is helpful to improve the processing efficiency and segmentation accuracy of the network.
[0089] The original U-Net network adopts an encoding-decoding architecture, and the network architecture is left-right symmetric, which is composed of an encoder (down-sampling path) on the left side of the network, a decoder (up-sampling path) on the right side of the network, and a skip connection path. The transverse skip connection transmits high-resolution image features from the down-sampling path to the up-sampling path to maintain high-frequency image features and obtain clear segmentation output. The down-sampling path and the up-sampling path use 3*3 convolution layers for convolution, each down-sampling column is connected in turn through a 2*2 pooling layer, each up-sampling column is connected in turn through a 2*2 up-sampling layer, and the encoder and the decoder are connected through a 2*2 pooling layer, two 3*3 convolution layers and a 2*2 up-sampling layer.
[0090] In this embodiment, the original U-Net network is improved. Specifically, compared with the original U-Net network, 10*10 filter convolution layers are used for convolution processing in the up-sampling and down-sampling paths.
[0091] Further, in order to keep the size of the feature map unchanged after the convolution operation of the convolution layer, the improved U-Net network can also set the padding of the convolution layer to 5.
[0092] Further, in the improved U-Net network, convolution with a step size of 2 (2*2 filter kernel size) is used for down-sampling, and transposed convolution is used for up-sampling. Compared with other down-sampling and up-sampling methods that the original U-Net network can use (such as pooling and de-pooling), this method saves more memory, because in the down-sampling process, the convolution operation itself reduces the size of the feature map while not needing to store additional pooling index information; in the up-sampling process, transposed convolution can directly learn the optimal up-sampling filter without relying on stored index information as de-pooling does, reducing the memory occupation. And transposed convolution for up-sampling can learn the optimal up-sampling filter, and convolution with a step size of 2 for down-sampling can also learn the optimal down-sampling filter, which enables the model to automatically adjust the parameters of the filter according to the training data to better adapt to the eye segmentation task. For example, when processing eye images under different resolutions and different imaging conditions, the model can better perform feature map down-sampling and up-sampling operations through the learned optimal filter, thereby improving the generalization ability of the model.
[0093] In this embodiment, the improved U-Net network is used for image segmentation processing, and specifically, the improved U-Net network uses a 10*10 convolution layer for convolution processing. Compared with the smaller 3*3 filter used in the original U-Net network, the larger 10*10 filter can cover more pixels, so that more context information in a larger range can be extracted at each stage. As the down-sampling stage progresses, the size of the feature map decreases, and the receptive field of the convolution filter increases, allowing more complex features in a larger context to be extracted. For example, in iris segmentation, the structure information around the iris, such as eyelashes and eyelids, is important for accurate iris segmentation. This improvement can better capture the relationship between these structures and the iris. Because more complex features in a larger context can be extracted, the model's ability to learn features such as shape and texture of different parts of the eye (such as the pupil and the iris) will be stronger. For example, the texture features of the iris are complex and diverse, and this improvement helps the model to more accurately capture global context information, thereby improving the accuracy of segmentation. And single-layer large convolution can replace the stacking of multiple layers of small convolution, thereby simplifying the network structure and avoiding the problem of gradient disappearance or unstable training in deep networks.
[0094] In a possible implementation, the gaze estimation model is a lightweight Densenet network, and the step of inputting the eye feature map into the pre-trained gaze estimation model to output a gaze estimation result comprises:
[0095] Step C10, inputting the eye feature map into the pre-trained lightweight Densenet network to output a gaze estimation result;
[0096] The lightweight Densenet network includes five dense blocks, and the dense blocks are connected through transition layers.
[0097] The Densenet network is also called a dense convolutional network, which mainly includes a DenseBlock (dense block) and a Transition layer (transition layer). In this embodiment, a lightweight Densenet network including five DenseBlocks is used for gaze estimation to reduce the calculation amount and parameter quantity of the model. Through this optimization measure, the lightweight Densenet network has lower calculation and storage costs while maintaining high accuracy, can quickly process eye feature maps, and output accurate gaze estimation results, achieving lightweight gaze estimation.
[0098] For example, to help understand the technical concept or technical principle of the gaze estimation method combined with the above first embodiment, a specific embodiment is listed. In this specific embodiment, referring to FIG. 1, the gaze estimation process includes: first, obtaining a near-eye image (i.e., an original image), detecting the eye region and the corresponding eye state in the near-eye image using an eye detection model (i.e., a target detection model), and cropping the corresponding eye region image based on the eye region. If the eye state is an open-eye state, the eye region image is processed for image segmentation (i.e., eye segmentation) to obtain an eye feature map, then the eye feature map is input into a gaze estimation model, and finally the gaze direction of the eye is output, i.e., the gaze estimation result. Figure 2
[0099] Specifically: 1. The target detection model uses a MobileNet-SSD model, which is a target detection network model combining MobileNet and SSD. By using deep separable convolution and feature pyramid network, MobileNet-SSD has lower calculation and storage costs while maintaining high accuracy. The MobileNet-SSD network model is used for eye detection, and in the training process, the background, open-eye image, and closed-eye image in the training data are divided into three categories, so that the network model can output the position of the eye and the open-closed state of the eye at the same time.
[0100] 2. Image segmentation is performed using an improved U-Net network, which is a special convolutional neural network designed to solve image segmentation problems. Referring to FIG. 2, the U-Net network includes an encoder and a decoder. The encoder is used to extract features from the input image, and the decoder is used to reconstruct the output image from the extracted features. The U-Net network is able to learn the context information of the image and generate a precise segmentation result. Figure 3 As shown, the network architecture of the U-Net network is left-right symmetrical, which is composed of an encoder (down-sampling path) on the left side of the network, a decoder (up-sampling path) on the right side of the network, and a skip connection path 3, the transverse skip connection transmits high-resolution image features from the down-sampling path to the up-sampling path to maintain high-frequency image features and obtain clear segmentation output. Compared with the original U-Net network, the improved U-Net network uses a convolutional layer with a 10*10 filter at each stage of the up-sampling and down-sampling paths, and outputs a feature map with the same size as the input by appropriate padding (padding = 5). The down-sampling path reduces the size of the feature map and increases the size of the receptive field of the convolutional filter at each stage, so that more complex features in a larger context can be extracted. Down-sampling is performed using convolution with a step size of 2 (2*2 filter kernel size), and up-sampling is performed using transposed convolution, which saves more memory and enables learning of optimal down-sampling / up-sampling filters. The final output layer outputs an eye feature map with the same size as the input layer, corresponding to the pupil, iris, and background, respectively.
[0101] 3. The gaze estimation is performed using a Densenet network, also known as a dense convolutional network, which mainly includes two parts of DenseBlock and Transition layer, wherein the DenseBlock is composed of multiple convolutional layers, the size of the feature map of each convolutional layer is the same, and the convolutional layers are connected in a dense connection manner. The Transition layer is connected between two adjacent DenseBlock, mainly including a 1*1 convolutional layer and a 2*2 average pooling layer, and the purpose is to reduce the size of the feature map. The specific embodiment adopts a lightweight Densenet network, specifically, the lightweight Densenet network includes five DenseBlock, each block includes eight convolutional layers, through the combination of five DenseBlock and Transition, finally 62 feature maps are generated at the end of DenseNet, the global average pooling is used to convert these feature maps into 62 feature vectors, and a linear layer is used to map the 62 feature vectors to the final three-dimensional gaze direction.
[0102] It should be noted that the above examples are only used to assist in understanding the embodiment, and do not constitute a limitation on the gaze estimation process of the embodiment, and more forms of simple transformation based on this technical concept are within the protection scope of the present application.
[0103] Based on the first embodiment and / or the second embodiment of the present application, in the third embodiment of the present application, the same or similar contents as the above embodiments one and two can be referred to the above introduction, and will not be described hereinafter. On this basis, after the step of obtaining the original image to be estimated, the method further comprises:
[0104] Step D10, obtaining an ambient light intensity;
[0105] The ambient light intensity of the current environment can be collected in real time by the ambient light sensor, and the ambient light intensity collected synchronously with the original image can be obtained, so that the ambient light intensity when the original image is collected can be obtained.
[0106] Step D20, if the ambient light intensity is less than or equal to a first preset threshold, performing image enhancement processing on the original image, and performing the steps of extracting the eye region image in the original image and the eye state corresponding to the eye region image based on the original image after the image enhancement processing;
[0107] If the ambient light intensity is less than or equal to the first preset light threshold, the original image is subjected to image enhancement processing. Specifically, the image enhancement processing can include one or more of image brightness enhancement processing, contrast enhancement processing, and local detail enhancement processing, so as to improve the recognition degree of the eye feature in the dark light environment through the image enhancement processing, and then improve the accuracy of subsequent image segmentation. The first preset threshold can be a threshold representing a low light condition.
[0108] Step D30, if the ambient light intensity is greater than a second preset threshold, performing desaturation processing on the original image, and performing the steps of extracting the eye region image in the original image and the eye state corresponding to the eye region image based on the original image after the desaturation processing; wherein the second preset threshold is greater than or equal to the first preset threshold.
[0109] If the ambient light intensity is greater than the second preset threshold, the original image is subjected to desaturation processing, so as to suppress the color deviation and overexposure interference caused by strong light reflection by reducing the color saturation, while retaining the key gray scale information, and then improving the accuracy of subsequent image segmentation. The second preset threshold can be a threshold representing a high light condition. The second preset threshold is greater than the first preset threshold.
[0110] In one possible implementation, after the step of inputting the eye feature map into a preset gaze estimation model and outputting a gaze estimation result, the method further includes:
[0111] Step E10, obtaining a preset calibration coefficient, wherein the calibration coefficient is used to calibrate the deviation between the model estimated gaze direction and the actual gaze direction of the user;
[0112] The calibration coefficient is obtained in this embodiment, so as to calibrate the possible deviation of the model estimated gaze direction by using the calibration coefficient.
[0113] It should be noted that the calibration coefficient can be measured by a calibration experiment, and can be in the form of a numerical value, a matrix, or a function mapping, without specific limitation. The calibration experiment refers to a process of establishing an error compensation parameter by comparing the estimated line-of-sight direction output by the model with the actual known real line-of-sight direction in a controlled environment.
[0114] Step E20, calibrating the line-of-sight estimation result based on the preset calibration coefficient to obtain a calibrated line-of-sight direction.
[0115] After obtaining the preset calibration coefficient, the line-of-sight estimation result output by the model is calibrated based on the calibration coefficient to obtain a calibrated line-of-sight direction, and the calibrated line-of-sight direction can be output as the final line-of-sight estimation result or passed to an upper application.
[0116] In this embodiment, by obtaining the preset calibration coefficient, the model estimation deviation caused by individual physiological differences and other factors is quantified as a compensation parameter through the calibration coefficient, so that the model estimation deviation caused by factors such as corneal reflection point offset and lens refractive index change can be eliminated in the line-of-sight estimation process through the calibration system, the general line-of-sight estimation method has the ability to adapt to different eye movement characteristics, and finally more accurate line-of-sight estimation results are obtained, providing more reliable line-of-sight direction data for upper interactive applications.
[0117] In one possible implementation, the step of obtaining the preset calibration coefficient includes:
[0118] Step F10, obtaining a device identifier of the line-of-sight estimation device, obtaining a calibration coefficient corresponding to the device identifier based on a first preset mapping relationship, and determining that the calibration coefficient corresponding to the device identifier is the preset calibration coefficient, wherein the first preset mapping relationship is a mapping relationship between different device identifiers and calibration coefficients; or,
[0119] Step F20, obtaining an account identifier of a user login account in the line-of-sight estimation device, obtaining a calibration coefficient corresponding to the account identifier based on a second preset mapping relationship, and determining that the calibration coefficient corresponding to the account identifier is the preset calibration coefficient, wherein the second preset mapping relationship is a mapping relationship between different account identifiers and calibration coefficients.
[0120] The line-of-sight estimation device can be a device that collects original images.
[0121] The device identifier refers to a credential identifier of device identity, such as a device MAC (Media Access Control) address, a device serial number, etc. It should be noted that if the calibration coefficient corresponding to the device identifier is not found in the first mapping relationship, the user can be guided to complete the corresponding calibration experiment to obtain the calibration coefficient of the current device, and store it in the first mapping relationship for subsequent use.
[0122] The account identifier refers to a credential identifier of user identity, such as an account name, an iris hash value, a face ID (Identifier), etc. Similarly, if the calibration coefficient corresponding to the account identifier is not found in the second mapping relationship, the user can be guided to complete the corresponding calibration experiment to obtain the calibration coefficient of the current account identifier, and store it in the second mapping relationship for subsequent use.
[0123] The embodiment finds the corresponding calibration coefficient through the device identifier or the account identifier, thereby converting individual physiological differences and / or device hardware errors into dynamically loadable compensation parameters through the establishment of a mapping relationship, automatically matching the optimal calibration coefficient when the user wears different devices or shares a device, systematically eliminating multi-source errors such as corneal reflection point offset, lens refractive index change, device assembly tolerance, etc., ensuring that cross-device and cross-user gaze estimation is always based on personalized compensation parameters, making the general model accurately adapt to the coupling differences of specific hardware and ergonomics while retaining universality, so that the final gaze direction data stream can be provided to the upper application without being disturbed by device replacement and user switching.
[0124] In a possible implementation, before the step of obtaining the preset calibration coefficient, the method further includes:
[0125] Step G10, obtaining an actual gaze direction of the user when gazing at a preset reference point and a reference image when gazing at the preset reference point;
[0126] It should be noted that the preset reference point can be one or more. In order to improve the accuracy of the obtained calibration coefficient, preferably, the preset reference point is multiple. For example, in a specific implementation, the preset reference point is 4 corner points and a center point of the screen when the user faces the screen.
[0127] Further, when the preset reference point is multiple, the actual gaze direction of the user when gazing at each preset reference point and the reference image when gazing at the preset reference point can be obtained respectively. The reference image refers to an image collected when the user gazes at the preset reference point.
[0128] Step G20, performing gaze estimation on the reference image based on the gaze estimation model to obtain an estimated gaze direction;
[0129] The line-of-sight estimation model is used to estimate the line-of-sight direction of the reference image, and it is to be noted that the line-of-sight estimation of the reference image is achieved in the same manner as the original image, which will not be described here.
[0130] In step G30, the deviation between the estimated line-of-sight direction and the actual line-of-sight direction is fitted to obtain a preset calibration coefficient.
[0131] After the estimated line-of-sight direction of the model and the actual line-of-sight direction of the user are obtained, the deviation between the two can be fitted to obtain a preset calibration coefficient.
[0132] For fitting the deviation between the estimated line-of-sight direction and the actual line-of-sight direction, a mature deviation fitting method (such as least squares method) can be used to obtain the calibration coefficient, which will not be described here.
[0133] In this embodiment, the actual line-of-sight direction and the reference image are synchronously collected by gazing at the reference point, the same line-of-sight estimation model is used to process the reference image to obtain the estimated direction, and then the system deviation between the estimated value and the actual value is fitted to generate the calibration coefficient. Thus, the nonlinear optical characteristics of the eyeball rotation are captured through the spatially distributed reference points, the generated calibration coefficient can compensate for the composite errors caused by factors such as corneal curvature variation, pupil center offset, and lens distortion, the general line-of-sight estimation model is upgraded to a personalized system that adapts to individual eye movement characteristics, and the accuracy of the final estimated line-of-sight direction is improved without changing the model structure.
[0134] In addition, the embodiment of the present application also provides a line-of-sight estimation system, as shown in Figure 4 The line-of-sight estimation system comprises:
[0135] The extraction module 10 is configured to acquire an original image to be subjected to line-of-sight estimation, extract an eye region image in the original image, and extract an eye state corresponding to the eye region image.
[0136] The segmentation processing module 20 is configured to, if the eye state is an open-eye state, perform image segmentation processing on the eye region image to obtain an eye feature map, wherein the eye feature map is a classification representation feature map of a pupil, an iris, and a background in the eye region image.
[0137] The line-of-sight estimation module 30 is configured to input the eye feature map into a pre-trained line-of-sight estimation model, and output a line-of-sight estimation result.
[0138] In an embodiment, the line-of-sight estimation system further comprises an image processing module configured to:
[0139] acquire an ambient light intensity;
[0140] if the ambient light intensity is less than or equal to a first preset threshold, performing image enhancement processing on the original image;
[0141] if the ambient light intensity is greater than a second preset threshold, performing desaturation processing on the original image; wherein the second preset threshold is greater than or equal to the first preset threshold.
[0142] In an embodiment, the line-of-sight estimation system further comprises a calibration module for:
[0143] obtaining a preset calibration coefficient, wherein the calibration coefficient is used to calibrate the deviation between the line-of-sight direction estimated by the model and the actual line-of-sight direction of the user;
[0144] calibrating the line-of-sight estimation result based on the preset calibration coefficient to obtain the calibrated line-of-sight direction.
[0145] In an embodiment, the calibration module is further configured to:
[0146] obtain a device identifier of the line-of-sight estimation device, obtain the calibration coefficient corresponding to the device identifier based on a first preset mapping relationship, and determine that the calibration coefficient corresponding to the device identifier is the preset calibration coefficient, wherein the first preset mapping relationship is a mapping relationship between different device identifiers and calibration coefficients; or
[0147] obtain an account identifier of a user login account in the line-of-sight estimation device, obtain the calibration coefficient corresponding to the account identifier based on a second preset mapping relationship, and determine that the calibration coefficient corresponding to the account identifier is the preset calibration coefficient, wherein the second preset mapping relationship is a mapping relationship between different account identifiers and calibration coefficients.
[0148] In an embodiment, the calibration module is further configured to:
[0149] obtain the actual line-of-sight direction when the user gazes at a preset reference point and the reference image when the user gazes at the preset reference point;
[0150] perform line-of-sight estimation on the reference image based on the line-of-sight estimation model to obtain an estimated line-of-sight direction;
[0151] fit the deviation between the estimated line-of-sight direction and the actual line-of-sight direction to obtain the preset calibration coefficient.
[0152] In an embodiment, the extraction module 10 is further configured to:
[0153] input the original image into a preset target detection model to output the eye region position and the eye prediction state;
[0154] Crop the original image based on the eye region position to obtain an eye region image, and determine the eye prediction state as an eye state corresponding to the eye region image.
[0155] In an embodiment, the segmentation processing module 20 is further configured to:
[0156] input the eye region image into an improved U-Net network, perform image segmentation processing on the eye region image through the improved U-Net network, and output an eye feature map;
[0157] In an embodiment, the improved U-Net network uses a 10x10 convolution layer for convolution processing in the up-sampling path and the down-sampling path.
[0158] In an embodiment, the line-of-sight estimation model is a lightweight Densenet network, and the line-of-sight estimation module 30 is further configured to:
[0159] input the eye feature map into the pre-trained lightweight Densenet network, and output a line-of-sight estimation result;
[0160] In an embodiment, the lightweight Densenet network includes five dense blocks, and each of the dense blocks is connected through a transition layer.
[0161] In addition, an embodiment of the present application further proposes a line-of-sight estimation device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the line-of-sight estimation method as described above.
[0162] Reference Figure 5 which shows a structural schematic diagram of a line-of-sight estimation device suitable for being used to implement an embodiment of the present application. The line-of-sight estimation device in the embodiment of the present application can also include, but is not limited to, mobile terminals such as servers, notebook computers, virtual reality glasses, virtual reality headsets, extended reality glasses, extended reality headsets, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), and the like, as well as fixed terminals such as digital TVs, desktop computers, and the like. Figure 5 The line-of-sight estimation device shown is only an example, and should not bring any limitation to the functions and use range of the embodiments of the present application.
[0163] As Figure 5As shown, the line-of-sight estimation device can include a processing apparatus 1001 (e.g., a central processing unit, a graphics processing unit, etc.) that can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 1002 or programs loaded from a storage apparatus 1003 into a random access memory (RAM) 1004. Various programs and data required for operation of the line-of-sight estimation device are also stored in the RAM 1004. The processing apparatus 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems can be connected to the I / O interface 1006: input apparatuses 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; output apparatuses 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; the storage apparatus 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication apparatus 1009. The communication apparatus 1009 can allow the line-of-sight estimation device to communicate wirelessly or wired with other devices to exchange data. Although the line-of-sight estimation device with various systems is shown in the figure, it should be understood that all of the shown systems are not required to be implemented or possessed. More or less systems can be alternatively implemented or possessed.
[0164] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network by a communication apparatus, or installed from the storage apparatus 1003, or installed from the ROM 1002. When the computer program is executed by the processing apparatus 1001, the above-mentioned functions defined in the methods of embodiments of the present disclosure are performed.
[0165] The line-of-sight estimation device provided by the embodiments of the present disclosure adopts the line-of-sight estimation method in the above-mentioned embodiments, and can solve the technical problem of how to reduce the computational complexity of line-of-sight estimation to achieve lightweight line-of-sight estimation. Compared with the prior art, the line-of-sight estimation device provided by the present disclosure has the same beneficial effects as the line-of-sight estimation method provided by the above-mentioned embodiments, and other technical features in the line-of-sight estimation device are the same as the features disclosed in the previous embodiment method, which will not be repeated here.
[0166] It should be understood that various aspects of the disclosure can be implemented in hardware, software, firmware, or a combination thereof. In the description of the embodiments above, specific features, structures, materials or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0167] The above description is merely illustrative of the application and is not intended to limit the scope of the application. Any variations and modifications that can be made by any person skilled in the art within the spirit and scope of the application are intended to be encompassed by the application. The scope of the application is defined by the appended claims.
[0168] In addition, to achieve the above object, the embodiments of the present application further provide a readable storage medium having computer readable program instructions (i.e. computer programs) stored thereon, the computer readable program instructions being used to execute the line-of-sight estimation method in the above embodiments.
[0169] The computer readable storage medium provided by the embodiments of the present application may, for example, be a U disk, but is not limited to an electric, magnetic, optical, electromagnetic, infrared, or semiconductor system, system or device, or any combination of the above. More specific examples of the computer readable storage medium can include, but are not limited to, an electric connection with one or more conductive wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the embodiments, the computer readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system or device. The program code contained on the computer readable storage medium can be transmitted by any suitable medium, including but not limited to an electric wire, an optical cable, an RF (Radio Frequency), etc., or any suitable combination of the above.
[0170] The above computer readable storage medium can be contained in the line-of-sight estimation device; or can exist separately without being assembled into the line-of-sight estimation device.
[0171] The computer readable storage medium described above carries one or more programs, when the one or more programs are executed by the line-of-sight estimation device, the line-of-sight estimation device is caused to: acquire an original image to be line-of-sight estimated, extract an eye region image in the original image and an eye state corresponding to the eye region image; if the eye state is an open-eye state, perform image segmentation processing on the eye region image to obtain an eye feature map, wherein the eye feature map is a classification representation feature map of the pupil, the iris and the background in the eye region image; input the eye feature map into a pre-trained line-of-sight estimation model to output a line-of-sight estimation result.
[0172] Computer program code for carrying out operations of the present application can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0173] The flow diagrams and the block diagrams in the drawings are illustrations of architectures, functionalities, and operations of possible implementations of systems, methods, and computer program products according to various embodiments of present application. In this regard, each block in the flow diagrams or block diagrams can represent a module, a procedure, or a portion of code, which comprises one or more executable instructions for implementing the specific logical functions specified for the block. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may
[0174] The modules described in the embodiments of the present application can be implemented in the form of software or in the form of hardware. In some cases, the names of the modules do not constitute a limitation on the modules themselves.
[0175] The readable storage medium provided in the present application is a computer readable storage medium, which stores computer readable program instructions (i.e., a computer program) for executing the above-mentioned line-of-sight estimation method, and can solve the technical problem of how to reduce the computational complexity of line-of-sight estimation to achieve lightweight line-of-sight estimation. Compared with the prior art, the computer readable storage medium provided in the present application has the same beneficial effects as the line-of-sight estimation method provided in the above-mentioned embodiments, and will not be described here.
[0176] In addition, the embodiments of the present application also provide a computer program product, which includes a computer program, and the computer program is executed by a processor to implement the steps of the line-of-sight estimation method as described above.
[0177] The computer program product of the present application has substantially the same implementation as the above-mentioned line-of-sight estimation method, and will not be described here.
[0178] It should be noted that in this paper, the term "including", "containing" or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, article or system including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or system. Without more limitations, the element defined by the statement "including a" does not exclude the presence of another identical element in the process, method, article or system including the element.
[0179] The serial numbers of the above-mentioned embodiments of the present application are only for description, and do not represent the advantages and disadvantages of the embodiments.
[0180] Through the above description of the embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment method can be realized by software and a necessary general hardware platform, of course, it can also be realized by hardware, but in many cases, the former is a better implementation. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software sensor, which is stored in a storage medium (such as a ROM / RAM, a magnetic disk, an optical disk) and includes a plurality of instructions for making a terminal device (which can be a mobile phone, a computer, a server or a network device, etc.) execute the methods described in the embodiments of the present application.
[0181] The above merely preferred embodiments of the present application and are not intended to limit the patent scope of the present application, any equivalent structure or equivalent process transformation using the content of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A line of sight estimation method, characterized in that: The sight line estimation method comprises the following steps: Acquire an original image for which sight line estimation is to be performed, and extract an eye region image and an eye state corresponding to the eye region image from the original image; If the eye state is an open eye state, performing image segmentation processing on the eye region image to obtain an eye feature map, wherein the eye feature map is a classification representation feature map of the pupil, iris, and background in the eye region image; The eye feature map is input into a pre-trained gaze estimation model, and a gaze estimation result is output.
2. The sight line estimation method according to claim 1, wherein: After the step of obtaining the original image to be estimated, the method further includes: Get the ambient light intensity; If the ambient light intensity is less than or equal to a first preset threshold, performing image enhancement processing on the original image, and performing the step of extracting the eye region image and the eye state corresponding to the eye region image from the original image based on the original image after image enhancement processing; If the ambient light intensity is greater than a second preset threshold, desaturating the original image, and extracting the eye region image and the eye state corresponding to the eye region image from the original image based on the desaturated original image; The second preset threshold is greater than or equal to the first preset threshold.
3. The sight line estimation method according to claim 1, wherein: After the step of inputting the eye feature map into a preset gaze estimation model and outputting a gaze estimation result, the method further includes: Obtaining a preset calibration coefficient, wherein the calibration coefficient is used to calibrate the deviation between the gaze direction estimated by the calibration model and the actual gaze direction of the user; The line of sight estimation result is calibrated based on the preset calibration coefficient to obtain a calibrated line of sight direction.
4. The sight line estimation method according to claim 3, wherein: The step of obtaining a preset calibration coefficient includes: Obtaining a device identifier of a line of sight estimation device, obtaining a calibration coefficient corresponding to the device identifier based on a first preset mapping relationship, and determining that the calibration coefficient corresponding to the device identifier is a preset calibration coefficient, wherein the first preset mapping relationship is a mapping relationship between different device identifiers and calibration coefficients; or Obtain the account identifier of the user login account in the line of sight estimation device, obtain the calibration coefficient corresponding to the account identifier based on a second preset mapping relationship, and determine the calibration coefficient corresponding to the account identifier and the preset calibration coefficient, wherein the second preset mapping relationship is a mapping relationship between different account identifiers and calibration coefficients.
5. The sight line estimation method according to claim 3, wherein: Before the step of obtaining the preset calibration coefficient, the method further includes: Acquire the actual sight line direction of the user when gazing at a preset reference point and the reference image when gazing at the preset reference point; performing line of sight estimation on the reference image based on the line of sight estimation model to obtain an estimated line of sight direction; The deviation between the estimated sight line direction and the actual sight line direction is fitted to obtain a preset calibration coefficient.
6. The sight line estimation method according to any one of claims 1 to 5, wherein: The step of extracting the eye region image from the original image and the eye state corresponding to the eye region image comprises: Input the original image into a preset object detection model, and output the eye area position and eye prediction state; The original image is cropped based on the eye region position to obtain an eye region image, and the predicted eye state is determined to be the eye state corresponding to the eye region image.
7. The sight line estimation method according to any one of claims 1 to 5, wherein: The step of performing image segmentation processing on the eye region image to obtain an eye feature map comprises: Inputting the eye region image into an improved U-Net network, performing image segmentation processing on the eye region image through the improved U-Net network, and outputting an eye feature map; The improved U-Net network uses 10×10 convolution layers for convolution processing in the upsampling path and the downsampling path.
8. The sight line estimation method according to any one of claims 1 to 5, wherein: The gaze estimation model is a lightweight DenseNet network. The step of inputting the eye feature map into the pre-trained gaze estimation model and outputting a gaze estimation result comprises: Input the eye feature map into the pre-trained lightweight Densenet network, and output a gaze estimation result; The lightweight Densenet network includes 5 dense blocks and each dense block is connected by a transition layer.
9. The sight line estimation method according to any one of claims 1 to 5, wherein: Before the step of inputting the eye feature map into a pre-trained gaze estimation model, the method further includes: Obtaining a training data set and a gaze estimation model to be trained, wherein the training data set includes a training eye feature map obtained by image segmentation processing and a gaze direction corresponding to the training eye feature map; The gaze estimation model is trained with the training eye feature map as input and the gaze direction as a label to obtain the trained gaze estimation model.
10. A sight line estimation device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the computer program is executed by the processor, the line of sight estimation method according to any one of claims 1 to 9 is implemented.
11. A readable storage medium, characterized in that: The readable storage medium includes a computer-readable storage medium, on which a sight line estimation program is stored. When the sight line estimation program is executed by a processor, the steps of the sight line estimation method according to any one of claims 1 to 9 are implemented.
Citation Information
Patent Citations
Eyeball tracking method, terminal and computer storage medium
CN116149461A
Driver sight line estimation method and system based on face three-dimensional reconstruction
CN117935334A
Eye movement tracking interaction method and device
CN119225528A
Sight line position processing apparatus, image capturing apparatus, training apparatus, sight line position processing method, training method, and storage medium
US20220026985A1
Gaze capturing method and apparatus, storage medium, and terminal
WO2022193809A1