Deep learning-based eye tracking glint detection method and apparatus

Through deep learning of eye tracking spot detection method, the primary neural network model is used to perform semantic segmentation of eyeball images, solving the problem of low spot detection accuracy in the existing technology, achieving accurate confirmation of spot sequence numbers, and improving the accuracy of eye tracking.

WO2025146022A1PCT designated stage expired Publication Date: 2025-07-10NANCHANG VIRTUAL REALITY RES INST CO LTD

Patent Information

Application Number
PCT/CN2024/143932
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-02
Filing Date
2024-12-30
Publication Date
2025-07-10

AI Technical Summary

Technical Problem

In the prior art, the eye movement spot detection accuracy is not high, and the spot number cannot be effectively confirmed.

Method used

Using a deep learning-based method, single-channel eyeball images are semantically segmented through a primary neural network model, multi-channel label images are generated, and the neural network model is iteratively optimized through loss function to determine the spot center and spot sorting.

Benefits of technology

The accuracy of eye movement spot detection is improved, the accuracy of spot sequence numbers is ensured, and effective guarantees are provided for subsequent eye movement tracking and eye movement posture estimation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024143932_10072025_PF_FP_ABST
    Figure CN2024143932_10072025_PF_FP_ABST
Patent Text Reader

Abstract

The present application provides a deep learning-based eye tracking glint detection method and apparatus. The method comprises: processing data sets of single-channel sample eye images having glints, and storing the processed data sets in a txt file; generating first multi-channel label images of the single-channel sample eye images; performing semantic segmentation on the data sets corresponding to the single-channel sample eye images by means of a primary neural network model and outputting second multi-channel label images; determining a loss function on the basis of the first multi-channel label images and the second multi-channel label images; iteratively optimizing the primary neural network model by means of the loss function to obtain a final neural network model; and processing, by means of the final neural network model, the single-channel eye images to be detected, and performing reasoning to obtain glint centers of the single-channel eye images to be detected and a glint sequence. The present application can implement accurate eye glint detection to confirm glint serial numbers.
Need to check novelty before this filing date? Find Prior Art

Description

A method and device for spot detection based on deep learning eye tracking

[0001] Cross-references to related documents

[0002] This application claims priority to the Chinese patent application filed with the Patent Office of the State Intellectual Property Office of China on January 2, 2024, with application number 2024100036613, and invention name “A Digital Human System Application Method and Device”, all contents of which are incorporated herein by reference. Technical Field

[0003] The present application belongs to the field of deep learning technology, and in particular relates to a method and device for detecting eye-tracking spots based on deep learning. Background Art

[0004] With the development of science and technology, eye tracking technology has become a research hotspot. Eye tracking is a technology used to study the movement trajectory of the human eye during visual tasks. It can record the position and duration of the human eye's gaze point when viewing visual information, and further infer the perception, cognition, and decision-making process of the human eye in visual tasks, helping scientists understand the mechanism of human visual information processing. Eye tracking can be applied in many fields, such as human-computer interaction design, psychology, neuroscience, advertising and marketing. In eye tracking, gaze estimation is the key, but visual estimation requires eye spot detection to confirm the spot number for visual estimation. The accuracy of eye spot detection in existing technologies is not high, so a new solution is needed to solve the problems in existing technologies. Technical Solutions

[0005] In order to solve or alleviate the problems in the prior art, a method and device for eye tracking spot detection based on deep learning are proposed.

[0006] In a first aspect, an embodiment of the present application provides a method for spot detection based on deep learning eye tracking, comprising:

[0007] The data set of single-channel sample eyeball images with light spots is processed and stored in a txt file;

[0008] Read the data set whose first digit is not 0 in the data set of single-channel sample eyeball images in the txt file;

[0009] Use the OpenCV image vision library to generate a floating-point image with all pixel values ​​1, where the size of the floating-point image is the same as that of the single-channel sample eyeball image;

[0010] The value obtained by multiplying the last two values ​​in each data group by the width and height of the single-channel sample eye image is used as the center of the circle, the first digit of each data group is used as the pixel, and the preset pixel value is used as the radius to draw a circle on the floating-point image to obtain a first multi-channel label image corresponding to the single-channel sample eye image;

[0011] Performing semantic segmentation on the data group corresponding to the single-channel sample eye image using a primary neural network model to output a second multi-channel label image;

[0012] Determine a loss function based on the first multi-channel label image and the second multi-channel label image;

[0013] Iteratively optimizing the primary neural network model through the loss function to obtain a final neural network model;

[0014] The final neural network model is used to process the single-channel eye image to be tested with light spots, and the light spot center and light spot order of the single-channel eye image to be tested are obtained by inference.

[0015] Compared with the prior art, the embodiment of the present application provides a method for spot detection based on deep learning eye tracking, which processes a data set of a single-channel sample eye image with a spot and stores it in a txt file; reads a data set in which the first digit of the single-channel sample eye image data set in the txt file is not 0; uses the opencv image vision library to generate a floating-point image with all pixel values ​​1, and the size of the floating-point image is the same as that of the single-channel sample eye image; takes the last two values ​​in each of the data sets multiplied by the width and height of the single-channel sample eye image as the center of a circle, takes the first digit of each of the data sets as a pixel, and draws a circle on the floating-point image with a preset pixel value as the radius. circle, obtain a first multi-channel label image corresponding to the single-channel sample eye image; perform semantic segmentation on the data group corresponding to the single-channel sample eye image through the primary neural network model to output a second multi-channel label image; determine the loss function according to the first multi-channel label image and the second multi-channel label image; iteratively optimize the primary neural network model through the loss function to obtain a final neural network model; process the single-channel eye image to be tested with a light spot through the final neural network model, and infer the light spot center and light spot order of the single-channel eye image to be tested. The technical solution provided by the present application can more accurately perform eye movement light spot detection to confirm the light spot sequence number.

[0016] In a second aspect, an embodiment of the present application further provides a deep learning-based eye tracking spot detection device, comprising:

[0017] A processing module is used to process the single-channel sample eyeball image with light spots and store it in a txt file;

[0018] A generation module is configured to read a data group whose first digit is not 0 from a data group of a single-channel sample eye image in a txt file; generate a floating-point image whose pixel values ​​are all 1 using the OpenCV image vision library, wherein the floating-point image has the same size as the single-channel sample eye image; draw a circle on the floating-point image with the first digit of each data group as a pixel and a preset pixel value as a radius, using the value obtained by multiplying the last two values ​​in each data group by the width and height of the single-channel sample eye image as the center of the circle, and obtaining a first multi-channel label image corresponding to the single-channel sample eye image;

[0019] A semantic segmentation module, configured to perform semantic segmentation on the data group corresponding to the single-channel sample eye image using a primary neural network model to output a second multi-channel label image;

[0020] a determination module, configured to determine a loss function based on the first multi-channel label image and the second multi-channel label image;

[0021] An optimization module, configured to iteratively optimize the primary neural network model through the loss function to obtain a final neural network model;

[0022] The inference module is used to process the single-channel eye image to be tested with the light spot through the final neural network model, and infer the light spot center and light spot order of the single-channel eye image to be tested.

[0023] Compared with the prior art, the embodiment of the present application provides an eye tracking spot detection device based on deep learning, which has the same beneficial effects as the technical solution provided in the first aspect and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. Some specific embodiments of the present application will be described in detail in an illustrative and non-restrictive manner with reference to the drawings. The same reference numerals in the drawings indicate the same or similar components or parts. It should be understood by those skilled in the art that these drawings are not necessarily drawn to scale. In the drawings:

[0025] FIG1 is a flow chart of a method for spot detection based on deep learning eye tracking provided by an embodiment of the present application;

[0026] Figure 2 is a structural schematic diagram of a deep learning-based eye tracking spot detection device provided in an embodiment of the present application. Best Mode for Carrying Out the Invention

[0027] In order to enable those skilled in the art to better understand the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.

[0028] Referring to FIG1 , in a first aspect, an embodiment of the present application provides a method for detecting a spot of light based on deep learning eye tracking, comprising:

[0029] Step S01, processing the data set of the single-channel sample eyeball image with light spots and storing it in a txt file;

[0030] Step S01 specifically includes: collecting a single-channel sample eye image with a light spot;

[0031] On the collected single-channel sample eyeball images, the spot center of each single-channel sample eyeball image is marked in sequence, and the spot center of each single-channel sample eyeball image is normalized;

[0032] The single-channel sample eyeball image after normalization of the spot center is saved in a txt file.

[0033] It should be noted that a single-channel sample eye image with a light spot is collected using relevant equipment (the equipment can be a VR headset with a circle of lights and a camera installed at the corresponding positions of the left and right eye corners, and the left and right eye images are collected through the camera). On the collected single-channel sample eye image, the center position of the light spot is manually marked in sequence and normalized. The location where the light spot is not captured is marked with a label and coordinate of 0. The single-channel sample eye image after the light spot center normalization is saved in a txt file.

[0034] The data saved in the txt file is similar to the following:

[0035] 1 0.834609 0.384967 (serial number 1); 1 0.864758 0.784047 (serial number 2); 1 0.794779 0.567892 (serial number 3); 0 0.000000 0.000000 (serial number 4); 1 0.694934 0.749345 (serial number 5); 0 0.000000 0.000000 (serial number 6); 0 0.000000 0.000000 (serial number 7); 1 0.479966 0.397679 (serial number 8);

[0036] Starting from the corner of the eye, the left eye is marked in clockwise order, and the right eye is marked in counterclockwise order. The first integer 1 of each number above indicates that there is a light spot, and the integer 0 indicates that there is no light spot. The two decimals behind it represent the relative position of the center of the light spot to the center of the image. For example, the first three values ​​1 0.834609 0.384967, 1 indicates that there is a light spot at the corner of the eye. If the pixel coordinates of the center of the light spot are (x, y), and the width and height of the image are H and W respectively, then x / W = 0.834609, y / H = 0.384967, and 0 0.000000 0.000000 means that no light spot is detected. The above data means that there are 8 light spots in total, of which 5 are detected.

[0037] Step S02: After processing the content stored in the txt file, a first multi-channel label image corresponding to the single-channel sample eye image is generated;

[0038] Step S02 specifically includes: reading a data group whose first digit is not 0 in a data group of a single-channel sample eye image in a txt file;

[0039] Use the OpenCV image vision library to generate a floating-point image with all pixel values ​​1, where the size of the floating-point image is the same as that of the single-channel sample eye image;

[0040] The value obtained by multiplying the last two values ​​in each of the data groups by the width and height of the single-channel sample eye image is used as the center of the circle, the first digit of each of the data groups is used as the pixel, and a circle is drawn on the floating-point image with a preset pixel value as the radius to obtain the first multi-channel label image corresponding to the single-channel sample eye image.

[0041] It should be noted that the data groups whose first digit of the label is not 0 in the above txt file are read (serial number 1) 1 0.834609 0.384967; (serial number 2) 1 0.864758 0.784047; (serial number 3) 1 0.794779 0.567892; (serial number 5) 1 0.694934 0.749345; (serial number 8) 1 0.479966 0.397679; (each three data are a group), and the label data is changed to the corresponding serial number plus 1, such as the data group in the above txt becomes:

[0042] [[2 0.834609 0.384967] [3 0.864758 0.784047] [4 0.794779 0.567892] [6 0.694934 0.749345] [9 0.479966 0.397679]]

[0043] Use the OpenCV image vision library to generate a floating-point image with all pixel values ​​​​as 1. The image size of the floating-point image is consistent with the size of the original image when the camera is captured. The width and height of the floating-point image are H and W respectively. Then, take the value obtained by multiplying the last two values ​​of each set of data by the image width and height as the center, take the first digit as the pixel, and draw a circle (i.e., a solid circle) on the floating-point image in a filled manner with a radius of R (R=4 pixels). The solid circle is the light spot in the area with the same pixel points.

[0044] For example, the data set [2 0.834609 0.384967]: Draw a solid circle with a radius of 4 pixels using 2 as the center coordinate and the first digit 2 in the data group as the pixel.

[0045] In this way, each single-channel sample eye image generates a first multi-channel label image with the same name as the original image.

[0046] Step S03, performing semantic segmentation on the data group corresponding to the single-channel sample eye image using a primary neural network model to output a second multi-channel label image;

[0047] It should be noted that the input of the primary neural network model is designed to be batch*m*W*H, and the output of the primary neural network model is batch*n*W*H, wherein batch is the number of label images corresponding to the single-channel sample eye image used in each iteration, m and n represent the number of channels, and W and H represent the width and height of the label image corresponding to the single-channel sample eye image.

[0048] It should be noted that after semantic segmentation through a neural network, the single-channel image will be converted into a multi-channel image label. In the embodiment of the present application, if there are 9 pixels on a single-channel image, the single-channel image is a grayscale image, and the pixel value of each pixel is one of 1 to 9. Then, converting the single-channel image into a multi-channel image label is actually converting the single-channel image into 9 single-channel binary images, in which each pixel value is 0 or 1. For example, in the first image label, except for the pixel with a pixel value of 1, the pixel values ​​of other areas are all 0. For another example, in the second image, except for the pixel value of 2 in the single-channel image, the pixel value of other areas is 0. And so on to obtain a 9-channel image label.

[0049] In a specific application, in the first channel, all pixel values ​​of the circular area drawn above are 0, and if there is no circular area drawn above in the first channel, the pixel value is 1; in the second channel, if the pixel value of the pixel point in the circle drawn above is 1, and there is no circular area drawn above in the second channel, the pixel value is 0; in the third channel, if the pixel value of the pixel point in the circle drawn above is 1, and there is no circular area drawn above in the second channel, the pixel value is 0, and so on to obtain the second multi-channel graphic label l.

[0050] In an embodiment of the present application, the primary neural network model is a Net network model, and the Net network model may be a Le-Net network model.

[0051] In an embodiment of the present application, the first multi-channel label image and the second multi-channel label image are both multiple binary images in which each pixel value is 0 or 1.

[0052] Step S04, determining a loss function according to the first multi-channel label image and the second multi-channel label image;

[0053] Step S04 specifically includes: obtaining a loss value loss1 between the first channel label image in the first multi-channel label image and the first channel label image in the second multi-channel label image, and a loss value loss2 between other channel label images in the first multi-channel label image and other channel label images in the second multi-channel label image, and determining a loss function according to the following formula:

[0054]

[0055] Among them, W1 and W2 represent the weight values ​​of loss value loss1 and loss value loss2 respectively.

[0056] It should be noted that the loss function is divided into two parts. One part is the loss value loss1 between the first first multi-channel label image of the single-channel sample eye image and the first second multi-channel label image output by the primary neural network, and the other part is the loss value loss2 between the other channel label images of the single-channel sample eye image and the other channel label images output by the primary neural network.

[0057] Step S05, iteratively optimizing the primary neural network model through the loss function to obtain a final neural network model;

[0058] It should be noted that the above loss value is used to continuously iteratively optimize the primary neural network model until the primary neural network model fully converges and the final neural network model is output.

[0059] Step S06: Processing the single-channel eye image with the light spot through the final neural network model, and inferring the light spot center and light spot order of the single-channel eye image.

[0060] Step S06 specifically includes: inputting the collected single-channel eye image to be tested into the final neural network model to obtain a third multi-channel label image of the single-channel eye image to be tested;

[0061] Inputting the collected single-channel eye image to be tested into the final neural network model to obtain a third multi-channel label image of the single-channel eye image to be tested;

[0062] sequentially polling the third multi-channel label image of the single-channel eye image to be tested to determine a single-channel image, wherein the pixel value of each pixel coordinate point on the single-channel image is a channel sequence number corresponding to the maximum pixel value of the third multi-channel label image having the same pixel coordinate point;

[0063] Obtaining a binary image having the same resolution as the pixel values ​​of each pixel coordinate point in the single-channel image;

[0064] The center position of each connected domain of each channel of the binary image is determined by the findContours function in the opencv image vision library. The connected domain corresponds to the spot number, and the center position of the spot and the spot order are obtained according to the spot number.

[0065] It should be noted that a single-channel eye image to be tested is collected, input into the final neural network model for inference, and outputs a third multi-channel label image output1.

[0066] Polling each channel of the third multi-channel label image output1, obtaining the channel where the maximum pixel value is located to determine the single-channel image output2, wherein the pixel value of each pixel coordinate point in the single-channel image output2 is the channel sequence number corresponding to the maximum pixel value of the same pixel coordinate point in the third multi-channel label image;

[0067] If there are 9 light spots, the first channel is channel 0, and the channels of the third multi-channel label image output1 are in the order of 0, 1, 2, 3, 4, 5, 6, 7, 8, that is, 9 channels. For example, the pixel values ​​of each channel of output1 at the pixel coordinate (0, 0) are [0.034554 0.05459 0.000000 0.000000 0.007462 0.934712 0.000000 0.0034401 0.000000], the maximum pixel value at this position is 0.934712, and the channel ordinal number is 5. Then the pixel value at the pixel coordinate (0, 0) of the single-channel image output2 is 5, and all output1s are polled in turn to obtain the pixel values ​​of each pixel coordinate point of the single-channel image output2.

[0068] A binary image output3 having the same resolution as the single-channel image output2 is obtained according to the pixel value of each pixel coordinate point of the single-channel image output2, and the pixel value of the binary image output3 is 255.

[0069] The findContours function in the OpenCV image vision library is used to determine the center position of each connected domain in the binary image output3, that is, to infer the center of the light spot through the final neural network model.

[0070] The connected domain of the binary image Output3 corresponds to the pixel value of the single-channel image Output2, which is the light spot sequence number. From this, the center position of the light spot can be obtained, and the order of the light spots can be obtained, providing effective data for subsequent eye tracking.

[0071] The embodiment of the present application transforms the eye movement spot detection problem into a semantic segmentation problem by processing the spot into a spot region, i.e., a point-to-surface sample label generation method, thereby effectively and rapidly achieving eye movement spot detection. At the same time, the semantic segmentation concept is borrowed and applied to spot detection in eye tracking. Because the semantic segmentation concept can remove natural light and tear points, it more effectively overcomes the interference of natural light and tear points in the eyes. The results of deep learning reasoning are post-processed to effectively extract eye movement spots and ensure the accuracy of the spot sequence number, which can provide a strong guarantee for subsequent eye tracking and eye posture estimation.

[0072] Referring to FIG2 , in a second aspect, an embodiment of the present application further provides a deep learning-based eye tracking spot detection device, comprising:

[0073] The processing module 21 is used to process the data set of the single-channel sample eyeball image with light spots and store it in a txt file;

[0074] A generating module 22 is configured to read a data group whose first digit is not 0 from a data group of a single-channel sample eye image in a txt file; generate a floating-point image whose pixel values ​​are all 1 using the OpenCV image vision library, wherein the floating-point image has the same size as the single-channel sample eye image; draw a circle on the floating-point image using the value obtained by multiplying the last two values ​​in each data group by the width and height of the single-channel sample eye image as the center of a circle, using the first digit of each data group as a pixel, and using a preset pixel value as a radius, to obtain a first multi-channel label image corresponding to the single-channel sample eye image;

[0075] A semantic segmentation module 23 is configured to perform semantic segmentation on the data group corresponding to the single-channel sample eye image using a primary neural network model to output a second multi-channel label image;

[0076] a determination module 24, configured to determine a loss function based on the first multi-channel label image and the second multi-channel label image;

[0077] An optimization module 25 is configured to iteratively optimize the primary neural network model using the loss function to obtain a final neural network model;

[0078] The inference module 26 is used to process the single-channel eye image to be tested with the light spot through the final neural network model, and infer the light spot center and light spot order of the single-channel eye image to be tested.

[0079] Compared with the prior art, the embodiment of the present application provides an eye tracking spot detection device based on deep learning, which has the same beneficial effects as the technical solution provided in the first aspect and will not be repeated here.

[0080] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some or all of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for detecting eye-tracking light spots based on deep learning, characterized in that, Including: Processing the data group of the single-channel sample eyeball image with a light spot and storing it in a txt file; Reading the data groups in the txt file whose first digit is not 0 in the data group of the single-channel sample eyeball image; Generating a floating-point image with all pixel values being 1 using the opencv image vision library, and the size of the floating-point image is the same as that of the single-channel sample eyeball image; Taking the value obtained by multiplying the last two values in each data group by the width and height of the single-channel sample eyeball image as the center of the circle, taking the first digit in each data group as the pixel, and drawing a circle on the floating-point image with a preset pixel value as the radius to obtain the first multi-channel label image corresponding to the single-channel sample eyeball image; Performing semantic segmentation on the data group corresponding to the single-channel sample eyeball image through a primary neural network model to output a second multi-channel label image; Determining a loss function according to the first multi-channel label image and the second multi-channel label image; Iteratively optimizing the primary neural network model through the loss function to obtain a final neural network model; Processing the single-channel eyeball image with a light spot to be measured through the final neural network model, and inferring the light spot center and light spot sorting of the single-channel eyeball image to be measured.

2. The method for detecting the light spot of eye tracking based on deep learning according to claim 1, wherein, The processing the data group of the single-channel sample eyeball image with a light spot and storing it in a txt file includes: Collecting a single-channel sample eyeball image with a light spot; On the collected single-channel sample eyeball image, sequentially marking the light spot center of each single-channel sample eyeball image, and normalizing the light spot center of each single-channel sample eyeball image; Saving the data group of the single-channel sample eyeball image after normalizing the light spot center in a txt file.

3. The method for detecting an eye-tracking light spot based on deep learning according to claim 1, characterized in that The determining a loss function according to the first multi-channel label image and the second multi-channel label image includes: Obtaining the loss value loss1 between the first channel label image in the first multi-channel label image and the first channel label image in the second multi-channel label image, and the loss value loss2 between the other channel label images in the first multi-channel label image and the other channel label images in the second multi-channel label image, and determining the loss function according to the following formula: where W1 and W2 respectively represent the weight values of the loss value loss1 and the loss value loss2.

4. The method for detecting the eye tracking light spot based on deep learning according to claim 3, wherein Both the first multi-channel label image and the second multi-channel label image are multiple binary images with each pixel value being 0 or 1.

5. The method for detecting an eye tracking light spot based on deep learning according to claim 4, characterized in that, The processing the single-channel eyeball image with a light spot to be measured through the final neural network model, and inferring the light spot center and light spot sorting of the single-channel eyeball image to be measured includes: Inputting the collected single-channel eyeball image to be measured into the final neural network model to obtain a third multi-channel label image of the single-channel eyeball image to be measured; Sequentially polling the third multi-channel label image of the single-channel eyeball image to be measured to determine a single-channel image, and the pixel value of each pixel coordinate point on the single-channel image is the channel number corresponding to the maximum pixel value with the same pixel coordinate point in the third multi-channel label image; Obtain a binary image with the same resolution as the pixel values of each pixel coordinate point in the single-channel image; Determine the central positions of each connected component in each channel of the binary image through the findContours function in the opencv image vision library. The connected component corresponds to a spot serial number, and obtain the spot center position and spot sorting according to the spot serial number.

6. An eyeball tracking spot detection device based on deep learning, characterized in that, It includes: A processing module, configured to process a single-channel sample eyeball image with spots and store it in a txt file; A generation module, configured to read the data groups in the txt file whose first digit in the data group of the single-channel sample eyeball image is not 0; generate a floating-point image with all pixel values being 1 using the opencv image vision library, and the size of the floating-point image is the same as the size of the single-channel sample eyeball image; draw circles on the floating-point image with the value obtained by multiplying the last two values in each data group by the width and height of the single-channel sample eyeball image as the center, the first digit in each data group as the pixel, and a preset pixel value as the radius, to obtain a first multi-channel label image corresponding to the single-channel sample eyeball image; A semantic segmentation module, configured to perform semantic segmentation on the data group corresponding to the single-channel sample eyeball image through a primary neural network model and output a second multi-channel label image; A determination module, configured to determine a loss function according to the first multi-channel label image and the second multi-channel label image; An optimization module, configured to iteratively optimize the primary neural network model through the loss function to obtain a final neural network model; An inference module, configured to process a single-channel test eyeball image with spots through the final neural network model, and infer the spot center and spot sorting of the single-channel test eyeball image.

Citation Information

Patent Citations

  • Target extraction method and device

    CN112070793A

  • Metal spot detection method and system based on neural network

    CN115082428A

  • Light spot labeling method and system

    CN116051631A

  • Eyeball tracking light spot detection method and device based on deep learning

    CN117496584A

  • Category learning neural networks

    US20200027002A1

Cited By

  • Glint numbering method, system, and device, and storage medium

    US12693741B2