Infrared near-eye pupil detection method based on pulse neural network
By reconstructing and lightweighting the spiking neural network of the YOLOv8 architecture, and combining the P-Block1 module and A-LIF neurons, the problem of balancing accuracy, model size and power consumption in pupil detection in head-mounted AR/VR devices was solved, achieving high-precision and low-power pupil detection results.
Patent Information
- Application Number
- CN202511883198.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-13
- Publication Date
- 2026-02-10
AI Technical Summary
Existing pupil detection solutions struggle to achieve a good balance between detection accuracy, model lightweighting, and power consumption in head-mounted AR/VR devices, especially exhibiting poor robustness under complex eye structures and uneven lighting conditions.
Based on the YOLOv8 architecture, a spiking neural network was reconstructed and lightweighted. A pupil detection model was constructed using the lightweight spiking feature extraction module P-Block1 and integer spiking neurons A-LIF. Through multi-scale feature extraction and cross-scale feature fusion, the model size and computational overhead were reduced, and the detection accuracy and robustness were improved.
It achieves high-precision, low-power pupil detection in resource-constrained near-eye interaction scenarios, with a model parameter count of no more than 0.5M and an average accuracy of no less than 95%, making it suitable for head-mounted AR/VR devices.
Smart Images

Figure CN121505680A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of neural network technology, and in particular to an infrared near-eye pupil detection method based on a spiking neural network. Background Technology
[0002] In recent years, the rapid development of Augmented Reality (AR) and Virtual Reality (VR) technologies has profoundly changed the mode of human-computer interaction and promoted the establishment of an immersive virtual-real fusion environment (Reference [1]). In head-mounted AR / VR devices, eye-tracking-based interaction has attracted widespread attention due to its natural, non-contact, and intuitive characteristics (Reference [2]). Among them, pupil detection, as the core technology of eye-tracking interaction, directly affects the accuracy and stability of interaction.
[0003] Currently, mainstream pupil detection algorithms are mainly divided into two categories: Traditional image processing methods rely on low-level visual features such as edge detection, ellipse fitting, and threshold segmentation (references [3, 4]). Their advantage is low computational complexity. However, in real near-eye scenarios faced by head-mounted devices, such as complex eye structures, uneven lighting, and reflections from eyeglass lenses, these methods are less robust, and their detection accuracy and stability are difficult to guarantee.
[0004] Deep learning-based methods: These methods utilize convolutional neural networks (CNNs) for end-to-end training, significantly improving the accuracy and generalization ability of pupil detection (References [5, 6]). However, high-performance CNN models typically have a large number of parameters and are computationally intensive, resulting in high power consumption. Although compression can be achieved through network pruning, quantization, or the use of lightweight architectures (References [11-14]), detection accuracy is often sacrificed while compressing the model.
[0005] Therefore, on head-mounted AR / VR edge devices where computing resources, storage space, and power consumption are strictly limited, existing pupil detection solutions struggle to achieve a good balance between detection accuracy, model lightweighting, and operating power consumption.
[0006] Spiking Neural Networks (SNNs), as the third generation of neural network models, use discrete pulse sequences for information encoding and transmission. They have the characteristics of event-driven and computationally sparsity, and theoretically can significantly reduce computational power consumption (references [7-9, 15]), providing a new approach for low-power visual perception tasks. Recently, some studies have attempted to apply SNNs to the field of object detection (reference
[16] ), for example by replacing the activation function in the YOLO model or constructing a general detection model. However, these works are mainly aimed at general object detection scenarios, and their model designs are not optimized for the specific task of near-eye pupil detection. When directly transferring the general model to near-eye pupil detection, there are often problems such as model redundancy, insufficient extraction of eye detail features, and failure to adapt to the characteristics of near-infrared images, resulting in an unsatisfactory balance between accuracy and power consumption.
[0007] In summary, there is an urgent need to propose a near-eye pupil detection method specifically designed for resource-constrained near-eye interaction scenarios, which can simultaneously achieve high accuracy, ultra-low power consumption, and extremely small model size.
[0008] The references are as follows: [1] H.-C. Jetter et al., 'Transitional Interfaces in Mixed and Cross-Reality: A new frontier?', 2021. [2] E. Campos-Castillo and EG Del Hierro-Gutierrez, 'Eye Control and Accessibility in Computers', in XLVII Mexican Conference on Biomedical Engineering, vol. 116, JDJA Flores Cuautle, B. Benitez-Mata, JJReyes-Lagos, HY Hernandez Acosta, G. Ames Lastra, E. Zuñiga-Aguilar, E. Del Hierro-Gutierrez, and RA Salido-Ruiz, Eds., in IFMBE Proceedings, vol. , Cham : Springer Nature Switzerland , 2025 , pp. 100 - 100 . 219-2 doi: 10.1007 / 978-3-031-82123-3_21. [3] U. Wildenmann and F. Schaeffel, 'Variations of pupil centrationand their effects on video eye tracking', Ophthalmic Physiologic Optic, vol.33, no. 6, pp. 634–641, Nov. 1999; 2013, doi: 10.1111 / opo.12086. [4] RS Kothari, AK Chaudhary, RJ Bailey, JB Pelz, and GJ Diaz, 'EllSeg: An Ellipse Segmentation Framework for Robust GazeTracking', IEEE Trans. Visually. Computing. Graphics, vol. 27, no. 5, pp. 2757–2767, May 2021, doi:10.1109 / TVCG.2021.3067765. [5] S. Eivazi, T. Santini, A. Keshavarzi, T. Kübler, and A. Mazzei,‘Improving real-time CNN-based pupil detection through domain-specific dataaugmentation’, in Proceedings of the 11th ACM Symposium on Eye TrackingResearch & Applications, Denver Colorado: ACM, June 2019, pp. 1–6. doi:10.1145 / 3314111.3319914. [6] Y. W. Lee, K. W. Kim, T. M. Hoang, M. Arsalan, and K. R. Park,‘Deep Residual CNN-Based Ocular Recognition Based on Rough Pupil Detection inthe Images by NIR Camera Sensor’, Sensors, vol. 19, no. 4, p. 842, Feb. 2019,doi: 10.3390 / s19040842. [7] Y. W. Lee, K. W. Kim, T. M. Hoang, M. Arsalan, and K. R. Park,‘Deep Residual CNN-Based Ocular Recognition Based on Rough Pupil Detection inthe Images by NIR Camera Sensor’, Sensors, vol. 19, no. 4, p. 842, Feb. 2019,doi: 10.3390 / s19040842. [8] G. Verma, N. Bindal, A. Nisar, S. Dhull, and B. K. Kaushik,‘Advances in Neuromorphic Spin-Based Spiking Neural Networks: A review’, IEEENanotechnology Mag., vol. 15, no. 5, pp. 33-44, Oct. 2021, doi: 10.1109 / MNANO.2021.3098219. [9] S. Zheng, L. Qian, P. Li, C. He, X. Qin, and X. Li, ‘AnIntroductory Review of Spiking Neural Network and Artificial Neural Network:From Biological Intelligence to Artificial Intelligence’, in ArtificialIntelligence Trends, Academy and Industry Research Collaboration Center(AIRCC), June 2022, pp. 12–145. doi: 10.5121 / csit.2022.121010.
[10] M. Yaseen, ‘What is YOLOv8: An In-Depth Exploration of theInternal Features of the Next-Generation Object Detector’, Aug. 28, 2024,arXiv: arXiv:2408.15857. doi: 10.48550 / arXiv.2408.15857.
[11] M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen,‘MobileNetV2: Inverted Residuals and Linear Bottlenecks’, Mar. 21, 2019,arXiv: arXiv:1801.04381. doi:
[12] M. Tan and Q. V. Le, ‘EfficientNet: Rethinking Model Scaling forConvolutional Neural Networks’, Sept. 11, 2020, arXiv: arXiv:1905.11946. doi:10.48550 / arXiv.1905.11946.
[13] L.-D. Quach, K. N. Quoc, A. N. Quynh, H. T. Ngoc, and N. Thai-Nghe, ‘Tomato Health Monitoring System: Tomato Classification, Detection, andCounting System Based on YOLOv8 Model With Explainable MobileNet Models UsingGrad-CAM++’, IEEE Access, vol. 12, pp. 9719–9737, 2024, doi: 10.1109 / ACCESS.2024.3351805.
[14] L. Jia et al., ‘MobileNet-CA-YOLO: An Improved YOLOv7 Based onthe MobileNetV3 and Attention Mechanism for Rice Pests and DiseasesDetection’, Agriculture, vol. 13, no. 7, p. 1285, June 2023, doi: 10.3390 / agriculture13071285.
[15] W. Maass, ‘Networks of spiking neurons: The third generation ofneural network models’, Neural Networks, vol. 10, no. 9, pp. 1659–1671, Dec.1997, doi: 10.1016 / S0893-6080(97)00011-7.
[16] X. Luo, M. Yao, Y. Chou, B. Xu, and G. Li, 'Integer-ValuedTraining and Spike-Driven Inference Spiking Neural Network for High-performance and Energy-efficient Object. Summary of the Invention
[0009] This invention proposes an infrared near-eye pupil detection method based on a spiking neural network, which solves the problem in existing pupil detection schemes that are difficult to achieve a good balance between detection accuracy, model lightweighting, and power consumption.
[0010] The technical solution of this invention is implemented as follows: The first aspect of this invention provides an infrared near-eye pupil detection method based on a spiking neural network, comprising the following steps: Acquire the infrared near-eye image to be detected; The infrared near-eye image is input into a pre-trained pupil detection model, and the pupil position detection result is output. The pupil detection model is based on the macroscopic topology of YOLOv8, namely "backbone network-neck network-detection head," and undergoes pulsed reconstruction and lightweight adaptation, specifically including: In the backbone network, multiple cascaded lightweight pulse feature extraction modules P-Block1 and downsampling operations are used to construct a feature extraction network for extracting multi-scale spatiotemporal features of infrared near-eye images. In the neck network, a pulse convolutional network is constructed using the P-Block1 module combined with upsampling and concatenation operations to replace the feature pyramid network structure in the original YOLOv8, thereby achieving cross-scale feature fusion.
[0011] Specifically, the P-Block1 module includes the following components connected in sequence: A 1×1 convolutional layer is used to increase the dimensionality of the input feature map's channels; The spiking neuron activation layer performs spiking activation on the upgraded features. A 3×3 depth convolutional layer is used to extract global spatial features; The spiking neuron activation layer performs spiking activation on the features after deep convolution. A 1×1 convolutional layer is used to reduce the dimensionality of the feature map's channels; The spiking neuron activation layer performs spiking activation on the dimensionality-reduced features; The P-Block1 module also includes a residual connection for fusing the module's input with the output after dimensionality reduction by the second 1×1 convolutional layer.
[0012] Furthermore, the operation process of the P-Block1 module is defined by the following formula: ; ; in, Indicates at time step t Input characteristics after pulse activation; This indicates that the P-Block1 module is at time step t The output; X This represents the input feature sequence after pulse activation. T For time steps, C For the number of channels, H , W These are the height and width of the feature map, respectively; The pulse-inverted bottleneck convolution operation, used to capture temporal and spatial features, is expressed as follows: ; in, This represents the dimensionality increase operation of the first 1×1 convolutional layer; This represents a depthwise convolution operation in a 3×3 convolutional layer. This represents the dimensionality reduction operation of the second 1×1 convolutional layer; S This represents the activating layer of spiking neurons.
[0013] Preferably, the spiking neuron is an integer spiking neuron (A-LIF), and its dynamic characteristics are described by the following formula: ; ; ; in, Indicates the current time step t The membrane potential is used to integrate the previous time step. t -1 time information Spatial input at the current moment ; Indicates the current time step t The output pulse, the value of which is determined by... The ratio to the membrane potential threshold is quantitatively determined; Indicates integer operation. express x The amplitude limit range is [0,4]; This is the membrane potential decay factor.
[0014] Specifically, the A-LIF neurons use fixed integers from 0 to 4 for pulse intensity quantization to distinguish the pupil, iris, sclera, and background regions with significant grayscale differences in infrared near-eye images.
[0015] Specifically, after training, the total number of parameters of the pupil detection model is no more than 0.5M, and the average accuracy in the infrared near-eye pupil detection task is no less than 95%.
[0016] A second aspect of the present invention provides an electronic device, including a memory and a processor, wherein the memory stores a computer program executable on the processor, and the processor executes the computer program to implement the steps of the infrared near pupil detection method.
[0017] A third aspect of the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the infrared near-eye pupil detection method.
[0018] Compared with the prior art, the beneficial effects of the present invention are as follows: (1) This invention uses the YOLOv8 architecture as a benchmark and performs a systematic spiking neural network reconstruction and lightweight design. It not only inherits the powerful feature extraction and target detection capabilities of the original framework, but also introduces the characteristics of event-driven and sparse computing through spiking design, achieving an excellent balance between detection accuracy, model efficiency and power consumption. It is particularly suitable for near-eye interaction scenarios with limited resources. (2) The lightweight pulse feature extraction module P-Block1 designed in this invention can fully extract multi-scale spatiotemporal features with very few parameters by combining multiple convolutional layers with spiking neuron activation layers and residual connections, which greatly reduces the model size and computational cost, and lays the foundation for the deployment of the model at the edge. (3) The A-LIF (integer pulse) neuron used in this invention expands the pulse output from "whether to fire" to "fire intensity level", realizing more refined encoding of analog information into pulse sequence, significantly reducing the quantization error generated when quantizing continuous values into pulses, enabling the network to more accurately characterize the subtle grayscale differences between regions such as pupil and iris in infrared images, thereby improving the detection accuracy and robustness of the model under complex eye imaging conditions. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is a schematic diagram of the network topology of the pupil detection model in an embodiment of the present invention.
[0021] Figure 2 This is a schematic diagram of the network structure of the P-Block1 module in an embodiment of the present invention.
[0022] Figure 3 This is a schematic diagram comparing the pulse quantization mechanism of A-LIF neurons with that of traditional LIF neurons and I-LIF neurons in an embodiment of the present invention; Figure 3 In the diagram, (a) represents a schematic diagram of the pulse quantization mechanism of a traditional binary LIF neuron; (b) represents a schematic diagram of the pulse quantization mechanism of a traditional general integer I-LIF neuron; and (c) represents a schematic diagram of the pulse quantization mechanism of the integer pulse neuron A-LIF of the present invention. Detailed Implementation
[0023] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0024] Reference Figure 1 The first aspect of this invention provides an infrared near pupil detection method based on a spiking neural network, comprising the following steps: Acquire the infrared near-eye image to be detected; The infrared near-eye image is input into a pre-trained pupil detection model, and the pupil position detection result is output. The pupil detection model is based on the macroscopic topology of YOLOv8, namely "backbone network-neck network-detection head," and undergoes pulsed reconstruction and lightweight adaptation, specifically including: In the backbone network, multiple cascaded lightweight pulse feature extraction modules P-Block1 and downsampling operations are used to construct a feature extraction network for extracting multi-scale spatiotemporal features of infrared near-eye images. Specifically, the backbone network comprises multiple stages, each consisting of at least one P-Block1 module followed by a downsampling operation. Through this cascaded structure, the network progressively compresses the spatial size of the feature maps while increasing the number of channels and the receptive field, thereby efficiently extracting multi-scale feature representations from the input image, from coarse to fine.
[0025] In the neck network, a pulse convolutional network is constructed using the P-Block1 module combined with upsampling and concatenation operations to replace the feature pyramid network (FPN / PAFPN) structure in the original YOLOv8, thereby achieving cross-scale feature fusion.
[0026] Specifically, the neck network receives feature maps from different stages of the backbone network. Through upsampling, the semantically rich feature maps from deeper layers are enlarged to the same spatial size as the shallower feature maps. Then, through concatenation, features from different depths are fused along the channel dimension. The fused features are then fed into the P-Block1 module for further feature extraction and integration. This series of operations (upsampling-concatenation-P-Block1) constitutes the lightweight pulse feature fusion path of this invention, effectively achieving the complementarity and enhancement of multi-scale information, replacing the more computationally complex feature pyramid module in the original YOLOv8.
[0027] Specifically, such as Figure 2 As shown, the P-Block1 module includes the following components connected in sequence: A 1×1 convolutional layer is used to increase the dimensionality of the input feature map's channels; The spiking neuron activation layer performs spiking activation on the upgraded features. A 3×3 depth convolutional layer is used to extract global spatial features; The spiking neuron activation layer performs spiking activation on the features after deep convolution. A 1×1 convolutional layer is used to reduce the dimensionality of the feature map's channels; The spiking neuron activation layer performs spiking activation on the dimensionality-reduced features; The P-Block1 module also includes a residual connection for fusing the module's input with the output after dimensionality reduction by the second 1×1 convolutional layer.
[0028] The P-Block1 module structure is inspired by the inverted residual bottleneck structure of MobileNetV2, employing an "expansion-convolution-compression" approach. First, a 1×1 convolution increases the channel dimension to enhance feature representation. Then, a 3×3 depthwise convolution extracts lightweight spatial features in a higher-dimensional space, with its larger kernel helping to capture broader contextual information. Finally, a 1×1 convolution compresses the number of channels back to the target dimension. Each convolution operation is followed by a spiking neuron activation layer, converting continuous activation values into sparse spiking sequences. The residual connections at the end of the module help mitigate information loss caused by the sparsity and discreteness of spiking neurons, promoting gradient flow and stabilizing network training.
[0029] Furthermore, the operation process of the P-Block1 module is defined by the following formula: ; ; in, Indicates at time step t Input characteristics after pulse activation; This indicates that the P-Block1 module is at time step t The output; X This represents the input feature sequence after pulse activation. T For time steps, C For the number of channels, H , W These are the height and width of the feature map, respectively; The pulse-inverted bottleneck convolution operation, used to capture temporal and spatial features, is expressed as follows: ; in, This represents the dimensionality increase operation of the first 1×1 convolutional layer; This represents a depthwise convolution operation in a 3×3 convolutional layer. This represents the dimensionality reduction operation of the second 1×1 convolutional layer; S This represents the activating layer of spiking neurons.
[0030] Preferably, the spiking neuron is an integer spiking neuron (A-LIF), and its dynamic characteristics are described by the following formula: ; ; ; in, Indicates the current time step tThe membrane potential is used to integrate the previous time step. t -1 time information Spatial input at the current moment ; Indicates the current time step t The output pulse, the value of which is determined by... The ratio to the membrane potential threshold is quantitatively determined; Indicates integer operation. express x The amplitude limit range is [0,4]; This is the membrane potential decay factor.
[0031] Figure 3 This invention demonstrates the evolution of the quantization mechanism of the integer spiking neuron A-LIF, as shown in the following example. Figure 3 As shown in (a), traditional binary LIF neurons use binary output (0 or 1) to indicate whether a pulse is fired. Although this enhances sparsity, its expressive power is limited, making it difficult to represent the differences in complex inputs, thus resulting in the largest quantization error. Figure 3 As shown in (b), traditional general-purpose integer I-LIF neurons expand the output into multiple discrete integer values, thereby improving network performance without significantly increasing computational complexity. This is suitable for general scenarios, but still results in significant errors when applied to the scenario described in this invention. Figure 3 As shown in (c), the integer spiking neuron A-LIF used in this invention optimizes the fixed range [0,4] integer quantization for the grayscale characteristics of infrared images. This mechanism directly maps the ratio of membrane potential to threshold to a finite integer intensity during the inference stage, rather than performing binary judgment, which significantly reduces quantization error. This allows the network to distinguish grayscale differences in regions such as pupil and iris more precisely, thereby greatly improving detection accuracy while maintaining low power consumption.
[0032] The working principle of A-LIF neurons is as follows: At each time step, the neuron first integrates the temporal information from the previous time step. and the spatial input at the current moment Update to obtain the membrane potential at the current moment. Then, for the current membrane potential... Perform rounding operation The result is then passed through a limiting function. Constrained within the integer interval [0, 4], the pulse firing intensity at the current moment is directly obtained. This output pulse is no longer a binary value (0 or 1) like that of a traditional LIF neuron, but rather an integer between 0 and 4, representing different firing intensities. Finally, from the current membrane potential... Subtract the emitted pulse value from the middle Multiply the difference by the attenuation factor The result serves as the time information for the next time step. Throughout the process, the intensity of the pulse firing is determined by direct integer quantization of the membrane potential, rather than by the Boolean judgment of "whether it exceeds the threshold" in traditional LIF neurons. This mechanism based on integer quantization within a fixed range of [0,4] significantly reduces the quantization error in the information encoding process of traditional binary spiking neurons.
[0033] Specifically, under near-infrared illumination, the pupil region absorbs a large amount of infrared light and appears deep black. The iris, sclera, and background skin have different reflective properties, creating a clear gradient in image grayscale. The A-LIF neuron of this invention uses fixed integers from 0 to 4 for pulse intensity quantization to distinguish the pupil, iris, sclera, and background regions with significant grayscale differences in near-eye infrared images. The integer-level pulse output of the A-LIF neuron can more finely encode this grayscale level information. Compared to traditional binary pulse neurons, it can more effectively preserve and transmit the contrast details between the target and the background, thereby improving the model's robustness in complex near-eye scenes (such as those with localized reflections or shadows).
[0034] Specifically, after training, the total number of parameters of the pupil detection model is no greater than 0.5M, and the average accuracy (mAP@50) in the infrared near-eye pupil detection task is no less than 95%.
[0035] The training dataset for the pupil detection model of this invention consists of a large number of infrared near-eye images labeled with pupil positions, and data augmentation techniques for near-infrared highlights and occlusion scenarios can be applied to improve the model's generalization ability. Through the aforementioned global lightweight architecture design, efficient P-Block1 module, and low-quantization-error A-LIF neurons, the model achieves extremely high detection accuracy on specialized datasets while greatly compressing the parameter scale, making it perfectly suited for head-mounted AR / VR devices that are extremely sensitive to power consumption and size.
[0036] A second aspect of the present invention provides an electronic device, including a memory and a processor, wherein the memory stores a computer program executable on the processor, and the processor executes the computer program to implement the steps of the infrared near pupil detection method.
[0037] A third aspect of the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the infrared near-eye pupil detection method.
[0038] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for detecting near-eye pupils based on a spiking neural network, characterized in that, Includes the following steps: Acquire the infrared near-eye image to be detected; The infrared near-eye image is input into a pre-trained pupil detection model, and the pupil position detection result is output. The pupil detection model is based on the macroscopic topology of YOLOv8, namely "backbone network-neck network-detection head," and undergoes pulsed reconstruction and lightweight adaptation, specifically including: In the backbone network, multiple cascaded lightweight pulse feature extraction modules P-Block1 and downsampling operations are used to construct a feature extraction network for extracting multi-scale spatiotemporal features of infrared near-eye images. In the neck network, a pulse convolutional network is constructed using the P-Block1 module combined with upsampling and concatenation operations to replace the feature pyramid network structure in the original YOLOv8, thereby achieving cross-scale feature fusion.
2. The infrared near-eye pupil detection method based on a spiking neural network as described in claim 1, characterized in that, The P-Block1 module comprises the following components connected in sequence: A 1×1 convolutional layer is used to increase the dimensionality of the input feature map's channels; The spiking neuron activation layer performs spiking activation on the upgraded features. A 3×3 depth convolutional layer is used to extract global spatial features; The spiking neuron activation layer performs spiking activation on the features after deep convolution. A 1×1 convolutional layer is used to reduce the dimensionality of the feature map's channels; The spiking neuron activation layer performs spiking activation on the dimensionality-reduced features; The P-Block1 module also includes a residual connection for fusing the module's input with the output after dimensionality reduction by the second 1×1 convolutional layer.
3. The infrared near-eye pupil detection method based on a spiking neural network as described in claim 2, characterized in that, The operation process of the P-Block1 module is defined by the following formula: ; ; in, Indicates at time step t Input characteristics after pulse activation; This indicates that the P-Block1 module is at time step t The output; X This represents the input feature sequence after pulse activation. T For time steps, C For the number of channels, H , W These are the height and width of the feature map, respectively; The pulse-inverted bottleneck convolution operation, used to capture temporal and spatial features, is expressed as follows: ; in, This represents the dimensionality increase operation of the first 1×1 convolutional layer; This represents a depthwise convolution operation on a 3×3 convolutional layer. This represents the dimensionality reduction operation of the second 1×1 convolutional layer; S This represents the activating layer of spiking neurons.
4. The infrared near-eye pupil detection method based on a spiking neural network as described in claim 2, characterized in that, The spiking neuron is an integer spiking neuron A-LIF, and its dynamic characteristics are described by the following formula: ; ; ; in, Indicates the current time step t The membrane potential is used to integrate the previous time step. t -1 time information Spatial input at the current moment ; Indicates the current time step t The output pulse, the value of which is determined by... The ratio to the membrane potential threshold is quantitatively determined; Indicates integer operation. express x The amplitude limit range is [0,4]; This is the membrane potential decay factor.
5. The infrared near-eye pupil detection method based on a spiking neural network as described in claim 4, characterized in that, The A-LIF neurons use fixed integers from 0 to 4 for pulse intensity quantization to distinguish the pupil, iris, sclera, and background regions with significant grayscale differences in infrared near-eye images.
6. The infrared near-eye pupil detection method based on a spiking neural network as described in claim 1, characterized in that, After training, the total number of parameters of the pupil detection model is no more than 0.5M, and the average accuracy in the infrared near-eye pupil detection task is no less than 95%.
7. An electronic device comprising a memory and a processor, wherein the memory stores a computer program executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the infrared near-eye pupil detection method as described in any one of claims 1 to 6.
8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the infrared near-eye pupil detection method as described in any one of claims 1 to 6.