Pupil positioning method and device, electronic equipment and storage medium

By acquiring the event stream within the preset time interval collected by the event camera and adjusting the initial signal according to the polarity, coordinates and timestamp of the event, the problems of low pupil positioning efficiency and low accuracy in the prior art are solved, and fast and accurate pupil positioning is achieved.

CN119992633APending Publication Date: 2025-05-13BEIJING 7INVENSUN TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311499243.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-10
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art uses event cameras to collect data for pupil positioning, and the existing conversion methods are cumbersome, making it easy to filter out effective information.

Method used

By obtaining the event stream within the preset time interval, adjusting the value of the initial signal according to the polarity, coordinates and timestamp of the event, obtaining the initial eye signal, and pupil positioning is performed based on the signal.

Benefits of technology

It realizes the rapid and accurate conversion of event stream data, avoids information loss, and improves the efficiency and accuracy of pupil positioning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992633A_ABST
    Figure CN119992633A_ABST
Patent Text Reader

Abstract

The invention discloses a pupil positioning method and device, electronic equipment and a storage medium. The method comprises the steps that an event stream in a preset time interval is acquired, events in the event stream comprise polarity, coordinates and timestamps, and the event stream comprises eye features; adjusting the value of the initial signal according to the polarity, the coordinate and the timestamp of each event to obtain an initial eye signal; the pupil positioning is performed based on the initial eye signal, the position information of the pupil is determined, the problems of low efficiency and low accuracy when the pupil positioning is performed according to event stream data are solved, the problem of high data volume is avoided by acquiring the eye features by acquiring the event stream, the numerical value of the signal is determined through the polarity, the coordinate and the timestamp recorded by the event, and the accuracy of the pupil positioning is improved. According to the method, the initial eye signal is obtained, all information is reserved in the process of converting the event stream into the signal, information loss is avoided, event stream data are quickly and accurately converted, the implementation process is simple, and accurate positioning of pupils is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a pupil positioning method, device, electronic equipment and storage medium. Background Art

[0002] With the development of VR, AR, eye control and other technologies, it is becoming increasingly important to accurately locate pupils and sight lines. When locating and tracking pupils, eye tracking technology usually collects eye image data through ordinary cameras. This method requires a lot of redundant data to be processed, has high requirements on the processor, and has a low dynamic range rate.

[0003] At present, in order to improve data processing efficiency, event cameras can be used to collect data when performing pupil positioning, which can greatly reduce redundant data and thus improve the computational efficiency of the post-processing algorithm. However, event cameras do not have the concept of "frames". When changes occur in the real scene, the event camera will generate some pixel-level outputs (i.e., events). An event specifically includes (t, x, y, p), where x and y are the pixel coordinates of the event in 2D space, t is the timestamp of the event, and p is the polarity of the event (if the brightness is enhanced, p = +1 indicates a positive event, otherwise, P = -1 indicates a negative event). When using event camera data, it is usually necessary to convert it into an image. The existing conversion method will count and denoise the event stream data of the event camera over a period of time, and generate images in a gradient manner. The process is cumbersome and it is easy to filter out effective information. Therefore, how to quickly and accurately perform pupil positioning based on event stream data has become a problem to be solved. Summary of the invention

[0004] The present invention provides a pupil positioning method, device, electronic device and storage medium to solve the problems of low efficiency and low accuracy when pupil positioning is performed based on event stream data.

[0005] According to one aspect of the present invention, there is provided a pupil positioning method, comprising:

[0006] Acquire an event stream within a preset time interval, wherein the events in the event stream include polarity, coordinates, and timestamps, and the event stream contains eye features;

[0007] Adjusting the value of the initial signal according to the polarity, coordinates and timestamp of each event to obtain an initial eye signal;

[0008] Pupil positioning is performed based on the initial eye signal to determine the position information of the pupil.

[0009] According to another aspect of the present invention, there is provided a pupil positioning device, comprising:

[0010] An event stream acquisition module, used to acquire an event stream within a preset time interval, wherein the events in the event stream include polarity, coordinates and timestamps, and the event stream contains eye features;

[0011] An initial eye signal generating module, used for adjusting the value of the initial signal according to the polarity, coordinates and timestamp of each event to obtain an initial eye signal;

[0012] The pupil positioning module is used to perform pupil positioning based on the initial eye signal to determine the position information of the pupil.

[0013] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0014] at least one processor; and

[0015] a memory communicatively connected to the at least one processor; wherein,

[0016] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can perform the pupil positioning method described in any embodiment of the present invention.

[0017] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the pupil positioning method described in any embodiment of the present invention when executed.

[0018] The technical solution of the embodiment of the present invention is to obtain an event stream within a preset time interval, wherein the events in the event stream include polarity, coordinates and timestamps, and the event stream contains eye features; adjust the value of the initial signal according to the polarity, coordinates and timestamp of each event to obtain an initial eye signal; locate the pupil based on the initial eye signal to determine the position information of the pupil, thereby solving the problems of low efficiency and low accuracy when locating the pupil according to the event stream data, and obtaining eye features by collecting the event stream to avoid the problem of high data volume, and determine the value of the signal by the polarity, coordinates and timestamp recorded by the event to obtain the initial eye signal, and retain all information in the process of converting the event stream into the signal to avoid information loss, and quickly and accurately convert the event stream data, and the implementation process is simple, thereby achieving accurate positioning of the pupil.

[0019] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0021] Figure 1 is a flow chart of a pupil positioning method provided according to Embodiment 1 of the present invention;

[0022] Figure 2 is a flow chart of a pupil positioning method provided according to Embodiment 2 of the present invention;

[0023] Figure 3 is a network structure diagram of a deep learning model provided according to Embodiment 2 of the present invention;

[0024] Figure 4 is a network structure diagram of a convolution block provided according to Embodiment 2 of the present invention;

[0025] Figure 5 is a schematic diagram of the structure of a pupil positioning device provided according to Embodiment 3 of the present invention;

[0026] Figure 6 It is a schematic diagram of the structure of an electronic device for implementing the pupil positioning method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0027] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0028] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0029] Embodiment 1

[0030] Figure 1 This is a flow chart of a pupil positioning method provided in the first embodiment of the present invention. This embodiment is applicable to the case where the pupil is quickly and accurately positioned according to event stream data. The method can be performed by a pupil positioning device. The pupil positioning device can be implemented in the form of hardware and / or software. The pupil positioning device can be configured in an electronic device. Figure 1 As shown, the method includes:

[0031] S101. Acquire an event stream within a preset time interval, where the events in the event stream include polarity, coordinates, and timestamps, and the event stream contains eye features.

[0032] In this embodiment, the preset time interval can be specifically understood as a preset time interval, and the size of the preset time interval can be set according to the requirements for data processing accuracy in the actual application process; if the preset time interval is too large, data redundancy will occur, and the accuracy of the data will also be affected; if the preset time interval is too small, insufficient accuracy will occur due to too little data, for example, it can be set to 1ms, 2ms, etc. The polarity can be a positive number or a negative number, which is used to indicate brightness enhancement and brightness reduction, respectively. Eye features can be specifically understood as feature data used to represent the eyes, for example, pupil information (such as pupil position, pupil shape), iris position, iris shape, eyelid position, eye corner position, light spot (also called Purkin spot) information (such as light spot position), etc. Specifically in this embodiment, pupil information can be pupil center, ellipse major and minor axis lengths, and major axis angles, etc.

[0033] The event stream is collected by an event camera, which can be installed on eye-tracking devices, AR / VR glasses and other devices. Users can wear eye-tracking devices and other devices, and then the event camera can collect the user's eye features while wearing them; or the event camera is set in a fixed position, and the user needs to get close to the event camera to complete the collection of eye features. The event camera is a sensor that "only senses moving objects". There is an independent photoelectric sensor module at each pixel of the event camera. When the brightness change at the pixel exceeds the set threshold, event data (sometimes also called pulse data) will be generated and output; because all pixels work independently, the data output of the event camera is asynchronous and sparse in space.

[0034] The event camera collects event streams in real time and records the events that have changed. The preset time interval is set in advance, and the event stream within the preset time interval is obtained. The event stream within the preset time interval is used to determine a signal. The event stream includes events that record changes. The event includes polarity, coordinates and timestamps. The coordinates are used to indicate the location where the event occurred, the polarity is used to indicate whether the brightness of the event increases or decreases, and the timestamp is used to indicate the time when the event occurred.

[0035] S102 , adjusting the value of the initial signal according to the polarity, coordinates and timestamp of each event to obtain an initial eye signal.

[0036] In this embodiment, the initial signal can be understood as the default signal or data used when forming the signal, and the initial signal includes different coordinate positions and the values ​​corresponding to each coordinate position; the initial eye signal can be specifically understood as the signal or data used for pupil positioning, and the signal or data is the processed value. Each event stream acquired within a preset time interval can form an initial eye signal, and each initial eye signal is formed on the basis of the initial signal to describe or represent the eye features.

[0037] The initial signal is pre-set, and the position of the coordinate point that needs to be adjusted in the initial signal is determined according to the coordinates of each event in the event stream, and the order of each event used to adjust the value of the coordinate point is determined according to the timestamp. The value of the coordinate point at the corresponding position in the initial signal is adjusted according to the polarity of the event. The adjustment of the value of the coordinate point can be increased, decreased or unchanged. When adjusting the value of a certain coordinate point, the polarity and the number of occurrences of all events occurring at the coordinate point can be counted, and the calculation is performed according to the number of event occurrences. The number of occurrences can be the sum of the number of times the polarity is positive and the number of times the polarity is negative, or the sum of the positive number accumulated when the polarity is positive and the negative number accumulated when the polarity is negative, etc. The value of the initial signal at the adjusted coordinate point is obtained by polarity calculation. The initial eye signal is obtained by adjusting the value of the initial signal. The initial eye signal obtained in this step is a one-time signal. By processing the event stream within a preset time interval, an initial eye signal is obtained. Since the initial eye signal includes event information within the preset time interval, the data volume is large, and the changes in the eye can be accurately described.

[0038] S103: Perform pupil positioning based on the initial eye signal to determine pupil position information.

[0039] In this embodiment, the position information of the pupil may be the center point, width, height, or coordinates of the four vertices of the minimum circumscribed rectangle of the pupil, etc. Any information that can describe the position of the pupil and uniquely determine a position will suffice.

[0040] Since the event stream contains eye features, the initial eye signal formed according to the event stream also contains eye features, which can reflect the condition of the eye. For the initial eye signal, it can be identified through data processing, algorithms, models, etc., to determine the pupil in the signal, and further determine the position information of the pupil. When the pupil is located according to the initial eye signal, the initial eye signal can be directly located, or the initial eye signal can be processed, for example, denoising, normalization, etc., and then the processed signal is located to obtain the position information of the pupil. A condition can be set in advance to determine whether to process the initial eye signal. After determining the initial eye signal, it is determined whether the initial eye signal meets the condition. If so, the initial eye signal is processed, otherwise, it is not processed. The condition can be whether the maximum value or the minimum value corresponding to each coordinate point in the initial signal is within the threshold range, etc.

[0041] An embodiment of the present invention provides a pupil positioning method, which obtains an event stream within a preset time interval, wherein the events in the event stream include polarity, coordinates and timestamps, and the event stream contains eye features; the value of an initial signal is adjusted according to the polarity, coordinates and timestamps of each event to obtain an initial eye signal; pupil positioning is performed based on the initial eye signal to determine the position information of the pupil, thereby solving the problems of low efficiency and low accuracy when performing pupil positioning based on event stream data, and obtaining eye features by collecting event streams to avoid the problem of high data volume. The value of the signal is determined by the polarity, coordinates and timestamps recorded by the events to obtain the initial eye signal. In the process of converting the event stream into a signal, all information is retained to avoid information loss, and the event stream data is converted quickly and accurately. The implementation process is simple, thereby achieving accurate positioning of the pupil.

[0042] Embodiment 2

[0043] Figure 2 This is a flow chart of a pupil positioning method provided in Embodiment 2 of the present invention. This embodiment is refined on the basis of the above embodiment. Figure 2 As shown, the method includes:

[0044] S201. Acquire an event stream within a preset time interval, where the events in the event stream include polarity, coordinates, and timestamps, and the event stream contains eye features.

[0045] S202 , sorting the timestamps corresponding to the events, selecting the first timestamp as the current timestamp, and taking the initial signal as the signal to be processed.

[0046] In this embodiment, the current timestamp can be specifically understood as a timestamp used to determine the currently processed data during data processing. The signal to be processed can be specifically understood as a signal for currently adjusting the value corresponding to the coordinate point.

[0047] Sorting the timestamps corresponding to each event can be done from front to back in chronological order, or from back to front. The embodiment of the present application takes sorting from front to back as an example, that is, the event that occurred first is in front. Sorting each timestamp can be achieved by a sorting algorithm, and the sorting algorithm can be a bubble sort, a selection sort, a quick sort, etc. The first timestamp is selected as the current timestamp, and the initial signal is used as the signal to be processed, so that the signal to be processed is processed according to the event corresponding to the current timestamp.

[0048] Alternatively, the embodiment of the present application can also sort each event according to the timestamp. Taking the sorting from front to back as an example, all timestamps corresponding to each event are determined, each timestamp is sorted, and then all events corresponding to each timestamp are determined, and all events corresponding to each timestamp are sorted in parallel to obtain the sorting of all events.

[0049] Optionally, the values ​​of each coordinate of the initial signal are all 0.

[0050] In the embodiment of the present application, the numerical value corresponding to each coordinate in the initial signal is assigned to 0, so that all events are recorded based on the initial signal. The size of the initial signal is the resolution of the event camera. The numerical value of each coordinate of the initial signal is set to 0. When the numerical value of the initial signal is processed based on the event stream data, only the changes in the eye collected are recorded according to the polarity of the event, without paying attention to the original data, thereby reducing the amount of data processing.

[0051] S203: Take the event corresponding to the current timestamp as the event to be processed.

[0052] In this embodiment, the pending event can be specifically understood as the event that needs to be processed currently. For each moment, the number of events generated can be 0, 1, or more. Therefore, after determining the current timestamp, the event corresponding to the current timestamp is determined. The number of events is not unique, and the event corresponding to the current timestamp is determined as the pending event.

[0053] S204, analyzing the polarity of each event to be processed, and determining the coordinates without polarity according to the coordinates of the event to be processed with positive polarity and the coordinates of the event to be processed with negative polarity.

[0054] The polarity of each event to be processed is analyzed to determine the coordinates of the event to be processed with positive polarity and the coordinates of the event to be processed with negative polarity, and the coordinates for which there are neither the event to be processed with positive polarity nor the event to be processed with negative polarity are determined as coordinates without polarity.

[0055] S205, adding up the values ​​of the coordinates of each event to be processed with positive polarity at the corresponding position in the signal to be processed, subtracting the values ​​of the coordinates of each event to be processed with no polarity at the corresponding position in the signal to be processed, and not processing the values ​​of the coordinates of each event to be processed with negative polarity at the corresponding position in the signal to be processed, to form a converted signal.

[0056] For each coordinate point in the signal to be processed, determine the coordinate point position corresponding to the coordinate of each event to be processed with positive polarity in the signal to be processed, and accumulate the values ​​at the corresponding positions. For example, event 1 (t1, x1, y1, p+) and event 2 (t1, x2, y2, p+) occurred at time t1, where p+ represents positive polarity, and the preset value is 1. If the values ​​of the coordinates (x1, y1) and (x2, y2) in the signal to be processed are both 0, after accumulation, the values ​​of the coordinates (x1, y1) and (x2, y2) are both 1; determine the coordinate point position corresponding to the coordinate of each event to be processed with negative polarity in the signal to be processed, and do not process the value at this position; determine the coordinate point position corresponding to the coordinate of the event to be processed without polarity in the signal to be processed, and subtract the value at this position. In this embodiment, the accumulation of values ​​may be adding a number of preset values, and the accumulation subtraction may be subtracting a number of preset values, the preset value may be 1, 2 or other values, and the accumulated preset value and the accumulated preset value may also be set to different values. The converted signal is obtained by performing three different processings, namely, adding, subtracting and not processing, on the values ​​of each coordinate in the signal to be processed.

[0057] For each signal to be processed, there is only one numerical processing method corresponding to each coordinate of the current timestamp, that is, each coordinate can only generate one event or no event at the same time.

[0058] S206: Determine whether there is a next timestamp after the current timestamp. If so, execute S207; otherwise, execute S208.

[0059] After the processing of the pending event corresponding to the current timestamp is completed, the remaining events are processed. It is determined whether there is a next timestamp after the current timestamp. If so, it is determined that there are still events that have not been processed, and S207 is executed to continue processing; otherwise, it is determined that all events have been processed, and S208 is executed to determine the initial eye signal.

[0060] S207: Taking the next timestamp of the current timestamp as the new current timestamp, taking the converted signal as the new signal to be processed, and repeating step S203.

[0061] The next timestamp of the current timestamp is used as a new current timestamp, and the converted signal is used as a new signal to be processed, and the updating of the current timestamp and the signal to be processed is completed, so as to continue data processing according to the updated current timestamp and the signal to be processed.

[0062] S208: Use the converted signal as the initial eye signal.

[0063] In the embodiment of the present application, the numerical values ​​corresponding to the coordinate points in the initial signal are adjusted through S202-S208, that is, the numerical values ​​corresponding to the coordinate points in the initial signal are adjusted in turn according to the polarity of the events at different times, and the final initial eye signal is obtained, so that the eye features recorded by the events within the preset time interval are merged into a signal through processing, which can avoid the problem of inaccurate recognition caused by too little data. And by accumulating the numerical values ​​of the positions corresponding to the coordinates of the pending events with positive polarity, not processing the numerical values ​​of the positions corresponding to the coordinates of the pending events with negative polarity, and subtracting the numerical values ​​of the positions corresponding to the coordinates without polarity, the gap between the numerical values ​​is expanded, thereby improving the pupil positioning accuracy. In the embodiment of the present application, the conversion of the event stream retains all information, including the position information and time information of each point, and the information is retained to a certain extent when the numerical values ​​of the positions corresponding to the coordinates without polarity are subtracted, so that the information loss is less.

[0064] S209: normalize the initial eye signal to obtain a target eye signal.

[0065] In this embodiment, the target eye signal can be specifically understood as a signal for pupil positioning. Since the initial eye signal is obtained by increasing and decreasing the values ​​of each coordinate point on the basis of the initial signal, the values ​​corresponding to each coordinate point in the initial eye signal may be positive or negative. Therefore, the embodiment of the present application normalizes the initial eye signal, normalizes each value in the initial eye signal to 0-1 or other range values, and obtains the target eye signal. When normalizing the initial eye signal, it can be processed by calling a normalization function or the like.

[0066] As an optional embodiment of this embodiment, this optional embodiment further normalizes the initial eye signal to obtain the target eye signal, which is optimized as follows:

[0067] A1. Compare the values ​​of the coordinates in the initial eye signal to obtain the maximum value and the minimum value.

[0068] Determine the value of each coordinate in the initial eye signal, compare the values, and determine the maximum and minimum values. For example, sort the values ​​using a sorting algorithm to obtain the maximum and minimum values. The sorting algorithm may be bubble sort, selection sort, quick sort, etc.

[0069] A2. Calculate the first difference between the maximum value and the minimum value.

[0070] In this embodiment, the first difference is the difference between the maximum value and the minimum value. The difference obtained by subtracting the minimum value from the maximum value is used as the first difference.

[0071] A3. For the value of each coordinate in the initial eye signal, calculate the second difference between the value and the minimum value, calculate the ratio of the second difference to the first difference, and use the product of the ratio and a preset normalization coefficient as the normalized value of the coordinate.

[0072] In this embodiment, the second difference is the difference between any value and the minimum value in the initial eye signal; the normalization coefficient can be specifically understood as the coefficient for scaling the value when performing normalization processing; the normalized value can be specifically understood as the value after normalization processing.

[0073] A normalization coefficient is preset, and the normalization coefficient can be 1 or any other appropriate value. For the value of each coordinate in the initial eye signal, normalization processing is performed in the following manner: the value is subtracted from the minimum value, and the difference obtained is the second difference; the second difference is divided by the first difference to obtain the ratio of the second difference to the first difference, and the ratio is multiplied by the preset normalization coefficient, and the product obtained is the normalized value corresponding to the coordinate.

[0074] Exemplarily, the present application embodiment provides a method for calculating a normalized value:

[0075] P ij =k*(p ij -p min)*(p max-p min), where P ij is the normalized value corresponding to the coordinate of the i-th row and j-th column, p ij is the value corresponding to the coordinate of the i-th row and j-th column in the initial eye signal, p min is the minimum value, p max is the maximum value, k is the normalization coefficient, and k can be 1 or any other appropriate value.

[0076] A4. forming a target eye signal based on each normalized value.

[0077] Based on the normalized values ​​corresponding to each coordinate, a target eye signal is formed.

[0078] S210, inputting the target eye signal into a pre-trained deep learning model to obtain a center point feature map, an offset feature map, and a width and height feature map output by the deep learning model.

[0079] In this embodiment, the deep learning model is a network model that can be identified based on experience after training. The center point feature map can be specifically understood as a feature map containing the corresponding information of the center point of the pupil; the offset feature map can be specifically understood as a feature map containing the corresponding information of the offset of the pupil; the width and height feature map can be specifically understood as a feature map containing the corresponding information of the width and height of the pupil.

[0080] Pre-train the deep learning model. To reduce the amount of data processing, a model with fewer convolution layers can be used. Obtain training samples, which include the signal to be trained formed according to the event flow within a preset time interval and the label value of the signal to be trained. Input the training samples into the deep learning model for training in sequence, and obtain the predicted value after processing by each convolution layer of the deep learning model. The predicted value is compared with the label value, and the loss function is calculated. The parameters of the deep learning model are adjusted by back propagation through the loss function until a deep learning model that meets the requirements is obtained. In the embodiment of the present application, the deep learning model can output three feature maps, which respectively represent the center point, offset, and width and height. Since the three feature maps represent different types of data respectively, compared with containing different types of data in one map at the same time, the model parameters can be accurately adjusted during the model training process, the model accuracy is higher, and the recognition result is more accurate and reliable. After the model is trained, the target eye signal is input into the pre-trained deep learning model. The deep learning model recognizes the target eye image signal according to the experience learned during the training process, and obtains and outputs the center point feature map, the offset feature map, and the width and height feature map.

[0081] For example, Figure 3A network structure diagram of a deep learning model is provided, taking 2 convolution layers and 4 convolution blocks as an example, the convolution kernel of the first convolution layer 31 is 3*3, the step size S=2, the convolution kernel of the second convolution layer 36 is 3*3, the step size S=1, the convolution kernel of each convolution block is 3*3, the step size S=2, and the number of convolution layers is N. A 1*h*w target eye signal 30 is input to the first convolution layer 31 for processing, wherein h is the height of the target eye signal, w is the width of the target eye signal, and 1 is the dimension of the target eye signal. The output result of the convolution layer 31 is input to the first convolution block 32, where N=2. The convolution block 32 performs convolution processing and inputs the output result to the convolution block 33, where N=3. The convolution block 33 performs convolution processing and inputs the output result to the convolution block 34, where N=3. The convolution block 34 performs convolution processing and inputs the output result to the convolution block 35, where N=2. The convolution block 35 performs convolution processing and multiplies the output result by 2. Sampling; adding the 2x upsampling result to the output result of the convolution block 34; upsampling the added result by 2x, adding the 2x upsampling result to the output result of the convolution block 33, repeatedly upsampling the added result by 2x, and then continuing to add the 2x upsampling result to the output result of the convolution block 32, and inputting the added result to the second convolution layer 36, the convolution layer 36 processes the input data and outputs a center point feature map 37, an offset feature map 38 and a width and height feature map 39.

[0082] For example, Figure 4 A network structure diagram of a convolution block is provided. The convolution block contains N layers, that is, the convolution block is composed of N convolution layers. The convolution kernel of each convolution layer is 3*3, the step size of the first convolution layer is S=2, and the step size of the convolution layers from the second layer to the Nth layer is S=1.

[0083] This application sets up a deep learning model for pupil detection, and its detection accuracy is related to the preset time interval. When the preset time interval is long, the probability of each coordinate point generating an event is higher, the encoded matrix is ​​non-sparse, and the amount of calculation is small, which is more conducive to the subsequent deep learning model to detect the target and improve the detection accuracy.

[0084] S211. Determine the position information of the pupil based on the center point feature map, the offset feature map, and the width and height feature map.

[0085] The center point feature map contains the position information of the corresponding center point. The center point feature map can be identified to determine the position information of the center point. For example, the coordinates corresponding to the maximum value in the center point feature map are determined as the coordinates of the pupil center point; or the center point feature map is processed, for example, by maximum pooling, mean pooling, etc., and the coordinates of the pupil center point are identified according to the processed results. The pupil center point is further offset corrected according to the offset feature map, and the width and height of the pupil center point are obtained from the width and height feature map according to the coordinates of the pupil center point. The position information of the pupil is determined based on the coordinates, offset, width and height of the pupil center point.

[0086] As an optional embodiment of this embodiment, this optional embodiment further optimizes the pupil position information determined based on the center point feature map, the offset feature map and the width and height feature map as follows:

[0087] B1. Perform a maximum pooling operation on the center point feature map, and determine the coordinates corresponding to the maximum value in the center point feature map after pooling as the coordinates of the pupil center point.

[0088] Perform a maximum pooling operation on the center point feature map. The area size of the maximum pooling operation can be set according to needs. For example, 3*3, that is, perform a maximum pooling operation on the center point feature map according to 3*3, and take the maximum value in the 3*3 area as the value of this area. Process each area in turn to obtain the pooled center point feature map, compare the size of the value of each coordinate point in the pooled center point feature map, obtain the maximum value, and determine the coordinates corresponding to the maximum value as the coordinates of the center point of the pupil.

[0089] B2. Correct the coordinates of the pupil center point according to the offset feature map to obtain the target pupil center point coordinates.

[0090] In this embodiment, the target pupil center coordinates can be specifically understood as the corrected pupil center coordinates. According to the pupil center coordinates, the corresponding value in the offset feature map is determined, and this value is used as the offset. The pupil center coordinates are corrected according to the offset to obtain the target pupil center coordinates.

[0091] As an optional embodiment of this embodiment, this optional embodiment further corrects the coordinates of the pupil center point according to the offset feature map to obtain the target pupil center point coordinates, which are optimized as follows:

[0092] B21. Determine the value of the position corresponding to the coordinates of the pupil center point in the first channel of the offset feature map as the horizontal coordinate offset, and determine the sum of the horizontal coordinate offset and the horizontal coordinate in the coordinates of the pupil center point as the corrected horizontal coordinate.

[0093] In this embodiment, the horizontal axis offset can be specifically understood as the offset of the horizontal axis; the first channel can be specifically understood as a channel in the offset feature map, and the offset feature map includes multiple channels to store offsets in different directions.

[0094] The coordinate axis directions corresponding to the values ​​stored in different channels in the offset feature map are predefined. The embodiment of the present application takes the offset of the horizontal coordinate stored in the first channel as an example. The value of the corresponding position in the first channel of the offset feature map is determined according to the coordinates of the pupil center point, and this value is determined as the horizontal coordinate offset. The horizontal coordinate in the coordinates of the pupil center point is determined, and the horizontal coordinate plus the horizontal coordinate offset is obtained. The sum is the corrected horizontal coordinate.

[0095] B22. Determine the value of the position corresponding to the coordinates of the pupil center point in the second channel of the offset feature map as the ordinate offset, and determine the sum of the ordinate offset and the ordinate in the coordinates of the pupil center point as the corrected ordinate.

[0096] In this embodiment, the vertical coordinate offset can be specifically understood as the offset of the vertical coordinate; the second channel can be specifically understood as another channel in the offset feature map, which stores the offset in the vertical axis direction.

[0097] According to the coordinates of the pupil center point, the value of the corresponding position in the second channel of the offset feature map is determined, and this value is determined as the ordinate offset. The ordinate in the coordinates of the pupil center point is determined, and the ordinate offset is added to the ordinate, and the sum obtained is the corrected ordinate.

[0098] B23. Determine the target pupil center coordinates based on the corrected horizontal coordinate and the corrected vertical coordinate.

[0099] The corrected abscissa and the corrected ordinate are used as the abscissa and ordinate in the target pupil center coordinates to obtain the target pupil center coordinates.

[0100] B3. The value of the position corresponding to the coordinates of the pupil center point in the first channel of the width-height feature map is determined as the pupil width, and the value of the position corresponding to the coordinates of the pupil center point in the second channel of the width-height feature map is determined as the pupil height.

[0101] In this embodiment, the width and height feature map also includes two channels, which respectively store the width and height of the pupil. For example, the first channel stores the width and the second channel stores the height. The pupil width is the width of the pupil, and the pupil height is the height of the pupil. The pupil width and the pupil height can be used to determine the size of the pupil. For example, a rectangle is determined as the pupil according to the pupil height and the pupil width, or the pupil height and the pupil width are determined as the short axis and the long axis of the pupil, and an ellipse is determined as the pupil according to the short axis and the long axis.

[0102] According to the coordinates of the pupil center point, the value of the corresponding position in the first channel of the width and height feature map is determined, and this value is determined as the pupil width; according to the coordinates of the pupil center point, the value of the corresponding position in the second channel of the width and height feature map is determined, and this value is determined as the pupil height.

[0103] B4. Determine the pupil position information based on the target pupil center coordinates, pupil width and pupil height.

[0104] Directly use the target pupil center coordinates, pupil width and pupil height as pupil position information; or, determine the coordinates of the vertex or boundary point of the pupil according to the target pupil center coordinates, pupil width and pupil height as pupil position information.

[0105] The embodiment of the present invention provides a pupil positioning method, which improves the pupil positioning accuracy by accumulating the values ​​of the positions corresponding to the coordinates of the to-be-processed events with positive polarity, not processing the values ​​of the positions corresponding to the coordinates of the to-be-processed events with negative polarity, and subtracting the values ​​of the positions corresponding to the coordinates without polarity, thereby expanding the gap between the values ​​and solving the problems of low efficiency and low accuracy when performing pupil positioning based on event stream data. By processing and merging the eye features recorded by the events within a preset time interval into a single signal, the problem of inaccurate recognition caused by too little data can be avoided. In the process of converting the event stream into a signal, all information is retained to avoid information loss, and the event stream data is converted quickly and accurately. The implementation process is simple, thereby achieving high-precision positioning of the pupil.

[0106] Embodiment 3

[0107] Figure 5 This is a schematic diagram of the structure of a pupil positioning device provided in Embodiment 3 of the present invention. Figure 5 As shown, the device includes: an event stream acquisition module 41, an initial eye signal generation module 42 and a pupil positioning module 43.

[0108] The event stream acquisition module 41 is used to acquire an event stream within a preset time interval, wherein the events in the event stream include polarity, coordinates and timestamps, and the event stream contains eye features;

[0109] An initial eye signal generating module 42 is used to adjust the value of the initial signal according to the polarity, coordinates and timestamp of each event to obtain an initial eye signal;

[0110] The pupil positioning module 43 is used to perform pupil positioning based on the initial eye signal to determine the position information of the pupil.

[0111] The embodiment of the present invention provides a pupil positioning device, which solves the problems of low efficiency and low accuracy when pupil positioning is performed based on event stream data. Eye features are acquired by collecting event streams to avoid the problem of high data volume. The value of the signal is determined by the polarity, coordinates and timestamp recorded by the event to obtain an initial eye signal. All information is retained in the process of converting the event stream into a signal to avoid information loss. Event stream data is converted quickly and accurately, and the implementation process is simple, thereby achieving accurate positioning of the pupil.

[0112] Optionally, the initial eye signal generating module 42 includes:

[0113] A sorting unit, used to sort the timestamps corresponding to the events, select the first timestamp as the current timestamp, take the initial signal as the signal to be processed, and take the event corresponding to the current timestamp as the event to be processed;

[0114] A polarity analysis unit, used for analyzing the polarity of each of the to-be-processed events, and determining the coordinates without polarity according to the coordinates of the to-be-processed events with positive polarity and the coordinates of the to-be-processed events with negative polarity;

[0115] A numerical processing unit, used for accumulating the numerical values ​​of the coordinates of each of the to-be-processed events with positive polarity at the corresponding positions in the to-be-processed signal, accumulating the numerical values ​​of the coordinates of each of the to-be-processed events with non-polarity at the corresponding positions in the to-be-processed signal, and not processing the numerical values ​​of the coordinates of each of the to-be-processed events with negative polarity at the corresponding positions in the to-be-processed signal, to form a converted signal;

[0116] A judgment unit is used to judge whether there is a next timestamp after the current timestamp. If so, the next timestamp of the current timestamp is used as the new current timestamp, the converted signal is used as the new signal to be processed, and the step of taking the event corresponding to the current timestamp as the event to be processed is repeated; otherwise, the converted signal is used as the initial eye signal.

[0117] Optionally, the values ​​of each coordinate of the initial signal are all 0.

[0118] Optionally, the pupil positioning module 43 includes:

[0119] A normalization processing unit, configured to perform normalization processing on the initial eye signal to obtain a target eye signal;

[0120] A model processing unit, used to input the target eye signal into a pre-trained deep learning model to obtain a center point feature map, an offset feature map, and a width and height feature map output by the deep learning model;

[0121] A position determination unit is used to determine the position information of the pupil based on the center point feature map, the offset feature map and the width and height feature map.

[0122] Optionally, a normalization processing unit is specifically used to: compare the values ​​of each coordinate in the initial eye signal to obtain a maximum value and a minimum value; calculate a first difference between the maximum value and the minimum value; for each coordinate in the initial eye signal, calculate a second difference between the value and the minimum value, calculate a ratio of the second difference to the first difference, and use the product of the ratio and a preset normalization coefficient as a normalized value of the coordinate; and form a target eye signal based on each of the normalized values.

[0123] Optionally, the position determination unit comprises:

[0124] A pupil coordinate determination subunit, used to perform a maximum pooling operation on the center point feature map, and determine the coordinates corresponding to the maximum value in the center point feature map after pooling as the coordinates of the pupil center point;

[0125] A pupil coordinate correction subunit, used for correcting the coordinates of the pupil center point according to the offset feature map to obtain the target pupil center point coordinates;

[0126] a pupil width and height determination subunit, configured to determine the value of the position corresponding to the coordinates of the pupil center point in the first channel of the width and height feature map as the pupil width, and determine the value of the position corresponding to the coordinates of the pupil center point in the second channel of the width and height feature map as the pupil height;

[0127] The position determination subunit is used to determine the position information of the pupil based on the target pupil center point coordinates, pupil width and pupil height.

[0128] Optionally, a pupil coordinate correction subunit is specifically used to determine the value of the position corresponding to the coordinates of the pupil center point in the first channel of the offset feature map as the horizontal coordinate offset, and determine the sum of the horizontal coordinate offset and the horizontal coordinate in the coordinates of the pupil center point as the corrected horizontal coordinate; determine the value of the position corresponding to the coordinates of the pupil center point in the second channel of the offset feature map as the vertical coordinate offset, and determine the sum of the vertical coordinate offset and the vertical coordinate in the coordinates of the pupil center point as the corrected vertical coordinate; and determine the target pupil center point coordinates based on the corrected horizontal coordinate and the corrected vertical coordinate.

[0129] The pupil positioning device provided in the embodiment of the present invention can execute the pupil positioning method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0130] Embodiment 4

[0131] Figure 6 A schematic diagram of an electronic device 50 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.

[0132] like Figure 6 As shown, the electronic device 50 includes at least one processor 51, and a memory connected to the at least one processor 51 in communication, such as a read-only memory (ROM) 52, a random access memory (RAM) 53, etc., wherein the memory stores a computer program that can be executed by at least one processor, and the processor 51 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 52 or the computer program loaded from the storage unit 58 to the random access memory (RAM) 53. In the RAM 53, various programs and data required for the operation of the electronic device 50 can also be stored. The processor 51, the ROM 52, and the RAM 53 are connected to each other via a bus 54. An input / output (I / O) interface 55 is also connected to the bus 54.

[0133] A number of components in the electronic device 50 are connected to the I / O interface 55, including: an input unit 56, such as a keyboard, a mouse, etc.; an output unit 57, such as various types of displays, speakers, etc.; a storage unit 58, such as a disk, an optical disk, etc.; and a communication unit 59, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 59 allows the electronic device 50 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0134] The processor 51 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the processor 51 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 51 executes the various pupil positioning methods and processes described in Embodiment 1 and / or Embodiment 2 above.

[0135] In some embodiments, the pupil localization method may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 58. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 50 via the ROM 52 and / or the communication unit 59. When the computer program is loaded into the RAM 53 and executed by the processor 51, one or more steps of the pupil localization method described in the above embodiment 1 and / or embodiment 2 may be performed. Alternatively, in other embodiments, the processor 51 may be configured to perform the pupil localization method in any other appropriate manner (e.g., by means of firmware).

[0136] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0137] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the computer program is executed by the processor, the functions / operations specified in the flow chart and / or block diagram are implemented. The computer program may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0138] In the context of the present invention, a computer-readable storage medium may be a tangible medium that may contain or store a computer program for use by or in combination with an instruction execution system, device or equipment. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0139] To provide interaction with a user, the systems and techniques described herein may be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).

[0140] The systems and techniques described herein may be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0141] A computing system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The client and server relationship is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services.

[0142] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps described in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of the present invention can be achieved, and this document does not limit this.

[0143] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A pupil positioning method, characterized in that: include: Acquire an event stream within a preset time interval, wherein the events in the event stream include polarity, coordinates, and timestamps, and the event stream contains eye features; Adjusting the value of the initial signal according to the polarity, coordinates and timestamp of each event to obtain an initial eye signal; Pupil positioning is performed based on the initial eye signal to determine the position information of the pupil.

2. The method according to claim 1, characterized in that The step of adjusting the value of the initial signal according to the polarity, coordinates and timestamp of each event to obtain the initial eye signal includes: Sort the timestamps corresponding to the events, select the first timestamp as the current timestamp, use the initial signal as the signal to be processed, and use the event corresponding to the current timestamp as the event to be processed; Analyzing the polarity of each of the events to be processed, and determining the coordinates without polarity according to the coordinates of the events to be processed with positive polarity and the coordinates of the events to be processed with negative polarity; The values ​​of the coordinates of the events to be processed with positive polarity at the corresponding positions in the signal to be processed are accumulated, the values ​​of the coordinates of the events to be processed with no polarity at the corresponding positions in the signal to be processed are accumulated, and the values ​​of the coordinates of the events to be processed with negative polarity at the corresponding positions in the signal to be processed are not processed, so as to form a converted signal; Determine whether there is a next timestamp after the current timestamp. If so, take the next timestamp of the current timestamp as the new current timestamp, take the converted signal as the new signal to be processed, and repeat the step of taking the event corresponding to the current timestamp as the event to be processed; otherwise, take the converted signal as the initial eye signal.

3. The method according to claim 2, characterized in that The values ​​of each coordinate of the initial signal are all 0.

4. The method according to claim 1, characterized in that The performing pupil positioning based on the initial eye signal to determine the position information of the pupil includes: Normalizing the initial eye signal to obtain a target eye signal; Inputting the target eye signal into a pre-trained deep learning model to obtain a center point feature map, an offset feature map, and a width and height feature map output by the deep learning model; The position information of the pupil is determined based on the center point feature map, the offset feature map and the width and height feature map.

5. The method according to claim 4, characterized in that The normalizing the initial eye signal to obtain a target eye signal includes: Comparing the values ​​of the coordinates in the initial eye signal to obtain the maximum value and the minimum value; Calculate a first difference between the maximum value and the minimum value; For the value of each coordinate in the initial eye signal, calculate a second difference between the value and the minimum value, calculate a ratio of the second difference to the first difference, and use the product of the ratio and a preset normalization coefficient as the normalized value of the coordinate; A target eye signal is formed based on each of the normalized values.

6. The method according to claim 4, characterized in that The determining the pupil position information based on the center point feature map, the offset feature map and the width and height feature map includes: Performing a maximum pooling operation on the center point feature map, and determining the coordinates corresponding to the maximum value in the center point feature map after pooling as the coordinates of the pupil center point; Correcting the coordinates of the pupil center point according to the offset feature map to obtain the coordinates of the target pupil center point; Determine the value of the position corresponding to the coordinates of the pupil center point in the first channel of the width-height feature map as the pupil width, and determine the value of the position corresponding to the coordinates of the pupil center point in the second channel of the width-height feature map as the pupil height; The pupil position information is determined based on the target pupil center point coordinates, pupil width and pupil height.

7. The method according to claim 6, characterized in that The step of correcting the coordinates of the pupil center point according to the offset feature map to obtain the coordinates of the target pupil center point includes: Determine the value of the position corresponding to the coordinate of the pupil center point in the first channel of the offset feature map as the horizontal coordinate offset, and determine the sum of the horizontal coordinate offset and the horizontal coordinate in the coordinate of the pupil center point as the corrected horizontal coordinate; Determine the value of the position corresponding to the coordinate of the pupil center point in the second channel of the offset feature map as the ordinate offset, and determine the sum of the ordinate offset and the ordinate in the coordinate of the pupil center point as the corrected ordinate; The target pupil center point coordinates are determined according to the corrected abscissa and the corrected ordinate.

8. A pupil locating device, characterized in that: include: An event stream acquisition module, used to acquire an event stream within a preset time interval, wherein the events in the event stream include polarity, coordinates and timestamps, and the event stream contains eye features; An initial eye signal generating module, used for adjusting the value of the initial signal according to the polarity, coordinates and timestamp of each event to obtain an initial eye signal; The pupil positioning module is used to perform pupil positioning based on the initial eye signal to determine the position information of the pupil.

9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can perform the pupil positioning method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the pupil positioning method according to any one of claims 1 to 7 when executed.