Pupil positioning method and device and head-mounted display equipment

By acquiring gradient information from multiple image frames and comprehensively evaluating pupil motion status, the problem of inaccurate pupil positioning in head-mounted display devices is solved, improving the accuracy and stability of pupil positioning and enhancing the user experience.

CN120954082APending Publication Date: 2025-11-14VIVO MOBILE COMM CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511202027.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-26
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Due to individual differences in users and different wearing methods, head-mounted display devices have large differences in pupil characteristics when performing pupil localization, resulting in inaccurate localization and poor stability, which affects the accuracy of gaze tracking.

Method used

By acquiring gradient information from multiple image frames, the motion state of the pupil is comprehensively evaluated, and the pupil contour points are located in the current frame using the pupil temporal information, thereby improving the accuracy of the positioning.

Benefits of technology

It improves the accuracy and stability of pupil positioning in head-mounted displays, thus enhancing the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120954082A_ABST
    Figure CN120954082A_ABST
Patent Text Reader

Abstract

The invention discloses a pupil positioning method and device and head-mounted display equipment, and belongs to the technical field of artificial intelligence. The method comprises the steps that first information corresponding to first N image frames of a first image frame is acquired, and the first information comprises a first gradient value of pupil contour points of eyeballs of a user in each image frame of the first N image frames and a second gradient value of eyelid contour points of the user in each image frame, a third gradient value of a mapping point, in the first image frame, of the pupil contour point in each image frame, and a fourth gradient value of a mapping point, in the first image frame, of the eyelid contour point in each image frame; determining a motion state of a pupil in the eyeballs based on the first information; a pupil contour point of the pupil is located in the first image frame based on the motion state.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of artificial intelligence technology, specifically relating to a pupil positioning method, device, and head-mounted display device. Background Technology

[0002] Currently, with the development of head-mounted display devices, these devices can track the user's eye movements to perform corresponding operations. For example, head-mounted display devices can perform gaze-based rendering and gaze-based interaction based on the user's gaze direction and movement trajectory.

[0003] In related technologies, in order to identify a user's gaze, a head-mounted display device can capture an image containing the user's eyeballs using an infrared (IR) camera within the device, and obtain pupil feature information from the image. Based on this pupil feature information, the pupil can be located to determine the user's gaze.

[0004] However, during actual use of head-mounted display devices, due to various differences in the eye structure of different users, the design of different head-mounted display devices, and the way different users wear head-mounted display devices, the state of the pupils of the user's eyes captured by the IR camera varies in the captured images. There may also be situations where eyelashes obscure the pupils, resulting in significant differences in the pupil features obtained by the head-mounted display device. Therefore, it is quite difficult to accurately locate pupils of various shapes. Summary of the Invention

[0005] The purpose of this application is to provide a pupil positioning method, apparatus, and head-mounted display device that can improve the accuracy of pupil positioning in electronic devices.

[0006] In a first aspect, embodiments of this application provide a pupil localization method, which includes: acquiring first information corresponding to the preceding N image frames of a first image frame, the first information including a first gradient value of the pupil contour point of the user's eyeball in each of the preceding N image frames, a second gradient value of the eyelid contour point of the user in each of the preceding N image frames, a third gradient value of the mapping point of the pupil contour point in the first image frame in each of the preceding N image frames, and a fourth gradient value of the mapping point of the eyelid contour point in the first image frame in each of the preceding N image frames, where N∈[1,3] and N is an integer; determining the motion state of the pupil in the eyeball based on the first information; and locating the pupil contour point of the pupil in the first image frame based on the motion state.

[0007] Secondly, embodiments of this application provide a pupil positioning device, which includes: an acquisition module, a determination module, and a positioning module. The acquisition module is used to acquire first information corresponding to the preceding N image frames of a first image frame. The first information includes a first gradient value of the pupil contour point of the user's eyeball in each of the preceding N image frames, a second gradient value of the eyelid contour point of the user in each of the preceding N image frames, a third gradient value of the mapping point of the pupil contour point in the first image frame, and a fourth gradient value of the mapping point of the eyelid contour point in the first image frame, where N ∈ [1,3] and N is an integer. The determination module is used to determine the motion state of the pupil in the eyeball based on the first information acquired by the acquisition module. The positioning module is used to locate the pupil contour point in the first image frame based on the motion state determined by the determination module.

[0008] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, wherein the memory stores programs or instructions executable on the processor, and the programs or instructions, when executed by the processor, implement the steps of the method described in the first aspect.

[0009] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.

[0010] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect.

[0011] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the method described in the first aspect.

[0012] In this embodiment, first information corresponding to the previous N image frames of the first image frame is obtained. The first information includes the first gradient value of the pupil contour point of the user's eyeball in each of the previous N image frames, the second gradient value of the eyelid contour point of the user in each of the previous N image frames, the third gradient value of the mapping point of the pupil contour point in the first image frame, and the fourth gradient value of the mapping point of the eyelid contour point in the first image frame, where N∈[1,3] and N is an integer. Then, based on the first information, the motion state of the pupil in the eyeball is determined. Finally, based on the motion state, the pupil contour point of the pupil is located in the first image frame. In this solution, the user's pupil motion state is comprehensively evaluated by combining the current frame (the first image frame) with the inter-frame information of the previous N image frames (the first information). Based on the pupil motion state, the user's pupil outline point is located in the first image frame. In other words, by making full use of the pupil temporal information, the gradient value of the historical pupil outline point in the current frame is calculated. Based on this gradient value, the pupil motion state is comprehensively evaluated, and the user's pupil outline point can be accurately located. Thus, the accuracy of pupil positioning in head-mounted display devices is improved. Attached Figure Description

[0013] Figure 1 This is one of the schematic diagrams of a pupil localization method provided in the embodiments of this application;

[0014] Figure 2 This is a second schematic diagram of a pupil localization method provided in an embodiment of this application;

[0015] Figure 3A This is one of the schematic diagrams of a pupil light spot provided in the embodiments of this application;

[0016] Figure 3B This is a second schematic diagram of a pupil light spot provided in an embodiment of this application;

[0017] Figure 4A This is the third schematic diagram of a pupil light spot provided in the embodiments of this application;

[0018] Figure 4B This is the fourth schematic diagram of a pupil light spot provided in the embodiments of this application;

[0019] Figure 4C This is the fifth schematic diagram of a pupil light spot provided in the embodiments of this application;

[0020] Figure 5 This is a third schematic diagram of a pupil localization method provided in an embodiment of this application;

[0021] Figure 6This is a fourth schematic diagram of a pupil localization method provided in an embodiment of this application;

[0022] Figure 7A This is one of the schematic diagrams of a pupil provided in the embodiments of this application;

[0023] Figure 7B This is a second schematic diagram of a pupil provided in an embodiment of this application;

[0024] Figure 7C This is the third schematic diagram of a pupil provided in the embodiments of this application;

[0025] Figure 8 This is the fifth schematic diagram of a pupil localization method provided in the embodiments of this application;

[0026] Figure 9 This is a schematic diagram of a pupil search area provided in an embodiment of this application;

[0027] Figure 10 This is a schematic diagram of the structure of a buffer provided in an embodiment of this application;

[0028] Figure 11 This is the seventh schematic diagram of a pupil localization method provided in the embodiments of this application;

[0029] Figure 12 This is the eighth schematic diagram of a pupil localization method provided in the embodiments of this application;

[0030] Figure 13 This is a schematic diagram of the structure of a pupil positioning device provided in an embodiment of this application;

[0031] Figure 14 This is one of the hardware structure diagrams of an electronic device provided in the embodiments of this application;

[0032] Figure 15 This is a second schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0033] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0034] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0035] The terms "at least one," "at least one of," etc., used in the specification and claims of this application refer to any one, any two, or a combination of two or more of the included items. For example, at least one of a, b, and c can mean: "a," "b," "c," "a and b," "a and c," "b and c," and "a, b, and c," where a, b, and c can be single or multiple. Similarly, "at least two" refers to two or more items, and its meaning is similar to that of "at least one."

[0036] The following provides a detailed explanation of some of the technical terms used in this application.

[0037] Virtual Reality (VR): Virtual reality technology is a computer simulation system that can create and experience virtual worlds. It uses computers to generate a simulated environment, which is an interactive three-dimensional dynamic visual and physical behavior system simulation that integrates multi-source information, immersing users in the environment.

[0038] Augmented Reality (AR): AR is a human-computer interaction technology that seamlessly overlays virtual information onto the real world. This technology enhances the user's real-world experience by capturing environmental information in real time and overlaying digital images, audio, and other sensory information onto the user's field of vision.

[0039] Eye Tracking: Eye tracking is a technique that detects and records eye characteristics to obtain the direction and movement trajectory of the human eye during viewing. Eye tracking technology typically uses eye trackers to record the position and speed of eye movement. By analyzing this eye-tracking data, researchers can obtain information about the subject's focus point, fixation duration, saccade path, and number of saccades during viewing. This data can be used to help understand the subject's level of interest in different stimuli, responses to illusions and attentional guidance, and performance in cognitive tasks. IR Camera: An IR camera is a camera that uses infrared technology to capture and detect infrared radiation emitted by objects and convert it into electrical signals to generate thermal images. IR cameras have a wide range of applications, including night vision surveillance, thermal imaging, scientific research, industrial inspection, medical diagnostics, agricultural applications, and remote sensing. These applications demonstrate the importance and versatility of IR cameras in modern technology and society.

[0040] Light spot: A light spot refers to a spot of light formed when light, moonlight or other light sources are projected onto the surface of water or the ground. In this application, the light spot refers to the light spot formed when light emitted by a light-emitting diode (LED) in a head-mounted display device is reflected on the surface of the eyeball.

[0041] Pupil contour points: Pupil contour points are specific points on the edge of the pupil, playing a crucial role in pupil localization and eye-tracking technologies. In pupil localization methods, a trained eye semantic segmentation model is used to obtain the pupil semantic contour map corresponding to the image to be detected. Then, the target contour is determined from the pupil semantic contour map. An ellipse fitting is performed on the target contour, and the coordinates of the center point of the fitted ellipse are used as the coordinates of the center point of the pupil region. In other words, the edge of the pupil region, after fitting, yields contour points on the ellipse. These contour points are essential for accurately identifying the position and shape of the pupil, thereby ensuring the accuracy and stability of gaze tracking.

[0042] The pupil localization method provided in this application will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.

[0043] The pupil localization method provided in this application can be applied to scenarios involving gaze tracking when a user wears a head-mounted display device. Alternatively, the pupil localization method provided in this application can also be applied to scenarios involving the localization and detection of object contours in a video sequence. For example: the iris of the eye, or moving spherical, circular, or elliptical objects in the video.

[0044] Currently, with the development of head-mounted display devices, these devices can track the user's eye movements to perform corresponding operations. For example, a head-mounted display device can perform page-turning operations based on the user's eye movements from left to right.

[0045] In head-mounted displays based on pupil-corneal reflection technology for gaze tracking, pupil feature extraction is a crucial step. In practical applications, differences in individual eye structure, head-mounted display design, and user wearing habits lead to variations in the pupil's state captured by the IR camera, potentially including eyelash obstruction. These factors can cause the IR camera to capture images of obscured or partially open pupils, resulting in inaccurate and unstable pupil feature extraction, severely impacting the user experience.

[0046] To avoid the above situation, there are currently many methods for pupil localization and detection in eye-tracking systems. For example, in order to accurately locate the pupil coordinates and thus determine the user's gaze information, after obtaining the face image frame corresponding to the target user, the eye region can be initially identified in the face image frame. Then, by cropping the face region in the face image frame, the eye image frame can be obtained, and pupil localization can be achieved in the image frame of the eye region. However, due to the influence of image noise and localization methods, there will still be some fluctuation in pupil localization when the eye is stationary. In high-precision eye-tracking systems, such fluctuations will lead to inaccurate gaze tracking.

[0047] In the pupil localization method provided in this application embodiment, the user's pupil motion state is comprehensively evaluated by combining the current frame, i.e., the first image frame, with the inter-frame information of the previous N image frames, i.e., the first information. Based on the pupil motion state, the user's pupil contour point is located in the first image frame. In other words, by making full use of the pupil temporal information, the gradient value of the historical pupil contour point in the current frame is statistically calculated. Based on the gradient value, the pupil motion state is comprehensively evaluated, and the user's pupil contour point can be accurately located. Thus, the accuracy of pupil localization by the head-mounted display device is improved.

[0048] The pupil positioning method provided in this application can be implemented by a pupil positioning device, which can be an electronic device or a functional module within an electronic device. The following description uses a head-mounted display device as an example to illustrate the technical solution provided in this application.

[0049] This application provides a pupil localization method. Figure 1 A flowchart of a pupil localization method provided in an embodiment of this application is shown. Figure 1As shown, the pupil localization method provided in this application embodiment may include the following steps 201 to 203.

[0050] Step 201: The head-mounted display device acquires the first information corresponding to the first N frames of the first image frame.

[0051] In this embodiment of the application, the first information includes the first gradient value of the pupil contour point of the user's eyeball in each of the first N image frames, the second gradient value of the eyelid contour point of the user in each of the first N image frames, the third gradient value of the mapping point of the pupil contour point in the first image frame in each of the first N image frames, and the fourth gradient value of the mapping point of the eyelid contour point in the first image frame in each of the first N image frames, where N∈[1,3] and N is an integer.

[0052] Optionally, in this embodiment, the head-mounted display device can be an AR device or a VR device. The specific device can be determined based on actual usage needs, and this embodiment does not impose any limitations.

[0053] It should be noted that the aforementioned first image frame is any frame in the continuous video stream sequence other than the first image frame.

[0054] Optionally, in the embodiments of this application, each image frame in the above-mentioned continuous video stream sequence includes the user's eyeball and eyelid.

[0055] Optionally, in this embodiment of the application, the above-mentioned continuous video stream sequence may be captured by an IR camera in a head-mounted display device.

[0056] Optionally, in the embodiments of this application, the aforementioned first N image frames are at least one image frame before the first image frame and at most three image frames before the first image frame.

[0057] For example, when the first image frame is the second image frame in a continuous video stream sequence, the aforementioned first N image frames are the first image frames in the continuous video stream sequence; when the first image frame is the third image frame in a continuous video stream sequence, the aforementioned first N image frames are the first image frame and the second image frame in the continuous video stream sequence; when the first image frame is the fourth image frame in a continuous video stream sequence, the aforementioned first N image frames are the first image frame, the second image frame, and the third image frame in the continuous video stream sequence.

[0058] It should be noted that the pupil contour points and eyelid contour points involved in the embodiments of this application are all a group of pixels, and this group of pixels constitutes the pupil contour points or eyelid contour points. Therefore, it can be understood that the aforementioned first gradient value, second gradient value, third gradient value, and fourth gradient value are all a set. Each set contains the gradient value of each contour point in the corresponding pupil contour points or eyelid contour points.

[0059] Optionally, in this embodiment of the application, for obtaining pupil contour points in the first frame of a continuous video stream sequence by a head-mounted display device, the head-mounted display device can input the first frame of the image frame into a first model to obtain the pupil contour points in the first frame of the image frame.

[0060] Optionally, in this embodiment, the first model can be any of the following: a neural network model or an artificial intelligence (AI) model, etc. The specific model can be determined according to actual usage requirements, and this embodiment does not impose any limitations.

[0061] For example, taking a neural network model as an example, the head-mounted display device can input the first image frame into the neural network model. The neural network model can first initially identify the eye region in the face image frame, and then obtain the eye image frame by cropping the face region in the face image frame. Then, the eye image frame is convolved to obtain the feature information of the user's pupil and the feature information of the eyelid. Then, the pupil contour point is determined by the pupil feature information, and the eyelid contour point is determined by the eyelid feature information.

[0062] Optionally, in the embodiments of this application, the specific process of acquiring pupil contour points and eyelid contour points in image frames other than the first image frame in a continuous video stream sequence by a head-mounted display device can be found in the following embodiments, and will not be repeated here to avoid repetition.

[0063] Optionally, in this embodiment of the application, the head-mounted display device may use a gradient algorithm to calculate the first gradient value of the pupil contour point in each of the first N image frames.

[0064] For example, the gradient algorithm described above can be any of the following: Batch Gradient Descent, Stochastic Gradient Descent, Momentum optimization, etc. The specific algorithm can be determined according to actual usage requirements, and this application embodiment does not impose any limitations.

[0065] It should be noted that the gradient values ​​involved in the embodiments of this application are all gradient magnitudes, and each pixel in the above group of pixels corresponds to a gradient magnitude.

[0066] For example, taking the batch gradient descent algorithm as an example, taking one image frame out of the previous N image frames as an example, after the head-mounted display device obtains the coordinate information of the pupil contour points in that image frame, it can use the Sobel operator to calculate the gradient value on the horizontal axis and the gradient value on the vertical axis of each contour point in the pupil contour points. Then, the gradient value on the horizontal axis and the gradient value on the vertical axis of each contour point are processed to obtain the gradient magnitude corresponding to each contour point.

[0067] Optionally, in this embodiment of the application, the coordinate information may include horizontal axis coordinates and vertical axis coordinates.

[0068] For example, taking one of the pupil contour points as an example, the gradient value of the contour point on the horizontal axis and the gradient value on the vertical axis can be calculated by the Sobel operator as shown in formula (1).

[0069]

[0070] Among them, grad x This represents the gradient value on the horizontal axis corresponding to a contour point. For Sobel operators, grad represents the x-axis coordinate of a contour point. y This represents the gradient value on the vertical axis corresponding to a contour point. This represents the ordinate of a contour point.

[0071] For example, taking one of the pupil contour points as an example, the gradient values ​​of the contour point on the horizontal axis and the gradient values ​​on the vertical axis are processed to obtain the gradient magnitude corresponding to the contour point as shown in formula (2).

[0072]

[0073] Where magnitude is the gradient magnitude corresponding to a contour point, grad x grad is the gradient value on the horizontal axis corresponding to a contour point. y This represents the gradient value on the vertical axis corresponding to a contour point.

[0074] It should be noted that the above explanation uses only one contour point as an example. For the pupil contour point in each of the first N image frames, the first gradient value corresponding to each of the first N image frames can be obtained through the above formulas (1) and (2). To avoid repetition, this will not be elaborated here.

[0075] Optionally, in this embodiment, the second gradient value corresponding to the eyelid contour point in each of the first N image frames can also be obtained using the above formulas (1) and (2). To avoid repetition, this will not be elaborated further here.

[0076] Optionally, in this embodiment, the third gradient value of the pupil contour point mapped to the first image frame in each of the first N image frames, and the fourth gradient value of the eyelid contour point mapped to the first image frame in each of the first N image frames, can also be obtained using the above formulas (1) and (2). To avoid repetition, this will not be elaborated further here.

[0077] Step 202: The head-mounted display device determines the motion state of the pupil in the eyeball based on the first information.

[0078] In this embodiment of the application, the above-mentioned motion state refers to the motion state of the pupil from the first N image frames to the first image frame.

[0079] Optionally, in the embodiments of this application, the above-mentioned motion state can be any of the following: a relatively static state, a first motion state, i.e. a minute motion state, or a second motion state, i.e. a violent motion state.

[0080] It should be noted that the specific process of step 202 above can be found in the following embodiments, and will not be repeated here to avoid repetition.

[0081] Step 203: The head-mounted display device locates the pupil outline point in the first image frame based on the motion state.

[0082] It should be noted that the specific process of step 203 above can be found in the following embodiments, and will not be repeated here to avoid repetition.

[0083] In the pupil localization method provided in this application embodiment, the head-mounted display device can acquire first information corresponding to the previous N image frames of the first image frame. The first information includes the first gradient value of the pupil contour point of the user's eyeball in each of the previous N image frames, the second gradient value of the eyelid contour point of the user in each of the previous N image frames, the third gradient value of the mapping point of the pupil contour point in the first image frame, and the fourth gradient value of the mapping point of the eyelid contour point in the first image frame, where N∈[1,3] and N is an integer. Then, based on the first information, the motion state of the pupil in the eyeball is determined. Finally, based on the motion state, the pupil contour point of the pupil is located in the first image frame. In this solution, the user's pupil motion state is comprehensively evaluated by combining the current frame (the first image frame) with the inter-frame information of the previous N image frames (the first information). Based on the pupil motion state, the user's pupil outline point is located in the first image frame. In other words, by making full use of the pupil temporal information, the gradient value of the historical pupil outline point in the current frame is calculated. Based on this gradient value, the pupil motion state is comprehensively evaluated, and the user's pupil outline point can be accurately located. Thus, the accuracy of pupil positioning in head-mounted display devices is improved.

[0084] Optionally, in this embodiment of the application, the first information further includes: the first position information of the light spot in the eyeball in each of the first N image frames.

[0085] For example, combined Figure 1 ,like Figure 2 As shown, step 202 above can be specifically implemented through steps 202a to 202e below.

[0086] Step 202a: The head-mounted display device acquires the second position information of the light spot in the eyeball in the first image frame.

[0087] Optionally, in the embodiments of this application, the light spot in the eyeball in each of the first N image frames and the light spot in the eyeball in the first image frame can be one or more.

[0088] In this embodiment of the application, the first position information and the second position information can both include horizontal axis coordinates and vertical axis coordinates.

[0089] Optionally, in this embodiment of the application, the number of light spots in the eyeball in each of the first N image frames is the same as the number of light spots in the eyeball in the first image frame.

[0090] Optionally, in this embodiment of the application, after the head-mounted display device obtains the light spot in the eyeball in each of the first N image frames and the light spot in the eyeball in the first image frame, it can uniquely identify the light spot in the eyeball in each of the first N image frames, with one identifier corresponding to one light spot; and it can also uniquely identify the light spot in the eyeball in the first image frame, with one identifier corresponding to one light spot, wherein the identifier of the light spot in the eyeball in the first image frame is the same as the identifier of the light spot in the eyeball in each of the first N image frames.

[0091] Optionally, in the embodiments of this application, the above-mentioned spot marking can be any of the following: a digital marking, such as an identity document (ID) marking, a special symbol marking, or a color marking.

[0092] Preferably, the above-mentioned spot identifier can be an ID identifier.

[0093] For example, suppose that each of the first N image frames and the first image frame includes 5 light spots, and the light spot labels in each of the first N image frames are 1-5, and the light spot labels in the first image frame are also 1-5.

[0094] Optionally, in this embodiment of the application, the head-mounted display device can determine the second position information of the light spot in the eyeball in the first image frame by the pixel value of the pixel point in the aforementioned local area of ​​the eye.

[0095] For example, a head-mounted display device can identify pixels in the local area of ​​the eye whose pixel value is greater than or equal to a preset pixel value as light spot pixels, and determine the second position information of the light spot based on the coordinates of the light spot pixel in the local area of ​​the eye.

[0096] For example, a head-mounted display device can identify pixels with a pixel value greater than or equal to 225 in the local area of ​​the eye as light spot pixels, and determine the second position information of the light spot based on the coordinates of the light spot pixel in the local area of ​​the eye.

[0097] It should be noted that, for the light spot in the eyeball in each of the first N image frames, the head-mounted display device can obtain the first position information of the light spot in the eyeball in each of the first N image frames through the above embodiments.

[0098] Step 202b: The head-mounted display device calculates the first value based on the first gradient value corresponding to each of the first N image frames and the third gradient value corresponding to the first image frame.

[0099] In this embodiment of the application, the first value is used to characterize the range of motion of the pupil between the first N image frames and the first image frame.

[0100] Optionally, in the embodiments of this application, step 202b can be implemented by steps 202b1 to 202b4 as described below.

[0101] Step 202b1: The head-mounted display device performs an average calculation on the first gradient value corresponding to each of the first N image frames to obtain the first average value corresponding to the first N image frames.

[0102] In this embodiment of the application, since there are multiple first gradient values ​​corresponding to each image frame in the first N image frames, the head-mounted display device can perform an average operation on the multiple first gradient values ​​in each image frame in the first N image frames to obtain a third average value corresponding to each image frame in the first N image frames; then, it continues to perform an average operation on the third average value corresponding to the first N image frames to obtain a first average value corresponding to the first N image frames, that is, the first average value is one.

[0103] For example, assuming the first N image frames are 3 image frames, the head-mounted display device performs an average operation on the multiple first gradient values ​​corresponding to each of the 3 image frames to obtain 3 third average values, and then performs an average operation on the 3 third average values ​​again to obtain the first average value.

[0104] Step 202b2: The head-mounted display device performs a mean operation on the third gradient value corresponding to the first image frame to obtain the second mean value corresponding to the first image frame.

[0105] In this embodiment of the application, since there are also multiple third gradient values ​​corresponding to the first image frame, the head-mounted display device can perform an average operation on the multiple third gradient values ​​to obtain a second average value corresponding to the first image frame, that is, the second average value is one.

[0106] Step 202b3: The head-mounted display device performs a difference operation on the first mean and the second mean to obtain the first difference value.

[0107] In this embodiment of the application, the head-mounted display device can subtract the second mean from the first mean to obtain the aforementioned first difference.

[0108] Step 202b4: The head-mounted display device performs an absolute value operation on the first difference to obtain the first value.

[0109] For example, the head-mounted display device can obtain the first value through the following formula (3), which is as follows:

[0110] Gdiff = ABS(G i -G' i (3)

[0111] Where Gdiff is the first value, ABS is the absolute value operator, and G... iG' is the first mean. i It is the second mean.

[0112] Step 202c: The head-mounted display device calculates the second value based on the second gradient value corresponding to each of the first N image frames and the fourth gradient value corresponding to the first image frame.

[0113] In this embodiment of the application, the second value is used to characterize the range of motion of the eyelid between the first N image frames and the first image frame.

[0114] Optionally, in the embodiments of this application, step 202c can be implemented by steps 202c1 to 202c4 as described below.

[0115] Step 202c1: The head-mounted display device performs an average calculation on the second gradient value corresponding to each of the first N image frames to obtain the fourth average value corresponding to the first N image frames.

[0116] In this embodiment, since there are multiple second gradient values ​​corresponding to each of the first N image frames, the head-mounted display device can perform an average operation on the multiple second gradient values ​​in each of the first N image frames to obtain a fifth average value corresponding to each of the first N image frames; then, it continues to perform an average operation on the fifth average value corresponding to the first N image frames to obtain a fourth average value corresponding to the first N image frames, i.e., the fourth average value is one.

[0117] For example, assuming the first N image frames are 3 image frames, the head-mounted display device performs an average operation on the multiple second gradient values ​​corresponding to each of the 3 image frames to obtain 3 fifth average values, and then performs an average operation on the 3 fifth average values ​​again to obtain a fourth average value.

[0118] Step 202c2: The head-mounted display device performs a mean operation on the fourth gradient value corresponding to the first image frame to obtain the sixth mean value corresponding to the first image frame.

[0119] In this embodiment of the application, since there are also multiple fourth gradient values ​​corresponding to the first image frame, the head-mounted display device can perform an average operation on the multiple fourth gradient values ​​to obtain the sixth average value corresponding to the first image frame, that is, the sixth average value is one.

[0120] Step 202c3: The head-mounted display device performs a difference operation on the fourth mean and the sixth mean to obtain the second difference value.

[0121] In this embodiment of the application, the head-mounted display device can subtract the sixth mean from the fourth mean to obtain the aforementioned second difference.

[0122] Step 202b4: The head-mounted display device performs an absolute value operation on the second difference to obtain the second value.

[0123] For example, the head-mounted display device can obtain the first value through the following formula (4), where formula (3) is specifically:

[0124] GLdiff = ABS(GL i -GL' i (4)

[0125] Where GLDiff is the second value, ABS is the absolute value operator, and GL i The fourth mean, GL' i It is the sixth mean.

[0126] Step 202d: The head-mounted display device calculates the average displacement of the light spot in the first image frame relative to the light spot in the first image frame based on the first position information corresponding to each image frame in the first N image frames and the second position information corresponding to the first image frame.

[0127] In this embodiment of the application, the head-mounted display device can subtract the second position information corresponding to the first image frame from the first position information corresponding to each of the previous N image frames to obtain the third difference value corresponding to each image frame; and perform an average operation on the third difference value corresponding to each image frame to obtain the seventh average value corresponding to each image frame, and then continue to perform an average operation on the seventh average value corresponding to each image frame to obtain the above-mentioned displacement average value.

[0128] For example, the head-mounted display device can subtract the second position information corresponding to each of the five light spots in the first image frame from the first position information corresponding to the five light spots in each of the three image frames to obtain five third differences for each image frame; then, it can perform an average operation on the five third differences for each image frame to obtain three seventh averages; finally, it can perform an average calculation on the three seventh averages to obtain the above-mentioned displacement average value, wherein the five light spots in the first image frame correspond one-to-one with the five light spots in each of the three image frames.

[0129] For example, such as Figure 3A As shown, Figure 3A The image in the middle contains six light spots from one of the first N image frames. Figure 3A Black dots 1-6 indicate, for example Figure 3B As shown, Figure 3B The image in the middle is a light spot in the first image frame.

[0130] For example, such as Figure 4A As shown, the head-mounted display device can input the aforementioned image frame into the aforementioned first model, which segments the human eye image region corresponding to the human eye region from the image frame, such as... Figure 4B As shown, the human eye image region is binarized to obtain the binarized human eye image region, as shown. Figure 4C As shown, the binarized human eye image region is convolved to obtain the light spot in the human eye, and the light spot is then labeled. Figure 4C The numbers 3, 4, 5, and 6 are used to represent them.

[0131] Step 202e: The head-mounted display device determines the motion state of the pupil in the eyeball based on the first value, the second value, and the average displacement.

[0132] It should be noted that the specific implementation process of step 202e above can be found in the following embodiments, and will not be repeated here to avoid repetition.

[0133] In this embodiment, the head-mounted display device can comprehensively determine the motion state of the pupil by combining the gradient values ​​corresponding to the pupil contour points in historical image frames and the gradient values ​​corresponding to the eyelid contour points in historical image frames with the gradient values ​​of the pupil contour points and the eyelid contour points in the current image frame, thereby improving the accuracy of the head-mounted display device in determining the motion state of the pupil.

[0134] Optionally, in the embodiments of this application, combined with Figure 2 ,like Figure 5 As shown, step 202e can be implemented through step 202e1, step 202e2, or step 202e3 as described below.

[0135] Step 202e1: When the second value is less than the first threshold and both the first value and the average displacement are less than the second threshold, the head-mounted display device determines that the motion state of the pupil is a relatively stationary state.

[0136] Optionally, in this embodiment, the first threshold and the second threshold can be user-defined or preset by the electronic device. The specific threshold can be determined according to actual usage needs, and this embodiment does not impose any limitations.

[0137] For example, the second threshold mentioned above can be 1.

[0138] Optionally, in this embodiment of the application, the first threshold can be obtained based on the pixel value of the scleral region in the user's eyeball.

[0139] For example, the head-mounted display device can obtain the pixel value corresponding to the sclera region in the user's eyeball through the first model described above, and then perform an average operation on the pixel value corresponding to the sclera region to obtain the average pixel value corresponding to the sclera region. Finally, the average pixel value is divided by a preset coefficient to obtain the first threshold. Specifically, this can be achieved through the following formula (5).

[0140] Thresh2 = (average brightness of the sclera region) / 2 (5)

[0141] Among them, Thresh2 is the first threshold, and 2 is a preset coefficient.

[0142] Exemplarily, the above-mentioned second value being less than the first threshold, and both the first value and the average displacement being less than the second threshold can be expressed as: GLDiff < Thresh2, and GDiff < 1 & diffFacuals < 1.

[0143] Among them, GLDiff is the above-mentioned second value, GDiff is the above-mentioned first value, and diffFacuals is the above-mentioned average displacement.

[0144] Step 202e2: When the second value is less than the first threshold, and the first value is greater than the third threshold and less than or equal to the fourth threshold, and the average displacement is less than the second threshold, the head-mounted display device determines that the motion state of the pupil is the first motion state.

[0145] In the embodiments of the present application, the motion amplitude of the pupil corresponding to the above-mentioned first motion state is less than or equal to the first motion amplitude threshold.

[0146] It can be understood that the above-mentioned first motion state is the above-mentioned micro motion state.

[0147] Optionally, in the embodiments of the present application, the above-mentioned third threshold and fourth threshold can be user-defined or preset by the electronic device. Specifically, it can be determined according to actual usage requirements, and the embodiments of the present application do not make restrictions.

[0148] Exemplarily, the above-mentioned third threshold can be 15.

[0149] Optionally, in the embodiments of the present application, the above-mentioned fourth threshold can be obtained according to the pixel values of the pupil region and the iris region in the user's eye.

[0150] Exemplarily, the head-mounted display device can obtain the pixel values corresponding to the pupil region, the pixel values corresponding to the iris region, and the pixel values corresponding to the sclera region in the user's eye through the above-mentioned first model, and then perform a mean operation on the pixel values corresponding to the pupil region to obtain the pixel mean corresponding to the pupil region; perform a mean operation on the pixel values corresponding to the iris region to obtain the pixel mean corresponding to the iris region. Finally, the fourth threshold is obtained according to the pixel mean corresponding to the pupil region, the pixel mean corresponding to the iris region, and the above-mentioned preset coefficient, and it can be specifically implemented through the following formula (6).

[0151] Thresh1 = (average brightness of the pupil region + average brightness of the iris region) / 2 (6)

[0152] Among them, Thresh1 is the fourth threshold, and 2 is a preset coefficient.

[0153] Exemplarily, the above-mentioned second value being less than the first threshold, and the first value being greater than the third threshold and less than or equal to the fourth threshold can be expressed as: GLDiff < Thresh2, and 15 < GDiff <= Thresh1, diffFacuals < 1.

[0154] Among them, GLDiff is the second value, Thresh2 is the first threshold, GDiff is the first value, Thresh1 is the fourth threshold, and diffFacuals is the average displacement value.

[0155] Step 202e3: When the second value is greater than the first threshold and the first value is greater than the fourth threshold, the head-mounted display device determines that the movement state of the pupil is the second movement state.

[0156] In the embodiment of the present application, the movement amplitude of the pupil corresponding to the above-mentioned second movement state is greater than the first movement amplitude threshold.

[0157] It can be understood that the above-mentioned second movement state is the above-mentioned violent movement state.

[0158] Exemplarily, the above-mentioned second value being greater than the first threshold and the first value being greater than the fourth threshold can be expressed as: GLDiff > Thresh2, and GDiff > Thresh1.

[0159] Among them, GLDiff is the second value, Thresh2 is the first threshold, GDiff is the first value, and Thresh1 is the fourth threshold.

[0160] Optionally, in the embodiment of the present application, when the second value is greater than the first threshold and the average displacement value is greater than the second threshold, the head-mounted display device may determine that the movement state of the pupil is the second movement state.

[0161] Exemplarily, the above-mentioned second value being greater than the first threshold and the average displacement value being greater than the second threshold can be expressed as: GLDiff > Thresh2, and diffFacuals > 1.

[0162] Among them, GLDiff is the second value, Thresh2 is the first threshold, and diffFacuals is the average displacement value.

[0163] In the embodiment of the present application, the head-mounted display device can judge different movement states of the pupil in the eyeball according to the feature information in the user's eyeball, that is, the gradient value of the pupil contour point, the gradient value of the eyelid contour point, and the average spot displacement value, and then can adopt different pupil positioning strategies through different movement states, which improves the flexibility of pupil positioning of the head-mounted display device.

[0164] Optionally, in the embodiments of this application, combined with Figure 5 ,like Figure 6 As shown, step 203 can be implemented through step 203a, step 203b, or step 203c as described below.

[0165] Step 203a: When the motion state is relatively stationary, the head-mounted display device uses the pupil contour point corresponding to the second image frame as the pupil contour point of the pupil in the first image frame.

[0166] In this embodiment of the application, the second image frame is the image frame preceding the first image frame.

[0167] It is understandable that when the head-mounted display device determines that the pupil's motion state in the current image frame, i.e. the first image frame, is relatively stationary, the head-mounted display device can directly reuse the pupil outline points in the previous frame of the first image frame as the pupil outline points in the first image frame.

[0168] Optionally, in this embodiment of the application, when the motion state is relatively static, the head-mounted display device uses the eyelid contour point corresponding to the second image frame as the eyelid contour point of the pupil in the first image frame.

[0169] Step 203b: When the motion state is the first motion state, the head-mounted display device predicts the pupil contour point in the first image frame based on the center position information of the pupil contour point in the previous N image frames.

[0170] It should be noted that the specific implementation process of step 203b above can be found in the following embodiments, and will not be repeated here to avoid repetition.

[0171] Optionally, in this embodiment of the application, when the motion state is the first state, the head-mounted display device uses the eyelid contour point corresponding to the second image frame as the eyelid contour point of the pupil in the first image frame.

[0172] Step 203c: When the motion state is the second motion state, the head-mounted display device locates the pupil outline point of the pupil in the first image frame based on the first image frame.

[0173] In this embodiment of the application, when the motion state is the second motion state, the head-mounted display device can input the first image frame into the first model to obtain the pupil contour points and eyelid contour points of the pupil in the first image frame.

[0174] For example, such as Figure 7AAs shown, the head-mounted display device can input the first image frame into a neural network model, and use the Convolutional Neural Network (CNN) algorithm in the neural network model to segment the human eye region from the first image frame, such as... Figure 7B As shown, the human eye region is binarized using a CNN algorithm to obtain a binarized image of the human eye region. Then, convolution is performed on the binarized image to obtain the iris region 11, sclera region 12, eyelid contour region 13, and pupil contour region 14 in the human eye. Figure 7C As shown, Figure 7C The dotted line in the diagram represents the pupil outline.

[0175] In this embodiment, the head-mounted display device can adopt different positioning strategies according to different motion states, thereby improving the flexibility of pupil positioning of the head-mounted display device.

[0176] Optionally, in the embodiments of this application, combined with Figure 6 ,like Figure 8 As shown, step 203b can be implemented through steps 301 to 303 as described below.

[0177] Step 301: When the motion state is the first motion state, the head-mounted display device calculates the motion speed of the pupil between the first N image frames and the first image frame based on the center position information of the pupil contour points in the previous N image frames.

[0178] For example, the aforementioned center position information can be the center coordinates of the pupil contour point, which includes horizontal axis coordinates and vertical axis coordinates.

[0179] Optionally, in this embodiment of the application, the aforementioned motion speed can be the average motion speed of the pupil between the first N image frames and the first image frame.

[0180] Optionally, in this embodiment of the application, taking the first N image frames as two image frames as an example, the head-mounted display device can subtract the horizontal coordinate of the center coordinate of the pupil contour point in the first image frame from the horizontal coordinate of the center coordinate of the pupil contour point in the second image frame to obtain a first center coordinate difference; then, it can subtract the vertical coordinate of the center coordinate of the pupil contour point in the first image frame from the vertical coordinate of the center coordinate of the pupil contour point in the second image frame to obtain a second center coordinate difference; finally, it can divide the first center coordinate difference from the second center coordinate difference to obtain the above-mentioned average motion speed.

[0181] Step 302: The head-mounted display device performs a division operation on the second and third values ​​to obtain the motion intensity coefficient.

[0182] In this embodiment of the application, the second value is the sum of the average values ​​of the first gradient values ​​corresponding to each of the first N image frames, and the third value is the average value of the third gradient values ​​corresponding to the first image frame.

[0183] For example, the motion intensity coefficient of a head-mounted display device can be obtained by the following formula (7), which is specifically:

[0184]

[0185] Where Coef is the exercise intensity coefficient, and pullilGrad is the exercise intensity coefficient. N It is the third value. It is the second value.

[0186] Thus, by adding the Coef coefficient, the problem of large deviation in the prediction of the eye center in the current frame due to motion inertia when the eye moves suddenly or stops can be effectively solved.

[0187] Step 303: The head-mounted display device predicts the pupil outline point in the first image frame based on the motion speed, the center position information of the pupil in the second image frame, and the motion intensity coefficient.

[0188] It should be noted that the specific process of step 303 above can be found in the above embodiments, and will not be repeated here to avoid repetition.

[0189] In this embodiment, when the motion state is the first motion state, the head-mounted display device can predict the pupil outline point in the first image frame by using the motion speed, the center position information of the pupil in the second image frame, and the motion intensity coefficient. Furthermore, the motion intensity coefficient avoids the problem of large deviation in the prediction of the eyeball center in the current frame due to motion inertia. In addition, selecting different pupil outline positioning strategies can improve the temporal stability of pupil precision positioning and improve the accuracy of pupil positioning performed by the head-mounted display device.

[0190] Optionally, in the embodiments of this application, step 303 above can be specifically implemented by steps 303a to 303d below.

[0191] Step 303a: The head-mounted display device calculates the first pupil contour point in the first image frame based on the motion speed, the center position information of the pupil contour point in the second image frame, and the motion intensity coefficient.

[0192] It can be understood that the first pupil outline point mentioned above is a set of pupil outline points.

[0193] For example, the center position information of the pupil contour point in the second image frame can be the center coordinates of the pupil contour point.

[0194] For example, the head-mounted display device can specifically obtain the first pupil contour point of the pupil in the first image frame using the following formula (8), specifically:

[0195] pupilCenter N =pupilCenter N-1 +velocity(x,y)*Coef (8)

[0196] Among them, pupilCenter N PupilCenter is the first pupil outline point. N-1 Here are the center coordinates of the pupil in the second image frame, velocity(x,y) is the motion velocity, and Coef is the motion intensity coefficient.

[0197] Step 303b: The head-mounted display device takes the center point of the first pupil contour point as the center position, expands outward by N pixels in the second image frame to obtain a wide band ring region containing the pupil contour point, and shrinks inward by N pixels to obtain a narrow band ring region containing the pupil contour point.

[0198] Optionally, in the embodiments of this application, the above-mentioned N is a preset value.

[0199] For example, N can be 5Pixel, 10Pixel, 15Pixel, 20Pixel, etc. The specific value can be determined based on actual usage requirements, and this application embodiment does not impose any limitations.

[0200] It should be noted that the above description is only for one image frame out of the previous N image frames. For each image frame out of the previous N image frames, the head-mounted display device can obtain the corresponding broadband ring region and narrowband ring region in the same way.

[0201] Step 303c: The head-mounted display device determines the pixel closest to the first pupil contour point in the first pixel group within the first annular region as the second pupil contour point in the second image frame, so as to obtain the second pupil contour point corresponding to each image frame in the first N image frames.

[0202] In this embodiment, the first annular region is the union of the narrow-band annular region and the wide-band annular region, and the first pixel group is determined based on the center position of the first pupil contour point.

[0203] It can be understood that the above-mentioned second pupil outline point is a set of pupil outline points.

[0204] For example, such as Figure 9As shown, the head-mounted display device can take the center point of the first pupil contour point as the center position, expand 5 pixels outward in the second image frame to obtain a wide band ring region containing the pupil contour point, and shrink 5 pixels inward in the second image frame to obtain a narrow band ring region containing the pupil contour point. Figure 9 The wide band ring region is represented by an expanding search area, while the narrow band ring region is represented by an inward-shrinking search area. Then, taking a contour point from the second pupil contour point as an example, the head-mounted display device can randomly mark a pixel P within the narrow band ring region, using the center point of the first pupil contour point as the origin. Figure 9 In the diagram, denoted by O, a ray is drawn in the direction corresponding to a pixel to obtain at least one pixel in both the narrow-band ring region and the wide-band ring region. Then, each of these at least one pixel is used as a contour point in the second pupil contour point by connecting it to the pixel closest to the first pupil contour point. Figure 9 The point is represented by Pi.

[0205] It should be noted that each contour point in the second pupil contour points can be obtained in the same way as described above. To avoid repetition, it will not be repeated here.

[0206] It can be understood that each pixel in the first pixel group is obtained by drawing rays from the center point of the first pupil outline point in each direction of the head-mounted display device.

[0207] Step 303d: The head-mounted display device obtains the third value corresponding to the first image frame and the fourth value corresponding to each image frame in the previous N image frames. From each image frame in the previous N image frames, the pixel point that meets the predetermined condition is determined as the pupil contour point of the pupil in the first image frame.

[0208] In this embodiment, the third value is the sum of pixel values ​​of the second pupil contour point in a preset neighborhood in the first image frame, the fourth value is the sum of pixel values ​​of the second pupil contour point in a preset neighborhood in each of the first N image frames, and the predetermined condition is that the difference between the third value and the fourth value corresponding to each of the first N image frames is the smallest.

[0209] It is understandable that a head-mounted display device can use the previous N image frames as reference frames to determine the pupil outline point in the first image frame.

[0210] Optionally, in this embodiment of the application, the aforementioned preset neighborhood refers to the preset neighborhood corresponding to each contour point in the second pupil contour points.

[0211] For example, the aforementioned preset neighborhood is a pixel region obtained by expanding X pixels outward from each contour point in the second pupil contour points.

[0212] Optionally, in this embodiment, the size of the preset neighborhood can be any of the following: 3×3, 5×5, or 7×7, etc. The specific size can be determined according to actual usage requirements, and this embodiment does not impose any limitations.

[0213] Optionally, in this embodiment of the application, the head-mounted display device can subtract the third value from the fourth value corresponding to each of the previous N image frames to determine the preset neighborhood with the smallest difference from the previous N image frames, and then determine the contour point corresponding to the preset neighborhood with the smallest difference as the pupil contour point of the pupil in the first image frame.

[0214] Optionally, in this embodiment of the application, the head-mounted display device can perform subtraction and absolute value operations on the third value and the fourth value corresponding to each of the previous N image frames, respectively, to determine the preset neighborhood with the smallest difference from the N image frames, and then determine the contour point corresponding to the preset neighborhood with the smallest difference as the pupil contour point of the pupil in the first image frame.

[0215] Optionally, in this embodiment of the application, the head-mounted display device stores the pupil contour points, spot positions, spot numbers, gradient information of pupil contour points, eyelid contour point positions, and eyelid contour point gradient information (all of the above information are hereinafter referred to as feature information) obtained in the first image frame in a cache.

[0216] Optionally, in this embodiment of the application, the buffer can store feature information corresponding to 18-36 image frames at a time.

[0217] For example, the head-mounted display device can calculate the frame rate (FrameRate) of the continuous video stream based on the image frame timestamp; and considering that the blinking process of the human eye takes 0.2-0.4s and the eyeball rotation process takes 0.1-0.4s; and that the IR camera frame rate is 90FPS, it can be estimated that the number of frames to be buffered for the above process is 18-36 frames.

[0218] For example, the above frame number 18 is obtained by multiplying 90 by the human eye blinking process of 0.2s, and the above frame number 36 is obtained by multiplying 90 by the human eye blinking process of 0.4s.

[0219] For example, such as Figure 10 As shown, the structure of the above-mentioned buffer can include 6 data segments. Figure 10 The data segments are represented by N to N-5, and each data segment stores the feature information corresponding to each image frame.

[0220] Optionally, in this embodiment of the application, the head-mounted display device can store the feature information corresponding to the first image frame in a buffer through a circular storage method.

[0221] In this embodiment, the head-mounted display device employs a pupil contour localization strategy based on a temporal reference frame. This strategy leverages the similarity of pupils in consecutive frame eye diagrams to enhance the stability of pupil center localization across frames. By selecting a reference frame in time sequence to guide subsequent frames in pupil contour localization, the accuracy of pupil localization by the head-mounted display device is improved.

[0222] The above-described method embodiments, or various possible implementations of the method embodiments, can be executed individually, or, provided there are no contradictions, they can be combined with each other. The specific implementation can be determined according to actual usage requirements, and this application embodiment does not impose any restrictions on this.

[0223] For example, such as Figure 11 As shown below, the pupil localization method provided in this application embodiment will be explained in detail through specific examples. Specifically, it can be implemented through the following steps 1 to 17.

[0224] Step 1: The head-mounted display device obtains the pupil contour points C1 / C2 / C3 corresponding to each of the first 3 image frames from the pupil tracking buff.

[0225] It should be noted that when the cached data is less than 3, the actual number of frames is used.

[0226] Step 2: The head-mounted display device acquires the gradient values ​​G1 / G2 / G3 of C1 / C2 / C3 in the corresponding image frame.

[0227] Step 3: The head-mounted display device calculates the gradient values ​​G1' / G2' / G3' of C1 / C2 / C3 in the current frame.

[0228] Step 4: The head-mounted display device calculates the absolute value of the corresponding gradient difference Di = ABS(Gi-Gi) and its mean value Davg.

[0229] Step 5: The head-mounted display device obtains the eyelid contour points L1 / L2 / L3 from the first 3 image frames in the pupil tracking buf.

[0230] It should be noted that when the cached data is less than 3, the actual number of frames will be used.

[0231] Step 6: The head-mounted display device obtains the gradient values ​​Gl1 / Gl2 / Gl3 corresponding to the eyelid contour points in the corresponding image frames for L1 / L2 / L3.

[0232] Step 7: The head-mounted display device calculates the gradient values ​​Gl1' / Gl2' / Gl3' of L1 / L2 / L3 corresponding to the eyelid contour points in the current frame.

[0233] Step 8: The head-mounted display device calculates the absolute value of the corresponding gradient difference Dli = ABS(Gli - Gli), and calculates its average value Dlavg.

[0234] Step 9: The head-mounted display device determines whether Dlavg is greater than Thresh1.

[0235] In the embodiment of the present application, when Dlavg < Thresh1, step 10 is executed; when Dlavg > Thresh1, step 16 is executed.

[0236] Step 10: The head-mounted display device determines whether Davg is greater than Thresh1.

[0237] In the embodiment of the present application, when Davg < Thresh1, step 11 is executed; when Davg > Thresh1, step 13 is executed.

[0238] Step 11: The head-mounted display device determines that the current frame is stationary relative to the historical frame.

[0239] Step 12: The head-mounted display device reuses the pupil features of the previous frame and reuses the eyelid contour points of the previous frame.

[0240] Step 13: The head-mounted display device determines whether Davg is greater than Thresh2.

[0241] In the embodiment of the present application, when Davg < Thresh2, step 14 is executed; when Davg > Thresh2, step 16 is executed.

[0242] Step 14: The head-mounted display device determines that a slight movement has occurred in the current frame relative to the historical frame.

[0243] Step 15: The head-mounted display device performs precise pupil edge localization and reuses the eyelid contour points of the previous frame.

[0244] Step 16: The head-mounted display device determines that a large movement has occurred in the current frame relative to the historical frame.

[0245] Step 17: The head-mounted display device performs precise AI segmentation of the pupil.

[0246] In the embodiment of the present application, through the current frame, that is, the above-mentioned first image frame, combined with the inter-frame information of the previous N image frames, that is, the above-mentioned first information, the movement state of the user's pupil is comprehensively evaluated, so as to locate the pupil contour points of the user in the first image frame according to the movement state of the pupil. That is to say, by making full use of the pupil time-domain information, the gradient values of the historical pupil contour points in the current frame are statistically obtained, and then the movement state of the pupil is comprehensively evaluated according to the gradient values, and then the pupil contour points of the user can be accurately located. In this way, the accuracy of the head-mounted display device in locating the pupil is improved.

[0247] For example, such as Figure 12 As shown below, a specific example is used to explain how, in the pupil localization method provided in this application, when the motion state is the first motion state, the head-mounted display device predicts the pupil contour point in the first image frame based on the center position information of the pupil contour point in the previous N image frames. This can be specifically implemented through the following steps 20 to 23.

[0248] Step 20: The head-mounted display device calculates the pupil motion velocity Vp based on the pupil center coordinates of the most recent 3 frames.

[0249] Step 21: The head-mounted display device calculates the motion intensity coefficient (Coef) of the current frame relative to the previous frame.

[0250] Step 22: The head-mounted display device predicts the pupil position in the current frame based on the pupil center point coordinates of the previous frame and the velocity Vp.

[0251] Step 23: The head-mounted display device obtains the pupil ROI eye map based on the pupil center coordinates.

[0252] In this embodiment of the application, the above-mentioned ROI eye map is the pupil contour point in the first image frame.

[0253] In this embodiment, the head-mounted display device employs a pupil contour localization strategy based on a temporal reference frame. This strategy leverages the similarity of pupils in consecutive frame eye diagrams to enhance the stability of pupil center localization across frames. By selecting a reference frame in time sequence to guide subsequent frames in pupil contour localization, the accuracy of pupil localization by the head-mounted display device is improved.

[0254] It should be noted that the pupil positioning method provided in this application can be executed by a pupil positioning device. This application uses a pupil positioning device executing the pupil positioning method as an example to illustrate the pupil positioning device provided in this application.

[0255] Figure 13 A schematic diagram of a possible structure of the pupil positioning device involved in an embodiment of this application is shown. For example... Figure 13 As shown, the pupil positioning device 70 may include: an acquisition module 71, a determination module 72, and a positioning module 73.

[0256] The acquisition module 71 is used to acquire first information corresponding to the first N image frames of the first image frame. This first information includes the first gradient value of the pupil contour point of the user's eyeball in each of the first N image frames, the second gradient value of the eyelid contour point of the user in each of the first N image frames, the third gradient value of the mapping point of the pupil contour point in each of the first image frames, and the fourth gradient value of the mapping point of the eyelid contour point in each of the first image frames, where N∈[1,3] and N is an integer. The determination module 72 is used to determine the motion state of the pupil in the eyeball based on the first information acquired by the acquisition module 71. The positioning module 73 is used to locate the pupil contour point in the first image frame based on the motion state determined by the determination module 72.

[0257] In one possible implementation, the first information further includes: first position information of the light spot in the eyeball in each of the first N image frames; the pupil positioning device 70 further includes: a calculation module. The acquisition module 71 is also used to acquire second position information of the light spot in the eyeball in the first image frame. The calculation module is used to calculate a first value based on the first gradient value corresponding to each of the first N image frames and the third gradient value corresponding to the first image frame, the first value being used to characterize the movement amplitude of the pupil between the first N image frames and the first image frame; and to calculate a second value based on the second gradient value corresponding to each of the first N image frames and the fourth gradient value corresponding to the first image frame, the second value being used to characterize the movement amplitude of the eyelid between the first N image frames and the first image frame; and to calculate the average displacement of the light spot in the first N image frames and the light spot in the first image frame based on the first position information corresponding to each image frame and the second position information corresponding to the first image frame. The aforementioned determining module 72 is specifically used to determine the motion state of the pupil in the eyeball based on the first value, the second value, and the average displacement obtained by the calculation module.

[0258] In one possible implementation, the determining module 72 is specifically used to determine the pupil's motion state as a relatively stationary state when the second value is less than the first threshold and both the first value and the average displacement are less than the second threshold; or, when the second value is less than the first threshold, and the first value is greater than the third threshold and less than or equal to the fourth threshold, and the average displacement is less than or equal to the second threshold, the pupil's motion state is determined to be a first motion state, and the pupil's motion amplitude corresponding to the first motion state is less than or equal to the first motion amplitude threshold; or, when the second value is greater than the first threshold and the first value is greater than the fourth threshold, the pupil's motion state is determined to be a second motion state, and the pupil's motion amplitude corresponding to the second motion state is greater than the first motion amplitude threshold.

[0259] In one possible implementation, the positioning module 73 is specifically used to, when the motion state is relatively static, take the pupil contour point corresponding to the second image frame as the pupil contour point of the pupil in the first image frame, where the second image frame is the previous image frame of the first image frame; or, when the motion state is a first motion state, predict the pupil contour point of the pupil in the first image frame based on the center position information of the pupil contour points in the previous N image frames; or, when the motion state is a second motion state, locate the pupil contour point of the pupil in the first image frame based on the first image frame.

[0260] In one possible implementation, the aforementioned calculation module is specifically used to calculate the pupil's motion velocity between the first image frame and the previous N image frames, based on the center position information of the pupil contour points in the previous N image frames, when the motion state is the first motion state; and to perform a division operation on the second and third values ​​to obtain the motion intensity coefficient, where the second value is the sum of the average values ​​of the first gradient values ​​corresponding to each of the previous N image frames, and the third value is the average value of the third gradient values ​​corresponding to the first image frame. The aforementioned positioning module 73 is specifically used to predict the pupil contour points in the first image frame based on the motion velocity obtained by the calculation module, the center position information of the pupil in the second image frame, and the motion intensity coefficient.

[0261] In one possible implementation, the pupil positioning device 70 provided in this application embodiment further includes: a processing module. The calculation module is specifically used to calculate the first pupil contour point of the pupil in the first image frame based on the motion speed, the center position information of the pupil contour point in the second image frame, and the motion intensity coefficient. The processing module is used to expand outward by N pixels in the second image frame to obtain a wideband ring region containing the pupil contour point, and shrink inward by N pixels to obtain a narrowband ring region containing the pupil contour point, with the center point of the first pupil contour point as the center position. The determination module 72 is specifically used to determine the pixel closest to the first pupil contour point in the first pixel group within the first ring region as the second pupil contour point of the pupil in the second image frame, so as to obtain the second pupil contour point corresponding to each image frame in the first N image frames. The first ring region is the union region of the narrowband ring region and the wideband ring region, and the first pixel group is determined based on the center position of the first pupil contour point. The aforementioned positioning module 73 is specifically used to obtain the third value corresponding to the first image frame and the fourth value corresponding to each image frame in the previous N image frames, and to determine the pixel points that meet the predetermined conditions from each image frame in the previous N image frames as the pupil contour points of the pupil in the first image frame; wherein, the third value is the sum of the pixel values ​​of the second pupil contour points in the preset neighborhood of the first image frame, and the fourth value is the sum of the pixel values ​​of the second pupil contour points in the preset neighborhood of each image frame, and the predetermined condition is that the difference between the third value and the fourth value corresponding to each image frame is the smallest.

[0262] This application provides a pupil positioning device that comprehensively evaluates the user's pupil movement state by combining the current frame (the first image frame) with the inter-frame information of the previous N image frames (the first information). Based on the pupil movement state, the device locates the user's pupil outline point in the first image frame. In other words, it fully utilizes the pupil temporal information to statistically calculate the gradient value of historical pupil outline points in the current frame. Based on this gradient value, the device comprehensively evaluates the pupil movement state and can accurately locate the user's pupil outline point. This improves the accuracy of the pupil positioning device in locating the pupil.

[0263] The pupil positioning device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, a mobile electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television set (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the device.

[0264] The pupil positioning device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit it.

[0265] The pupil positioning device provided in this application embodiment can realize the various processes implemented in the above embodiments. To avoid repetition, it will not be described again here.

[0266] Optionally, such as Figure 14As shown, this application embodiment also provides an electronic device 90, including a processor 91 and a memory 92. The memory 92 stores a program or instructions that can run on the processor 91. When the program or instructions are executed by the processor 91, they implement the various steps of the above-described pupil positioning method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.

[0267] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.

[0268] Figure 15 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application.

[0269] The electronic device 100 includes, but is not limited to, components such as: radio frequency unit 101, network module 102, audio output unit 103, input unit 104, sensor 105, display unit 106, user input unit 107, interface unit 108, memory 109, and processor 110.

[0270] Those skilled in the art will understand that the electronic device 100 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 110 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 15 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0271] The processor 110 is configured to acquire first information corresponding to the first N image frames of the first image frame. The first information includes the first gradient value of the pupil contour point of the user's eyeball in each of the first N image frames, the second gradient value of the eyelid contour point of the user in each of the first N image frames, the third gradient value of the mapping point of the pupil contour point in each of the first image frames, and the fourth gradient value of the mapping point of the eyelid contour point in each of the first image frames, where N∈[1,3] and N is an integer; and based on the first information, determine the motion state of the pupil in the eyeball; and based on the motion state, locate the pupil contour point of the pupil in the first image frame.

[0272] Optionally, in this embodiment, the first information further includes: the first position information of the light spot in the eyeball in each of the first N image frames. The processor 110 is specifically configured to: obtain the second position information of the light spot in the eyeball in the first image frame; calculate a first value based on the first gradient value and the third gradient value corresponding to each of the first N image frames, the first value representing the amplitude of pupil movement between the first and second image frames; calculate a second value based on the second gradient value and the fourth gradient value corresponding to each of the first and second image frames, the second value representing the amplitude of eyelid movement between the first and second image frames; calculate the average displacement between the light spot in the first and second image frames and the light spot in the first image frame based on the first position information and the second position information corresponding to each of the first and second image frames; and determine the motion state of the pupil in the eyeball based on the first value, the second value, and the average displacement.

[0273] Optionally, in this embodiment of the application, the processor 110 is specifically configured to determine the pupil's motion state as a relatively stationary state when the second value is less than the first threshold and both the first value and the average displacement are less than the second threshold; or, when the second value is less than the first threshold, and the first value is greater than the third threshold and less than or equal to the fourth threshold, and the average displacement is less than the second threshold, determine the pupil's motion state as a first motion state, wherein the pupil's motion amplitude corresponding to the first motion state is less than or equal to the first motion amplitude threshold; or, when the second value is greater than the first threshold and the first value is greater than the fourth threshold, determine the pupil's motion state as a second motion state, wherein the pupil's motion amplitude corresponding to the second motion state is greater than the first motion amplitude threshold.

[0274] Optionally, in this embodiment of the application, the processor 110 is specifically configured to, when the motion state is a relatively static state, use the pupil contour point corresponding to the second image frame as the pupil contour point of the pupil in the first image frame, wherein the second image frame is the previous image frame of the first image frame; or, when the motion state is a first motion state, predict the pupil contour point of the pupil in the first image frame based on the center position information of the pupil contour points in the previous N image frames; or, when the motion state is a second motion state, locate the pupil contour point of the pupil in the first image frame based on the first image frame.

[0275] Optionally, in this embodiment of the application, the processor 110 is specifically configured to, when the motion state is a first motion state, calculate the motion velocity of the pupil between the first N image frames and the first image frame based on the center position information of the pupil contour points in the first N image frames; perform a division operation on the second value and the third value to obtain a motion intensity coefficient, wherein the second value is the sum of the average values ​​of the first gradient values ​​corresponding to each image frame in the first N image frames, and the third value is the average value of the third gradient values ​​corresponding to the first image frame; and predict the pupil contour points of the pupil in the first image frame based on the motion velocity, the center position information of the pupil in the second image frame, and the motion intensity coefficient.

[0276] Optionally, in this embodiment, the processor 110 is specifically configured to calculate the first pupil contour point of the pupil in the first image frame based on the motion speed, the center position information of the pupil contour point in the second image frame, and the motion intensity coefficient; using the center point of the first pupil contour point as the center position, expand outward by N pixels in the second image frame to obtain a wide band ring region containing the pupil contour point, and shrink inward by N pixels to obtain a narrow band ring region containing the pupil contour point; determine the pixel point in the first pixel group within the first ring region that is closest to the first pupil contour point as the second pupil contour point of the pupil in the second image frame, so as to obtain the second pupil contour point corresponding to each image frame in the first N image frames. The first annular region is the union of the narrow-band annular region and the wide-band annular region. The first pixel group is determined based on the center position of the first pupil contour point. The third value corresponding to the first image frame and the fourth value corresponding to each of the previous N image frames are obtained. From each of the previous N image frames, pixels that meet predetermined conditions are determined as the pupil contour points of the pupil in the first image frame. The third value is the sum of the pixel values ​​of the second pupil contour point in the preset neighborhood of the first image frame, and the fourth value is the sum of the pixel values ​​of the second pupil contour point in the preset neighborhood of each image frame. The predetermined condition is that the difference between the third value and the fourth value corresponding to each image frame is the smallest.

[0277] This application provides an electronic device that comprehensively evaluates the motion state of a user's pupil by combining the current frame (the first image frame) with the inter-frame information of the previous N image frames (the first information). Based on the pupil's motion state, the device locates the user's pupil outline point in the first image frame. In other words, by fully utilizing the pupil's temporal information, the device statistically calculates the gradient value of historical pupil outline points in the current frame, and comprehensively evaluates the pupil's motion state based on this gradient value. This allows for accurate location of the user's pupil outline point, improving the accuracy of pupil positioning in head-mounted display devices.

[0278] The electronic device provided in this application embodiment can implement the various processes implemented in the above method embodiments and achieve the same technical effect. To avoid repetition, it will not be described again here.

[0279] For details on the beneficial effects of the various implementation methods in this embodiment, please refer to the beneficial effects of the corresponding implementation methods in the above method embodiments. To avoid repetition, these will not be repeated here.

[0280] It should be understood that, in this embodiment, the input unit 104 may include a graphics processing unit (GPU) 1041 and a microphone 1042. The GPU 1041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 106 may include a display panel 1061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 107 includes at least one of a touch panel 1071 and other input devices 1072. The touch panel 1071 is also called a touch screen. The touch panel 1071 may include a touch detection device and a touch controller. Other input devices 1072 may include, but are not limited to, a physical keyboard, function keys (such as volume control buttons, power buttons, etc.), a trackball, a mouse, and a joystick, which will not be described in detail here.

[0281] The memory 109 can be used to store software programs and various data. The memory 109 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 109 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 109 in the embodiments of this application includes, but is not limited to, these and any other suitable types of memory.

[0282] Processor 110 may include one or more processing units; optionally, processor 110 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 110.

[0283] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0284] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0285] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above method embodiments and achieve the same technical effect. To avoid repetition, it will not be described again here.

[0286] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0287] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the pupil positioning method embodiment described above, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0288] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0289] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0290] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A pupil localization method, characterized in that, The method includes: Obtain first information corresponding to the first N frames of the first image frame. The first information includes the first gradient value of the pupil contour point of the user's eyeball in each of the first N frames, the second gradient value of the eyelid contour point of the user in each of the first N frames, the third gradient value of the mapping point of the pupil contour point in each of the first N frames, and the fourth gradient value of the mapping point of the eyelid contour point in each of the first N frames, where N∈[1,3] and N is an integer. Based on the first information, the motion state of the pupil in the eyeball is determined; Based on the motion state, the pupil outline point is located in the first image frame.

2. The method according to claim 1, characterized in that, The first information also includes: the first position information of the light spot in the eyeball in each of the first N image frames; Determining the motion state of the pupil in the eyeball based on the first information includes: Obtain the second position information of the light spot in the eyeball in the first image frame; Based on the first gradient value corresponding to each of the first N image frames and the third gradient value corresponding to the first image frame, a first value is calculated. The first value is used to characterize the movement amplitude of the pupil between the first N image frames and the first image frame. Based on the second gradient value corresponding to each of the first N image frames and the fourth gradient value corresponding to the first image frame, a second value is calculated. The second value is used to characterize the movement amplitude of the eyelid between the first N image frames and the first image frame. Based on the first position information corresponding to each image frame and the second position information corresponding to the first image frame, the average displacement of the light spot in the first N image frames and the light spot in the first image frame is calculated. Based on the first value, the second value, and the average displacement, the motion state of the pupil in the eyeball is determined.

3. The method according to claim 2, characterized in that, Determining the motion state of the pupil in the eyeball based on the first value, the second value, and the average displacement includes: If the second value is less than the first threshold, and both the first value and the average displacement are less than the second threshold, the motion state of the pupil is determined to be a relatively stationary state; or, If the second value is less than the first threshold, and the first value is greater than the third threshold and less than or equal to the fourth threshold, and the average displacement is less than the second threshold, then the pupil's motion state is determined to be a first motion state, and the pupil's motion amplitude corresponding to the first motion state is less than or equal to a first motion amplitude threshold; or... If the second value is greater than the first threshold and the first value is greater than the fourth threshold, the movement state of the pupil is determined to be the second movement state, and the movement amplitude of the pupil corresponding to the second movement state is greater than the first movement amplitude threshold.

4. The method according to claim 1, characterized in that, The step of locating the pupil contour point in the first image frame based on the motion state includes: When the motion state is the relatively static state, the pupil contour point corresponding to the second image frame is used as the pupil contour point of the pupil in the first image frame, and the second image frame is the image frame preceding the first image frame; or... When the motion state is the first motion state, based on the center position information of the pupil contour points in the previous N image frames, predict the pupil contour points in the first image frame; or, When the motion state is the second motion state, the pupil outline point of the pupil in the first image frame is located based on the first image frame.

5. The method according to claim 4, characterized in that, When the motion state is the first motion state, predicting the pupil contour point in the first image frame based on the center position information of the pupil contour points in the previous N image frames includes: When the motion state is the first motion state, the motion speed of the pupil between the first N image frames and the first image frame is calculated based on the center position information of the pupil contour points in the previous N image frames. The motion intensity coefficient is obtained by dividing the second value and the third value. The second value is the sum of the average values ​​of the first gradient values ​​corresponding to each of the first N image frames, and the third value is the average value of the third gradient values ​​corresponding to the first image frame. Based on the motion speed, the center position information of the pupil in the second image frame, and the motion intensity coefficient, the pupil contour point in the first image frame is predicted.

6. The method according to claim 5, characterized in that, The step of predicting the pupil contour points in the first image frame based on the motion speed, the center position information of the pupil in the second image frame, and the motion intensity coefficient includes: Based on the motion speed, the center position information of the pupil contour point in the second image frame, and the motion intensity coefficient, the first pupil contour point of the pupil in the first image frame is calculated. With the center point of the first pupil contour point as the center position, expand outward by N pixels in the second image frame to obtain a wideband ring region containing the pupil contour point, and shrink inward by N pixels to obtain a narrowband ring region containing the pupil contour point. The pixel closest to the first pupil contour point in the first pixel group within the first annular region is determined as the second pupil contour point of the pupil in the second frame image frame, so as to obtain the second pupil contour point corresponding to each frame image frame in the first N frames image frames. The first annular region is the union region of the narrow band annular region and the wide band annular region. The first pixel group is determined based on the center position of the first pupil contour point. Obtain the third value corresponding to the first image frame and the fourth value corresponding to each image frame in the preceding N image frames. From each image frame in the preceding N image frames, determine the pixel points that meet the predetermined conditions as the pupil contour points of the pupil in the first image frame. Wherein, the third value is the sum of the pixel values ​​of the second pupil contour point in a preset neighborhood in the first image frame, the fourth value is the sum of the pixel values ​​of the second pupil contour point in a preset neighborhood in each image frame, and the predetermined condition is that the difference between the third value and the fourth value corresponding to each image frame is the smallest.

7. A pupil positioning device, characterized in that, The pupil positioning device includes: an acquisition module, a determination module, and a positioning module; The acquisition module is used to acquire first information corresponding to the first N image frames of the first image frame. The first information includes the first gradient value of the pupil contour point of the user's eyeball in each of the first N image frames, the second gradient value of the eyelid contour point of the user in each of the first N image frames, the third gradient value of the mapping point of the pupil contour point in each of the first image frames in the first image frame, and the fourth gradient value of the mapping point of the eyelid contour point in each of the first image frames in the first image frame, where N∈[1,3] and N is an integer. The determining module is used to determine the motion state of the pupil in the eyeball based on the first information obtained by the acquiring module. The positioning module is used to locate the pupil outline point in the first image frame based on the motion state determined by the determining module.

8. The apparatus according to claim 7, characterized in that, The first information further includes: the first position information of the light spot in the eyeball in each of the first N image frames; the pupil positioning device further includes: a calculation module; The acquisition module is further configured to acquire second position information of the light spot in the eyeball in the first image frame; The calculation module is used to calculate a first value based on the first gradient value corresponding to each of the first N image frames and the third gradient value corresponding to the first image frame. The first value is used to characterize the movement amplitude of the pupil between the first N image frames and the first image frame. Based on the second gradient value corresponding to each of the first N image frames and the fourth gradient value corresponding to the first image frame, a second value is calculated. The second value is used to characterize the movement amplitude of the eyelid between the first N image frames and the first image frame. And based on the first position information corresponding to each image frame and the second position information corresponding to the first image frame, the average displacement of the light spot in the first N image frames and the light spot in the first image frame is calculated. The determining module is specifically used to determine the motion state of the pupil in the eyeball based on the first value, the second value, and the average displacement obtained by the calculation module.

9. The apparatus according to claim 8, characterized in that, The determining module is specifically used to determine that the motion state of the pupil is a relatively stationary state when the second value is less than the first threshold and both the first value and the average displacement are less than the second threshold; or, If the second value is less than the first threshold, and the first value is greater than the third threshold and less than or equal to the fourth threshold, and the average displacement is less than or equal to the second threshold, the motion state of the pupil is determined to be the first motion state, and the motion amplitude of the pupil corresponding to the first motion state is less than or equal to the first motion amplitude threshold. or, If the second value is greater than the first threshold and the first value is greater than the fourth threshold, the movement state of the pupil is determined to be the second movement state, and the movement amplitude of the pupil corresponding to the second movement state is greater than the first movement amplitude threshold.

10. A head-mounted display device, characterized in that, Includes the pupil positioning device as described in any one of claims 7 to 9.