Gaze point rendering method, near-to-eye display device and storage medium
By dividing the human eye image into grid cells in a near-eye display device, determining the position information of the gaze point, and rendering it to the user interface, the problem of gaze point jitter is solved, and the user's visual experience is improved.
Patent Information
- Application Number
- CN202510878805.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-11-18
AI Technical Summary
In existing near-eye display devices, the jitter of the gaze point is relatively large, resulting in a poor visual experience for users.
The human eye image is divided into multiple grid cells according to a preset number of rows and columns. The position information of the human eye's gaze point is determined based on the target grid cell and rendered to the user interface, thereby reducing the amplitude and frequency of gaze point jitter.
By dividing the grid cells and determining their location information, the amplitude and frequency of jitter at the gaze point are reduced, thus improving the user's visual experience.
Smart Images

Figure CN120976384A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of eye-tracking technology, and more particularly to a gaze point rendering method, a near-eye display device, and a storage medium. Background Technology
[0002] Near-eye display devices (NIVs) have eye-tracking capabilities. When using eye-tracking, the user's gaze point needs to be rendered in real-time onto the user interface displayed on the NIV. Currently, gaze point rendering in eye tracking primarily involves directly calculating the pixel coordinates of the user's gaze point in each frame of the eye image, and then rendering the gaze point onto the NIV's user interface in real-time based on these coordinates. However, the displayed gaze point can exhibit significant jitter due to NIV device fluctuations, lighting changes, or algorithm noise, resulting in a poor user visual experience. Therefore, reducing the jitter of the gaze point to improve the user's visual experience is a pressing issue that needs to be addressed. Summary of the Invention
[0003] This invention provides a gaze point rendering method, a near-eye display device, and a storage medium, aiming to reduce the jitter amplitude of the gaze point and improve the user's visual experience.
[0004] In a first aspect, embodiments of the present invention provide a gaze-point rendering method applied to a near-eye display device, the method comprising:
[0005] The near-eye display device acquires an image of the wearer's eye.
[0006] The human eye image is divided into multiple grid units according to a preset number of rows and columns;
[0007] Based on the plurality of grid cells, the target grid cell that the wearer's eye gazes at is determined;
[0008] Based on the target grid cell, the position information of the human eye gaze point is determined, and based on the position information of the human eye gaze point, the human eye gaze point is rendered and displayed in the user interface currently displayed by the near-eye display device.
[0009] In a second aspect, embodiments of the present invention also provide a near-eye display device, the near-eye display device including a processor, a memory, a computer program stored in the memory and executable by the processor, and a data bus for implementing communication between the processor and the memory, wherein when the computer program is executed by the processor, it implements the foveation rendering method as described in the first aspect.
[0010] Thirdly, embodiments of the present invention also provide a storage medium for computer-readable storage, wherein the storage medium stores one or more programs that can be executed by one or more processors to implement the foveated rendering method as described in the first aspect.
[0011] This invention provides a gaze point rendering method, a near-eye display device, and a storage medium. In this invention, the grid cells are used to divide the human eye image according to a preset number of rows and columns. Therefore, each grid cell includes multiple pixels. Thus, based on the multiple grid cells obtained, the range of the target grid cell that the human eye gazes at is wider than that of a single pixel. Even if the human eye gaze point drifts within the same target grid cell, the position of the human eye gaze point rendered using the position information of the human eye gaze point determined based on the target grid cell can remain unchanged in the user interface. This reduces the jitter amplitude of the gaze point and also reduces the frequency of gaze point jitter, improving the user's visual experience. Attached Figure Description
[0012] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 This is a schematic diagram of a scene implementing the foveated rendering method provided in the embodiments of the present invention;
[0014] Figure 2 This is a flowchart illustrating a foveated rendering method provided in an embodiment of the present invention;
[0015] Figure 3 This is a flowchart illustrating another foveated rendering method provided in an embodiment of the present invention;
[0016] Figure 4 This is a flowchart illustrating another foveated rendering method provided in an embodiment of the present invention;
[0017] Figure 5 This is a flowchart illustrating another foveated rendering method provided in an embodiment of the present invention;
[0018] Figure 6 This is a flowchart illustrating another foveated rendering method provided in an embodiment of the present invention;
[0019] Figure 7 This is a schematic block diagram of a near-eye display device provided in an embodiment of the present invention. Detailed Implementation
[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0021] The flowchart shown in the attached diagram is for illustrative purposes only and does not necessarily include all content and operations / steps, nor does it necessarily have to be performed in the order described. For example, some operations / steps can be broken down, combined, or partially merged, so the actual execution order may change depending on the actual situation.
[0022] It should be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0023] Currently, the main method for gazing point rendering in eye tracking is to directly calculate the pixel coordinates of the human eye's gaze point in each frame of the human eye image, and then render the gaze point in real time onto the user interface displayed on the near-eye display device based on these pixel coordinates. However, the pixel coordinates of the gaze point can fluctuate significantly due to the jitter of the near-eye display device, changes in lighting, or algorithm noise, resulting in noticeable jitter in the displayed gaze point and a poor visual experience for the user.
[0024] To address the aforementioned problems, embodiments of the present invention provide a gaze point rendering method, a near-eye display device, and a storage medium. In these embodiments, the grid cells are used to divide the human eye image according to a preset number of rows and columns. Therefore, each grid cell includes multiple pixels. This means that the range of the target grid cell determined by the human eye's gaze point is wider than that of a single pixel. Even if the human eye's gaze point drifts within the same target grid cell, the position of the human eye's gaze point rendered using the position information determined based on the target grid cell remains unchanged in the user interface. This reduces the amplitude and frequency of gaze point jitter, improving the user's visual experience.
[0025] The following detailed description of some embodiments of the present invention is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0026] Please see Figure 1 , Figure 1 This is a schematic diagram of a scene implementing the foveated rendering method provided in the embodiments of the present invention.
[0027] like Figure 1 As shown, the near-eye display device 100 is worn on the user's head 11. The near-eye display device 100 includes a waveguide lens 101, a gaze point detection device, and an optical engine. Figure 1 (Not shown). The optical engine (microdisplay) is used to guide the projection of images, videos, pages, or user interfaces onto the optical waveguide lens, so that the optical waveguide lens 101 directs the images, videos, pages, or user interfaces to the user's eyes. The gaze point detection device is used to detect the position information of the user's gaze point.
[0028] In some embodiments, the gaze detection device includes a detection light source and an imaging unit. For example, the near-eye display device 100 can use the gaze detection device to determine the position of the human eye's gaze point on the display screen 101 using eye tracking. For instance, when light emitted from the detection light source shines on the human eye, the imaging unit can capture an image of the human eye's response to the current light, obtaining an eye image. Then, the eye image is divided into multiple grid cells according to a preset number of rows and columns. Based on the multiple grid cells, the target grid cell that the wearer's gaze point is focused on is determined. Based on the target grid cell, the position information of the human eye's gaze point is determined. Based on the position information of the human eye's gaze point, the human eye's gaze point is rendered and displayed on the currently displayed user interface of the near-eye display device 100's display screen 101.
[0029] In some implementations, the imaging unit may include a lens for adjusting the field of view and converging light, and an image sensor for sensing light. For example, the image sensor may be a complementary metal-oxide-semiconductor (CMOS) based image sensor or a charge-coupled device (CCD) based image sensor. It should be understood that the above examples illustrate several specific implementations of the imaging unit's functions, and in other implementations of the present invention, the imaging unit may also achieve its functions through other components, which will not be elaborated here.
[0030] In some embodiments, the near-eye display device 100 includes a sensor for acquiring attitude data of the near-eye display device 100. Figure 1(Not shown). For example, the sensor that acquires the attitude data of the near-eye display device 100 may include an inertial measurement unit (IMU), which may include an accelerometer, a gyroscope, and / or a magnetometer. For example, the inertial measurement unit includes a three-axis accelerometer, a three-axis gyroscope, and a three-axis magnetometer. The near-eye display device 100 may include augmented reality (AR) glasses, AR helmets, mixed reality (MR) glasses, and MR helmets, etc.
[0031] The following will combine Figure 1 The following scenario provides a detailed description of the foveated rendering method provided by embodiments of the present invention. It should be noted that... Figure 1 The scenarios described are only used to explain the foveated rendering method provided in the embodiments of the present invention, but do not constitute a limitation on the application scenarios of the foveated rendering method provided in the embodiments of the present invention.
[0032] Please see Figure 2 , Figure 2 This is a flowchart illustrating a foveated rendering method provided in an embodiment of the present invention.
[0033] like Figure 2 As shown, the foveated rendering method includes steps S101 to S104.
[0034] Step S101: Obtain the human eye image obtained by the near-eye display device from the wearer's human eye.
[0035] In this embodiment, the eye image can be a current eye image obtained by the near-eye display device capturing the wearer's eye at the current moment. Specifically, this eye image is obtained by capturing the wearer's eye after the near-eye display device has enabled eye-tracking functionality.
[0036] Step S102: Divide the human eye image into multiple grid units according to the preset number of rows and columns.
[0037] In this embodiment, the preset number of rows and columns can be set based on different application scenarios, and this embodiment does not impose specific limitations on them. For example, the current eye image of the wearer is obtained by the near-eye display device at the current moment. The current eye image is divided into n*m grid units according to the preset number of rows n and the preset number of columns m, where m and n are integers greater than or equal to 3. For example, the preset number of rows n = 5, 8, 10, 15, or 20, and the preset number of columns m = 5, 8, 10, 15, or 20, etc.
[0038] It should be noted that the size of the grid cell is related to the size of the human eye image and the number of rows and columns used to divide the human eye image. For example, when the size of the human eye image is fixed, the size of the grid cell is negatively correlated with the number of rows and columns used to divide the human eye image. That is, the smaller the number of rows and columns used to divide the human eye image, the larger the size of the grid cell; and the larger the number of rows and columns used to divide the human eye image, the smaller the size of the grid cell. Specifically, let the width of the human eye image be W, the height be H, and the number of rows and columns of the grid cell be n and m, then the width of the grid cell w = W / m, and the height of the grid cell h = H / n.
[0039] Step S103: Based on multiple grid cells, determine the target grid cell that the wearer's eye gazes at.
[0040] In this embodiment, the pixel coordinates of the pixels included in each grid cell can be obtained to obtain the pixel coordinate set of each grid cell. The pixel coordinates of the human eye gaze point in the human eye image can be determined. Then, the pixel coordinates of the human eye gaze point in the human eye image are matched with the pixel coordinates in the pixel coordinate set of each grid cell to determine the pixel coordinate set containing the pixel coordinates of the human eye gaze point. The grid cell corresponding to the pixel coordinate set containing the pixel coordinates of the human eye gaze point is determined as the target grid cell gazed at by the human eye gaze point.
[0041] In some embodiments, determining the target grid cell that the wearer's eye gazes at, based on a plurality of grid cells, may include: determining the pixel coordinates of the eye gaze point in the eye image; determining a target row index based on the ordinate of the pixel coordinates of the eye gaze point, the height of the grid cell, and the maximum row index of the grid cell; determining a target column index based on the abscissa of the pixel coordinates of the eye gaze point, the width of the grid cell, and the maximum column index of the grid cell; and determining the grid cell among the plurality of grid cells corresponding to the target row index and the target column index as the target grid cell that the eye gazes at. Wherein, let the width of the eye image be W, the height be H, and the number of rows and columns of the grid cell be n and m, then the width w of the grid cell = W / m, the height h of the grid cell = H / n, the range of the row index of the grid cell is [0, n-1], the maximum row index of the grid cell is n-1, the range of the column index of the grid cell is [0, m-1], and the maximum column index of the grid cell is m-1.
[0042] In some embodiments, determining the target row index based on the ordinate of the pixel coordinates of the eye's gaze point, the height of the grid cell, and the maximum row index of the grid cell may include: dividing the ordinate of the pixel coordinates of the eye's gaze point by the height of the grid cell and rounding down to obtain an initial row index; and determining the smaller of the initial row index and the maximum row index as the target row index. For example, this can be achieved using a formula. The target row index is calculated, where y t It is the ordinate of the pixel coordinates of the point of human eye fixation, i t is the target row index, h is the height of the grid cell, and n is the number of rows in the grid cell.
[0043] In some embodiments, determining the target column index based on the x-coordinate of the pixel coordinates of the eye's gaze point, the width of the grid cell, and the maximum column index of the grid cell may include: dividing the x-coordinate of the pixel coordinates of the eye's gaze point by the width of the grid cell and rounding down to obtain an initial column index; and determining the smaller of the initial column index and the maximum column index as the target column index. For example, this can be achieved using a formula. The target column index is calculated, where x t j is the x-coordinate of the pixel coordinates of the point of human eye fixation. t is the target column index, w is the width of the grid cell, and m is the number of columns in the grid cell.
[0044] Step S104: Determine the position information of the human eye gaze point based on the target mesh cell, and render and display the human eye gaze point in the user interface currently displayed on the near-eye display device based on the position information of the human eye gaze point.
[0045] In this embodiment, the grid cells are used to divide the human eye image according to a preset number of rows and columns. Therefore, each grid cell includes multiple pixels. As a result, the range of the target grid cell that the human eye gazes at is determined by the multiple grid cells is wider than that of a single pixel. Even if the human eye gazes at the same target grid cell, the position of the human eye gazes at the user interface can remain unchanged by using the position information of the human eye gazes at the target grid cell. This reduces the amplitude of gaze jitter and the frequency of gaze jitter, thus improving the user's visual experience.
[0046] In some embodiments, determining the location information of the human eye's gaze point based on the target grid cell may include: determining the pixel coordinates of the center point of the target grid cell; and determining the pixel coordinates of the center point of the target grid cell as the location information of the human eye's gaze point. For example, determining the pixel coordinates of the center point of the target grid cell may include: obtaining the row index, column index, width, and height of the target grid cell, and determining the pixel coordinates of the center point of the target grid cell based on the row index, column index, width, and height of the target grid cell. Wherein, the range of the row index of the grid cell is [0, n-1], the range of the column index of the grid cell is [0, m-1], n is the number of rows of the grid cell, and m is the number of columns of the grid cell.
[0047] For example, when the n*m grid cells obtained are of the same size, the x-coordinate of the pixel coordinates of the center point of the target grid cell can be obtained by formula x. t ′=(j t The result is obtained by calculating (+0.5)*w, where x t ′ is the x-coordinate of the center point of the target mesh cell in the pixel coordinates, j t is the column index of the target grid cell, and w is the width of the target grid cell. When the resulting n*m grid cells are of the same size, the ordinate of the pixel coordinates of the center point of the target grid cell can be determined using the formula y. t ′=(i t The result is calculated as +0.5)*h, where y t ′ is the ordinate of the pixel coordinates of the center point of the target mesh cell, i t is the row index of the target grid cell, and h is the height of the target grid cell.
[0048] In some embodiments, such as Figure 3 As shown, after step S102, the following steps are also included:
[0049] Step S105: Determine the fixation heat of each grid cell in the multiple grid cells. The fixation heat of the grid cell is related to the frequency and / or duration of human eye fixation on the grid cell.
[0050] In this embodiment, the fixation intensity of a grid cell is positively correlated with the frequency and / or duration of fixation on the grid cell by the human eye's fixation point. Specifically, the higher the frequency of fixation on the grid cell, the higher the fixation intensity; conversely, the lower the frequency of fixation, the lower the fixation intensity. Furthermore, the higher the frequency and the longer the fixation duration, the higher the fixation intensity; and the lower the frequency and the shorter the fixation duration, the lower the fixation intensity. Finally, the longer the fixation duration, the higher the fixation intensity; and the shorter the fixation duration, the lower the fixation intensity.
[0051] In some embodiments, determining the gaze heat of each grid cell in a plurality of grid cells may include: acquiring multiple historical location information of human eye gaze points, and determining the location distribution information of human eye gaze points based on the multiple historical location information of human eye gaze points; generating a gaze heat map based on the location distribution information of human eye gaze points, wherein the size of the gaze heat map is the same as the size of the human eye image; and determining the gaze heat of each grid cell in the plurality of grid cells based on the gaze heat map. The multiple historical location information of human eye gaze points is determined before the current time, and the difference between the time when the historical location information was determined and the current time is less than or equal to a preset time difference. For example, the preset time difference is 0.5 seconds or 1 second, meaning that the gaze heat map is generated based on multiple historical location information within 0.5 seconds or 1 second before the current time.
[0052] In some embodiments, generating a gaze heatmap based on the location distribution information of human eye gaze points may include: dividing a blank image of the same size as the human eye image into multiple gaze heat units in the same manner as dividing the human eye image; determining the number of human eye gaze points falling into each gaze heat unit based on the location distribution information of human eye gaze points, determining the gaze heat of each gaze heat unit based on the number of human eye gaze points falling into each gaze heat unit, and marking the corresponding gaze heat in each gaze heat unit to obtain a gaze heatmap.
[0053] In some embodiments, determining the gaze heat of a gaze heat unit based on the number of human eye gaze points falling into the gaze heat unit may include: determining the gaze heat of the gaze heat unit based on the number of human eye gaze points falling into the gaze heat unit and the mapping relationship between the number of human eye gaze points and gaze heat. The mapping relationship between the number of human eye gaze points and gaze heat is pre-established, and the gaze heat can be quickly determined using the number of human eye gaze points falling into the gaze heat unit and this mapping relationship.
[0054] In some embodiments, determining the gaze heat of each of a plurality of grid cells based on the gaze heatmap may include: determining a corresponding gaze heat cell from the gaze heatmap based on the row and column index of each of the plurality of grid cells, and determining the gaze heat marked within the gaze heat cell as the gaze heat of the corresponding grid cell, wherein the gaze heat cell has the same row and column index as the corresponding grid cell.
[0055] Step S106: Determine the size adjustment coefficient of each grid cell based on the gaze heat of each grid cell. The size adjustment coefficient of the grid cell is negatively correlated with the gaze heat of the grid cell.
[0056] In this embodiment, the size adjustment coefficient of the grid cell is negatively correlated with the gaze intensity of the grid cell. That is, the higher the gaze intensity of the grid cell, the smaller the size adjustment coefficient of the grid cell, and the lower the gaze intensity of the grid cell, the larger the size adjustment coefficient of the grid cell. The size adjustment coefficient is within a preset value range, which can be set based on actual conditions. This embodiment does not specifically limit this range. For example, the preset value range includes [0.5, 2].
[0057] In some embodiments, determining the size adjustment coefficient of each grid cell based on the gaze heat of each grid cell may include: determining the size adjustment coefficient of each grid cell based on the gaze heat of each grid cell and the width and height of the human eye image, such that the sum of the heights of each grid cell after being sized based on the size adjustment coefficient is the same as the height of the human eye image, and the sum of the widths of each grid cell after being sized is the same as the width of the human eye image. This embodiment determines the size adjustment coefficient of each grid cell based on the gaze heat of each grid cell and the width and height of the human eye image, thereby ensuring that the sum of the heights of each grid cell after being sized is the same as the height of the human eye image, and the sum of the widths of each grid cell after being sized is the same as the width of the human eye image, thus avoiding overall size distortion.
[0058] Step S107: Multiply the size of each grid cell by the corresponding size adjustment factor to update the size of each grid cell. Then execute steps S103 and S104.
[0059] This embodiment dynamically adjusts the size of each grid cell based on the gaze intensity of each grid cell, so that the higher the gaze intensity, the smaller the size of the grid cell, and the lower the gaze intensity, the larger the size of the grid cell. This makes the size of the target grid cell that the human eye gazes at smaller. In this way, the position information of the human eye gaze point can be determined more accurately based on the smaller target grid cell, thus improving the accuracy of the position information of the human eye gaze point.
[0060] In some embodiments, determining the location information of the human eye's gaze point based on the target mesh cell may include: determining the pixel coordinates of the center point of the target mesh cell; and determining the location information of the human eye's gaze point using the pixel coordinates of the center point of the target mesh cell. Where, in the case that some or all of the n*m mesh cells after size adjustment have different sizes, the abscissa of the pixel coordinates of the center point of the target mesh cell can be determined using the formula... Calculations show that i t j t It is the row and column index of the target grid cell, w pq It is the width of the grid cell corresponding to row and column index pq. It is the width of the target mesh cell. The ordinate of the pixel coordinates of the center point of the target mesh cell can be obtained by formula. The calculation yields h pq It is the height of the grid cell corresponding to row and column index pq. It is the height of the target mesh cell.
[0061] In some embodiments, dividing a human eye image into multiple grid units according to a preset number of rows and a preset number of columns may include: obtaining the cumulative number of times the first preset number of rows and the first preset number of columns have been used in the historical division of the human eye image, when the preset number of rows and the preset number of columns are both a first preset number; in response to the cumulative number of uses being less than a preset threshold, dividing the human eye image into multiple grid units according to the first preset number of rows and the first preset number of columns, and incrementing the cumulative number of uses by 1; in response to the cumulative number of uses being greater than or equal to the preset threshold, dividing the human eye image into multiple grid units according to a second preset number of rows and the second preset number of columns, and resetting the cumulative number of uses to zero, wherein the first preset number of rows is less than the second preset number of rows, and the first preset number of columns is less than the second preset number of columns. In this embodiment, the size of the grid cells obtained by dividing the human eye image according to the first preset number of rows and the first preset number of columns is larger than the size of the grid cells obtained by dividing the human eye image according to the second preset number of rows and the second preset number of columns. The speed of determining the position information of the human eye gaze point by using the larger grid cells is greater than the speed of determining the position information of the human eye gaze point by using the smaller grid cells. However, the accuracy of the determined position information of the human eye gaze point is less than the accuracy of the determined position information of the human eye gaze point by using the smaller grid cells. Therefore, after using the larger grid cells multiple times to determine the position information of the human eye gaze point, the system switches to using the smaller grid cells to determine the position information of the human eye gaze point in order to improve the accuracy of the position information of the human eye gaze point.
[0062] It should be noted that the preset number of rows and columns used for historically dividing the human eye image refers to the preset number of rows and columns used to divide the human eye image acquired before the current moment. The first preset number of rows, the second preset number of rows, the second preset number of columns, and the second preset number of columns can be set based on different application scenarios, and this embodiment does not impose specific limitations on them. For example, if the first preset number of rows and the second preset number of columns are both 5, the human eye image is divided into 5*5=25 grid units; if the second preset number of rows and the second preset number of columns are both 10, the human eye image is divided into 10*10=100 grid units; or if the second preset number of rows and the second preset number of columns are both 20, the human eye image is divided into 20*20=400 grid units.
[0063] In some embodiments, such as Figure 4As shown, step S102 includes sub-step S1021.
[0064] Sub-step S1021: When the preset number of rows and preset number of columns used in the historical division of the human eye image are the second preset number, the human eye image is divided into multiple grid cells according to the second preset number of rows and the second preset number of columns. Then, steps S103 and S104 are executed. This embodiment uses smaller grid cells to determine the location information of the human eye gaze point, resulting in higher accuracy in determining the location information of the human eye gaze point.
[0065] In some embodiments, such as Figure 4 As shown, after step S103, the following steps are also included:
[0066] Step S108: Obtain the historical grid cell that the human eye is looking at, and determine the deviation value between the row and column indices of the historical grid cell and the row and column indices of the target grid cell.
[0067] In this embodiment, the historical grid cell that the human eye fixates on is the grid cell that the human eye fixates on before the current moment. The deviation between the row and column indices of the historical grid cell and the target grid cell includes a first deviation between the row index of the historical grid cell and the row index of the target grid cell, and a second deviation between the column index of the historical grid cell and the column index of the target grid cell.
[0068] Step S109: When the deviation value is greater than or equal to the preset deviation threshold, the human eye image is divided into multiple grid units according to the first preset number of rows and the first preset number of columns.
[0069] In this embodiment, after executing step S109, the process returns to step S103, which determines the target grid cell that the wearer's eye is fixating on based on multiple grid cells, and step S104, which determines the position information of the eye's fixation point based on the target grid cell. Based on the position information of the eye's fixation point, the eye's fixation point is rendered and displayed in the user interface currently displayed on the near-eye display device. In this embodiment, if the determined target grid cell differs significantly from the historical grid cells after dividing the eye image using a higher number of rows and columns, causing significant jitter in the fixation point, the process switches to using a lower number of rows and columns to divide the eye image, thereby increasing the size of the grid cells and reducing fixation point jitter.
[0070] In some embodiments, the preset deviation threshold is related to the size of the grid cells obtained by dividing the human eye image according to a first preset number of rows and a first preset number of columns, and the size of the grid cells obtained by dividing the human eye image according to a second preset number of rows and a second preset number of columns. For example, the preset deviation threshold includes a first deviation threshold and a second deviation threshold. The first deviation threshold is obtained by dividing the height of the grid cells obtained by dividing the human eye image according to the first preset number of rows and the first preset number of columns by the height of the grid cells obtained by dividing the human eye image according to the second preset number of rows and the second preset number of columns, and then rounding down or up. The second deviation threshold is obtained by dividing the width of the grid cells obtained by dividing the human eye image according to the first preset number of rows and the first preset number of columns by the width of the grid cells obtained by dividing the human eye image according to the second preset number of rows and the second preset number of columns, and then rounding down or up.
[0071] In some embodiments, when the deviation value is greater than or equal to a preset deviation threshold, dividing the human eye image into multiple grid units according to a first preset number of rows and a first preset number of columns may include: when the first deviation value is greater than or equal to a first deviation threshold and / or the second deviation value is greater than or equal to a second deviation threshold, dividing the human eye image into multiple grid units according to a first preset number of rows and a first preset number of columns.
[0072] If the deviation value is less than a preset deviation threshold, step S104 is executed: the position information of the human eye's gaze point is determined based on the target mesh cell, and the human eye's gaze point is rendered and displayed on the user interface currently displayed on the near-eye display device based on the position information of the human eye's gaze point. For example, if the first deviation value is less than a first deviation threshold and the second deviation value is less than a second deviation threshold, step S104 is executed: the position information of the human eye's gaze point is determined based on the target mesh cell, and the human eye's gaze point is rendered and displayed on the user interface currently displayed on the near-eye display device based on the position information of the human eye's gaze point.
[0073] Please see Figure 5 , Figure 5 This is a flowchart illustrating another foveated rendering method provided in an embodiment of the present invention.
[0074] like Figure 5 As shown, the foveated rendering method includes steps S201 to 205.
[0075] Step S201: Acquire multiple frames of human eye images obtained by the near-eye display device from the wearer's eyes.
[0076] In this embodiment, the multiple frames of eye images acquired by the near-eye display device from the wearer's eyes are continuous, and these multiple frames include the current eye image and at least two historical eye images. The at least two historical eye images are the historical eye images most recently acquired. The current eye image is acquired by the near-eye display device at the current moment, and the historical eye images are acquired by the near-eye display device before the current moment. Furthermore, the near-eye display device acquires each historical eye image at a different time. For example, the near-eye display device acquires k most recently acquired eye images from the wearer, where k is an integer greater than or equal to 3, such as k = 3, 4, or 5.
[0077] Step S202: Divide each frame of human eye image into multiple grid units according to the preset number of rows and columns.
[0078] In this embodiment, the preset number of rows and columns can be set based on different application scenarios, and this embodiment does not impose specific limitations on them. For example, the current eye image of the wearer is obtained by the near-eye display device at the current moment. The current eye image is divided into n*m grid units according to the preset number of rows n and the preset number of columns m, where m and n are integers greater than or equal to 3. For example, the preset number of rows n = 5, 8, 10, 15, or 20, and the preset number of columns m = 5, 8, 10, 15, or 20, etc.
[0079] In some embodiments, dividing each frame of human eye image into multiple grid units according to a preset number of rows and a preset number of columns may include: when the preset number of rows used in the historical division of multiple frames of human eye image is a first preset number of rows and the preset number of columns is a first preset number of columns, obtaining the cumulative number of uses of the first preset number of rows and the first preset number of columns; in response to the cumulative number of uses being less than a preset number of uses threshold, dividing each frame of human eye image into multiple grid units according to the first preset number of rows and the first preset number of columns, and incrementing the cumulative number of uses by 1; in response to the cumulative number of uses being greater than or equal to the preset number of uses threshold, dividing each frame of human eye image into multiple grid units according to a second preset number of rows and the second preset number of columns, and resetting the cumulative number of uses to zero, wherein the first preset number of rows is less than the second preset number of rows and the first preset number of columns is less than the second preset number of columns. In this embodiment, the size of the grid cells obtained by dividing the human eye image according to the first preset number of rows and the first preset number of columns is larger than the size of the grid cells obtained by dividing the human eye image according to the second preset number of rows and the second preset number of columns. The speed of determining the position information of the human eye gaze point by using the larger grid cells is greater than the speed of determining the position information of the human eye gaze point by using the smaller grid cells. However, the accuracy of the determined position information of the human eye gaze point is less than the accuracy of the determined position information of the human eye gaze point by using the smaller grid cells. Therefore, after using the larger grid cells multiple times to determine the position information of the human eye gaze point, the system switches to using the smaller grid cells to determine the position information of the human eye gaze point in order to improve the accuracy of the position information of the human eye gaze point.
[0080] Step S203: Determine the grid cell that the wearer's eye gazes at from the multiple grid cells corresponding to each frame of human eye image, and obtain multiple candidate grid cells.
[0081] In this embodiment, one frame of human eye image corresponds to one candidate grid cell. It should be noted that the specific implementation method for determining the grid cell that the wearer's eye gazes at from the multiple grid cells corresponding to each frame of human eye image can refer to the corresponding embodiment for determining the grid cell that the eye gazes at in the foregoing embodiments, and will not be repeated here.
[0082] In some embodiments, determining the grid cell that the wearer's eye gazes at from a plurality of grid cells corresponding to each frame of human eye image to obtain a plurality of candidate grid cells may include: for each frame of human eye image, determining the pixel coordinates of the eye gaze point in the human eye image; determining a target row index based on the ordinate of the pixel coordinates of the eye gaze point, the height of the grid cell, and the maximum row index of the grid cell; determining a target column index based on the abscissa of the pixel coordinates of the eye gaze point, the width of the grid cell, and the maximum column index of the grid cell; and determining the grid cell among the plurality of grid cells corresponding to the human eye image that corresponds to the target row index and the target column index as the grid cell that the eye gaze point gazes at, thereby obtaining a plurality of candidate grid cells.
[0083] In some embodiments, determining the grid cell gazed at by the wearer's eye fixation point from a plurality of grid cells corresponding to each frame of human eye image to obtain a plurality of candidate grid cells includes: determining the fixation heat of each grid cell in the plurality of grid cells corresponding to each frame of human eye image, wherein the fixation heat of the grid cell is related to the frequency and / or duration of the human eye fixation point gazing at the grid cell; determining a size adjustment coefficient for each grid cell based on the fixation heat of each grid cell, and multiplying the size of each grid cell by the corresponding size adjustment coefficient to update the size of each grid cell; and determining the grid cell to which the wearer's eye fixation point belongs from the plurality of grid cells corresponding to each frame of human eye image to obtain a plurality of candidate grid cells. This embodiment dynamically adjusts the size of each grid cell based on the fixation heat of each grid cell to achieve the effect that the higher the fixation heat, the smaller the size of the grid cell, and the lower the fixation heat, the larger the size of the grid cell. This results in a smaller size for the candidate grid cells gazed at by the human eye fixation point, thus making the size of the target grid cell subsequently selected from the plurality of candidate grid cells smaller, thereby more accurately determining the position information of the human eye fixation point and improving the accuracy of the position information of the human eye fixation point.
[0084] In some embodiments, determining the gaze heat of each grid cell among the multiple grid cells corresponding to each frame of human eye image may include: acquiring multiple historical position information of human eye gaze points, and determining the positional distribution information of human eye gaze points based on the multiple historical position information of human eye gaze points; generating a gaze heat map based on the positional distribution information of human eye gaze points, and determining the gaze heat of each grid cell among the multiple grid cells corresponding to each frame of human eye image based on the gaze heat map. The multiple historical position information of human eye gaze points is determined before the current time, and the difference between the determination time of the historical position information and the current time is within a preset time difference. For example, the preset time difference is 0.5 seconds or 1 second, that is, the gaze heat map is generated based on multiple historical position information within 0.5 seconds or 1 second before the current time. It should be noted that the specific implementation process of determining the gaze heat of grid cells can be referred to the corresponding description in the foregoing embodiments, and will not be repeated here.
[0085] Step S204: Determine the number of identical candidate mesh cells among multiple candidate mesh cells, and select the candidate mesh cell with the largest number of identical candidate mesh cells as the target mesh cell.
[0086] In this embodiment, candidate grid cells with the same row and column indices can be counted as the same candidate grid cells, thereby obtaining the number of different candidate grid cells, and selecting the candidate grid cell with the largest number as the target grid cell.
[0087] In some embodiments, selecting the candidate grid cell with the largest number of candidate grid cells as the target grid cell may include: in response to at least two candidate grid cells with the largest number of candidate grid cells among the multiple candidate grid cells, selecting the candidate grid cell belonging to the current human eye image as the target grid cell from the at least two candidate grid cells with the largest number of candidate grid cells.
[0088] Step S205: Determine the position information of the human eye gaze point based on the target mesh cell, and render and display the human eye gaze point in the user interface currently displayed on the near-eye display device based on the position information of the human eye gaze point.
[0089] In this embodiment, the target grid cells are divided into human eye images according to a preset number of rows and columns. Therefore, the range of the target grid cells that the human eye gazes over is wider than that of a single pixel. Even if the human eye gazes over the same target grid cell, the position of the human eye gaze in the user interface can remain unchanged when rendered using the position information of the human eye gaze determined by the target grid cells. This reduces the jitter amplitude and frequency of the gaze. Furthermore, the target grid cells are obtained by majority voting on multiple candidate grid cells that the human eye gazes over in the most recent frames. Therefore, when using the target grid cells to determine the position information of the human eye gaze, sudden jumps in the position information of the human eye gaze can be further eliminated, thereby further reducing the frequency of jitter in the human eye gaze and making the display of the human eye gaze smoother and more stable. Moreover, the calculation latency of majority voting is low, which can meet the real-time rendering requirements of the human eye gaze.
[0090] In some embodiments, selecting the candidate grid cell with the largest number from multiple candidate grid cells as the target grid cell may include: selecting the candidate grid cell with the largest number from multiple candidate grid cells as the grid cell to be optimized; performing sliding weighted filtering on the row and column indices of the grid cell to be optimized based on the frame offset of the human eye image to which each candidate grid cell belongs, to obtain the target row and column index; and determining the candidate grid cell corresponding to the target row and column index as the target grid cell. Here, the frame offset of the human eye image refers to the absolute value of the difference between the frame number of each frame of the human eye image in multiple frames and the current human eye image. In this embodiment, after obtaining the grid cell to be optimized by majority voting on multiple candidate grid cells gazed at by the human eye in the most recent multiple frames, the row and column indices of the grid cell to be optimized are further subjected to sliding weighted filtering. This makes the position information determined based on the target grid cell corresponding to the target row and column index obtained by sliding weighted filtering smoother, reducing the jitter amplitude of the human eye gaze point and further reducing the frequency of the human eye gaze point, making the display of the human eye gaze point smoother and more stable.
[0091] In some embodiments, performing sliding weighted filtering on the row and column indices of the grid cell to be optimized based on the frame offset of the human eye image to which each candidate grid cell belongs, to obtain the target row and column index, may include: determining the weighting coefficient of each candidate grid cell based on the frame offset of the human eye image to which each candidate grid cell belongs; accumulating the weighting coefficients of each candidate grid cell to obtain a total weighting coefficient value; multiplying the weighting coefficient of each candidate grid cell by the row index of the grid cell to be optimized and accumulating the results to obtain a total row index value; dividing the total row index value by the total weighting coefficient value to obtain a weighted row index; multiplying the weighting coefficient of each candidate grid cell by the column index of the grid cell to be optimized and accumulating the results to obtain a total column index value; dividing the total column index value by the total column weighting coefficient value to obtain a weighted column index; and rounding the weighted row index and the weighted column index to obtain the target row and column index.
[0092] For example, through formula The weighted row index can be calculated using the formula. The weighted column index can be calculated, where, It is a weighted row index. It is a weighted column index, where l is the frame offset, k is the number of frames in the human eye image, and w l It is the weighting coefficient of the candidate grid cell in the l-th frame of the human eye image out of k frames, i t-l j is the row index of the grid cell to be optimized. t-l It is the column index of the grid cell to be optimized.
[0093] In some embodiments, the row index weighting coefficient and column index weighting coefficient of each candidate grid cell are determined based on the frame offset of the human eye image to which each candidate grid cell belongs: A linear attenuation method is used to determine the row index weighting coefficient and column index weighting coefficient of each candidate grid cell based on the frame offset of the human eye image to which each candidate grid cell belongs. For example, the row index weighting coefficient and column index weighting coefficient of each candidate grid cell are obtained by subtracting the frame offset of the human eye image to which each candidate grid cell belongs from the number of frames of the human eye image. This embodiment uses a linear attenuation method to determine the row index weighting coefficient and column index weighting coefficient, which has a smaller impact on latency and higher real-time performance.
[0094] In some embodiments, the row index weighting coefficient and column index weighting coefficient of each candidate grid cell are determined based on the frame offset of the human eye image to which each candidate grid cell belongs: The row index weighting coefficient and column index weighting coefficient of each candidate grid cell are determined using an exponential decay method, based on the frame offset of the human eye image to which each candidate grid cell belongs. For example, the row index weighting coefficient and column index weighting coefficient of each candidate grid cell are obtained by multiplying the frame offset of the human eye image to which each candidate grid cell belongs by a preset decay coefficient. The preset decay coefficient can be set based on actual conditions, and this embodiment does not specifically limit it. For example, the preset decay coefficient is 0.7. This embodiment uses an exponential decay method to determine the row index weighting coefficient and column index weighting coefficient, resulting in better smoothing.
[0095] In some embodiments, such as Figure 6 As shown, step S202 includes sub-step S2021.
[0096] Sub-step S2021: When the preset number of rows and preset number of columns used for historically dividing multiple frames of human eye images are the second preset number, divide each frame of human eye image into multiple grid cells according to the second preset number of rows and the second preset number of columns. Then execute steps S203, S204, and S205. This embodiment uses smaller grid cells to determine the location information of the human eye gaze point, resulting in higher accuracy in determining the location information of the human eye gaze point.
[0097] In some embodiments, such as Figure 6 As shown, after step S204, the following steps are also included:
[0098] Step S206: Obtain the historical grid cell that the human eye is focused on, and determine the deviation value between the row and column indices of the historical grid cell and the row and column indices of the target grid cell.
[0099] In this embodiment, the historical grid cell that the human eye fixates on is the grid cell that the human eye fixates on before the current moment. The deviation between the row and column indices of the historical grid cell and the target grid cell includes a first deviation between the row index of the historical grid cell and the row index of the target grid cell, and a second deviation between the column index of the historical grid cell and the column index of the target grid cell.
[0100] Step S207: When the deviation value is greater than or equal to the preset deviation threshold, divide each frame of human eye image into multiple grid units according to the first preset number of rows and the first preset number of columns.
[0101] In this embodiment, after executing step S207, the process returns to step S204, determining the target grid cell that the wearer's eye is fixating on based on multiple grid cells, and step S205, determining the position information of the eye's fixation point based on the target grid cell. Based on the position information of the eye's fixation point, the eye's fixation point is rendered and displayed in the user interface currently displayed on the near-eye display device. In this embodiment, after dividing each frame of the human eye image using the higher number of rows and columns, if the determined target grid cell differs significantly from the historical grid cells, causing significant jitter in the fixation point, the process switches to using the lower number of rows and columns to divide each frame of the human eye image, thereby making the grid cell size larger and reducing fixation point jitter.
[0102] In some embodiments, the preset deviation threshold is related to the size of the grid cells obtained by dividing the human eye image according to a first preset number of rows and a first preset number of columns, and the size of the grid cells obtained by dividing the human eye image according to a second preset number of rows and a second preset number of columns. For example, the preset deviation threshold includes a first deviation threshold and a second deviation threshold. The first deviation threshold is obtained by dividing the height of the grid cells obtained by dividing the human eye image according to the first preset number of rows and the first preset number of columns by the height of the grid cells obtained by dividing the human eye image according to the second preset number of rows and the second preset number of columns, and then rounding down or up. The second deviation threshold is obtained by dividing the width of the grid cells obtained by dividing the human eye image according to the first preset number of rows and the first preset number of columns by the width of the grid cells obtained by dividing the human eye image according to the second preset number of rows and the second preset number of columns, and then rounding down or up.
[0103] In some embodiments, when the deviation value is greater than or equal to a preset deviation threshold, dividing the human eye image into multiple grid units according to a first preset number of rows and a first preset number of columns may include: when the first deviation value is greater than or equal to a first deviation threshold and / or the second deviation value is greater than or equal to a second deviation threshold, dividing the human eye image into multiple grid units according to a first preset number of rows and a first preset number of columns.
[0104] If the deviation value is less than a preset deviation threshold, step S205 is executed: the position information of the human eye's gaze point is determined based on the target mesh cell; and the human eye's gaze point is rendered and displayed on the user interface currently displayed on the near-eye display device based on the position information of the human eye's gaze point. For example, if the first deviation value is less than a first deviation threshold and the second deviation value is less than a second deviation threshold, step S205 is executed: the position information of the human eye's gaze point is determined based on the target mesh cell; and the human eye's gaze point is rendered and displayed on the user interface currently displayed on the near-eye display device based on the position information of the human eye's gaze point.
[0105] Please see Figure 7 , Figure 7This is a schematic block diagram of the structure of a near-eye display device provided in an embodiment of the present invention.
[0106] like Figure 7 As shown, the near-eye display device 100 includes a processor 101 and a memory 102, which are connected via a bus 103, such as an I2C (Inter-integrated Circuit) bus.
[0107] Specifically, processor 101 provides computing and control capabilities to support the operation of the entire near-eye display device 100. Processor 101 can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.
[0108] Specifically, the memory 102 can be a Flash chip, a read-only memory (ROM) disk, an optical disk, a USB flash drive, or a portable hard drive, etc.
[0109] Those skilled in the art will understand that Figure 7 The structure shown is merely a block diagram of a portion of the structure related to the embodiments of the present invention, and does not constitute a limitation on the near-eye display device 100 to which the embodiments of the present invention are applied. The specific near-eye display device 100 may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0110] The processor 101 is used to run a computer program stored in the memory 102, and implements any of the foveated rendering methods provided in the embodiments of the present invention when executing the computer program.
[0111] In one embodiment, the processor 101 is configured to run a computer program stored in the memory 102, and to perform the following steps when executing the computer program:
[0112] The near-eye display device acquires an image of the wearer's eye.
[0113] The human eye image is divided into multiple grid units according to a preset number of rows and columns;
[0114] Based on the plurality of grid cells, the target grid cell that the wearer's eye gazes at is determined;
[0115] Based on the target grid cell, the position information of the human eye gaze point is determined, and based on the position information of the human eye gaze point, the human eye gaze point is rendered and displayed in the user interface currently displayed by the near-eye display device.
[0116] In one embodiment, the human eye image includes multiple frames, and the plurality of grid cells include a plurality of grid cells corresponding to each frame of the human eye image. When the processor 101 determines the target grid cell gazed at by the wearer's eye fixation point based on the plurality of grid cells, it is configured to:
[0117] From the plurality of grid cells corresponding to each frame of the human eye image, determine the grid cell that the wearer's eye gazes at, and obtain a plurality of candidate grid cells;
[0118] Determine the number of identical candidate grid cells among the plurality of candidate grid cells, and select the candidate grid cell with the largest number of identical candidate grid cells as the target grid cell.
[0119] In one embodiment, when the processor 101 selects the candidate grid cell with the largest number from the plurality of candidate grid cells as the target grid cell, it is configured to:
[0120] Select the candidate grid cell with the largest number of candidates from the plurality of candidate grid cells as the grid cell to be optimized;
[0121] Based on the frame offset of the human eye image to which each of the multiple candidate grid cells belongs, the row and column indices of the grid cells to be optimized are subjected to sliding weighted filtering to obtain the target row and column indices;
[0122] The candidate grid cell corresponding to the target row and column index is determined as the target grid cell.
[0123] In one embodiment, when the processor 101 determines the grid cell that the wearer's eye gazes at from the plurality of grid cells corresponding to each frame of the human eye image, and obtains a plurality of candidate grid cells, it is configured to:
[0124] Determine the gaze heat of each of the plurality of grid cells corresponding to each frame of the human eye image, wherein the gaze heat of the grid cell is related to the frequency and / or duration of the human eye gaze at the grid cell;
[0125] Based on the gaze intensity of each grid cell, a size adjustment factor for each grid cell is determined, and the size of each grid cell is updated by multiplying the size of each grid cell by the corresponding size adjustment factor.
[0126] From the multiple grid cells corresponding to each frame of the human eye image, determine the grid cell to which the wearer's eye gaze point belongs, and obtain multiple candidate grid cells.
[0127] In one embodiment, when the processor 101 determines the gaze heat of each of the plurality of grid cells corresponding to each frame of the human eye image, it is configured to:
[0128] Acquire multiple historical location information of the human eye fixation point, and determine the location distribution information of the human eye fixation point based on the multiple historical location information of the human eye fixation point;
[0129] Based on the location distribution information of the human eye gaze point, a gaze heatmap is generated. Based on the gaze heatmap, the gaze heat of each of the multiple grid cells corresponding to each frame of the human eye image is determined.
[0130] In one embodiment, when the processor 101 divides the human eye image into multiple grid units according to a preset number of rows and columns, it is configured to:
[0131] When the number of preset rows used in historical segmentation of human eye images is the first preset number of rows and the number of preset columns is the first preset number of columns, the cumulative number of times the first preset number of rows and the first preset number of columns are used is obtained;
[0132] In response to the cumulative number of uses being less than a preset threshold, the human eye image is divided into multiple grid units according to the first preset number of rows and the first preset number of columns, and the cumulative number of uses is incremented by 1;
[0133] In response to the cumulative number of uses being greater than or equal to a preset threshold, the human eye image is divided into multiple grid units according to a second preset number of rows and a second preset number of columns, and the cumulative number of uses is reset to zero, wherein the first preset number of rows is less than the second preset number of rows, and the first preset number of columns is less than the second preset number of columns.
[0134] In one embodiment, when the processor 101 divides the human eye image into multiple grid units according to a preset number of rows and columns, it is configured to:
[0135] When the preset number of rows used for historically dividing the human eye image is the second preset number of rows and the preset number of columns is the second preset number of columns, the human eye image is divided into multiple grid units according to the second preset number of rows and the second preset number of columns.
[0136] In one embodiment, after determining the target grid cell that the wearer's eye gaze point is focused on based on the plurality of grid cells, the processor 101 is further configured to:
[0137] Obtain the historical grid cell that the human eye is focused on, and determine the deviation value between the row and column index of the historical grid cell and the row and column index of the target grid cell;
[0138] When the deviation value is greater than or equal to a preset deviation threshold, the human eye image is divided into multiple grid units according to the first preset number of rows and the first preset number of columns;
[0139] Return to the step of determining the target grid cell that the wearer's eye gazes at based on the plurality of grid cells;
[0140] When the deviation value is less than the preset deviation threshold, the position information of the human eye gaze point is determined according to the target grid cell, and the human eye gaze point is rendered and displayed in the user interface currently displayed by the near-eye display device according to the position information of the human eye gaze point.
[0141] It should be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the near-eye display device described above can be referred to the corresponding process in the aforementioned gaze point rendering method embodiment, and will not be repeated here.
[0142] This invention also provides a storage medium for computer-readable storage, wherein the storage medium stores one or more programs that can be executed by one or more processors to implement any of the foveated rendering methods provided in the specification of this invention.
[0143] The storage medium can be volatile or non-volatile. It can be an internal storage unit of the near-eye display device described in the foregoing embodiments, such as the hard drive or memory of the near-eye display device. Alternatively, it can be an external storage device of the near-eye display device, such as a plug-in hard drive, smart media card (SMC), secure digital card (SD), flash card, etc., provided on the near-eye display device.
[0144] It will be understood by those skilled in the art that all or some of the steps, systems, or apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware embodiments, the division between functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.
[0145] It should be understood that the term "and / or" as used in this specification and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations. It should be noted that, herein, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.
[0146] The sequence numbers of the above embodiments of the present invention are merely for descriptive purposes and do not represent the superiority or inferiority of the embodiments. The above descriptions are only specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A foveated rendering method, characterized in that, Applied to near-eye display devices, the method includes: The near-eye display device acquires an image of the wearer's eye. The human eye image is divided into multiple grid units according to a preset number of rows and columns; Based on the plurality of grid cells, the target grid cell that the wearer's eye gazes at is determined; Based on the target grid cell, the position information of the human eye gaze point is determined, and based on the position information of the human eye gaze point, the human eye gaze point is rendered and displayed in the user interface currently displayed by the near-eye display device.
2. The foveated rendering method according to claim 1, characterized in that, The human eye image includes multiple frames, and the plurality of grid cells include a plurality of grid cells corresponding to each frame of the human eye image. Determining the target grid cell that the wearer's gaze point is fixed on based on the plurality of grid cells includes: From the plurality of grid cells corresponding to each frame of the human eye image, determine the grid cell that the wearer's eye gazes at, and obtain a plurality of candidate grid cells; Determine the number of identical candidate grid cells among the plurality of candidate grid cells, and select the candidate grid cell with the largest number of identical candidate grid cells as the target grid cell.
3. The foveated rendering method according to claim 2, characterized in that, Selecting the candidate grid cell with the largest number from the plurality of candidate grid cells as the target grid cell includes: Select the candidate grid cell with the largest number of candidates from the plurality of candidate grid cells as the grid cell to be optimized; Based on the frame offset of the human eye image to which each of the multiple candidate grid cells belongs, the row and column indices of the grid cells to be optimized are subjected to sliding weighted filtering to obtain the target row and column indices; The candidate grid cell corresponding to the target row and column index is determined as the target grid cell.
4. The foveated rendering method according to claim 2, characterized in that, The step of determining the grid cell that the wearer's eye gazes at from the plurality of grid cells corresponding to each frame of the human eye image, thereby obtaining a plurality of candidate grid cells, includes: Determine the gaze heat of each of the plurality of grid cells corresponding to each frame of the human eye image, wherein the gaze heat of the grid cell is related to the frequency and / or duration of the human eye gaze at the grid cell; Based on the gaze intensity of each grid cell, a size adjustment factor for each grid cell is determined, and the size of each grid cell is updated by multiplying the size of each grid cell by the corresponding size adjustment factor. From the multiple grid cells corresponding to each frame of the human eye image, determine the grid cell to which the wearer's eye gaze point belongs, and obtain multiple candidate grid cells.
5. The foveated rendering method according to claim 4, characterized in that, Determining the gaze heat of each of the plurality of grid cells corresponding to each frame of the human eye image includes: Acquire multiple historical location information of the human eye fixation point, and determine the location distribution information of the human eye fixation point based on the multiple historical location information of the human eye fixation point; Based on the location distribution information of the human eye gaze point, a gaze heatmap is generated. Based on the gaze heatmap, the gaze heat of each of the multiple grid cells corresponding to each frame of the human eye image is determined.
6. The foveated rendering method according to any one of claims 1-5, characterized in that, The step of dividing the human eye image into multiple grid units according to a preset number of rows and columns includes: When the number of preset rows used in historical segmentation of human eye images is the first preset number of rows and the number of preset columns is the first preset number of columns, the cumulative number of times the first preset number of rows and the first preset number of columns are used is obtained; In response to the cumulative number of uses being less than a preset threshold, the human eye image is divided into multiple grid units according to the first preset number of rows and the first preset number of columns, and the cumulative number of uses is incremented by 1; In response to the cumulative number of uses being greater than or equal to a preset threshold, the human eye image is divided into multiple grid units according to a second preset number of rows and a second preset number of columns, and the cumulative number of uses is reset to zero, wherein the first preset number of rows is less than the second preset number of rows, and the first preset number of columns is less than the second preset number of columns.
7. The foveated rendering method according to any one of claims 1-5, characterized in that, The step of dividing the human eye image into multiple grid units according to a preset number of rows and columns includes: When the preset number of rows used for historically dividing the human eye image is the second preset number of rows and the preset number of columns is the second preset number of columns, the human eye image is divided into multiple grid units according to the second preset number of rows and the second preset number of columns.
8. The foveated rendering method according to claim 7, characterized in that, After determining the target grid cell that the wearer's eye gazes at based on the plurality of grid cells, the method further includes: Obtain the historical grid cell that the human eye is focused on, and determine the deviation value between the row and column index of the historical grid cell and the row and column index of the target grid cell; When the deviation value is greater than or equal to a preset deviation threshold, the human eye image is divided into multiple grid units according to the first preset number of rows and the first preset number of columns; Return to the step of determining the target grid cell that the wearer's eye gazes at based on the plurality of grid cells; When the deviation value is less than the preset deviation threshold, the position information of the human eye gaze point is determined according to the target grid cell, and the human eye gaze point is rendered and displayed in the user interface currently displayed by the near-eye display device according to the position information of the human eye gaze point.
9. A near-eye display device, characterized in that, The near-eye display device includes a processor, a memory, a computer program stored in the memory and executable by the processor, and a data bus for enabling communication between the processor and the memory, wherein when the computer program is executed by the processor, it implements the steps of the foveated rendering method as described in any one of claims 1 to 8.
10. A storage medium for computer-readable storage, characterized in that, The storage medium stores one or more programs, which can be executed by one or more processors to implement the steps of the foveated rendering method according to any one of claims 1 to 8.