Amblyopia training video display system based on eye tracking

By using an eye-tracking-based amblyopia training video display system, which utilizes an interactive processor and a visual display processor to update visual stimuli in a personalized and real-time manner, the system solves the problem of poor matching between training content and visual state in existing systems, thereby improving training effectiveness and resource utilization efficiency.

CN122363523APending Publication Date: 2026-07-10SHANGHAI RUISHI HEALTH TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI RUISHI HEALTH TECH CO LTD
Filing Date
2026-04-30
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

Existing amblyopia training video display systems are limited by fixed parameter preset modes, cannot respond in real time, and lack eye-tracking-based mechanisms. This results in poor matching between training content and the patient's visual state, visual jumps or stutters during image frame switching, and increases the training cycle, display bandwidth, and image processing resource consumption.

Method used

An eye-tracking-based amblyopia training video display system is adopted. The screen gaze coordinates are obtained through the interactive processor, the gaze area is determined by the central processing unit, and the visual display processor is used for differentiated update processing. Visual stimulus information with preset contrast and spatial frequency is inserted to improve personalization and real-time response capabilities.

Benefits of technology

It improves the personalization and real-time response capabilities of amblyopia training, reduces visual jumps and stuttering during image frame switching, shortens the training cycle, and reduces the consumption of display bandwidth and image processing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122363523A_ABST
    Figure CN122363523A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose an eye movement tracking-based amblyopia training video display system. A specific implementation of the system includes an interaction processor, a central processor, a visual display processor and a user-side screen; and the interaction processor is configured to obtain a screen gaze point coordinate of a target user; the central processor is configured to determine a target gaze point area corresponding to the screen gaze point coordinate; the visual display processor is configured to perform differential update processing on first display content and second display content in the target gaze point area; insert visual stimulation information in the first display content corresponding to a first eye in the target gaze point area, and perform corresponding update processing on the second display content corresponding to a second eye. The implementation can improve the personalization and real-time response capability of amblyopia training, reduce the training period of amblyopia training, and reduce the occupation of display bandwidth and image processing resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments disclosed herein relate to the field of computer technology, and more specifically to an eye-tracking-based amblyopia training video display system. Background Technology

[0002] In clinical practice and home rehabilitation for amblyopia treatment, visual function training, as an active intervention, has long been widely adopted and regarded as a core treatment method following refractive correction and occlusion therapy. Currently, when displaying amblyopia training videos for amblyopia training, the common approach is to use a binocular split-viewing method, where the amblyopic eye watches the training video, while the non-amblyopic eye watches the same content, with the visual input to the non-amblyopic eye undergoing simple blurring. Furthermore, the parameters of the training video (such as brightness, contrast, and frame rate) are fixed.

[0003] However, when using the above methods for visual function training in amblyopia, the following technical problems often arise: Existing amblyopia training video display systems are limited by fixed parameter preset modes, and can only serve as static video display terminals. They require pre-setting parameters such as brightness and contrast, and rely on fixed parameters for playback, which reduces the personalization and real-time response capability of amblyopia training. Furthermore, the system lacks an eye-tracking-based mechanism, making it unable to perform real-time fixation point detection and image frame replacement. This results in poor matching between training content and the patient's actual visual state, and visual jumps or stuttering during image frame switching. Consequently, a longer training cycle is required to compensate for the insufficient training effect caused by fixed display parameters and poor inter-frame fusion. This leads to a longer time required for the recovery of visual function and binocular balance reconstruction in amblyopia, and also consumes a lot of display bandwidth and image processing resources.

[0004] The information disclosed in this background section is only intended to enhance the understanding of the background of the inventive concept, and therefore may contain information that does not form prior art known to those skilled in the art. Summary of the Invention

[0005] The summary section of this disclosure is intended to provide a brief overview of concepts that will be described in detail in the detailed description section below. This summary section is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions. Some embodiments of this disclosure propose an eye-tracking-based amblyopia training system to address one or more of the technical problems mentioned in the background section above.

[0006] Some embodiments of this disclosure provide an eye-tracking-based amblyopia training video display system, including an interaction processor, a central processing unit, a visual display processor, and a user-side screen; the interaction processor is used to acquire the screen gaze coordinates of a target user on the user-side screen; the central processing unit is used to determine a target gaze region corresponding to the screen gaze coordinates; the visual display processor is used to, in response to determining that the amblyopia training method is a first display method, perform differential update processing on the first display content corresponding to the first eye and the second display content corresponding to the second eye in the target gaze region to obtain first video visual information as target video visual information; in response to determining that the amblyopia training method is a second display method, insert visual stimulus information with preset contrast and target spatial frequency into the first display content corresponding to the first eye in the target gaze region, and perform corresponding update processing on the second display content corresponding to the second eye to obtain second video visual information as target video visual information.

[0007] The above embodiments of this disclosure have the following beneficial effects: the eye-tracking-based amblyopia training video display system of some embodiments of this disclosure can improve the personalization and real-time response capability of amblyopia training, reduce visual jumps or stuttering when switching image frames, reduce the training cycle of amblyopia training, and reduce the occupation of display bandwidth and image processing resources. Specifically, the reasons for reduced personalization and real-time responsiveness in amblyopia training, visual jumps or stutters during image frame switching, increased training cycles, and increased consumption of display bandwidth and image processing resources are as follows: Existing amblyopia training video display systems are limited by fixed parameter preset modes, and can only function as static video display terminals. They require pre-setting parameters such as brightness and contrast, and rely on fixed parameters for playback, thus reducing the personalization and real-time responsiveness of amblyopia training. Furthermore, the system lacks an eye-tracking-based mechanism, making real-time fixation detection and image frame replacement impossible. This results in poor matching between training content and the patient's actual visual state, causing visual jumps or stutters during image frame switching. Consequently, longer training cycles are needed to compensate for insufficient training effects caused by fixed display parameters and poor inter-frame fusion, leading to longer time required for amblyopia visual function recovery and binocular balance reconstruction, and increased consumption of display bandwidth and image processing resources. Based on this, some embodiments of the eye-tracking-based amblyopia training video display system disclosed herein include an interactive processor, a central processing unit, a visual display processor, and a user-side screen. The interactive processor is used to acquire the screen fixation coordinates of the target user. Therefore, the screen gaze coordinates of the target user can be obtained. The central processing unit (CPU) is used to determine the target gaze region corresponding to the aforementioned screen gaze coordinates. The visual display processor, firstly, in response to determining that the amblyopia training method is the first display method, performs differential update processing on the first display content corresponding to the first eye and the second display content corresponding to the second eye in the target gaze region, obtaining first video visual information as target video visual information. Thus, the amblyopia training display content corresponding to the first training method can be obtained for amblyopia training. Then, in response to determining that the amblyopia training method is the second display method, it inserts visual stimulus information with preset contrast and target spatial frequency into the first display content corresponding to the first eye in the target gaze region, and performs corresponding update processing on the second display content corresponding to the second eye, obtaining second video visual information as target video visual information. Thus, the amblyopia training display content corresponding to the second training method can be obtained for amblyopia training.Because it doesn't rely on preset fixed parameters for training video display, but instead seamlessly integrates training into high-frequency, highly engaging everyday video content, and smoothly embeds Gabor grating stimuli through frame interval replacement strategies, it improves the smoothness and visual coherence of the training video display. Furthermore, by integrating eye-tracking and binocular split-view display hardware, it reduces visual jumps and stuttering caused by fixed parameters and poor inter-frame fusion during training, thus lowering the overall training time and the consumption of display bandwidth and image processing resources. Therefore, it enhances the personalization and real-time responsiveness of amblyopia training, reduces visual jumps or stuttering during image frame switching, shortens the training cycle for amblyopia training, and reduces the consumption of display bandwidth and image processing resources. Attached Figure Description

[0008] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.

[0009] Figure 1 This is a schematic diagram of the structure of an eye-tracking-based amblyopia training video display system according to some embodiments of this disclosure; Figure 2 This is an exemplary system architecture diagram of some embodiments of an eye-tracking-based amblyopia training video display system applying some embodiments of this disclosure; Figure 3 These are internal test images of a visual impairment training device suitable for implementing some embodiments of the present disclosure; Figure 4 These are internal test images of a visual impairment training device suitable for implementing other embodiments of the present disclosure; Figure 5 These are screenshots of the display content during internal testing and training of the eye-tracking-based amblyopia training video display system according to this disclosure. Detailed Implementation

[0010] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0011] It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.

[0012] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0013] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0014] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0015] Figure 1 The diagram shows structural schematics of some embodiments of the eye-tracking-based amblyopia training video display system that can be applied to this disclosure.

[0016] like Figure 1As shown, the eye-tracking-based amblyopia training video display system provided in this disclosure may include: an interaction processor 101, a central processing unit 102, a vision processor 103, and a user-side screen 104. The interaction processor 101, central processing unit 102, vision processor 103, and user-side screen 104 are communicatively connected. The interaction processor 101 may be a server for acquiring the screen gaze coordinates of the target user on the user-side screen. For example, the interaction processor 101 may be a depth camera module or a Tobii Pro Glasses 3. The central processing unit 102 may be a server for determining the target gaze region corresponding to the screen gaze coordinates. For example, the central processing unit 102 may be a Unity Engine. The aforementioned visual processor 103 can be a server that, in response to determining that the amblyopia training method is the first display method, performs differential update processing on the first display content corresponding to the first eye and the second display content corresponding to the second eye in the target fixation point area to obtain first video visual information as target video visual information; and, in response to determining that the amblyopia training method is the second display method, inserts visual stimulus information with preset contrast and target spatial frequency into the first display content corresponding to the first eye in the target fixation point area, and performs corresponding update processing on the second display content corresponding to the second eye to obtain the target video visual information. For example, the visual processor 103 can be an NVIDIA Maxine. The aforementioned user-side screen 104 can be a screen that can be displayed on the user side in the aforementioned amblyopia training device.

[0017] It should be noted that the aforementioned communication connections may include, but are not limited to, 3G / 4G connections, WiFi connections, Bluetooth connections, WiMAX connections, Zigbee connections, UWB (ultra wideband) connections, and other communication methods that are currently known or will be developed in the future.

[0018] The following is for reference. Figure 2 The diagram shows a timing diagram of an eye-tracking-based amblyopia training video display system disclosed herein.

[0019] like Figure 2 As shown, an eye-tracking-based amblyopia training video display system includes: an interaction processor, a central processing unit, a visual display processor, and a user-side screen. The interaction steps between the interaction processor, central processing unit, visual display processor, and user-side screen may include the following steps: Step 201: The interaction processor is used to obtain the screen gaze coordinates of the target user.

[0020] In some embodiments, the interaction processor can acquire the screen gaze coordinates of the target user. These screen gaze coordinates can represent the position coordinates of the gaze points of the target user's two eyes on the amblyopia training device. The target user can represent an amblyopic user undergoing amblyopia training through the amblyopia training device. Figure 3 , Figure 4 As shown, Figure 3 and Figure 4 This is a physical image of an internal test device used for amblyopia training.

[0021] In addressing the aforementioned technical problems in the application scenario—focusing on gaze point monitoring during amblyopia training using amblyopia training equipment—the following technical problem arises: Traditional eye-tracking methods typically employ only simple threshold segmentation and least-squares ellipse fitting. When there is physiological nystagmus or partial eyelid occlusion, the extracted pupil ellipse parameters become distorted, the conic surface obtained from inverse projection shifts, and the calculated monocular gaze vector fluctuates significantly. This leads to increased errors in the 3D eye center reconstruction and frequent jumps in screen gaze point coordinates, requiring multiple resampling and filtering compensations. This increases the response latency and computational power consumption for real-time monitoring of the target user's gaze point, thus increasing the system's computational resource consumption. Considering the following requirements for this application scenario—adapting to high precision, high real-time requirements, and limited computing power—we have decided to adopt the following solution: In some optional implementations, the interaction processor described above can obtain the screen gaze coordinates of the target user through the following steps: The first step is to acquire a set of binocular eye images captured by an infrared camera. This set of binocular eye images can represent multiple frames of eye images of the left and right eyes, respectively acquired by the infrared camera. The set of binocular eye images can include a set of left-eye eye images and a set of right-eye eye images. The number of eye images included in the multiple frames can be 300 frames. The set of left-eye eye images and the set of right-eye eye images each contain 150 frames. The left-eye eye images in the left-eye eye image set represent the left-eye eye image acquired by the infrared camera. The right-eye eye images in the right-eye eye image set represent the right-eye eye image acquired by the infrared camera. In practice, firstly, the interactive processor can control the infrared camera to continuously acquire the left-eye and right-eye eye images of the target user, obtaining the left-eye eye image set and the right-eye eye image set. Then, the left-eye eye image set and the right-eye eye image set are determined as the binocular eye image set.

[0022] The second step involves performing the following steps for each left-eye image in the left-eye image set included in the aforementioned binoculars image set and each right-eye image in the right-eye image set included in the aforementioned binoculars image set: The first sub-step involves performing pupil contour detection processing on the aforementioned left-eye or right-eye image to obtain pupil ellipse parameters. These pupil ellipse parameters characterize the geometric information describing the pupil contour. The pupil contour characterizes the contour of the pupil in the aforementioned left-eye or right-eye image. The geometric information characterizes the center coordinates, major axis length, minor axis length, and rotation angle of the ellipse's principal axis relative to the horizontal direction of the ellipse (pupil region). In practice, the aforementioned interactive processor first performs image preprocessing on the aforementioned left-eye or right-eye image through the following steps: First, convert the aforementioned left-eye or right-eye image into a grayscale image. Then, perform denoising processing on the grayscale image using Gaussian filtering to obtain a denoised image. Next, enhance the contrast between the pupil region and the background region in the denoised image using gamma transform to obtain a preprocessed image. The background region characterizes the region in the denoised image that is not the pupil region. Secondly, an adaptive threshold segmentation algorithm is used to binarize the preprocessed image to separate the pupil region from the surrounding tissue, resulting in a binarized image. Then, the binarized image undergoes morphological processing through the following steps: First, erosion is used to remove isolated noise points such as eyelashes and light spots. Then, a closing operation (dilation followed by erosion) is used to connect adjacent connected components within the pupil region. Erosion restores the original size of the pupil region to fill the small holes caused by eyelash obstruction or light spot interference. Finally, an opening operation (erosion followed by dilation) is used to remove irrelevant small regions connected to the pupil region, resulting in the processed pupil region. Next, the Canny edge detection algorithm is used to detect the edges of the processed pupil region, obtaining the boundary point set of the pupil region's contour. Then, a least-squares ellipse fitting method is used to calculate the center coordinates, major axis length, minor axis length, and rotation angle of the ellipse's principal axis relative to the horizontal direction based on the boundary point set. Finally, the pupil ellipse parameters are determined by the center coordinates of the ellipse (pupil region), the length of the major axis, the length of the minor axis, and the rotation angle of the principal axis of the ellipse relative to the horizontal direction.

[0023] The second sub-step involves inversely projecting the pupil ellipse parameters to obtain a conical surface with the focal point of the infrared camera as its vertex. This conical surface can represent a cone-shaped surface with the focal point of the infrared camera as its vertex and the pupil ellipse in either the left or right eye image as its base boundary. In practice, the interactive processor first obtains the intrinsic parameters of the infrared camera. These intrinsic parameters can represent the focal length, the coordinates of the principal point of the image, and the pixel size. The pixel size represents the physical size of each pixel on the image sensor of the infrared camera. Then, based on the aforementioned intrinsic parameters, the transformation relationship between the image coordinate system and the camera coordinate system is established through the following steps: First, for any pixel on the image plane, the difference between its x-coordinate and the x-coordinate of the image principal point is divided by the equivalent focal length of the infrared camera in the horizontal direction to obtain the ratio of the pixel's x-coordinate to its depth coordinate in the camera coordinate system; the difference between its y-coordinate and the y-coordinate of the image principal point is divided by the equivalent focal length of the infrared camera in the vertical direction to obtain the ratio of the pixel's y-coordinate to its depth coordinate in the camera coordinate system. The image plane can represent the two-dimensional plane corresponding to the left or right eye image in the aforementioned binocular eye image set. The depth coordinate can represent the focal length of the infrared camera. Next, the pupil ellipse is considered as the intersection of a cone in three-dimensional space and the image plane. Then, with the focal point of the infrared camera as the origin, rays connecting the focal point and each point on the boundary of the pupil ellipse are drawn to obtain a conical surface. The generatrix of the conical surface is all rays originating from the focal point and passing through the boundary points of the pupil ellipse. Finally, by substituting the coordinates obtained in the camera coordinate system after transformation into the general quadratic surface equation of the conic surface, the surface equation corresponding to the conic surface is obtained.

[0024] The third sub-step involves performing coordinate transformation on the aforementioned conic surface to obtain a monocular gaze vector. This monocular gaze vector represents the unit vector pointing from the center of the eyeball to the center of the pupil, i.e., the spatial direction of the target user's monocular gaze. In practice, firstly, the interaction processor performs eigenvalue decomposition on the quadratic matrix corresponding to the surface equation of the conic surface, obtaining three eigenvalues ​​and three corresponding eigenvectors. Then, these three eigenvectors are used as row vectors of a rotation matrix. Next, using the rotation matrix, a coordinate rotation transformation is performed on the surface equation of the conic surface to transform the surface equation in the original coordinate system into a standard conic equation containing only the squared terms of the horizontal, vertical, and depth coordinates, without any cross-product terms. The transformed coordinate system represents a new coordinate system with the principal axes of the conic surface as coordinate axes. Then, in this new coordinate system, a plane equation is defined, and the plane equation is combined with the standard conic equation to obtain the equation of the intersection line between the plane and the conic surface. The above plane equation can be expressed as the dot product of the plane normal vector and the coordinate points on the plane equal to a constant term. The above plane can represent the surface that intersects the above conic surface into a circle. Then, by setting the coefficient of the squared term of the abscissa in the above intersection equation to equal the coefficient of the squared term of the ordinate, we obtain the first equation; and by setting the coefficient of the product term of the abscissa and ordinate in the above intersection equation to zero, we obtain the second equation. Next, we combine the first equation, the second equation, and the third equation to form a system of equations. The third equation can represent the equation where the sum of the squares of the abscissa, ordinate, and depth coordinate components of the plane normal vector equals one. Next, we solve the system of equations to obtain eight numerical solutions corresponding to the above plane normal vector. The above plane normal vector can represent the normal vector of the plane that intersects the above conic surface into a circle. Then, using the above rotation matrix, we perform an inverse transformation on the eight numerical solutions to obtain eight plane normal vector solutions in the original coordinate system. Finally, taking the infrared camera focus as the origin and the screen direction as the positive Z-axis, the plane normal vector solution corresponding to the negative depth coordinate of the eye center and the distance from the eye center to the focus being closest to 40mm is selected from the above eight plane normal vector solutions to obtain the monocular gaze vector.

[0025] The third step involves constructing first constraint lines based on the monocular gaze vectors corresponding to the obtained left-eye image set, and constructing second constraint lines based on the monocular gaze vectors corresponding to the obtained right-eye image set. Each of the first constraint lines represents a spatial line containing the center of the left eye. Each of the second constraint lines represents a spatial line containing the center of the right eye. In practice, firstly, for each left-eye image in the left-eye image set, the interactive processor can determine a spatial line with the pupil center in the left-eye image as its starting point and the opposite direction of the monocular gaze vector corresponding to the left-eye image as its direction as the first constraint line, thus obtaining each first constraint line. Then, for each right-eye image in the right-eye image set, the interactive processor can determine a spatial line with the pupil center in the right-eye image as its starting point and the opposite direction of the monocular gaze vector corresponding to the right-eye image as its direction as the second constraint line, thus obtaining each second constraint line.

[0026] The fourth step involves determining the three-dimensional coordinates of the left eye center based on the aforementioned first constraint lines, and the three-dimensional coordinates of the right eye center based on the aforementioned second constraint lines. The three-dimensional coordinates of the left eye center represent the position of the center of the left eyeball in three-dimensional space. Similarly, the three-dimensional coordinates of the right eye center represent the position of the center of the right eyeball in three-dimensional space. In practice, the interactive processor can first use the least squares method to determine the three-dimensional coordinates of the left eye center by identifying the spatial point closest to each of the aforementioned first constraint lines. Then, it can use the least squares method to determine the three-dimensional coordinates of the right eye center by identifying the spatial point closest to each of the aforementioned second constraint lines.

[0027] Fifth, based on the aforementioned three-dimensional left eye center coordinates, the aforementioned three-dimensional right eye center coordinates, the left eye target monocular gaze vector, and the right eye target monocular gaze vector, determine the screen fixation point coordinates. The aforementioned left eye target monocular gaze vector represents the monocular gaze vector corresponding to the left eye image captured at the current moment. The aforementioned right eye target monocular gaze vector represents the monocular gaze vector corresponding to the right eye image captured at the current moment.

[0028] The above technical solution, as an inventive point of this disclosure, solves technical problem two: "Increased response latency and computational power consumption for real-time monitoring of the target user's gaze point, thereby increasing the computational resources consumed by the system." The reasons for increased response latency and computational power consumption for real-time monitoring of the target user's gaze point, and thus increased computational resources consumed by the system, are as follows: Traditional eye-tracking methods typically only use simple threshold segmentation and least-squares ellipse fitting. When there is physiological nystagmus or partial eyelid occlusion, the extracted pupil ellipse parameters become distorted, the conic surface obtained by inverse projection shifts, and the calculated monocular gaze vector fluctuates significantly. This leads to increased errors in the reconstruction of the three-dimensional eye center and frequent jumps in the screen gaze point coordinates, requiring the system to perform multiple resampling and filtering compensations. This increases the response latency and computational power consumption for real-time monitoring of the target user's gaze point, thereby increasing the computational resources consumed by the system. Solving these factors can reduce the response latency and computational power consumption for real-time monitoring of the target user's gaze point, thereby reducing the computational resources consumed by the system. To achieve this effect, the eye-tracking-based amblyopia training video display system disclosed herein first performs pupil contour detection on images of both eyes to obtain pupil ellipse parameters. Then, through inverse projection and coordinate transformation, the monocular gaze vector is calculated. Next, constraint lines are constructed using the gaze vectors corresponding to multiple frames, and the intersection of these constraint lines is determined as the three-dimensional eye center coordinates. Finally, based on the three-dimensional eye center coordinates and the gaze vector at the current moment, the screen gaze point coordinates are determined. This reduces the response latency and computational power consumption of real-time monitoring of the target user's gaze point, thereby reducing the computational resources consumed by the system.

[0029] In addressing the technical problems mentioned above, the application scenario—specifically, amblyopia training using amblyopia training equipment—often presents the following technical problem: unstable binocular fixation, eye deviation (such as esotropia or exotropia), or asymmetric binocular visual function in target users with binocular visual asymmetry. This leads to inconsistencies in the monocular gaze vectors calculated by the left and right eyes, and the inability to effectively fuse the binocular gaze equations in the same world coordinate system. Consequently, the calculated 3D gaze point fluctuates significantly, and the screen gaze point coordinates frequently change. This necessitates multiple resampling, filtering compensation, and repeated calibrations, increasing the response latency and computational power consumption for real-time monitoring and gaze point region determination, thus increasing the system's computational resource consumption. Considering the following requirements for this application scenario: consistent fusion of binocular gaze vectors and high-precision screen gaze point coordinate transformation, we have decided to adopt the following solution: In some optional implementations of certain embodiments, the above-mentioned interaction processor can be configured as follows: The first step involves correcting the aforementioned three-dimensional left and right eye center coordinates to obtain the target left and right eye center coordinates. The target left eye center coordinates represent the true position coordinates of the left eyeball center after corneal refraction correction. Similarly, the target right eye center coordinates represent the true position coordinates of the right eyeball center after corneal refraction correction. In practice, the interactive processor first obtains the pre-constructed correction polynomial through the following steps: First, using an open-source three-dimensional pseudo-eye model, the pseudo-eye model is controlled to change the eyeball center position and gaze direction within a preset range, while an infrared camera acquires images of the pseudo-eye model. The specific range of the preset range is not limited. For example, the preset range for the eyeball center position is -15mm to +15mm, and the preset range for the gaze direction can be -30° to +30° horizontally and -20° to +20° vertically. The aforementioned pseudo-eye model is used to simulate the human eyeball center, pupil position, and the optical characteristics of corneal refraction. Next, based on the acquired images, estimated values ​​of the three-dimensional left and right eye center coordinates are calculated, and the actual values ​​of the three-dimensional left and right eye center coordinates are recorded. It should be noted that the method for calculating the estimated values ​​of the three-dimensional left and right eye center coordinates based on the acquired images is the same as the method for determining the three-dimensional left and right eye center coordinates based on a set of binocular eye images; see step 101 for details. Then, polynomial regression analysis is performed on the obtained sets of estimated and actual values: First, three correction polynomials are established for the horizontal, vertical, and depth coordinates, respectively. Each of the three correction polynomials can be expressed as the linear term of the coordinate value to be corrected multiplied by the first coefficient, plus the square term of the coordinate value to be corrected multiplied by the second coefficient, plus the cube term of the coordinate value to be corrected multiplied by the third coefficient, plus a constant term. The coordinate values ​​to be corrected can represent the horizontal, vertical, or depth coordinates of the obtained estimates. Then, through regression analysis, the values ​​of the first, second, and third coefficients and the constant term in the three correction polynomials are determined. Finally, the x-coordinate, y-coordinate, and depth coordinates of the three-dimensional left eye center coordinates are substituted into the corresponding correction polynomials to calculate the target left eye center coordinates. Similarly, the x-coordinate, y-coordinate, and depth coordinates of the three-dimensional right eye center coordinates are substituted into the corresponding correction polynomials to calculate the target right eye center coordinates.

[0030] The second step involves combining the target's left-eye center coordinates and the target's left-eye gaze vector into a left-eye line-of-sight equation, and combining the target's right-eye center coordinates and the target's right-eye gaze vector into a right-eye line-of-sight equation. In practice, the executing entity can use a point-direction equation of a spatial straight line, taking the target's left-eye center coordinates as the starting point of the left-eye line-of-sight and the target's left-eye gaze vector as the direction vector of the left-eye line-of-sight, to obtain the left-eye line-of-sight equation. Then, using a point-direction equation of a spatial straight line, the target's right-eye center coordinates are taken as the starting point of the right-eye line-of-sight, and the target's right-eye gaze vector is taken as the direction vector of the right-eye line-of-sight, to obtain the right-eye line-of-sight equation.

[0031] The third step involves transforming the aforementioned left-eye and right-eye line-of-sight equations to the same world coordinate system, resulting in the target's left-eye and right-eye line-of-sight equations. The target's left-eye line-of-sight equation represents the equation transformed to the aforementioned world coordinate system. Similarly, the target's right-eye line-of-sight equation represents the equation transformed to the aforementioned world coordinate system. In practice, the interactive processor can obtain the translation vector and rotation matrix of the infrared camera in the world coordinate system from the product design parameters of the infrared camera. Through rigid body transformation, it can transform the aforementioned left-eye and right-eye line-of-sight equations to the same world coordinate system, thus obtaining the target's left-eye and right-eye line-of-sight equations.

[0032] The fourth step involves determining the first target point corresponding to the left-eye gaze equation and the second target point corresponding to the right-eye gaze equation, based on the aforementioned left-eye gaze equation and right-eye gaze equation. The first target point can be represented as the point on the line corresponding to the left-eye gaze equation that is closest to the right-eye gaze equation. Similarly, the second target point can be represented as the point on the line corresponding to the right-eye gaze equation that is closest to the left-eye gaze equation. In practice, the interaction processor can use the least squares method to obtain the first and second target points based on the left-eye and right-eye gaze equations.

[0033] The fifth step is to determine the three-dimensional gaze point based on the first and second target points mentioned above. This three-dimensional gaze point can represent the intersection or closest point of the target user's gaze in three-dimensional space, i.e., the three-dimensional spatial position where the target user is currently looking. In practice, the interaction processor can determine the midpoint between the first and second target points as the three-dimensional gaze point.

[0034] Step 6: Perform depth normalization on the aforementioned 3D gaze point to obtain normalized gaze point coordinates. These normalized gaze point coordinates represent the projected coordinates of the 3D gaze point on a plane with a depth of 1. In practice, firstly, the interactive processor divides the x-axis and y-axis coordinates of the 3D gaze point by the z-axis coordinate, respectively, to obtain the first target coordinate value and the second target coordinate value. Then, the first target coordinate value is determined as the x-axis coordinate value, and the second target coordinate value is determined as the y-axis coordinate value, resulting in the normalized gaze point coordinates. For example, if the spatial coordinates of the 3D gaze point are (x, y, z), then the normalized gaze point coordinates can be (x / z, y / z).

[0035] Step 7: Based on the obtained screen pixel coordinates of each preset calibration point and the corresponding depth-normalized gaze point coordinates, determine the coordinate mapping relationship between the screen pixel coordinates and the normalized gaze point coordinates. Here, each preset calibration point can represent a pre-defined 3D coordinate point. Each preset calibration point can correspond to 20 frames of binocular eye images. The depth-normalized gaze point coordinates can represent the normalized gaze point coordinates of the acquired binocular eye images corresponding to the preset calibration points. The coordinate mapping relationship represents the mapping relationship between the screen pixel coordinates and the normalized gaze point coordinates. In practice, the interaction processor can establish the mapping relationship between the screen pixel coordinates of each preset calibration point and the corresponding depth-normalized gaze point coordinates through quadratic polynomial regression. The quadratic polynomial regression can take the form of a complete quadratic polynomial.

[0036] Step 8: Based on the coordinate mapping relationship described above, transform the normalized gaze point coordinates to screen coordinates to obtain the screen gaze point coordinates. In practice, the interaction processor can use the normalized gaze point coordinates as input variables and substitute them into the coordinate mapping relationship to obtain the screen gaze point coordinates.

[0037] The above-described technical solution, as an inventive point of this disclosure, solves technical problem three: "The large fluctuations in the calculated three-dimensional gaze point and the frequent jumps in the screen gaze point coordinates lead to increased response latency and computational power consumption in real-time monitoring and gaze point region determination of the target user, thereby increasing the computational resources consumed by the system." The reasons for the large fluctuations in the calculated three-dimensional gaze point and the frequent jumps in the screen gaze point coordinates, leading to increased response latency and computational power consumption in real-time monitoring and gaze point region determination of the target user, and thus increasing the computational resources consumed by the system, are as follows: Unstable binocular gaze of the target user, eye deviation (such as esotropia or exotropia), or asymmetry in binocular visual function can cause inconsistencies in the monocular gaze vectors calculated by the left and right eyes, and the inability to effectively fuse the binocular gaze equations in the same world coordinate system. This results in large fluctuations in the calculated three-dimensional gaze point and frequent jumps in the screen gaze point coordinates, requiring the system to perform multiple resampling, filtering compensation, and repeated calibrations, thus increasing the response latency and computational power consumption in real-time monitoring and gaze point region determination of the target user, and consequently increasing the computational resources consumed by the system. If the above factors are addressed, the large fluctuations in the calculated 3D gaze point and frequent jumps in the screen gaze point coordinates can be reduced, thus reducing the response latency and computational power consumption for real-time monitoring and gaze point region determination of the target user, and consequently reducing the computational resources consumed by the system. To achieve this effect, the eye-tracking-based amblyopia training video display system disclosed herein first corrects the 3D left-eye center coordinates and 3D right-eye center coordinates to obtain the target eyeball center coordinates. Next, the target eyeball center coordinates of both eyes are combined with their respective monocular gaze vectors to form a binocular gaze equation, which is then transformed to the same world coordinate system. Then, based on the binocular gaze equation in the same world coordinate system, the 3D gaze point is calculated. Next, normalized gaze point coordinates are obtained through depth normalization. Finally, using a coordinate mapping relationship established by preset calibration points, the normalized gaze point coordinates are accurately transformed to the screen coordinates to obtain the screen gaze point coordinates. Therefore, through binocular gaze fusion calibration and depth normalization mapping, gaze point calculation deviations and jumps are reduced, lowering response latency and computational power consumption. This reduces the occurrence of large fluctuations in the calculated 3D gaze point and frequent changes in the screen gaze point coordinates, reduces the response delay and computational power consumption of real-time monitoring of the target user's gaze point and determination of the gaze point region, and thus reduces the computational resources consumed by the system.

[0038] Step 202: The central processing unit is used to determine the target gaze point region corresponding to the screen gaze point coordinates based on the screen gaze point coordinates.

[0039] In some embodiments, the central processing unit (CPU) can determine the target gaze point region corresponding to the screen gaze point coordinates. The target gaze point region can represent the area within a target range centered on the screen gaze point coordinates. The specific range of the target range is not limited here. For example, the target range can be 0 to 3 degrees. In practice, the CPU can determine the target gaze point region corresponding to the screen gaze point coordinates using a geometric mapping algorithm (which determines the radius of the target gaze point region by multiplying the tangent of the maximum angle corresponding to the target range by the distance from the user's eye to the user-side screen). The target gaze point region can be a circular region. Figure 5 As shown, Figure 5 This is a screenshot taken from the user's side screen during internal testing and training.

[0040] Step 203: In response to determining that the amblyopia training mode is the first display mode, the visual display processor performs differential update processing on the first display content corresponding to the first eye and the second display content corresponding to the second eye in the target fixation point area to obtain the first video visual information as the target video visual information.

[0041] In some embodiments, in response to determining the amblyopia training method as the first display method, the visual display processor can perform differential update processing on the first display content corresponding to the first eye and the second display content corresponding to the second eye in the target fixation point region to obtain first video visual information as target video visual information. The first display method can represent a training method achieved by adjusting visual information. The visual information can represent brightness and / or contrast information in the first or second display content. The brightness information can represent the brightness of the first or second display content. The contrast information can represent the contrast of the first or second display content. The first eye can represent the amblyopic eye of the target user. The first display content can represent the content viewed by the amblyopic eye. The second eye can represent the non-amblyopic eye of the target user. The second display content can represent the content viewed by the non-amblyopic eye. The first video visual information can represent the visual content obtained after updating the visual information of the first and second display content.

[0042] In some alternative implementations of some embodiments, the above-described visual display processor can be configured to: First, in response to detecting the first adjustment operation by the target user, the visual information of the second display content is adjusted to obtain modified visual information and modified second display content. The first adjustment operation can represent the target user's operation of increasing or decreasing the visual information. In practice, in response to determining that the first adjustment operation is an operation to increase the visual information, the executing entity increases the current visual information (brightness or contrast) by a preset step size to obtain modified visual information and modified second display content. In response to determining that the first adjustment operation is an operation to decrease the visual information, the executing entity can decrease the current visual information by a preset step size to obtain modified visual information and modified second display content. Here, the specific adjustment range of the preset step size is not limited. For example, the preset step size can be five percent of the brightness or contrast value.

[0043] Then, in response to determining that the modified visual information and the visual information of the first display content are the same, the modified second display content and the first display content are determined as the first video visual information.

[0044] In addressing the aforementioned technical problems in the application scenario—specifically, the training assistance scenario where users with different binocular vision functions view content using amblyopia training equipment, exhibiting non-uniform differences in their perception capabilities at different spatial frequencies (e.g., smaller differences in low-frequency contour perception and larger differences in high-frequency detail perception)—often presents the following technical problem: Using traditional global contrast adjustment methods, when the user's eyes perceive different spatial frequencies inconsistently, adjusting contrast according to high-frequency requirements leads to excessive weakening of low-frequency components, making binocular fusion difficult; conversely, adjusting contrast according to low-frequency requirements results in the eye with stronger high-frequency detail perception still dominating, while the high-frequency channels corresponding to the weaker eye cannot be effectively activated. This results in difficulty in binocular coordination during content viewing, with high-frequency visual channels consistently suppressed, increasing the time and energy consumption for adjusting the user's binocular coordination state, and consequently increasing the system's computational resource consumption. Considering the following requirements for this application scenario—adapting to high precision, high real-time requirements, and limited computing power—we have decided to adopt the following solution: In some alternative implementations of some embodiments, the above-described visual display processor can be configured to: The first step involves, in response to the detection of the target user's second adjustment operation, constructing a raw measurement dataset containing the correspondence between spatial frequency, orientation, and contrast values ​​based on a pre-defined standard test stimulus library with multiple different spatial frequencies and orientations. The second adjustment operation can represent an operation that directly adjusts the contrast of the second displayed content. The standard test stimulus library can represent a pre-constructed set of test stimulus patterns containing multiple different spatial frequencies and orientations. The test stimulus patterns can be Gabor gratings or sinusoidal grating patterns. The raw measurement data in the raw measurement dataset can represent the correspondence between the contrast value of the test stimulus pattern presented by the second eye and the spatial frequency and orientation when the target user achieves binocular coherence through the second adjustment operation. In practice, the visual display processor can first sequentially retrieve test stimulus patterns with different spatial frequencies and orientations from the standard test stimulus library. Each test stimulus pattern is presented in a binocular split-view manner, where the first eye presents a reference pattern with a fixed contrast, and the second eye presents a test stimulus pattern with the contrast to be adjusted. Then, in response to determining that the contrast of the adjusted test stimulus pattern perceived by the second eye is consistent with the contrast of the reference pattern perceived by the first eye, the contrast value of the adjusted test stimulus pattern is recorded as the contrast value at that spatial frequency and orientation. Next, the spatial frequency of the test stimulus pattern, the orientation corresponding to the test stimulus pattern, and the contrast value corresponding to the test stimulus pattern are recorded as discrete data points. Finally, the obtained discrete data points are determined as the original measurement dataset.

[0045] The second step involves processing the discrete data points in the original measurement dataset to create a continuous curve, with spatial frequency as the independent variable and contrast ratio as the dependent variable. This personalized attenuation curve characterizes the degree of contrast attenuation relative to standard perceptual ability in the non-affected eye of the target user at different spatial frequencies. The function curve can be plotted with spatial frequency on the x-axis and contrast ratio on the y-axis. In practice, firstly, the visual display processor sorts the discrete data points in the original measurement dataset in ascending order of spatial frequency. Then, an exponential function fitting method is used, and the parameters of the exponential function are determined using the least squares method to minimize the sum of squared errors between the fitted continuous curve and each discrete data point, thus fitting the discrete data points into a continuous curve. Finally, the fitted continuous curve is determined as the personalized attenuation curve.

[0046] Third, for each frame of image included in the second display content mentioned above, perform the following steps: The first sub-step involves performing multi-scale, multi-directional wavelet transform decomposition on the image to obtain sub-band images. These sub-band images represent the high-frequency or low-frequency components obtained after wavelet transform decomposition. In practice, the visual display processor can perform a three-level wavelet transform decomposition on the image using wavelet basis functions to obtain multiple low-frequency or high-frequency sub-bands as sub-band images. Specifically, the first level decomposition divides the image into a first-level low-frequency sub-band, a horizontal high-frequency sub-band, a vertical high-frequency sub-band, and a diagonal high-frequency sub-band; the second level decomposition divides the first-level low-frequency sub-band into a second-level low-frequency sub-band, a horizontal high-frequency sub-band, a vertical high-frequency sub-band, and a diagonal high-frequency sub-band; and the third level decomposition divides the second-level low-frequency sub-band into a third-level low-frequency sub-band, a horizontal high-frequency sub-band, a vertical high-frequency sub-band, and a diagonal high-frequency sub-band. The wavelet basis functions can be Daubechies wavelets, Symlets wavelets, or Biorthogonal wavelets.

[0047] The second sub-step involves performing the following steps for each sub-band image included in the aforementioned sub-band images: First, based on the personalized attenuation curve, the contrast attenuation coefficient corresponding to the center spatial frequency of the sub-band image is determined. The center spatial frequency represents the center value of the spatial frequency range covered by the sub-band image. The center spatial frequency can be used to represent the main frequency components corresponding to the sub-band image. The contrast attenuation coefficient represents the factor used to scale the pixel amplitude of the sub-band image. In practice, firstly, the visual display processor can use the scal2frq function in MATLAB to determine the center spatial frequency corresponding to the current sub-band image based on the wavelet basis function and the number of layers in the wavelet transform decomposition. Next, the center spatial frequency is substituted as the independent variable into the personalized attenuation curve to obtain the contrast value corresponding to the center spatial frequency. Then, the ratio between the contrast value and the standard contrast value is determined as the contrast attenuation coefficient. The standard contrast value can represent a preset reference contrast value.

[0048] Then, based on the aforementioned contrast attenuation coefficient, the amplitude of each pixel in the sub-band image is linearly scaled to obtain the contrast-attenuated eye-friendly frequency band component. The pixel amplitude among these individual pixel amplitudes can represent the amplitude of the pixels included in the sub-band image. The eye-friendly frequency band component can represent the image after linearly scaling the pixel amplitudes in the sub-band image. In practice, firstly, for the pixel amplitude of each pixel in the sub-band image, the visual display processor can determine the scaled pixel amplitude by multiplying the contrast attenuation coefficient by the pixel amplitude. Then, the obtained scaled pixel amplitudes replace the pixel amplitudes of the corresponding pixels in the sub-band image to obtain an updated sub-band image. Finally, the updated sub-band image is determined as the contrast-attenuated eye-friendly frequency band component.

[0049] The third sub-step involves performing inverse wavelet transform reconstruction on the obtained healthy eye frequency band components to obtain the reconstructed healthy eye image. This reconstructed healthy eye image characterizes the image reconstructed by inverse wavelet transform on the obtained healthy eye frequency band components. In practice, firstly, the aforementioned visual display processor performs inverse wavelet transform on the healthy eye frequency band components corresponding to the three high-frequency and low-frequency sub-bands of the third layer to obtain the low-frequency sub-bands of the second layer. Next, it performs inverse wavelet transform on the healthy eye frequency band components corresponding to the three high-frequency sub-bands of the second layer and the reconstructed low-frequency sub-bands of the second layer to obtain the low-frequency sub-bands of the first layer. Finally, it performs inverse wavelet transform on the healthy eye frequency band components corresponding to the three high-frequency sub-bands of the first layer and the reconstructed low-frequency sub-bands of the first layer to obtain the reconstructed healthy eye image.

[0050] The fourth step is to replace the images of the healthy eye that have been reconstructed with the images of the corresponding timestamps in the second display content in chronological order, so as to obtain the modified second display content.

[0051] The fifth step is to determine the modified second display content and the modified first display content as the first video visual information.

[0052] The above-described technical solution, as an inventive point of this disclosure, solves technical problem four: "Increasing the operation time and energy consumption for adjusting the user's binocular coordination state, thereby increasing the computational resources consumed by the system." The reasons for this increased operation time and energy consumption for adjusting the user's binocular coordination state, and consequently the increased computational resources consumed by the system, are as follows: Using traditional global contrast adjustment methods, when the user's eyes perceive differences inconsistently at different spatial frequencies, adjusting the contrast according to high-frequency requirements leads to excessive attenuation of low-frequency components, making it difficult for the eyes to fuse. Conversely, adjusting the contrast according to low-frequency requirements allows the eye with stronger high-frequency detail perception to still dominate, while the high-frequency channel corresponding to the weaker eye cannot be effectively activated. This results in difficulty for the eyes to coordinate during binocular viewing, and the high-frequency visual channel is constantly suppressed, leading to increased operation time and energy consumption for adjusting the user's binocular coordination state, and consequently, increased computational resources consumed by the system. Solving these factors can reduce the operation time and energy consumption for adjusting the user's binocular coordination state, thereby reducing the computational resources consumed by the system. To achieve this effect, the eye-tracking-based amblyopia training video display system disclosed herein decomposes each image corresponding to the non-amblyopic eye into sub-bands of different spatial frequencies using wavelet transform. Based on the user's daily measured personalized attenuation curve, the contrast of each sub-band is scaled separately before the image is reconstructed and displayed, thus achieving precise contrast adjustment across frequency bands. This reduces the time and energy consumption required to adjust the user's binocular coordination, thereby reducing the system's computational resources.

[0053] Optionally, before step 204, the interactive processor may acquire amblyopia severity information of the target user. This amblyopia severity information characterizes the degree of amblyopia in the target user's eyes. The amblyopia severity information can be mild, moderate, or severe amblyopia. Mild amblyopia can indicate that the target user's best corrected visual acuity is between 0.8 and 0.6. Moderate amblyopia can indicate that the target user's best corrected visual acuity is between 0.2 and 0.5. Severe amblyopia can indicate that the target user's best corrected visual acuity is less than or equal to 0.1. All the specific values ​​for the best corrected visual acuity are decimal visual acuity values.

[0054] Optionally, before step 204, the interaction processor can first display a preset test visual information set to the first display content of the target user. The preset test visual information set includes preset test visual information with different spatial frequencies. The preset test visual information in the preset test visual information set can represent pre-set visual patterns used for testing. The preset test visual information can be a Gabor grating.

[0055] Then, based on each preset test visual information included in the aforementioned preset test visual information set, a contrast threshold corresponding to the preset test visual information is determined, resulting in a contrast threshold set. In practice, firstly, for each preset test visual information included in the aforementioned preset test visual information set, the interaction processor can adjust the contrast of the corresponding preset test visual information to determine the lowest contrast that the target user can recognize, thus obtaining a contrast threshold. Then, the obtained contrast thresholds are defined as a contrast threshold set.

[0056] Next, based on the aforementioned preset test visual information set and the aforementioned contrast threshold set, a contrast sensitivity function is generated. This contrast sensitivity function characterizes the mapping relationship between contrast thresholds and their corresponding spatial frequencies within the aforementioned contrast threshold set. In practice, the aforementioned interactive processor can plot the contrast sensitivity function by using the spatial frequencies corresponding to the aforementioned test visual information set as the abscissa and the set of reciprocals of the contrast thresholds (i.e., contrast sensitivity) corresponding to the aforementioned test visual information set as the ordinate, employing a log-normal distribution function.

[0057] Finally, based on the aforementioned contrast sensitivity function and preset contrast threshold, the target spatial frequency is determined. The preset contrast threshold can represent a pre-defined contrast threshold. For example, the contrast threshold could be 0.5. The target spatial frequency can represent the spatial frequency obtained by inputting the reciprocal of the preset contrast threshold into the aforementioned contrast sensitivity function. In practice, firstly, the aforementioned interaction processor can determine the preset contrast sensitivity as the reciprocal of the preset contrast threshold. Then, the preset contrast sensitivity is input into the aforementioned contrast sensitivity function to obtain the target spatial frequency.

[0058] Step 204: In response to determining that the amblyopia training method is the second display method, visual stimulus information with preset contrast and target spatial frequency is inserted into the first display content corresponding to the first eye in the target fixation point area, and the second display content corresponding to the second eye is updated accordingly to obtain the second video visual information as the target video visual information.

[0059] In some embodiments, in response to determining that the aforementioned amblyopia training method is the second display method, the visual display processor can insert visual stimulus information with a preset contrast and target spatial frequency into the first display content corresponding to the first eye in the target fixation point region, and perform corresponding update processing on the second display content corresponding to the second eye to obtain second video visual information as target video visual information. The aforementioned second display method can represent a method of amblyopia training by inserting visual stimulus information. The aforementioned visual stimulus information can represent a Gabor grating. The aforementioned preset contrast can represent the contrast of the aforementioned pre-set visual stimulus information. Here, the specific value of the aforementioned preset contrast is not limited and can be adjusted according to actual needs.

[0060] In some optional implementations of certain embodiments, in response to determining that the amblyopia training method is the second display method, the visual display processor may be further configured to perform the following steps: inserting visual stimulus information with preset contrast and target spatial frequency into the first display content corresponding to the first eye in the target fixation point region, and performing corresponding update processing on the second display content corresponding to the second eye to obtain second video visual information as target video visual information: The first step involves, in response to determining that the target user's amblyopia level is mild, inserting the visual stimulus information into the first display content at a first preset density, and inserting preset blank information into the second display content at the same first preset density, to obtain second video visual information. The first preset density represents the frequency at which the visual stimulus information is inserted into the first display content. This first preset density can be achieved by replacing or adding the visual stimulus information to a frame in the first display content within a first preset period. The preset blank information can represent a blank frame. The specific number of frames in the first preset period is not limited. For example, the first preset period can be 9 frames. In practice, the visual display processor can first insert the visual stimulus information into the first display content at the first preset density, or replace the visual stimulus information in a frame of the first display content, and then insert the preset blank information into the second display content at the first preset density, or replace the preset blank information in a frame of the second display content, to obtain second video visual information. For example, every 9 frames, one frame of visual stimulus information and one frame of preset blank information are replaced (or added). It should be noted that the frames selected for inserting or replacing the first displayed content and the second displayed content are at the same time.

[0061] The second step involves, in response to determining that the target user's amblyopia level is moderate, inserting the visual stimulus information into the first display content at a second preset density, and inserting the preset blank information into the second display content at the same second preset density, to obtain second video visual information. The second preset density characterizes the frequency at which the visual stimulus information is inserted into the first display content. This second preset density can be achieved by replacing or adding the visual stimulus information to a frame of the first display content within a second preset period. The specific number of frames in the second preset period is not limited. For example, the second preset period can be 7. In practice, the visual display processor can first insert the visual stimulus information into the first display content at the second preset density, or replace the visual stimulus information in a frame of the first display content, and then insert the preset blank information into the second display content at the second preset density, or replace the preset blank information in a frame of the second display content, to obtain second video visual information. For example, every 7 frames, one frame of visual stimulus information and one frame of preset blank information are replaced (or added).

[0062] Thirdly, in response to determining that the target user's amblyopia level is severe, the visual stimulus information is inserted into the first display content at a third preset density, and the preset blank information is inserted into the second display content at the same third preset density, to obtain second video visual information. The third preset density characterizes the frequency at which the visual stimulus information is inserted into the first display content. This third preset density can be a time interval during which the visual stimulus information is replaced or added to a frame of the first display content. The specific number of frames in the third preset interval is not limited. For example, the third preset interval can be 5. In practice, the visual display processor can first insert the visual stimulus information into the first display content at the third preset density, or replace a frame of the first display content with the visual stimulus information, and then insert the preset blank information into the second display content at the third preset density, or replace a frame of the second display content with the preset blank information, to obtain second video visual information. For example, every 5 frames, 1 frame of visual stimulus information and 1 frame of preset blank information are replaced (or added).

[0063] Optionally, after step 204, the visual display processor may first superimpose a preset first defocus amount onto the second display content corresponding to the second eye in the target gaze area of ​​the user-side screen, and superimpose a preset second defocus amount onto the first display content corresponding to the first eye in the target gaze area of ​​the user-side screen. The preset first defocus amount can represent a pre-set, positive defocus parameter. The specific value of the preset first defocus amount is not limited. For example, the preset first defocus amount can be 0.25D. The preset second defocus amount can represent a pre-set, negative defocus parameter. The specific value of the preset second defocus amount is not limited. For example, the preset second defocus amount can be -0.1D. In practice, the visual display processor can use convolutional blurring to superimpose the second display content and the first display content according to the preset first defocus amount and the preset second defocus amount, respectively.

[0064] The above embodiments of this disclosure have the following beneficial effects: the eye-tracking-based amblyopia training video display system of some embodiments of this disclosure can improve the personalization and real-time response capability of amblyopia training, reduce visual jumps or stuttering when switching image frames, reduce the training cycle of amblyopia training, and reduce the occupation of display bandwidth and image processing resources. Specifically, the reasons for reduced personalization and real-time responsiveness in amblyopia training, visual jumps or stutters during image frame switching, increased training cycles, and increased consumption of display bandwidth and image processing resources are as follows: Existing amblyopia training video display systems are limited by fixed parameter preset modes, and can only function as static video display terminals. They require pre-setting parameters such as brightness and contrast, and rely on fixed parameters for playback, thus reducing the personalization and real-time responsiveness of amblyopia training. Furthermore, the system lacks an eye-tracking-based mechanism, making real-time fixation detection and image frame replacement impossible. This results in poor matching between training content and the patient's actual visual state, causing visual jumps or stutters during image frame switching. Consequently, longer training cycles are needed to compensate for insufficient training effects caused by fixed display parameters and poor inter-frame fusion, leading to longer time required for amblyopia visual function recovery and binocular balance reconstruction, and increased consumption of display bandwidth and image processing resources. Based on this, some embodiments of the eye-tracking-based amblyopia training video display system disclosed herein include an interactive processor, a central processing unit, a visual display processor, and a user-side screen. The interactive processor is used to acquire the screen fixation coordinates of the target user. Therefore, the screen gaze coordinates of the target user can be obtained. The central processing unit (CPU) is used to determine the target gaze region corresponding to the aforementioned screen gaze coordinates. The visual display processor, firstly, in response to determining that the amblyopia training method is the first display method, performs differential update processing on the first display content corresponding to the first eye and the second display content corresponding to the second eye in the target gaze region, obtaining first video visual information as target video visual information. Thus, the amblyopia training display content corresponding to the first training method can be obtained for amblyopia training. Then, in response to determining that the amblyopia training method is the second display method, it inserts visual stimulus information with preset contrast and target spatial frequency into the first display content corresponding to the first eye in the target gaze region, and performs corresponding update processing on the second display content corresponding to the second eye, obtaining second video visual information as target video visual information. Thus, the amblyopia training display content corresponding to the second training method can be obtained for amblyopia training.Because it doesn't rely on preset fixed parameters for training video display, but instead seamlessly integrates training into high-frequency, highly engaging everyday video content, and smoothly embeds Gabor grating stimuli through frame interval replacement strategies, it improves the smoothness and visual coherence of the training video display. Furthermore, by integrating eye-tracking and binocular split-view display hardware, it reduces visual jumps and stuttering caused by fixed parameters and poor inter-frame fusion during training, thus lowering the overall training time and the consumption of display bandwidth and image processing resources. Therefore, it enhances the personalization and real-time responsiveness of amblyopia training, reduces visual jumps or stuttering during image frame switching, shortens the training cycle for amblyopia training, and reduces the consumption of display bandwidth and image processing resources.

[0065] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0066] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0067] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.

[0068] The above description is merely a selection of preferred embodiments of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in the embodiments of this disclosure.

Claims

1. An eye-tracking-based amblyopia training video display system, wherein, The amblyopia training video display system includes: an interactive processor, a central processing unit, a visual display processor, and a user-side screen that are interconnected. The interaction processor is used to obtain the screen gaze point coordinates of the target user on the user-side screen; The central processing unit is used to determine the target gaze point region corresponding to the screen gaze point coordinates based on the screen gaze point coordinates. The visual display processor is configured to, in response to determining that the amblyopia training mode is the first display mode, perform differential update processing on the first display content corresponding to the first eye and the second display content corresponding to the second eye in the target fixation point area to obtain first video visual information as target video visual information; and in response to determining that the amblyopia training mode is the second display mode, insert visual stimulus information with preset contrast and target spatial frequency into the first display content corresponding to the first eye in the target fixation point area, and perform corresponding update processing on the second display content corresponding to the second eye to obtain second video visual information as target video visual information.

2. The system according to claim 1, wherein, The interaction processor is configured to: Obtain the amblyopia severity result information of the target user, wherein the amblyopia severity result information is mild amblyopia, moderate amblyopia, or severe amblyopia.

3. The system according to claim 1, wherein, The visual display processor is configured to: In response to detecting a first adjustment operation by the target user, the visual information of the second displayed content is adjusted to obtain the changed visual information and the changed second displayed content; In response to determining that the modified visual information and the visual information of the first displayed content are the same, the modified second displayed content and the first displayed content are determined as the first video visual information.

4. The system according to claim 1, wherein, The central processing unit is configured to: A preset set of test visual information is displayed on the first display content, wherein the spatial frequencies of each preset test visual information included in the preset set of test visual information are different; Based on each preset test visual information included in the preset test visual information set, a contrast threshold corresponding to the preset test visual information is determined to obtain a contrast threshold set. The contrast sensitivity function is determined based on the preset test visual information set and the contrast threshold set. The target spatial frequency is determined based on the contrast sensitivity function and the preset contrast threshold.

5. The system according to claim 2, wherein, The visual display processor is configured to: In response to determining that the target user's amblyopia level is mild, the visual stimulus information is inserted into the first display content at a first preset density, and the preset blank information is inserted into the second display content at the first preset density to obtain the second video visual information; In response to determining that the target user's amblyopia level is moderate, the visual stimulus information is inserted into the first display content at a second preset density, and the preset blank information is inserted into the second display content at the second preset density to obtain the second video visual information; In response to determining that the target user's amblyopia level is severe, the visual stimulus information is inserted into the first display content at a third preset density, and the preset blank information is inserted into the second display content at the third preset density to obtain the second video visual information.

6. The system according to claim 1, wherein, The visual display processor is configured to: A preset first defocus amount is superimposed onto the second display content corresponding to the second eye in the target gaze point area of ​​the user's side screen, and a preset second defocus amount is superimposed onto the first display content corresponding to the first eye in the target gaze point area of ​​the user's side screen.