Vision training method and visual function recognition method and storage medium, device and system
Through visual function diagnosis and treatment equipment, using eye movement information and pupil waveform diagrams, the abnormal visual function is accurately diagnosed and personalized to treat abnormal visual function, which solves the problems of inaccurate diagnosis and untimely monitoring of treatment effects in the prior art, and achieves efficient and personalized visual function treatment.
Patent Information
- Application Number
- CN202411443357.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-16
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-10-16
AI Technical Summary
The prior art has the problem of diagnosing and treating abnormal visual functions, especially amblyopia and strabismus, which have strong subjectivity, difficulty in accurately assessing binocular visual function asymmetry and gaze preferences, and untimely monitoring of treatment effects.
It provides a visual function diagnosis and treatment device, which obtains eye movement information through the camera, uses pre-trained object detection network and pupil segmentation model to track the gaze coordinates and pupil waveform diagram of the user's eyes, determines the visual function information, and plays personalized preset objects on the display for visual training.
Accurate diagnosis and personalized treatment of abnormal visual function are achieved, the requirements for cooperation in children's patients are reduced, the accuracy of diagnosis and treatment efficiency are improved, and the improvement of abnormality in abnormal eyes is evaluated in real time.
Smart Images

Figure CN118963560B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of ophthalmic medical equipment, and in particular to a visual function diagnosis and treatment device and a visual training method, a visual function recognition method, a storage medium, a computer device and a system. Background Art
[0002] Abnormal visual function such as amblyopia and strabismus is a common visual development disorder, which is manifested as decreased visual function. Even under the best correction conditions, vision cannot reach normal levels, and is not caused by abnormal eye structure. Its pathogenesis is complex and is usually related to interference with visual experience in infancy, such as insufficient eye muscle function, refractive error or form deprivation, which leads to inhibition of signal processing of the affected eye by the brain's visual center.
[0003] Traditionally, the diagnosis of abnormal visual function mainly relies on methods such as visual acuity chart testing, cover test and stereoscopic vision assessment. Although these methods can preliminarily determine whether there is abnormal visual function and its approximate degree, they are often highly subjective, require high cooperation from pediatric patients, and it is difficult to accurately assess the asymmetry and gaze preference of binocular vision, which affects the accuracy and timeliness of diagnosis.
[0004] In terms of treatment, the most commonly used strategies include occlusion therapy, vision training, and glasses or contact lens correction. Occlusion therapy promotes visual development by forcing the eye with abnormal visual function to be used. However, this method may cause a temporary decrease in the function of the covered eye, and patients have significant compliance issues. Although vision training can improve the condition in a targeted manner, it lacks an individualized plan and an immediate feedback mechanism for effect monitoring. Summary of the invention
[0005] The present application mainly provides a visual function diagnosis and treatment device and its visual training method, visual function recognition method, storage medium, computer device and system to solve the above-mentioned deficiencies in existing diagnostic and treatment methods for the visual function of the eyes.
[0006] In order to solve the above technical problems, a technical solution adopted by the present application is: to provide a visual training method. The visual training method includes: tracking the eye movement information of the user when watching the display, the eye movement information includes a waveform diagram of the pupil area and a visual trajectory formed by the coordinates of the gaze point; based on the difference in the change of the waveform diagram between the user's eyes, and the difference between the visual trajectory and the motion trajectory of the preset object on the display, determining the visual function information of the user, the visual function information includes a normal eye, an abnormal eye and the degree of abnormality; combining the real-time eye movement information and visual function information of the user's eyes, playing a preset object corresponding to the visual function information on the display, so as to digitally cover the normal eye and perform visual training on the abnormal eye.
[0007] In some embodiments, tracking the eye movement information of the user when viewing the display includes:
[0008] Acquire an input image including the eye areas of the user;
[0009] Detecting the iris regions of the left and right eyes in the input image using a pre-trained object detection network;
[0010] Detecting a pupil area and light spot areas of at least two reflective points in the iris area, and calculating a pupil center point of the pupil area and a reflective center point corresponding to each of the light spot areas;
[0011] The coordinates of the user's gaze point on the display are calculated based on the pupil center point and each of the reflection center points.
[0012] In some embodiments, after detecting the iris region in the input image by using the pre-trained object detection network, the method further includes:
[0013] A search box based on the iris region is established at a preset magnification ratio, the iris region of the next time sequence is tracked and detected by the search box, and the position of the search box is updated after the iris region is detected.
[0014] In some embodiments, after the iris region is tracked and extracted using the search box, the method further includes:
[0015] In response to the iris region not being detected in the search box, the step of detecting the iris region in the input image by using the pre-trained object detection network is performed again.
[0016] In some embodiments, the calculating the coordinates of the user's gaze point on the display based on the pupil center point and each of the reflection center points includes:
[0017] The pupil center point and each of the reflection center points are used as inputs of a line of sight estimation model to obtain the coordinates of the user's gaze point on the display, wherein the line of sight estimation model is a deep learning network model trained using a data set established based on a pupil corneal reflection method.
[0018] In some embodiments, the types of abnormal eyes include amblyopia and strabismus;
[0019] The combining the real-time eye movement information of both eyes of the user and the visual function information to play a preset object corresponding to the visual function information on the display includes:
[0020] In response to the abnormal eye of the user's eyes being amblyopic or strabismic, a treatment area is established with a preset radius based on the real-time gaze point coordinates of the user on the display, wherein the object of the first color channel in the treatment area is degraded to form a degraded object corresponding to the normal eye, the object of the second color channel in the treatment area is not changed to form a clear object corresponding to the abnormal eye, and the projection of the treatment area on the user's retina covers at least the macular area of the retina.
[0021] In order to solve the above technical problems, another technical solution adopted by the present application is to provide a visual function recognition method. The visual function recognition method includes: obtaining an input image containing the binocular regions of the user; detecting the iris region in the input image to obtain the pupil region, the pupil center point and at least two reflection center points in the iris region; calculating the user's gaze point coordinates on the display based on the pupil center point and each of the reflection center points; confirming the user's visual function information based on the waveform diagram of the pupil region obtained in time sequence and the visual trajectory formed by the gaze point coordinates.
[0022] In some embodiments, detecting the iris region in the input image to obtain the pupil region, the pupil center point and at least two reflection center points in the iris region includes:
[0023] Detecting an iris region in the input image using a pre-trained object detection network;
[0024] The pupil area and at least two reflective spot areas in the iris area are detected, and the pupil center point of the pupil area and the reflective center point corresponding to each of the spot areas are calculated.
[0025] In some embodiments, after detecting the iris region in the input image by using the pre-trained object detection network, the method further includes:
[0026] A search box based on the iris region is established at a preset magnification ratio, the iris region of the next time sequence is tracked and detected by the search box, and the position of the search box is updated after the iris region is detected.
[0027] In some embodiments, after the iris region of the next time sequence is obtained by tracking and detecting with the search frame, the method further includes:
[0028] In response to the iris region not being detected in the search box, the step of detecting the iris region in the input image by using the pre-trained object detection network is performed again.
[0029] In some embodiments, the calculating the coordinates of the user's gaze point on the display based on the pupil center point and each of the reflection center points includes:
[0030] The pupil center point and each of the reflection center points are used as inputs of a preset line of sight estimation model to obtain the coordinates of the user's gaze point on the display, wherein the line of sight estimation model is a deep learning network model trained with a data set established based on the pupil corneal reflection method.
[0031] In some embodiments, the confirming the visual function information of the user based on the waveform diagram of the pupil area obtained in time sequence and the visual track formed by the gaze point coordinates includes:
[0032] The visual function information of the user is determined based on the difference in the change of the waveform between the two eyes of the user and the difference between the visual trajectory and the motion trajectory of the preset object on the display.
[0033] To solve the above technical problems, another technical solution adopted by the present application is to provide a storage medium having program data stored thereon, and when the program data is executed by a processor, the steps of the above visual training method or the above visual function recognition method are implemented.
[0034] In order to solve the above technical problems, another technical solution adopted by the present application is to provide a computer device. The computer device includes a processor and a memory connected to each other, the memory stores a computer program, and when the processor executes the computer program, the steps of the above visual training method or the above visual function recognition method are implemented.
[0035] To solve the above technical problems, another technical solution adopted by the present application is to provide a visual function diagnosis and treatment device, which includes a display, a camera disposed on the display, at least two fill lights, and a computer device as described above associated with the display, the camera, and the fill lights.
[0036] In order to solve the above technical problems, another technical solution adopted by the present application is to provide a visual function diagnosis and treatment system. The visual function diagnosis and treatment system includes filter glasses and the above visual function diagnosis and treatment device, the filter glasses include a first filter lens corresponding to a normal eye and a second filter lens corresponding to an abnormal eye, and the filter glasses are used for a user to wear to view the content displayed on the display.
[0037] The beneficial effects of the present application are as follows: Different from the prior art, the present application discloses a device for diagnosing and treating amblyopia and an operating method, a storage medium, a computer device and a system. The present application uses the data obtained by the camera to track the eye movement information of the user when watching the display in real time, and can accurately and efficiently diagnose the user's visual function information based on the eye movement information, which includes normal eyes, abnormal eyes and their abnormality, and formulates a personalized treatment plan corresponding to the user, and plays the preset object corresponding to the visual function information on the display for visual training; that is, by using computer vision technology, the diagnosis and treatment are accurate and personalized, and by integrating digital technology with medical knowledge in the field of ophthalmology, an advanced solution is provided for the diagnosis and treatment of abnormal visual function, which does not require the child patient to have a high degree of cooperation, the diagnosis is accurate and efficient, and the patient's compliance problem is also solved. At the same time, based on the real-time eye movement information obtained during the treatment process, the improvement of the abnormality of the abnormal eye can be evaluated in real time, with good feedback, and the method provided by the present application can effectively promote the progress of visual rehabilitation technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the prior art descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative work, among which:
[0039] Figure 1 It is a structural schematic diagram of an embodiment of a visual function diagnosis and treatment device provided by the present application;
[0040] Figure 2 It is a flowchart of an embodiment of a visual training method provided by the present application;
[0041] Figure 3 yes Figure 2 A schematic flow chart of an embodiment of step 10 in the visual training method shown;
[0042] Figure 4 yes Figure 3 Schematic diagram of the input image and the iris region bounding box therein;
[0043] Figure 5 yes Figure 4 Schematic diagram of a search box in an input image and an iris region in the search box;
[0044] Figure 6 yes Figure 5 Schematic diagram of the pupil area in the iris region;
[0045] Figure 7yes Figure 6 Schematic diagram of reflective spots in the iris area;
[0046] Figure 8 It is a flowchart of an embodiment of a visual function recognition method provided by the present application;
[0047] Fig. 9 yes Figure 8 A flow chart of an embodiment of step 120 in the visual function recognition method shown;
[0048] Fig.10 It is a structural schematic diagram of an embodiment of the storage medium provided by the present application;
[0049] Fig.11 It is a structural schematic diagram of an embodiment of a computer device provided by the present application;
[0050] Fig.12 It is a structural schematic diagram of an embodiment of the visual function diagnosis and treatment system provided by the present application. DETAILED DESCRIPTION
[0051] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0052] The terms "first", "second", "third" in the embodiments of the present application are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features. Thus, the features defined as "first", "second", "third" can expressly or implicitly include at least one of the features. In the description of the present application, the meaning of "multiple" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device comprising a series of steps or units is not limited to the listed steps or units, but optionally also includes steps or units that are not listed, or optionally also includes other steps or units inherent to these processes, methods, products or devices.
[0053] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0054] When one eye of a patient is a healthy non-amblyopic eye and the other eye is an amblyopic eye, if no treatment intervention is performed, the brain will tend to use the non-amblyopic eye and ignore the visual signals from the amblyopic eye, resulting in the obstruction of the visual function development of the neglected amblyopic eye. If no timely treatment intervention is performed, the amblyopic eye may cause permanent vision loss.
[0055] The present application provides a visual function diagnosis and treatment device 100, see Figure 1 , Figure 1 It is a structural schematic diagram of an embodiment of a visual function diagnosis and treatment device provided in the present application.
[0056] The visual function diagnosis and treatment device 100 includes a display 10, a camera 20 and at least two fill lights 30 arranged on the display 10, and a computer device 40 associated with the display 10, the camera 20 and the fill lights 30. The display 10 is used to display image content, the camera 20 is used to obtain image data of a user when viewing the display 10, and the fill lights 30 can be mapped in the user's eyes and form a light spot area when turned on. The computer device 40 is used to execute a computer program to implement a visual training method or a visual function recognition method run by the visual function diagnosis and treatment device 100, so as to realize the integration of diagnosis and treatment of visual function.
[0057] The display 10 can be used to play image content including preset objects, wherein the image content can be pre-stored, or the image content can also be content uploaded in advance or downloaded in real time. A certain object in the image content can be designated as a preset object through a program, and the object can be a swimming fish, a falling leaf, a moving train, etc.
[0058] The camera 20 can be located at the center of the upper frame or the center of the lower frame of the display 10, and is used to obtain image data of the user viewing the display 10. The image data includes the user's eye area, and a light spot area corresponding to the fill light 30 is formed in the user's eyes.
[0059] Camera 20 is an infrared camera with a near-infrared filter, wherein the near-infrared light can clearly show the details of the eye, such as the pupil, iris and corneal reflective spots (i.e., Purkinje spots) under low light conditions.
[0060] The fill light 30 is a light source that outputs near-infrared light. There can be two fill lights 30, which can be symmetrically arranged on both sides of the camera 20. There can also be three or four fill lights 30, etc., which can be evenly distributed on the frame of the display 10.
[0061] The fill light 30 may be a monochromatic near-infrared light in the range of 600nm to 1500nm, and the camera 20 is an infrared camera covered with a corresponding wavelength filter. For example, if the fill light 30 is a light source with a narrow-band wavelength of 650nm or 850nm, the camera 20 is an infrared camera covered with a corresponding 650nm or 850nm filter.
[0062] The input image obtained by the camera 20 in the near-infrared light spectrum range can obtain clearer pupil edges and light spot areas within the pupil area compared to the input images obtained under red, green and blue light, thereby reducing the difficulty of processing image data, effectively reducing the demand for computing resources, and greatly improving the accuracy and precision of image processing, thereby obtaining more accurate eye movement information, and further providing users with more accurate amblyopia diagnosis and treatment effects.
[0063] Among them, the camera 20 and the fill light 30 can be integrated on the display 10, that is, the camera 20, the fill light 30 and the display 10 are integrated; or, at least one of the camera 20 and the fill light 30 is independent of the display 10, for example, the display 10 is a liquid crystal display, and the camera 20 and the fill light 30 are electronic accessories attached to the display 10, that is, the camera 20 and the fill light 30 are both detachably connected to the display 10, or one of the camera 20 and the fill light 30 is an electronic accessory attached to the display 10, and the other is integrated on the display 10; or, the display 10 is a screen for carrying projection, and the camera 20 and the fill light 30 are both independently arranged on the screen.
[0064] The computer device 40 can obtain eye movement information, non-amblyopic eye, amblyopic eye and its amblyopia degree and other information based on the obtained image data after processing, and form a treatment plan accordingly.
[0065] Among them, the computer device 40 may include a general-purpose processor (Central Processing Unit, CPU), and may further include a graphics processing unit (Graphics Processing Unit, GPU) and / or a neural network processing unit (Neural network Processing Unit, NPU) and other parallel processors that can help improve the reasoning speed of the deep learning model.
[0066] The model in this embodiment is a lightweight model. On an ordinary mid-range computing chip, it can be calculated in the form of a pure general-purpose processor and with the help of multi-threading technology to complete the analysis and calculation of a frame of image within a few milliseconds.
[0067] In other words, the visual training method and visual function recognition method provided by the present application and available for operation by the visual function diagnosis and treatment device 100 can be operated on an ordinary general-purpose processor, greatly reducing the required hardware cost.
[0068] The present application also provides a visual training method that can be used by the visual function diagnosis and treatment device 100, see Figure 2 , Figure 2 1 is a flow chart of an embodiment of a visual training method provided by the present application. The visual training method of the visual function diagnosis and treatment device 100 includes:
[0069] Step 10: Track the eye movement information of the user when viewing the display, the eye movement information including a waveform diagram of the pupil area and a visual track formed by the coordinates of the gaze point.
[0070] The eye movement information includes eye state information and gaze point information. The eye state information includes a waveform diagram of eye pupil contraction and dilation, and the gaze point information includes the user's gaze point coordinates on the display 10.
[0071] By continuously tracking the eye movement information of the user when viewing the display 10 in a time sequence, a waveform diagram of the changes in pupil contraction and dilation can be formed, and the gaze point coordinates can be calculated when the user gazes at a preset object on the display 10, and a visual trajectory can be formed by continuous gaze point coordinates when tracking the preset object.
[0072] In this embodiment, the camera 20 located on the display 10 can obtain image data of the user viewing the display 10 in real time, and the computer device 40 processes the image data and tracks the eye movement information of the user viewing the display 10 in real time.
[0073] The camera 20 collects video streams at a frequency of not less than 20 Hz, and then transmits the image frame data in the video stream to the computing processor for image processing and analysis to obtain the user's eye movement information. For example, the sampling frequency of the camera 20 can be 20 Hz, 30 Hz, 60 Hz, 80 Hz, 100 Hz or 120 Hz, etc.
[0074] It should be noted that the computer device 40 tracks the eye movement information of the user's eyes in real time.
[0075] Furthermore, the data obtained by the computer device 40 through the camera can also track the user's interaction with the visual function diagnosis and treatment device 100, so as to form an interactive behavior of controlling the visual function diagnosis and treatment device 100 through gestures or blinking and other actions, thereby improving the convenience of the user in controlling the visual function diagnosis and treatment device 100.
[0076] Optionally, the eye movement information of the user when viewing the display 10 may be tracked by detecting a small potential difference on the surface of the eyeball.
[0077] When the visual function of the user's eyes is normal, the waveform of his pupil changes rapidly and accurately, and the coordinates of the gaze point and the visual trajectory are also accurate; if the visual function of the user's eyes is abnormal, the waveform of his pupil changes relatively sluggishly and is not agile enough and accurate enough, and the coordinates of the gaze point and the visual trajectory also have large deviations and are not agile enough to follow the movement trajectory of the object. Therefore, based on the tracked eye movement information, it can be confirmed whether the user's visual function information is abnormal, the type and degree of the abnormality.
[0078] See also Figure 3 , Figure 3 yes Figure 2 The flowchart of step 10 of the visual training method is shown in FIG. Step 10 may specifically include:
[0079] Step 11: Obtain an input image containing the user's eye areas.
[0080] See also Figure 4 , Figure 4 yes Figure 3 Schematic diagram of an input image and an iris region outer frame in the input image. The camera 20 acquires an input image A including the user's eye regions, wherein when the camera 20 acquires the input image A, the fill light 30 is turned on.
[0081] In this embodiment, the two fill lights 30 are symmetrically arranged on both sides of the camera 20 , so that two reflection center points corresponding to the fill lights 30 are formed in the iris area of the input image A.
[0082] The input image A includes both the left and right eye areas of the user, where both eyes of the user are open, and is a valid image. If both eyes of the user in the input image A are closed, the input image A is determined to be an invalid image, and the eye tracking of the input image A is abandoned, and the next input image A is acquired.
[0083] Therefore, after the computer device 40 obtains the input image A, it will first determine whether the input image A is a valid image. If it is a valid image, eye tracking will continue on the input image A. If it is an invalid image, the eye gaze point obtained in the previous frame will remain unchanged and the next input image A will continue to be obtained.
[0084] Step 12: Detect the iris region in the input image through the pre-trained object detection network.
[0085] The target detection network can be an R-CNN (Region-based Convolutional Neural Network) network and its optimized version or a YOLO (You Only Look Once) network and its optimized version, etc. After being trained with a training set and verified by a verification set, a pre-trained target detection network is obtained.
[0086] The pre-trained target detection network is used to perform binocular detection. The pre-trained target detection network can identify the iris area A1 in the input image A, that is, the iris area A1 of the left eye and the right eye, and detect the circumscribed frame of the iris area A1 of the left and right eyes. The circumscribed frame can be an circumscribed rectangular frame or an circumscribed circular frame, etc., and the present application does not limit its specific shape.
[0087] Combined with reference Figure 4 and Figure 5 ,in Figure 5 yes Figure 4 Schematic diagram of the search box and the iris region in the search box in the input image.
[0088] Furthermore, after the circumscribed frames of the left and right eye iris areas A1 are detected in step 12, the process also includes: establishing a search box A2 based on the iris area at a preset magnification ratio r, tracking and detecting the iris area A1 of the next time sequence with the search box A2, and updating the position of the search box A2 after the iris area A1 is detected.
[0089] In other words, after detecting and obtaining the circumscribed frames of the left and right eye iris areas A1, the circumscribed frames are enlarged according to a preset enlargement ratio r to obtain search frames A2 for the two eyes respectively. The preset enlargement ratio r is greater than 1, and the preset enlargement ratio r can be 1.5, 2.0, 2.5 or 3.0, etc.
[0090] The position information of the search box A2 in the current input image A is used as the region of interest in the input image A of the next time series, and the pre-trained target detection network is used to extract the iris region A1 of the input image A of the next time series in the region of interest, and the position of the search box A2 is updated after the iris region A1 is extracted, that is, the external frame is enlarged again according to the preset enlargement ratio r to form a new search box A2, so that the user's iris region A1 can be continuously tracked, which greatly improves the recognition and tracking efficiency of the iris region A1, that is, compared with processing the global area of the input image A, by establishing the region of interest, the image area that the computer needs to process can be effectively reduced, thereby improving the recognition efficiency of the target detection network by actively excluding most of the irrelevant data, and reducing the demand for computing resources.
[0091] It can be understood that when the user is concentrating on watching the content displayed on the display 10, the position of his eyes changes relatively little. Therefore, by forming a search box A2, it is equivalent to predicting the area of interest where the eyes are located in the next input image A. Directly using a pre-trained target detection network in the area of interest can more efficiently identify and extract the iris area on the current input image A, which greatly reduces the data processing amount of the target detection network, making the target detection network relatively faster, more efficient and less energy-consuming.
[0092] Furthermore, the user may also move his head so that the position of his eyes on the input image A changes significantly, that is, the iris area cannot be detected in the region of interest of the next time-series input image A, indicating that the tracking has failed and binocular detection needs to be performed again based on the current input image A.
[0093] That is, after the step of tracking and extracting the iris region A1 with the search box A2, the method further includes: in response to not detecting the iris region A1 in the search box A2, detecting the iris region A1 in the input image A again by using the pre-trained target detection network.
[0094] After confirming that the target tracking has failed, the complete input image A is used as the input of the pre-trained target detection network again, binocular detection is performed again and the iris area A1 in the input image A is identified. The circumscribed frames of the left and right eye iris areas A1 are detected again, and target tracking is performed again.
[0095] Therefore, the present application can more efficiently identify and extract the iris area A1 on the current input image A by using a pre-trained target detection network to continuously perform target detection and target tracking, so that the target detection network can operate relatively faster, more efficiently and consume less energy.
[0096] Step 13: Detect the pupil area and at least two reflective light spot areas in the iris area, and calculate the pupil center point of the pupil area and the reflective center point of each light spot area.
[0097] See also Figures 5 to 7 , Figure 6 yes Figure 5 Schematic diagram of the pupil area in the iris area, Figure 7 yes Figure 6 Schematic diagram of the reflective points in the iris area. The pupil segmentation model is used to obtain the corresponding pupil area A3 from the iris area A1 of the left eye and the right eye. The pupil area A3 is Figure 6 The area enclosed by the green circle in , wherein the pupil segmentation model can be a pre-trained target segmentation network, or can be obtained by using the binarization operation of traditional image processing, which is responsible for extracting the pupil area A3 in the iris area A1.
[0098] The iris is a colored ring-shaped structure in the eye. Its main function is to adjust the amount of light entering the eye by controlling the size of the pupil. When the light is strong, the iris muscles will shrink the pupil and reduce the amount of light entering the eye; when the light is weak, the iris muscles will dilate the pupil and increase the amount of light entering the eye, thus ensuring that the retina can receive the right amount of light and form a clear visual image.
[0099] Therefore, by continuously obtaining the pupil area A3 in time sequence, a changing waveform of the pupil area A3, i.e., the pupil waveform diagram, can be obtained. The difference in changes in the waveform diagram of the pupil area A3 when both eyes observe the same object can be used as an evaluation basis for subsequently determining the user's visual function information.
[0100] Furthermore, after image post-processing the pupil area A3, the pupil center point of the pupil area A3 is calculated, and the pupil center point is Figure 6 and use the preset Purchin spot detection model to detect the spot area corresponding to the fill light 30, and further calculate the reflection center point according to the regional information of the spot area, the reflection center point is as follows Figure 7 The points marked by red circles are shown.
[0101] like Figure 6 As shown, the pupil area is circular, and the pupil center point is the center of the circle. After image processing, the pupil center point of the pupil area can be calculated.
[0102] Purkinje image or Purkinje spot refers to an optical phenomenon observed in the human eye. It is a bright spot formed on the retina after light is reflected by the cornea. The Purkinje spot detection model can detect and locate the reflective spot area in the pupil area, and can assist in determining the gaze direction and pupil position of the user's eyes, and then analyze the coordinates of the user's gaze point on the display 10.
[0103] Step 14: Calculate the coordinates of the user's gaze point on the display based on the pupil center point and each reflection center point.
[0104] Optionally, the pupil center point and at least two reflection center points are obtained, and an eyeball model based on the pupil corneal reflection method is used, and a group of equations are established using the physical principles of light reflection and refraction. Through numerical optimization calculations, the position information of the user's corneal center point and pupil center point in three-dimensional space is obtained, and then the vector of the visual axis is obtained based on the obtained corneal center point and pupil center point, and then the coordinates of the user's gaze point on the display 10 are calculated based on the user calibration parameters and the mapped polynomial equation.
[0105] This method requires solving a large number of nonlinear equations for each frame of image data, which requires a large amount of calculation and requires the use of high-performance dedicated processors, resulting in high hardware costs.
[0106] In this embodiment, step 14 specifically includes: using the pupil center point and each reflection center point as input of the line of sight estimation model to obtain the coordinates of the user's gaze point on the display 10, wherein the line of sight estimation model is a deep learning network model trained with a data set established based on the pupil corneal reflection method.
[0107] Specifically, after collecting a large amount of image data, the above method is used to detect the coordinates of the pupil center and the reflection center in each image as the input feature data X. After establishing a group of equations according to the pupil corneal reflection method, the offline numerical optimization solution method is used to calculate the optimal solution corresponding to each group of data as the label Y, so as to supervise the training of the deep learning model.
[0108] Specifically, the hardware parameters (the spatial position of the light source and the internal parameters obtained by camera calibration) and the key features of the human eye in the image (the coordinates of the pupil center and the reflection center) are mapped and transformed to obtain the physical position information on the camera sensor, which is used as the input data X of the deep learning network model. The corresponding label Y is the position information of the cornea center and the pupil center.
[0109] The input data X and label Y form a training data set, and the fully connected deep learning network model is trained in a supervised learning manner. The pre-trained deep learning model is then used as a line of sight estimation model after further pruning, compression, and quantization. The model has low computational complexity and high speed, and can complete inference calculations within 1ms on general mid-to-high-end embedded processing chips.
[0110] Therefore, based on the pre-trained deep learning model, the pupil center point and the coordinates of at least two reflection center points are used as the input of the pre-trained deep learning model, so as to quickly and efficiently obtain the position information of the user's corneal center point and pupil center point in three-dimensional space, without consuming a large amount of computing resources to solve a set of equations based on the pupil center point and at least two reflection center points obtained each time. The line of sight estimation model can greatly reduce the demand for computing resources and obtain the gaze point coordinates efficiently and accurately; then, based on the obtained corneal center point and pupil center point, the vector of the visual axis is obtained, and then, based on the user calibration parameters and the mapped polynomial equation, the user's gaze point coordinates on the display 10 are calculated, and the gaze point coordinates obtained continuously in time sequence will form a visual trajectory.
[0111] Therefore, during this process, eye movement information consisting of pupil waveforms and gaze point coordinates can be continuously obtained. Based on the obtained eye movement information, the user's visual function diagnosis and corresponding personalized treatment can be completed.
[0112] This embodiment tracks the eye movement information of the user when viewing the display 10 through the input image A obtained by the camera 20, and obtains the pupil area A3 and the light spot area therein through the target detection network and the pupil segmentation model, and then calculates the pupil center point of the pupil area A3 and the reflection center point of the light spot area. The pupil center point and the reflection center point are used as the input of the line of sight estimation model, so that the gaze point coordinates with high accuracy can be obtained efficiently and quickly. These combined technical means can track the eye movement information very efficiently, and the eye movement information obtained is reliable and accurate. These models are pre-trained lightweight models with low requirements for computing resources. They can be carried by ordinary tablet computers and other devices, and the cost of hardware modification is low. In addition, this method of tracking eye movement information is non-invasive, does not cause discomfort to users, and is more user-friendly.
[0113] Step 20: Determine the visual function information of the user based on the difference in the waveform between the user's eyes and the difference between the visual trajectory and the motion trajectory of the preset object on the display.
[0114] The visual function information includes normal eyes, abnormal eyes and the degree of abnormality. The abnormal eyes include the abnormal type of the eyes. For example, the type of the abnormal eye may be amblyopia or strabismus.
[0115] When the visual function of one eye of the user is normal and the visual function of the other eye is abnormal, the brain will tend to use the normal eye with normal visual function and ignore the visual signals from the abnormal eye with abnormal visual function, resulting in the obstruction of the visual function development of the neglected abnormal eye.
[0116] The abnormal eye with abnormal visual function is slower in sensitivity than the normal eye with normal visual function, and it takes longer and is not quick and accurate enough to fixate on the position of the preset object on the display 10. In other words, the pupil waveform of the abnormal eye changes more slowly than that of the normal eye, and the visual track formed by the corresponding fixation point coordinates is not accurate enough compared to the motion track of the preset object. The normal eye is more sensitive, the pupil waveform changes more sensitively and accurately, and its fixation point coordinates correspond to the position coordinates of the preset object more quickly and accurately, that is, its visual track can more accurately follow the motion track of the preset object.
[0117] Based on the difference in eye movement information between the normal eye and the abnormal eye, the normal eye and the abnormal eye can be efficiently identified, and based on the visual sensitivity and visual accuracy shown by the pupil waveform, gaze point coordinates and visual trajectory of the abnormal eye, the degree of abnormality of the abnormal eye can also be confirmed.
[0118] The differences in pupil area changes and visual trajectories obtained in a non-invasive manner in this embodiment can intuitively and accurately reflect the user's normal eye, abnormal eye and the degree of abnormality, thereby ensuring the diagnostic efficiency and accuracy of the abnormal eye, reducing diagnostic errors, and establishing a better foundation for the generation of subsequent treatment plans.
[0119] Due to the difference in visual sensitivity between the normal eye and the abnormal eye, the pupil waveform of the abnormal eye lags behind the pupil waveform of the normal eye in responding to changes in the preset object. The abnormal eye's gaze point coordinates are not sensitive enough and not precise enough to follow the trajectory of the preset object, while the normal eye's gaze point coordinates are more sensitive and precise to follow the trajectory of the preset object. It can be seen that there is an obvious difference in visual sensitivity between the normal eye and the abnormal eye. Therefore, based on these differences, the normal eye and the abnormal eye in the user's eyes can be accurately identified, and the degree of abnormality of the abnormal eye can also be identified accordingly.
[0120] Based on the differences in changes in the pupil area and the differences in following the visual trajectory, the degree of abnormality of the normal eye and the abnormal eye can also be scored. For example, the sensitivity of the normal eye is 95%, and the sensitivity of the abnormal eye is 30%, thereby achieving real-time diagnosis of the user's eyes.
[0121] Specifically, a preset diagnostic analysis model can be used to confirm the user's normal eye, abnormal eye and the degree of abnormality. The preset diagnostic analysis model can be a pre-trained deep learning model, and a large amount of collected data is used to verify the deep learning model. These data include pupil waveforms and their changes, visual trajectories formed by gaze point coordinates and corresponding diagnostic results.
[0122] The types of abnormal eyes are different, and the deviations of their pupil waveforms, gaze point coordinates, visual trajectories and other data compared with normal data are also different. Therefore, based on these data deviations, the abnormal type of the abnormal eye can be further identified. For example, the abnormal eye type can be detected as amblyopic, or the abnormal eye type is strabismus, etc.
[0123] The preset diagnostic analysis model can quickly and efficiently identify the user's normal eye, abnormal eye type and degree of abnormality based on eye movement information.
[0124] During the diagnosis process, the preset object displayed on the display 10 can be transmitted to the user's eyes in the form of binocular dichography, and then the eye movement information of the user's eyes is recorded; or, the same preset object displayed on the display 10 is transmitted to the user's eyes, that is, the user's eyes observe the same preset object on the same screen, and the eye movement information of the user's eyes is recorded.
[0125] Step 30: Combine the real-time eye movement information and visual function information of the user's eyes, and play a preset object corresponding to the visual function information on the display to digitally cover the normal eye and perform visual training on the abnormal eye.
[0126] After obtaining the eye movement information and visual function information of the user's eyes and other information as described above, a personalized diagnosis and treatment plan related to the user can also be formed based on this information.
[0127] The visual function information may also be manually input information. When the user changes or manually inputs information such as the normal eye, the abnormal eye and the degree of abnormality, the input data shall prevail.
[0128] A degenerate object for the normal eye and a clear object for the abnormal eye are played on the display 10 to form a digital cover for the normal eye, thereby forcing the brain to use the abnormal eye to receive visual stimulation, promote its visual development, and reduce the brain's dependence on the normal eye to a certain extent.
[0129] The degraded object for the normal eye may be obtained by processing the clear object by a degradation method such as Gaussian filtering or narrow-band filtering, and the clear object for the abnormal eye is not subjected to degradation processing.
[0130] The degradation method includes at least one of reducing contrast, reducing brightness, blurring and reducing color saturation.
[0131] In this embodiment, step 30 specifically includes: in response to the type of the abnormal eye of the user's eyes being amblyopic or strabismic, a treatment area is established with a preset radius based on the real-time gaze point coordinates of the user on the display 10, wherein the object of the first color channel in the treatment area is degraded to form a degraded object corresponding to the normal eye, and the object of the second color channel in the treatment area is not changed to form a clear object corresponding to the abnormal eye, and the projection of the treatment area on the user's retina at least covers the macular area of the retina.
[0132] The first color channel may be a red channel, and the second color channel may be a blue-green channel. Alternatively, the first color channel may be a blue channel, and the second color channel may be a red-green channel. Alternatively, the first color channel may be a green channel, and the second color channel may be a red-blue channel.
[0133] By establishing a treatment area on the display 10 that tracks the coordinates of the user's binocular gaze point in real time, fatigue caused by the user's eyes being restricted to a fixed object can be avoided, so that visual training and treatment can be performed more flexibly and friendly, thereby improving the quality of treatment; the size setting of the treatment area can ensure that the formed object effectively stimulates the macular area of the abnormal eye; wherein, by performing different treatments on the objects of the first color channel and the second color channel in the treatment area, the same object can form a complementary color image pair in the normal eye and the abnormal eye, and the visual needs of the user can be taken into account while treating the abnormal eye, so that the user can complete the treatment of the abnormal eye in entertainment, and the treatment experience is better and easier for users to accept.
[0134] It can be understood that after obtaining the eye movement information of the user's eyes in real time, during the treatment process, a treatment area with a preset radius can be established based on the real-time gaze point coordinates. In other words, the treatment area will move with the movement of the user's line of sight, thereby increasing the user's effective treatment time and effectively improving the quality of treatment.
[0135] During the treatment process, the same object in the treatment area is processed to form a degraded object and a clear object respectively. The user needs to wear filter glasses (not shown) to cooperate with the treatment. The filter glasses include a first filter lens corresponding to the normal eye and a second filter lens corresponding to the abnormal eye. The filter glasses are used for the user to wear to view the content displayed in the treatment area of the display 10.
[0136] The normal eye wears the first filter lens and can only observe the degraded object of the first color channel, thereby forming a digital cover for the normal eye; the abnormal eye wears the second filter lens and can only observe the clear object of the second color channel, thereby forcing the brain to use the abnormal eye, so that the abnormal eye receives visual stimulation, promotes its visual development, and reduces the brain's dependence on the normal eye to a certain extent; the normal eye and the abnormal eye can form a complementary color image pair.
[0137] The degree of degeneration of the degenerated object can be adjusted based on the degree of abnormality of the abnormal eye, wherein the higher the degree of abnormality, the higher the degree of degeneration of the degenerated object, so as to form more effective visual stimulation for the abnormal eye and promote its visual development.
[0138] Furthermore, the treatment time and frequency can be adjusted based on the degree of abnormality of the abnormal eye, wherein the higher the degree of abnormality, the longer the treatment time and frequency, so as to form more effective visual stimulation for the abnormal eye and promote its visual development.
[0139] The present application tracks the eye movement information of the user when viewing the display, and based on the eye movement information, can accurately and efficiently diagnose the user's visual function information, and formulate a personalized treatment plan corresponding to the user, and play the degenerated objects for the normal eye and the clear objects for the abnormal eye on the display for visual training; that is, by utilizing computer vision technology, the diagnosis and treatment are precise and personalized, and by integrating digital technology with medical knowledge in the field of ophthalmology, an advanced solution is provided for the diagnosis and treatment of abnormal eyes. There is no need to require child patients to have a high degree of cooperation, the diagnosis is accurate and efficient, and the patient's compliance problem is also solved. At the same time, during the treatment process, the improvement of the abnormality of the abnormal eye can be evaluated in real time, with good feedback. The method provided by the present application can effectively promote the advancement of visual rehabilitation technology.
[0140] The present application also provides a visual function identification method that can be used by the visual function diagnosis and treatment device 100, see Figure 8 , Figure 8 : is a flow chart of an embodiment of a visual function recognition method provided by the present application, the visual function recognition method comprising:
[0141] Step 110: Obtain an input image including the user's eye regions.
[0142] The camera 20 continuously acquires the input image A including the eye area of the user in a time sequence, wherein when the camera 20 acquires the input image A, the fill light 30 is turned on.
[0143] Step 120: Detect the iris region in the input image to obtain the pupil region, the pupil center point and at least two reflection center points in the iris region.
[0144] The iris area A1 in the input image A is detected to obtain the pupil area A3 and at least two light spot areas in the iris area A1, the pupil center point is calculated based on the obtained pupil area A3, and the corresponding reflection center point is calculated based on the light spot area.
[0145] See also Fig. 9 , Fig. 9 yes Figure 8 The flowchart of step 120 of an embodiment of the visual function recognition method is shown. Step 120 may specifically include:
[0146] Step 121: Detect the iris region in the input image through the pre-trained object detection network.
[0147] The target detection network can be an R-CNN network and its optimized version or a YOLO network and its optimized version, etc. After being trained with a training set and verified by a verification set, a pre-trained target detection network is obtained.
[0148] The pre-trained target detection network is used to perform binocular detection. The pre-trained target detection network can identify the iris area A1 in the input image A, that is, the iris area A1 of the left eye and the right eye, and detect the circumscribed frames of the left and right eye iris areas A1.
[0149] Furthermore, after the iris area A1 is detected in step 121, the iris area A1 is tracked so that the iris area A1 can be tracked more quickly in the next time sequence.
[0150] Specifically, the visual function recognition method further includes: establishing a search box A2 based on the iris area A1 at a preset enlargement ratio r, tracking and detecting the iris area A1 of the next time sequence with the search box A2, and updating the position of the search box A2 after detecting the iris area A1.
[0151] Furthermore, the user may also move his head so that the position of his eyes on the input image A changes significantly, that is, the iris area A1 cannot be detected in the search box A2 of the input image A of the next time sequence, indicating that the tracking has failed and the iris area A1 needs to be re-detected globally based on the current input image A.
[0152] That is, after the step of tracking and extracting the iris region A1 with the search box A2, the method further includes: in response to not detecting the iris region A1 in the search box A2, detecting the iris region A1 in the input image A again by using the pre-trained target detection network.
[0153] Step 122: Detect and obtain the pupil area and at least two reflective light spot areas in the iris area, and calculate and obtain the pupil center point of the pupil area and the reflective center point corresponding to each light spot area.
[0154] The pupil region A3 can be detected from the iris region A1 using a pupil segmentation model, and the pupil segmentation model can be a pre-trained target segmentation network; or, the pupil region A3 can be obtained using a binarization operation of traditional image processing. The pupil center point of the pupil region A3 can be calculated based on the region information of the pupil region A3.
[0155] The preset Purkinje spot detection model may be used to detect the light spot area corresponding to the fill light 30 , and the reflection center point may be further calculated based on the area information of the light spot area.
[0156] Step 130: Calculate the user's gaze point coordinates on the display based on the pupil center point and each reflection center point.
[0157] The coordinates of the gaze point can be obtained by establishing a system of equations and solving the equations, or by training a deep learning network model to obtain inference results more efficiently.
[0158] In this embodiment, step 130 specifically includes: using the pupil center point and each reflection center point as input of a preset line of sight estimation model to obtain the coordinates of the user's gaze point on the display, wherein the line of sight estimation model is a deep learning network model trained with a data set established based on the pupil corneal reflection method.
[0159] Among them, the trained line of sight estimation model has small calculation amount and fast speed, and can complete inference calculation within 1ms on general mid-to-high-end embedded processing chips.
[0160] Step 140: Confirm the user's visual function information based on the waveform diagram of the pupil area obtained in time series and the visual trajectory formed by the gaze point coordinates.
[0161] The visual function information of the user's eyes includes whether both eyes are normal eyes or both eyes are abnormal eyes and the degree of abnormality; or the visual function information of the user's eyes includes the normal eye, the abnormal eye and the degree of abnormality, that is, one eye of the user is a normal eye and the other eye is an abnormal eye. The type of the abnormal eye can be amblyopia or strabismus.
[0162] Normal eyes and abnormal eyes can show obvious differences in data such as waveform diagrams of the pupil area, gaze point coordinates and visual trajectories, and different types of abnormal eyes also show differences in the aforementioned data. Therefore, based on the pupil waveform diagrams, gaze point coordinates and visual trajectories of the user's eyes, visual function information such as normal eyes, abnormal eyes, types of abnormal eyes and degree of abnormality can be identified.
[0163] Optionally, the visual function diagnosis and treatment device 100 pre-stores boundary values obtained from big data statistics and which can distinguish normal eyes, abnormal eyes, the type of abnormal eyes and the degree of abnormality, etc. By comparing the pupil waveform, gaze point coordinates and visual trajectory with the corresponding boundary values, it can be identified whether the eye is a normal eye or an abnormal eye, as well as the type and degree of abnormality of the abnormal eye.
[0164] In this embodiment, step 140 specifically includes: determining the visual function information of the user based on the difference in the waveform between the user's eyes and the difference between the visual trajectory and the motion trajectory of the preset object on the display.
[0165] Due to the difference in visual sensitivity between the normal eye and the abnormal eye, the pupil waveform of the abnormal eye lags behind the pupil waveform of the normal eye in responding to changes in the preset object. The abnormal eye's gaze point coordinates are not sensitive enough and not precise enough to follow the trajectory of the preset object, while the normal eye's gaze point coordinates are more sensitive and precise to follow the trajectory of the preset object. It can be seen that there is an obvious difference in visual sensitivity between the normal eye and the abnormal eye. Therefore, based on these differences, the normal eye and the abnormal eye in the user's eyes can be accurately identified, and the degree of abnormality of the abnormal eye can also be identified accordingly.
[0166] Based on the differences in changes in the pupil area and the differences in following the visual trajectory, the degree of abnormality of the normal eye and the abnormal eye can also be scored. For example, the sensitivity of the normal eye is 95%, and the sensitivity of the abnormal eye is 30%, thereby achieving real-time diagnosis of the user's eyes.
[0167] Specifically, a preset diagnostic analysis model can be used to confirm the user's normal eye, abnormal eye and the degree of abnormality. The preset diagnostic analysis model can be a pre-trained deep learning model, and a large amount of collected data is used to verify the deep learning model. These data include pupil waveforms and their changes, visual trajectories formed by gaze point coordinates and corresponding diagnostic results.
[0168] Furthermore, the present application also provides a storage medium 50, see Figure 8 , Figure 8 It is a structural diagram of an embodiment of the storage medium provided by the present application.
[0169] The storage medium 50 stores program data 51. When the program data 51 is executed by the processor, the following is achieved: Figure 2 and Figure 3 Visual training methods as described in Figure 8 and Fig. 9 The steps of the visual function recognition method described in .
[0170] The program data 51 is stored in a storage medium 50, and includes a number of instructions for enabling a network device (such as a router, a personal computer, a server, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present application.
[0171] Optionally, the storage medium 50 may be a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, an optical disk, or other medium that can store the program data 51 .
[0172] Furthermore, the present application also provides a computer device 40, see Fig. 9 , Fig. 9 40 is a schematic diagram of a computer device according to an embodiment of the present application. The computer device 40 includes a processor 42 and a memory 41 connected to each other. The memory 41 stores a computer program. When the processor 42 executes the computer program, the following is achieved: Figure 2 and Figure 3 Visual training methods as described in Figure 8 and Fig. 9 The steps of the visual function recognition method described in .
[0173] Based on this, the present application also provides a visual function diagnosis and treatment system, see Fig.10 , Fig.10 1 is a schematic diagram of the structure of an embodiment of the visual function diagnosis and treatment system provided by the present application. The visual function diagnosis and treatment system comprises filter glasses 101 and the visual function diagnosis and treatment device 100 as described above, the filter glasses 101 comprises a first filter lens corresponding to a normal eye and a second filter lens corresponding to an abnormal eye, and the filter glasses 101 are used for a user to wear to view the content displayed on a display.
[0174] The normal eye wears the first filter lens and can only observe the degraded objects of the first color channel, thus forming a digital cover for the normal eye; the abnormal eye wears the second filter lens and can only observe the clear objects of the second color channel, thus forcing the brain to use the abnormal eye, allowing the abnormal eye to receive visual stimulation, promoting its visual development, and reducing the brain's dependence on the normal eye to a certain extent.
[0175] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the storage medium embodiment and the electronic device embodiment, since they are basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0176] The present application can be used in many general or special computing system environments or configurations, such as personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc.
[0177] In the several embodiments provided in this application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation described above is only illustrative, for example, the division of the modules or units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.
[0178] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0179] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0180] The above descriptions are merely embodiments of the present application and are not intended to limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A visual training method, characterized in that: include: Tracking the eye movement information of the user when viewing the display, the eye movement information including a waveform of the pupil area and a visual track formed by the coordinates of the gaze point, wherein the waveform of the pupil area is a waveform formed by the contraction and expansion changes of the pupil area of the user's eyes; Based on the difference in the change of the waveform graph between the user's two eyes, and the difference between the visual trajectory and the motion trajectory of the preset object on the display, the visual function information of the user is determined, and the visual function information includes a normal eye, an abnormal eye, and a degree of abnormality; wherein the waveform graph of the abnormal eye lags behind the waveform graph of the normal eye in responding to the change of the preset object, the deviation between the visual trajectory of the abnormal eye and the motion trajectory of the preset object is greater than the deviation between the visual trajectory of the normal eye and the motion trajectory of the same preset object, and the agility of the gaze point coordinates of the abnormal eye following the motion trajectory of the preset object is also lower than the agility of the gaze point coordinates of the normal eye following the motion trajectory of the same preset object; Combined with the real-time eye movement information of the user's eyes and the visual function information, a preset object corresponding to the visual function information is played on the display to digitally cover the normal eye and perform visual training on the abnormal eye.
2. The visual training method according to claim 1, characterized in that: The tracking of eye movement information of the user when viewing the display includes: Acquire an input image including the eye areas of the user; Detecting an iris region in the input image using a pre-trained object detection network; Detecting a pupil area and at least two reflective light spot areas in the iris area, and calculating a pupil center point of the pupil area and a reflective center point corresponding to each of the light spot areas; The coordinates of the user's gaze point on the display are calculated based on the pupil center point and each of the reflection center points.
3. The visual training method according to claim 2, characterized in that: After detecting the iris region in the input image by the pre-trained target detection network, the method further includes: A search box based on the iris region is established at a preset magnification ratio, the iris region of the next time sequence is tracked and detected by the search box, and the position of the search box is updated after the iris region is detected.
4. The visual training method according to claim 3, characterized in that: After the iris region is tracked and extracted by the search box, the method further includes: In response to the iris region not being detected in the search box, the step of detecting the iris region in the input image by using the pre-trained object detection network is performed again.
5. The visual training method according to claim 2, characterized in that: The calculating the coordinates of the user's gaze point on the display based on the pupil center point and each of the reflection center points includes: The pupil center point and each of the reflection center points are used as inputs of a line of sight estimation model to obtain the coordinates of the user's gaze point on the display, wherein the line of sight estimation model is a deep learning network model trained using a data set established based on a pupil corneal reflection method.
6. The visual training method according to claim 5, characterized in that: The types of abnormal eyes include amblyopia and strabismus; The combining the real-time eye movement information of both eyes of the user and the visual function information to play a preset object corresponding to the visual function information on the display includes: In response to the abnormal eye of the user's eyes being amblyopic or strabismic, a treatment area is established with a preset radius based on the real-time gaze point coordinates of the user on the display, wherein objects of the first color channel in the treatment area are degraded to form degraded objects corresponding to the normal eye, and objects of the second color channel in the treatment area are not changed to form clear objects corresponding to the abnormal eye, and the projection of the treatment area on the user's retina covers at least the macular area of the retina.
7. A visual function recognition method, characterized in that: include: Acquire an input image including the user's eye areas; Detecting an iris region in the input image to obtain a pupil region, a pupil center point, and at least two reflection center points in the iris region; Calculating the coordinates of the user's gaze point on the display based on the pupil center point and each of the reflection center points; Based on the difference in changes in the waveform of the pupil area obtained in time sequence between the user's two eyes, and the difference between the visual trajectory formed by the gaze point coordinates and the motion trajectory of the preset object on the display, the user's visual function information is confirmed, and the function information includes a normal eye, an abnormal eye and the degree of abnormality; wherein, the waveform of the pupil area is a waveform formed by the contraction and expansion changes of the pupil area of the user's eye; the waveform of the abnormal eye lags behind the waveform of the normal eye in responding to the changes of the preset object in response to the changes of the same preset object, the deviation between the visual trajectory of the abnormal eye and the motion trajectory of the preset object is greater than the deviation between the visual trajectory of the normal eye and the motion trajectory of the same preset object, and the agility of the gaze point coordinates of the abnormal eye to follow the motion trajectory of the preset object is also lower than the agility of the gaze point coordinates of the normal eye to follow the motion trajectory of the same preset object.
8. The visual function recognition method according to claim 7, characterized in that: The detecting the iris region in the input image to obtain the pupil region, the pupil center point and at least two reflection center points in the iris region includes: Detecting an iris region in the input image using a pre-trained object detection network; The pupil area and at least two reflective spot areas in the iris area are detected, and the pupil center point of the pupil area and the reflective center point corresponding to each of the spot areas are calculated.
9. The visual function recognition method according to claim 8, characterized in that: After detecting the iris region in the input image by the pre-trained target detection network, the method further includes: A search box based on the iris region is established at a preset magnification ratio, the iris region of the next time sequence is tracked and detected by the search box, and the position of the search box is updated after the iris region is detected.
10. The visual function recognition method according to claim 9, characterized in that: After the iris region of the next time sequence is obtained by tracking and detecting with the search frame, the method further includes: In response to the iris region not being detected in the search box, the step of detecting the iris region in the input image by using the pre-trained object detection network is performed again.
11. The visual function recognition method according to claim 7, characterized in that: The calculating the coordinates of the user's gaze point on the display based on the pupil center point and each of the reflection center points includes: The pupil center point and each of the reflection center points are used as inputs of a preset line of sight estimation model to obtain the coordinates of the user's gaze point on the display, wherein the line of sight estimation model is a deep learning network model trained with a data set established based on the pupil corneal reflection method.
12. A storage medium having program data stored thereon, characterized in that: When the program data is executed by a processor, the steps of the visual training method according to any one of claims 1 to 6 or the visual function recognition method according to any one of claims 7 to 11 are implemented.
13. A computer device, characterized in that: The method comprises a processor and a memory connected to each other, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of the visual training method according to any one of claims 1 to 6 or the visual function recognition method according to any one of claims 7 to 11 are implemented.
14. A visual function diagnosis and treatment device, characterized in that: The visual function diagnosis and treatment equipment includes a display, a camera arranged on the display and at least two fill lights, and a computer device as described in claim 13 associated with the display, the camera and the fill lights.
15. A visual function diagnosis and treatment system, comprising filter glasses and the visual function diagnosis and treatment device as described in claim 14, the filter glasses comprising a first filter lens corresponding to a normal eye and a second filter lens corresponding to an abnormal eye, the filter glasses are used for a user to wear to view the content displayed on the display.
Citation Information
Patent Citations
Fixation point track description method and system based on video analysis
CN111443804A
Visual inspection and visual training equipment
CN113208884A