Vision adaptive detection method and system based on generative adversarial network
By employing a vision adaptive detection method based on generative adversarial networks, the system can perceive the subject's state in real time and generate pre-distorted optotype images, thus solving the accuracy drift and consistency problems of digital vision testing equipment and achieving efficient and reliable vision testing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-24
- Publication Date
- 2026-03-27
AI Technical Summary
Existing digital vision testing equipment relies on static physical calibration, which leads to easy accuracy drift, high maintenance costs, and poor consistency across devices, making it unable to cope with the challenges of hardware aging and changes in the posture of the test subject.
A vision-adaptive detection method based on generative adversarial networks is adopted. By sensing the subject's state in real time, a pre-distorted optotype image is dynamically generated to achieve retinal equivalent imaging. Through a closed-loop iterative process of perception-generation-response, the detection accuracy and consistency are ensured.
It achieves long-term reliability of testing accuracy and consistency of results across devices, simplifies maintenance processes, reduces costs, and improves the level of testing automation and clinical work efficiency.
Smart Images

Figure CN121730733A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of medical diagnostic equipment and artificial intelligence, in particular to a vision adaptive detection technology applied to a digital vision detection device, more particularly to a vision adaptive detection method and system based on a generative adversarial network. BACKGROUND
[0002] Vision, as one of the core indicators of visual function, its accurate measurement is of great significance for the diagnosis of ophthalmic diseases, the correction of ametropia and the census of public visual health. Traditional vision detection relies on printed standard logarithmic visual charts (LogMAR Chart), and its detection is based on the principle of strictly following "one angle of view", that is, the physical size of the visual target should accurately correspond to a specific angle of view at a standard detection distance. This principle means that the accuracy of detection is highly dependent on the high standardization of the two physical quantities of "visual target physical size" and "detection distance".
[0003] With the development of information technology, electronic and digital vision detection devices such as electronic visual chart modules built into comprehensive optometry instruments and integrated vision screening instruments have been widely used in professional medical institutions. These devices present digital generated visual targets on the display screen, replacing the traditional printed chart. However, in order to ensure that the pixel composed electronic image can be equivalent to the physical visual target with standard physical size, the existing technology generally adopts a technical mode of "static calibration combined with mechanical fixation". In this mode, the device manufacturer will use professional optical instruments to conduct one-time precise measurement and calibration of the physical parameters of the specific model display screen used by the device, such as accurate pixel pitch (and then converted to pixel density PPI), brightness, contrast, etc., and these calibration data are fixed in the device system. At the same time, the device designs precise head and chin supports and other mechanical structures to force the testee to fix his head and eyes at a strictly unchanged physical distance from the screen.
[0004] Although this technical mode can ensure the accuracy of detection at the beginning of the device factory, in the long-term practical application, its inherent technical defects are increasingly exposed. First, this mode has serious precision drift and long-term reliability risk. As an electronic component, the performance of the display will inevitably decay with the increase of use time, for example, the backlight brightness of liquid crystal display will gradually decrease, and the material of organic light emitting diode (OLED) display screen will age and cause color performance to deviate. At the same time, the parameters of electronic components such as driving circuit may drift due to changes in environmental temperature and humidity or component aging. These physical changes make the initial calibration data fixed in the device no longer match the actual performance of the hardware, and the detection error generated thereby will slowly accumulate over time without the user's awareness, causing the device to unknowingly lose its accuracy.
[0005] Secondly, the maintenance cost of the existing technical mode is high and the process is complex. Once the core components of the device, especially the display, need to be replaced due to failure or accidental damage, even if the same type of spare parts are replaced, it cannot be guaranteed that their physical parameters will be completely consistent with the original components. Small batch differences are enough to cause the original fixed calibration data to be completely invalid. At this time, the manufacturer's engineers who have professional technology and special equipment must be present to perform a set of complex and expensive re-calibration process, which not only brings high maintenance cost to the user, but also causes business downtime during device maintenance.
[0006] Furthermore, the existing technical mode leads to poor consistency and comparability of detection results across devices. In large-scale vision screening, multi-center clinical trials or epidemiological studies, etc. The application scene has very high requirements for data consistency. This drawback is particularly prominent. Even through mechanical structures such as chin rests and chin rests, the small movements and posture changes of the subject's head cannot be completely avoided, and any deviation from the standard viewing distance will directly affect the angle of the optotype formed on the retina, thereby interfering with the accuracy of the detection results. Due to the inherent differences in display screens, cameras and other core hardware of devices of different brands, models and even different production batches, it is difficult to effectively guarantee the consistency and comparability of their detection results, which greatly limits their application value in multi-center clinical research, large-scale epidemiological investigation and other scenes that require high consistency, thereby seriously weakening the data standardization advantage that digital detection should bring.
[0007] In summary, the existing professional digital visual detection technology is trapped in a rigid and static implementation mode. Its technical system lacks effective closed-loop feedback and dynamic adaptive capability, and cannot cope with the challenges brought by natural aging of hardware, component replacement and slight changes in the posture of the tested person, thus there are technical bottlenecks in precision retention, maintenance convenience and data consistency that need to be solved. SUMMARY
[0008] An aspect of the present application is to provide a visual adaptive detection method and system based on a generative adversarial network, aiming to solve the technical problems in the prior art that digital visual detection devices rely on static physical calibration, resulting in easy precision drift, high maintenance cost and poor cross-device consistency.
[0009] To solve the above technical problems, the present application provides a visual adaptive detection method based on a generative adversarial network. This method no longer follows the traditional technical path of measuring physical parameters and compensating calculations, but adopts a perception-generation technical approach. The method includes the following steps: first, the system perceives the state of the tested person in the current detection environment in real time through its perception unit, such as a camera, and processes and encodes these state information, such as the real-time distance between the tested person's eyes and the display device, the posture of the head, etc., into a real-time environmental feature vector that can be understood by the machine.
[0010] Next, the module responsible for process control inside the system will decide the target visual level that needs to be presented to the tested person in the current round of testing according to a set of pre-set adaptive visual detection algorithms based on psychophysical principles, such as the QUEST algorithm or the Best PEST algorithm.
[0011] Then, the real-time environmental feature vector and the target visual level obtained in the previous two steps are input into a pre-trained conditional generative adversarial network (cGAN) model. The core function of this cGAN model is to serve as an inverse optical solver, which can generate a unique pre-distorted target image in real time and deterministically according to the input conditions. The special feature of this image is that although its pixel pattern displayed on the screen may be slightly different from the standard target, it is precisely designed so that after passing through the current specific, non-ideal detection environment (including the current distance, viewing angle, etc.) and the tested person's eye optical system, the final optical imaging on the retina can achieve the same effect as the retina imaging formed by a standard target presented under perfect standard conditions.
[0012] Subsequently, this instantaneously generated pre-distorted target image is presented on the display device for the subject to recognize.
[0013] Finally, the system acquires the feedback from the subject (e.g. whether the recognized direction is correct) through the response device, and feeds this response information back to the adaptive vision detection algorithm. The algorithm updates its internal estimate of the subject's vision threshold based on this response, and decides on the next target vision level to be tested, thus initiating a new round of the "determine level - generate image - display image" iteration. This iteration loop continues until the algorithm decides that its estimate of the subject's vision threshold has reached sufficient accuracy, i.e. the algorithm converges, at which point the loop terminates and the system outputs the final vision detection result.
[0014] In a preferred embodiment, the step of real-time sensing the subject's detection environment is implemented by an image acquisition device, such as an industrial grade camera, which acquires a sequence of images containing the subject's face and eyes. The system then processes these image sequences by running efficient computer vision algorithms, such as face detection, facial landmark localization, 3D reconstruction, etc., to accurately solve for one or more of the subject's eye position in 3D space relative to the display device (especially the distance in depth, i.e. the Z-axis), the 3D pose angles of the head (pitch, yaw, roll), and the physical diameter of the pupil determined by lighting and physiological state. These quantified parameters are then organized and encoded into a structured data, which is the real-time environmental feature vector.
[0015] In a preferred embodiment, the conditional generative adversarial network model's ability to generate accurate pre-distorted target images is obtained through a rigorous offline training process. The model consists of a generator and a discriminator. The core of the training process lies in the construction and utilization of a retinal imaging simulator, which can simulate the physical optical process inside a computer. This simulator is a differentiable mathematical model that can simulate the blurring projection of a target image on the retina after it is emitted from the screen and passes through the human eye optical system (which includes various aberration effects). The goal of the training is to jointly optimize the network parameters of the generator through the adversarial game between the generator and the discriminator, combined with the constraint of an additional imaging consistency loss function, so that the pre-distorted target image generated by the generator can achieve high consistency with a standard target image in the pixel level after being processed by the retinal imaging simulator.
[0016] Further, in order to make the simulation closer to the real physiological conditions, the human eye optical system model in the retina imaging simulator, which is the core, is characterized by constructing a dynamic point spread function (PSF) that can change with conditions. This PSF can comprehensively reflect all the blurring and distortion effects of the human eye optical system. Specifically, the shape and parameters of the PSF are modeled as at least a function of the three-dimensional spatial position of the subject's eye relative to the display device (the distance determines the degree of defocus blur) and the pupil physical diameter (the pupil size affects diffraction and aberration), etc., thereby realizing dynamic and personalized simulation of the optical properties of the human eye under different detection conditions.
[0017] Another aspect of the present application is to provide a vision adaptive detection system based on a generative adversarial network, which can implement the above method. The system includes an image acquisition unit for real-time perception of the detection environment of the subject, a display unit for displaying the target image, and a processing unit connected with the former two.
[0018] In the system, the processing unit is the core, which is configured to run multiple functional modules. First, it runs the environment feature encoding module, which is responsible for processing the information collected by the image acquisition unit in real time and extracting the aforementioned real-time environment feature vector from it. Second, it runs the vision detection logic control module, which has a pre-set adaptive vision detection algorithm inside, responsible for deciding the target vision level to be tested according to the response history of the subject. Third, and most importantly, it runs the target synthesis generative adversarial network module, which integrates a pre-trained conditional generative adversarial network model inside. The core task of this module is to receive the target vision level from the logic control module and the real-time environment feature vector from the encoding module, and based on these two conditions, to instantly generate the aforementioned pre-distortion target image that can achieve retinal equivalent optical imaging in the current environment.
[0019] Subsequently, the processing unit will control the display unit to present the pre-distortion target image just generated to the subject. After obtaining the response of the subject, the processing unit will drive the vision detection logic control module and the target synthesis generative adversarial network module to perform the next round of iteration, and this closed-loop process will continue until the adaptive algorithm determines that the detection converges, and the processing unit finally outputs a stable and accurate vision detection result.
[0020] In a preferred embodiment, the environment feature encoding module of the system specifically functions to acquire and process image sequences containing the face and eyes of the subject, calculate one or more key parameters such as eye-screen three-dimensional spatial position, head posture angle, and pupil physical diameter through computer vision algorithms, and encode them into a real-time environment feature vector.
[0021] In a preferred embodiment, the conditional generative adversarial network model in the system is obtained through an offline training process containing a differentiable retinal imaging simulator. The simulator, as a mathematical model, can simulate the retinal projection of an object image after passing through a human eye optical system model, and through the joint optimization of the model generator by the adversarial loss and the imaging consistency loss, the model generator is enabled to generate pre-distorted object images.
[0022] Further, the core of the human eye optical system model is a point spread function model that can dynamically generate a corresponding point spread function according to the parameters contained in the real-time environmental feature vector.
[0023] In a particularly preferred embodiment, the system of the present application is completely integrated into an integrated professional vision detection device. In the device, the optical axis of the image acquisition unit for perception and the geometric center of the display unit for presentation have undergone precise coaxial calibration, and this configuration helps to simplify the geometric relationship calculation, thereby improving the accuracy of environmental feature perception.
[0024] In summary, the present application provides a vision adaptive detection method and system based on a generative adversarial network, aiming to replace the traditional "static calibration-compensation correction" mode with a dynamic perception-generation paradigm to solve the precision drift, maintenance difficulty and poor cross-device consistency of digital vision detection equipment in the prior art due to reliance on static physical calibration. The scheme first acquires image information containing the face and eyes of the subject in real time and high frequency through the image acquisition unit configured in the system. Then, the environmental feature encoding module performs in-depth computer vision analysis on the acquired image sequence to accurately calculate multi-dimensional key parameters including the three-dimensional spatial position of the subject's eyes relative to the display unit, the three-dimensional attitude angle of the head, and the real-time pupil physical diameter, and encodes these parameters into a structured real-time environmental feature vector. At the same time, the vision detection logic control module decides the target vision level to be tested according to adaptive algorithms such as QUEST and combines the historical response of the subject. The key technical point of the scheme is that the conditional generative adversarial network model is pre-trained offline, and the real-time environmental feature vector and the target vision level are received as its dual conditional input to generate a precisely calculated pre-distorted target image in real time and deterministically. The pre-distorted target image can form an optical image on the subject's retina that is completely equivalent to the standard target under standard detection conditions after passing through the human eye optical system in the current specific and non-ideal detection environment, thereby realizing integrated and dynamic reverse compensation of all environmental disturbance factors. To achieve this function, the network model is trained through an offline process that includes a differentiable retinal imaging simulator. The simulator mathematically models the optical properties of the human eye through a conditional point spread function that can dynamically change with environmental conditions, and through the joint optimization of adversarial loss and imaging consistency loss, the network learns this inverse optical solving ability. In actual detection, the system continuously iterates the perception, decision and generation steps in a closed loop until the detection algorithm converges, and finally outputs an accurate and highly reliable vision detection result. Compared with the prior art, the present application has significant beneficial effects.
[0025] Firstly, the present application realizes dynamic real-time self-calibration, which fundamentally guarantees the detection accuracy and long-term reliability. The technical paradigm relying on static factory calibration is completely abandoned. By dynamically generating an optically equivalent target through real-time perception of the subject's accurate state, the system can instantaneously and automatically compensate and calibrate potential errors caused by device aging, environmental changes, and especially the subject's slight head position changes. This fundamentally solves the precision drift problem of traditional devices due to the "calibration-reality" disconnection, ensuring the high accuracy of each detection result and the long-term reliability throughout the device's life cycle.
[0026] Secondly, the present application ensures high consistency and comparability of cross-device results. The present application takes a unified, physiological-optics-model-based "retina equivalent imaging" as the underlying gold standard for all tests, instead of relying on each device's own, different, physical calibration data. This makes the output of vision tests by different devices, different production batches, and even different manufacturers, as long as they are equipped with the technology of the present application, have extremely high consistency and comparability. This has great scientific and social value for establishing standardized visual health databases, carrying out large-scale epidemiological surveys and multi-center clinical studies, etc.
[0027] Thirdly, the present application simplifies the production and after-sales maintenance processes, and reduces the whole life cycle cost. In the production link, the present application has self-calibration capability, has higher tolerance to the selection requirements (batch differences) of core hardware such as display screens, and completely eliminates the complex, time-consuming and expensive manual precise calibration process in traditional devices, thereby improving the production efficiency. In the after-sales maintenance link, when the display screen and other components need to be replaced, there is no need for professional technicians to go on-site for re-calibration. The device can automatically adapt to the characteristics of the new hardware after replacement, thereby greatly reducing the maintenance cost and downtime caused by maintenance.
[0028] Finally, the present application improves the automation level of the test and the clinical work efficiency. The present application completely builds the environmental self-adaptation capability into the system, maximally reduces the dependence on the professional experience of the tester and the potential human operation error. The automatic process and the high tolerance to the natural state of the testee not only improve the efficiency of a single test, but also improve the test experience and cooperation degree of the testee, especially children or special groups. BRIEF DESCRIPTION OF DRAWINGS
[0029] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.
[0030] Figure 1 The present application provides a kind of visual self-adaptive detection system overall architecture schematic diagram for embodiment.
[0031] Figure 2 The present application provides a kind of visual self-adaptive detection method flow chart for embodiment.
[0032] Figure 3 The principle schematic diagram of the conditional generative adversarial network (cGAN) offline training process in the embodiment of the present application.
[0033] Figure 4 The figure is a schematic diagram of the working principle of the retinal imaging simulator in the embodiment of the present application.
[0034] Figure 5 The figure is a schematic diagram of the working process of the real-time detection cycle in the embodiment of the present application.
[0035] Figure 6 The figure is a schematic diagram of the functional modules of the vision adaptive detection system provided by another embodiment of the present application.
[0036] Figure 7 The figure is a schematic diagram of the technical effect comparison between the embodiment of the present application and the prior art under different disturbance conditions.
[0037] Figure 8 The figure is a schematic diagram of the convergence process of the QUEST adaptive vision detection algorithm in the embodiment of the present application.
[0038] Figure 9 The figure is a schematic diagram of the dynamic change of the point spread function with the detection condition in the embodiment of the present application. DETAILED DESCRIPTION
[0039] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0040] It should be noted that in the description of the present application, unless otherwise explicitly specified and limited, the terms "mounting", "connection", "connecting" should be understood in a broad sense, for example, it can be fixed connection, or detachable connection, or integrally connected; it can be mechanical connection, or electrical connection; it can be directly connected, or indirectly connected through an intermediate medium, or the internal communication of two elements. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.
[0041] Embodiment one The embodiment of the present application provides a vision adaptive detection method based on a generative adversarial network, which can be applied to the professional integrated vision detection equipment as shown in the figure. Figure 1 The core of the method is to realize the dynamic adaptation to the state of the subject through a closed-loop "perception-generation-response" process, so as to ensure the accuracy of the detection result. Please refer to the figure, which shows the detailed process of the method of the embodiment, including the following steps: Figure 2 Step S101: Real-time perception of the detection environment of the subject, and obtaining the real-time environment feature vector.
[0042] In a specific application scenario, when the subject is guided by the detection personnel to stably place the head on the forehead support and chin support of the integrated vision detection device shown in the figure, and is ready to start detection, this step is activated and continuously runs. This step is executed by the environment feature encoding module 106 in the system, which serves as the perception front end of the system. The purpose is to convert the complex and dynamic detection environment in the real world into a structured digital signal that can be understood and utilized by subsequent deep learning models, i.e., the real-time environment feature vector c_current. Figure 1
[0043] Specifically, the implementation of this step relies on high-precision image acquisition and complex computer vision analysis. The system is equipped with an industrial-grade high-frame-rate camera 102, which continuously acquires high-definition video streams containing the subject's 101 face, especially the eye region, at a frequency of, for example, 60 frames per second or higher.
[0044] Preferably, in order to improve the accuracy and robustness of subsequent calculations, the system can perform a series of image preprocessing operations before acquiring the image, including but not limited to: image graying to reduce color information interference, histogram equalization to enhance image contrast under different lighting conditions, and image denoising filtering to eliminate sensor noise.
[0045] After obtaining the video stream, the algorithm inside the environment feature encoding module 106 starts processing each frame of image. First, a lightweight but efficient face detection algorithm, such as the deep learning-based MTCNN (Multi-task Cascaded Convolutional Networks) or the mobile-optimized YOLO (You Only Look Once) variant, is run to quickly and accurately mark the bounding box of the face in the image.
[0046] After the face region is determined, the system then runs a more detailed face key point positioning algorithm, such as the MediaPipe Face Mesh or the key point detection model in the Dlib library. This algorithm can accurately and stably locate the two-dimensional pixel coordinates of dozens or even hundreds of key points within the face region, including eyebrows, eye contours, nose, mouth, and facial contours. Among them, the most critical for this invention is the ability to accurately and stably locate the center of the pupils, the corners of the eyes, and other eye feature points.
[0047] After obtaining the coordinates of these two-dimensional key points, the system enters the three-dimensional information calculation stage. Using a pre-established camera calibration model (containing the camera's intrinsic parameter matrix and distortion coefficients) and a general 3D Morphable Model (3DMM), the system can employ the classic Perspective-n-Point (PnP) algorithm or its variants to calculate the three-dimensional spatial coordinates of the facial key points in the camera coordinate system. Through this process, the system can accurately calculate the three-dimensional spatial position of the subject's pupil centers relative to the display device's 103-screen plane, where the Z-axis coordinate is the high-precision real-time eye-to-screen distance. Simultaneously, by analyzing the orientation of the facial key points in three-dimensional space, the three-dimensional head posture angles—pitch, yaw, and roll—can be calculated.
[0048] In addition, this module segments and fits the pupil region in the image, calculating the pixel diameter of the pupil in the image. Combined with the previously calculated real-time eye-to-screen distance, the true physical diameter of the pupil can be obtained through simple geometric optics conversion.
[0049] Finally, all these precisely quantified, dynamically changing parameters, such as [eye-to-screen distance Z, pitch angle P, yaw angle Y, roll angle R, pupil diameter D_pupil], are organized into a structured vector, namely the real-time environment feature vector c_current. This vector is refreshed at the beginning of each detection loop to ensure that subsequent target generation is always based on the latest and most accurate on-site environment information.
[0050] Step S102: Determine the target visual acuity level to be tested based on the adaptive vision detection algorithm.
[0051] This step is executed by the system's vision detection logic control module 108. The core of this module is a built-in adaptive vision detection algorithm program based on psychophysical principles, which is used to detect the visual threshold of the test subject in the most efficient and accurate way.
[0052] In a preferred embodiment, the module employs the QUEST (Quick Estimation by Sequential Testing) algorithm. The QUEST algorithm intelligently selects the next most informative test point (i.e., visual acuity level) by updating a posterior probability density function about the subject's visual acuity threshold after each test.
[0053] Specifically, at the beginning of the test, the algorithm initializes the visual acuity threshold with a prior probability distribution based on large population statistics. Then, it selects the mean or mode of the distribution as the target visual acuity level v_initial for the first test. Upon receiving the subject's response (correct or incorrect) to the previous optotype, the QUEST algorithm updates the posterior probability distribution of the visual acuity threshold according to its internal Weber function model and Bayes' theorem. For example, if the answer is correct, the distribution moves and narrows towards higher visual acuity (smaller LogMAR value); vice versa. Subsequently, the algorithm selects the mean or mode of the updated posterior probability distribution as the target visual acuity level v_next for the next test. This process makes the selection of test points always tend to the current most uncertain area, thereby greatly reducing the number of tests required.
[0054] As shown in Figure 8 , the convergence process of the QUEST adaptive visual acuity test algorithm includes an initialization phase, an iterative test phase, and a convergence determination phase. In the initialization phase, the system establishes a prior probability distribution based on large population statistics, with an initial standard deviation of about 0.30 LogMAR. In the iterative test phase, the algorithm dynamically updates the posterior probability distribution according to the response result (correct or incorrect) of each test: when the subject correctly identifies the optotype, the posterior probability distribution moves towards higher visual acuity; when the identification is incorrect, the distribution adjusts towards lower visual acuity. As the iteration proceeds, the standard deviation of the posterior probability distribution continues to decrease, from the initial 0.30 to less than 0.05. The right side of the figure shows the morphological changes of the posterior probability distribution at three typical moments: at iteration 1, the distribution is wide (standard deviation 0.30), indicating high uncertainty; at iteration 9, the distribution is significantly narrowed (standard deviation 0.10), indicating more accurate estimation of the threshold; at iteration 18, the distribution is the narrowest (standard deviation 0.042), having reached the convergence condition. The entire process usually completes within 15-20 iterations, significantly reducing the test time compared to traditional methods, while ensuring the detection accuracy.
[0055] Alternatively, other adaptive algorithms such as the Best PEST (Best Parameter Estimation by Sequential Testing) algorithm or the simplified up-and-down method can also be used. The common feature of these algorithms is that they can dynamically adjust the difficulty of subsequent tests based on the subject's historical responses to achieve rapid convergence of the visual acuity threshold.
[0056] As shown in Figure 5As shown, the real-time detection cycle adopts a closed-loop perception-generation-response architecture. After detection initiation, the system continuously captures images at 120 fps and calculates eye-screen distance, head pose, and other parameters, outputting the real-time environmental feature vector c current. The logic control module determines the initial level based on the prior distribution according to the QUEST algorithm, and in subsequent iterations, updates the posterior probability distribution according to the response, and selects a new target visual acuity level v next. The GAN module inputs c current and v next into the conditional generative adversarial network, and generates the pre-distorted target image I generated within about 30 ms. The display and response module renders I generated to the display screen, and the subject feeds back the identification result through the response device, and the system records the response data. The convergence judgment module checks the standard deviation of the posterior probability distribution, and if it has not converged, returns to the level decision to continue iteration, and if it has converged, outputs the final visual acuity level result. The environmental perception module and the target generation module are kept in real-time synchronization through a continuous updating mechanism. A typical test process has about 18 iterations, with a total time of about 35 seconds, realizing efficient closed-loop control of adaptive visual acuity detection.
[0057] In a preferred embodiment, the initialization stage loads a pre-trained model and sets the initial parameters of the QUEST algorithm (mean μ = 0.0, standard deviation σ = 0.3 LogMAR). The environmental perception stage acquires images at 120 fps, extracts 137 facial key points through MediaPipeFace Mesh, and calculates eye-screen distance (accuracy ± 1 mm) and head three-axis pose angle (accuracy ± 0.5°) using the EPnP algorithm. In the decision-making stage, the mean of the posterior probability distribution is selected as the next test level according to the QUEST algorithm, for example, when P(θ > 0.1) = 0.75, v next = 0.1 is selected. The generation stage (504) inputs v next and c current into the optimized TensorRT engine, and outputs the pre-distorted image within 30 ms, using 16-bit floating-point precision to ensure visual quality. The display stage uses OpenGL texture mapping to achieve pixel-to-pixel rendering without scaling, and the brightness is calibrated to 85 ± 5 cd / m. The response stage (506) records the subject's keystrokes, calculates the response time (normal range 200-1500 ms), and verifies the effectiveness. The convergence judgment stage calculates the standard deviation of the posterior distribution, and if σ < 0.05 LogMAR, the test is terminated. The entire cycle adopts a double buffering architecture, with perception and generation executed in parallel, ensuring an inter-frame delay of < 50 ms. The exception handling path includes: when the eye-screen distance < 0.4 m or > 6.0 m, visual cues are triggered to guide adjustment; when a single inference > 100 ms, a lightweight fallback model is enabled; when the head pose angle > ± 20°, the test is paused and a repositioning prompt is given.
[0058] Step S103: input the real-time environment feature vector and the target visual acuity level into the cGAN model to generate a pre-distorted optotype image.
[0059] This step is performed by the optotype synthesis GAN module 107. This module receives as dual inputs the real-time environment feature vector c current from step S101 and the target visual acuity level v next from step S102, and its task is to instantly generate a pre-distorted optotype image I generated that is capable of achieving retinal equivalent optical imaging under the current specific environment.
[0060] In one specific implementation, this module internally deploys a conditional generative adversarial network (cGAN) model that is thoroughly offline trained and embedded optimized. The Generator part of this model, its network structure preferably adopts the U-Net architecture. When receiving the inputs, the target visual acuity level v next (e.g. a LogMAR value of 0.1) is first converted into a standardized, corresponding-size ideal optotype template image I target, while the real-time environment feature vector c current is expanded to the same spatial dimension as I target. These two are concatenated in the channel dimension with an optional random noise vector to form a multi-channel input tensor, which is fed into the encoder part of the Generator for feature extraction and compression. Subsequently, the decoder part utilizes these condensed features and gradually reconstructs a high-resolution output image, i.e. the pre-distorted optotype image I generated, through deconvolution and skip connections.
[0061] For example, if the real-time environment feature vector c current indicates that the subject is currently 5% farther away than the standard distance, and has a slight left tilt of the head. Then, the Generator, in order to counteract these effects, might generate an optotype image that is slightly larger in physical size than the standard 0.1 optotype, and has a slight reverse tilt and barrel distortion at the pixel level. This image, although it might look "unstandard" on the screen, is designed to achieve a "standard" projection on the retina after passing through the current optical transmission.
[0062] To achieve real-time generation, the cGAN model deployed on the embedded high-performance computing unit 104 (such as the NVIDIA Jetson series), is usually subjected to specialized optimization, such as using NVIDIA TensorRT for model quantization and operator fusion, to ensure that the time for a single forward inference is controlled within tens of milliseconds, so as to be instantaneous for the subject and the operator.
[0063] In one preferred embodiment, the cGAN is tasked to solve as a high-precision inverse optical function solver. Its goal is to accurately and deterministically generate an input image that can reverse the impact of the human eye optical system, given a specific target (i.e. desired visual acuity level) and specific environmental conditions (i.e. current detection environment).
[0064] The entire cGAN module consists of two deep neural networks: the generator (G) and the discriminator (D). Both form a dynamic adversarial relationship during the training process, playing against each other, co-evolving, and ultimately reaching an ideal balance point. The detailed design and training process are as shown in Figure 3 and specifically as follows.
[0065] (I) Design of the Generator (G) 1. Network structure: The generator G preferably adopts a U-Net-based encoder-decoder architecture. Its unique "skip connection" mechanism allows the network to utilize both high-level semantic information and low-level fine textures when decoding and reconstructing images, which is crucial for generating sharp-edged target images.
[0066] 2. Input: The generator G receives a multi-modal conditional input, usually by concatenating the following three pieces of information in the channel dimension: Target visual acuity level image I_target: A standardized ideal target template image corresponding to the target LogMAR level.
[0067] Environmental feature vector c: A real-time environmental vector provided by the environmental feature encoding module, which is expanded to the same spatial dimension as I_target.
[0068] Random noise vector z (optional): To increase the randomness of the generation process and learn more robust features.
[0069] 3. Output: Output a "pre-distorted" target image I_generated = G(I_target, c, z) with the same size as the input image, which is used for direct display on the device screen.
[0070] Preferably, to enable those skilled in the art to implement it, an exemplary network structure and parameter configuration of the generator (Generator) in the conditional generative adversarial network model are given below. The U-Net architecture adopted by the generator can specifically include a symmetric encoder-decoder structure with a total of 8 convolutional layer levels.
[0071] Encoder part: The input image first goes through a convolutional layer, which increases the number of channels from the input channel number to 64. Subsequently, the encoder contains 4 down-sampling modules. Each down-sampling module consists of a convolutional layer with a stride of 2x2 (to halve the feature map size), a Batch Normalization layer, and a LeakyReLU activation function. The output channel number of each module doubles successively, e.g., from 64 -> 128 -> 256 -> 512.
[0072] Decoder part: The decoder correspondingly contains 4 up-sampling modules. Each up-sampling module consists of a Transposed Convolution with a stride of 2x2 (to double the feature map size), a Batch Normalization layer, and a ReLU activation function. The output channel number of each module halves successively, e.g., from 512 -> 256 -> 128 -> 64. Crucially, the input of each up-sampling module of the decoder is concatenated with the output feature map of the corresponding level of the encoder in the channel dimension via a skip connection, to fuse the low-level detail information.
[0073] Specific method of input concatenation: In the input stage, the target visual acuity level image I_target (e.g., a single-channel binary image of 64x64 pixels), the expanded real-time environment feature vector c_current (e.g., a 5-parameter vector is expanded into a 64x64x5 tensor), and the expanded random noise vector `z` (e.g., a 100-dimensional noise vector is expanded into a 64x64x1 tensor) are concatenated in the channel dimension into a unified, multi-channel input tensor, which can have dimensions of 64x64x(1+5+1=7).
[0074] Output layer: The last layer of the decoder is connected to a convolutional layer, which uses a Tanh activation function and maps the channel number to 1 (for grayscale images), outputting the final, pre-distorted target image I_generated with pixel values in the range [-1, 1].
[0075] Optimization strategies for models on embedded devices: To ensure real-time performance on embedded devices such as NVIDIA Jetson, the trained cGAN model, particularly the generator part, can employ a series of optimization strategies. Preferably, deployment optimization can be performed using NVIDIA TensorRT, which includes the following specific steps: First, convert the PyTorch or TensorFlow trained model to ONNX (Open Neural Network Exchange) format; then, TensorRT parses the ONNX model and performs a series of techniques including layer and tensor fusion (e.g., fusing convolution, batch normalization, and activation functions into a single CBR layer), weight and activation precision calibration (e.g., from FP32 to FP16 or INT8 quantization to significantly improve computing speed and reduce power consumption), kernel automatic adjustment, and dynamic tensor memory optimization. Through these optimizations, the single forward inference time of the generator can be controlled within, for example, 30 milliseconds, fully meeting the needs of real-time detection. (B) Design of Discriminator (D) 1. Network structure: The discriminator D preferably adopts the PatchGAN structure, which divides the input image into N x N small blocks and performs independent authenticity discrimination on each block, forcing the generator to focus on and optimize the local high-frequency details of the image.
[0076] 2. Input: The discriminator D receives a pair of images for conditional discrimination: Image to be discriminated: It can be a simulated retinal image S(G(I_target, c)) generated by the generator, or a real standard retinal image S(I_standard). Where S represents the retinal imaging simulator.
[0077] Conditional information: I_target and c are also input as conditions, enabling the discriminator to learn to determine: “Under the given condition c, does this retinal image correspond to the target imaging of the standard visual level v?” 3. Output: An N x N probability matrix indicating the probability of each image block being “true”.
[0078] (III) Adversarial training process and loss function The training process is an iterative and alternating optimization process, as shown in Figure 3 The specific steps are as follows: First, a batch of training samples is randomly selected from the pre-constructed dataset, each sample contains a pair of (standard optotype image I standard, and its corresponding environmental condition c). Then, the target visual acuity v corresponding to I standard, environmental condition c and a random noise z are input into the current generator G to obtain a batch of "fake" pre-distorted optotype images generated by the generator. Here, I fake is defined as the image generated by the generator in the current training iteration step, i.e. I fake = G(v, c, z).
[0079] 1. Train discriminator D: fix G, pass I fake and I standard through retinal imaging simulator S to get the corresponding simulated retinal images, i.e. S(I fake) and S(I standard). Send S(I fake) and S(I standard) obtained by simulator S to D, calculate the adversarial loss, and update the weights of D.
[0080] 2. Train generator G: fix D, calculate the adversarial loss of generator G. At the same time, introduce an important imaging consistency loss (Reconstruction Loss), usually L1 loss: Loss recon = || S(I fake) - S(I standard) ||_1. This loss directly supervises and forces I fake to be highly consistent with I standard in pixel level after simulation.
[0081] 3. Total loss: the total loss of G = adversarial loss + λ × imaging consistency loss. Update the weights of G according to the total loss.
[0082] 4. Iterative loop: repeat the above steps until the network converges.
[0083] Step S104: display the pre-distorted optotype image on the display device.
[0084] This step is performed by the display and response module 109. This module receives the pixel matrix of the pre-distorted optotype image I generated generated by step S103. A key technical detail is that this module will render this image directly to the central area of the medical-grade high-resolution display screen 103 without any scaling, interpolation or operating system rendering engine intervention in a pixel-to-pixel manner. This ensures that each pixel value calculated by the generator can be accurately converted to the light intensity of the corresponding pixel point on the display screen, thereby ensuring the fidelity of the visual stimulus.
[0085] Step S105: Acquire subject response.
[0086] After the target is displayed, the subject 101 recognizes it and presses the button on the handheld wireless response device 105 corresponding to the direction of the opening of the target (e.g. up, down, left, right) as perceived by the subject. The display and response module 109 captures this button press signal via the wireless receiver and decodes it into structured response data (e.g. {timestamp:..., response: left}) which is then transmitted to the vision detection logic control module 108.
[0087] Step S106: Determine if the adaptive algorithm has converged.
[0088] Upon receiving the subject response, the vision detection logic control module 108 first compares the response to the true direction of the target presented at the time of the response to determine if it is "correct" or "incorrect". Then, as described in step S102, the result is used to update the probability density function maintained by the module to reflect the subject's visual acuity threshold.
[0089] After each update, the module checks a pre-defined convergence condition. In a preferred embodiment, the convergence condition is defined as whether the standard deviation of the probability density function is less than a pre-defined threshold (e.g. 0.05 in LogMAR units). The smaller the standard deviation, the more accurate the estimate of the visual acuity threshold and the lower the uncertainty.
[0090] If the result of the check is "no", i.e. the standard deviation is still greater than the threshold, it means that the detection has not yet reached sufficient accuracy and the process jumps back to step S102 to start a new iteration of the test.
[0091] If the result of the check is "yes", it means that the algorithm has successfully converged.
[0092] Step S107: Obtain and output the final vision detection result.
[0093] When the algorithm has converged, the detection process terminates. The vision detection logic control module 108 outputs the mean or mode of the probability density function it maintains as the final visual acuity threshold result of the test. This result can be formatted in a variety of standard forms for output and display, such as a LogMAR value (e.g. -0.1), decimal notation (e.g. 1.2) or fraction notation (e.g. 20 / 16). At the same time, the system can also output quality control parameters of the test, such as the average response time throughout the process, head stability indicators (derived by analyzing the changes in the c_current vector over time), and the confidence interval of the result, to provide richer reference information for clinical diagnosis. Finally, the process ends.
[0094] Further, in the technical system of the present application, the task of the conditional generative adversarial network (cGAN) model is to generate a "pre-distorted" optotype image that can be displayed on a physical screen, so that after the real-world optical transmission process, it can form an imaging on the retina of the subject that is completely equivalent to the ideal standard optotype. In order to be able to effectively supervise the training of GAN, it is necessary to have a tool that can simulate this real-world optical transmission process inside the computer. This tool is the retinal imaging simulator.
[0095] It needs to be emphasized that this simulator is not a physical hardware device, but a differentiable software module that is completely based on mathematical modeling and computer construction. Its differentiable nature is crucial because it ensures that the error gradient produced by the discriminator can be smoothly back-propagated back to the generator during the training process of the GAN, thereby effectively optimizing the network weights of the generator.
[0096] The core purpose of the simulator is to determine what kind of optical projection will be observed on the retina of a person's eye with specific optical characteristics (such as pupil size, aberration) located at a specific three-dimensional spatial position, given a digital image displayed on a screen with specific parameters (such as PPI, brightness).
[0097] For this, the simulator simulates the full-link process of light from the screen pixel point to the final projection on the retina photoreceptor. This process can be decomposed and modeled as the following cascading steps: 1. Screen imaging model: simulates how a digital image is converted into the emission of light on a physical screen. This step mainly considers the pixel density (PPI) of the screen, which converts the input pixel matrix into a light intensity distribution function on a two-dimensional physical plane with real physical dimensions (unit: millimeter).
[0098] 2. Spatial light propagation: simulates the process of light propagating from the screen plane to the cornea surface of the human eye. In the geometric optics level, this process is mainly determined by the eye-screen distance (D), which directly affects the opening angle of the object at the human eye node.
[0099] 3. Modeling of human eye optical system - Point Spread Function (PSF): This is the most core and complex part of the entire simulator. Human eye is not a perfect optical system, it has various aberrations (such as spherical aberration, coma, astigmatism, etc.) and diffraction effects. Therefore, an ideal point light source from infinity, after focusing through the human eye, will not form an infinitely small perfect point on the retina, but will be diffused into a light spot with a specific shape and intensity distribution. The two-dimensional intensity distribution function of this light spot is defined as the point spread function (PSF) of the optical system.
[0100] In the present application, the PSF is not a fixed function, but a conditional PSF (Conditional PSF) dynamically affected by multiple parameters, which is modeled as a function of the following key parameters: Defocus blur: When the subject has myopia, hyperopia, or other refractive errors, or the observation distance is not appropriate, the image cannot be accurately focused on the retina, which is the main source of blurred vision. The degree of diffusion of the PSF is directly related to the defocus amount.
[0101] Preferably, the defocus blur can be approximated by a Gaussian blur kernel proportional to the defocus amount; the effect of pupil diameter on diffraction can be modeled by an Airy disk function; the high-order aberrations can be parameterized by the first few (e.g., the first 15) coefficients of the Zernike polynomial. In the training phase, the value range of these aberration coefficients can be randomly sampled based on the distribution of large-scale population ophthalmic data, so that the trained model has the generalization ability to the general human eye aberrations.
[0102] Pupil diameter: The pupil size directly affects the light flux entering the eye and the diffraction effect of the optical system. The larger the pupil, the less obvious the diffraction effect, but the impact of various aberrations will be intensified; vice versa. In the simulator, a corresponding pupil diameter will be estimated according to the preset environmental light conditions and the human eye physiological model (such as Watson & Yellott model), and it will be used as one of the parameters for generating PSF.
[0103] Higher-order aberrations: To further improve the simulation accuracy, a higher-order aberration model based on wavefront analysis (such as Zernike polynomial) can be introduced. In the training phase, an "average" higher-order aberration model can be used, which is statistically derived from large-scale population ophthalmic data, or in more advanced applications, personalized simulation can be performed for the aberration data of a specific subject.
[0104] Finally, these parameters are integrated into a complex mathematical model (which can be a Fourier transform model based on physical optics, or an approximate model based on machine learning fitting), generating a two-dimensional PSF convolution kernel representing the overall performance of the human eye optical system under current conditions.
[0105] To be able to implement the simulator, the core optical transfer process can be accurately described by a mathematical model based on Fourier optics. The entire simulation process can be expressed as a convolution operation. Specifically, in the spatial domain, the final simulated retinal imaging I_retinal(x', y') is a two-dimensional convolution of the physical light intensity distribution I_physical(x, y) and the conditional point spread function PSF_c(x, y): I_retinal = I_physical * PSF_c Where * represents the two-dimensional convolution operator. For computational efficiency, this convolution operation is usually done in the frequency domain through multiplication.
[0106] First, the physical light intensity distribution I_physical(x, y) can be obtained by scaling the input pixel matrix I_input(i,j), where x = i * p, y = j * p, and p is the pixel pitch calculated from the screen PPI.
[0107] Second, the conditional point spread function PSF_c is the key to this model. In a preferred embodiment, PSF_c can be constructed by the wavefront aberration theory of the human eye optics. The wavefront aberration function W(u, v) on the pupil plane can be represented by a series of weighted Zernike polynomials Z_n^m: W(u, v) = Σ [c_n^m * Z_n^m(ρ, θ)] Where (u, v) are the pupil coordinates, and c_n^m are the coefficients corresponding to each Zernike polynomial. In particular, the defocus blur is mainly determined by the c_2^0 term, whose value is directly related to the eye-screen distance D. Other coefficients can be obtained from large-scale population statistics or personalized measurements.
[0108] The pupil function P(u, v) containing aberrations can be represented as: P(u, v) = A(u, v) * exp[j * (2π / λ) * W(u, v)] Where A(u, v) is the aperture function determined by the physical diameter of the pupil (e.g., 1 within the pupil area and 0 outside the area), and λ is the wavelength of light.
[0109] Finally, according to Fraunhofer diffraction theory, the point spread function PSF is the modulus square of the Fourier transform of the pupil function P(u, v): PSF_c = |FFT{P(u, v)}|^2 where FFT represents the two-dimensional Fourier transform. This PSF_c is a two-dimensional convolution kernel. In this way, the macroscopic environmental parameters (distance, pupil diameter, etc.) and the microscopic optical effects (blurring, aberration) are linked through a rigorous mathematical model, forming a computable and differentiable simulator core.
[0110] Based on the above modeling, the specific calculation process of the simulator is as follows: 1. Input: An input image I_input (an M x N pixel matrix) to be displayed on the screen, provided by a GAN generator or a standard target library.
[0111] A set of condition parameters C describing the current detection environment and the physiological state of the subject, including: screen PPI, eye-screen distance D, pupil diameter, defocus amount, high-order aberration coefficients, etc.
[0112] 2. Step one: Physical size conversion. According to the input PPI, convert I_input from the pixel domain to the physical domain to obtain a two-dimensional physical light intensity distribution I_physical(x, y).
[0113] 3. Step two: Generate conditional PSF. According to the condition parameter set C, calculate the corresponding two-dimensional point spread function convolution kernel PSF_c(u, v) through the PSF generation model.
[0114] 4. Step three: Convolution operation. Perform convolution operation in the Fourier domain to simulate the imaging process of the image after passing through the human eye optical system. The calculation formula is: I_retinal_freq = FFT(I_physical) * FFT(PSF_c) where FFT represents the two-dimensional fast Fourier transform, and * represents element-wise multiplication.
[0115] 5. Step four: Inverse transform to get retinal imaging. Convert the frequency domain result back to the spatial domain through inverse Fourier transform to get the final simulated retinal imaging I_retinal: I_retinal = IFFT(I_retinal_freq) where IFFT represents the two-dimensional inverse fast Fourier transform.
[0116] 6. Output: The simulator finally outputs a blurred and distorted simulated retinal image I_retinal. This image is the direct basis for the GAN discriminator to judge the authenticity.
[0117] As shown in Figure 9 , the point spread function (PSF) is the core component of the retinal imaging simulator, which is used to represent the blurring and distortion effects of the human eye optical system. The PSF in this invention is conditional and will change dynamically with the optical parameters of the detection environment. Figure 9 The 3x2 grid layout shows the morphological changes of PSF under different conditions. The first row fixes the pupil diameter at 4mm, and from left to right, it shows the PSF morphology when the defocus amount is 0D, 1D, and 2D, respectively. It can be observed that as the defocus amount increases, the PSF dispersion degree gradually increases, and the spot range significantly expands, indicating that the imaging blur increases. The second row fixes the defocus amount at 0D, and from left to right, it shows the PSF morphology when the pupil diameter is 2mm, 4mm, and 8mm, respectively. The diffraction effect dominates when the pupil is small (2mm), and the PSF presents a regular Airy pattern, containing a central bright spot and a weak diffraction ring. When the pupil is medium (4mm), the PSF is relatively compact. When the pupil is large (8mm), the aberration effect is intensified, and the PSF morphology appears asymmetric deformation, and the light intensity distribution is no longer completely symmetric. Each subgraph uses a gray-scale graph to represent the two-dimensional intensity distribution of PSF, with the highest intensity in the center showing white, gradually decreasing to black background outward, which directly reflects the spatial distribution characteristics of light energy. Through this parameterized PSF modeling method, this invention can accurately simulate the retinal imaging effect under different refractive states and pupil size conditions, providing accurate conditional image input for subsequent vision adaptive detection.
[0118] Example Two The embodiment of the present invention also provides a vision adaptive detection system based on a generative adversarial network. Please refer to Figure 6 , which is a schematic block diagram of the functional modules of the system. The system can be Figure 1 as shown, which is an integrated professional vision detection device, and the internal hardware entity supports the operation of multiple functional units and modules. The system at least includes: an image acquisition unit, a display unit, and a processing unit.
[0119] The image acquisition unit is used to perceive the detection environment of the subject in real time and provide raw visual information for subsequent processing. Specifically, this unit is used to perform the image acquisition part in step S101 described in embodiment one. In a specific implementation, the image acquisition unit can be an industrial-grade high-frame-rate, global-shutter camera, such as Figure 1Camera, which is responsible for capturing the image of the subject's face and eyes. It is precisely installed in the device, with its optical axis preferably coaxially calibrated with the geometric center of the display unit, to ensure that the image of the subject's face and eyes can be captured clearly and without deviation.
[0120] Display unit, which is responsible for presenting visual stimuli, i.e., the target image, to the subject. Specifically, this unit is used to perform step S104 described in embodiment one. In a specific implementation, the display unit can be a medical-grade high-resolution display screen, such as Figure 1 Display screen in the device. It needs to have high brightness, high contrast, high color accuracy, and excellent brightness uniformity to ensure that the pre-distorted target image generated can be converted into an optical signal faithfully and without distortion.
[0121] Preferably, the optical axis of the image acquisition unit and the geometric center of the display unit are precisely coaxially calibrated. This configuration helps to simplify the geometric relationship calculation from two-dimensional image coordinates to three-dimensional world coordinates, as it can reduce or eliminate the complex correction caused by parallax, thereby improving the accuracy of the environmental feature encoding module in solving parameters such as eye-screen distance and head pose.
[0122] Processing unit, which is the brain and computing core of the entire system, and interacts with the image acquisition unit and the display unit for data exchange and control. The processing unit is configured to run the core logic of the method of the present application. In a specific implementation, the processing unit can be Figure 1 embedded high-performance computing unit in the device, such as a NVIDIA Jetson AGX Orin module that integrates a powerful GPU and NPU. Inside the processing unit, through software programming and hardware acceleration, multiple function modules work cooperatively, at least including: environmental feature encoding module, vision detection logic control module, and target synthesis GAN module.
[0123] Environmental feature encoding module, which is responsible for extracting real-time environmental feature vectors that can quantitatively describe the current detection environment from the raw image information obtained by the image acquisition unit. Specifically, this module is used to perform the computer vision analysis part in step S101 described in embodiment one. In a specific implementation, this module is a set of optimized algorithm programs running on the processing unit. It receives real-time video streams from the image acquisition unit, processes them through a series of algorithms (such as face detection, key point positioning, 3D pose estimation, etc.), and finally outputs structured real-time environmental feature vectors c_current.
[0124] Vision detection logic control module, which is responsible for leading the whole adaptive detection process. Specifically, this module is used to implement the core logic of steps S102, S105, S106 and S107 in embodiment one. In a specific implementation, this module is a state machine or control program running on the processing unit. It internally implements adaptive algorithms such as QUEST, etc., and is responsible for deciding the next target visual acuity level v_next according to the user response received from the display and response module (not explicitly shown in the figure, but the function is included in the interaction between the processing unit and the display unit), and judging whether the detection process meets the convergence condition.
[0125] Synthetic target GAN module, which is responsible for generating optically equivalent pre-distorted target images in real time according to the changing conditions. Specifically, this module is used to implement step S103 in embodiment one. In a specific implementation, this module is a conditional generative adversarial network model deployed on the GPU or NPU of the processing unit, which is optimized by tools such as TensorRT. It receives c_current from the environmental feature encoding module and v_next from the vision detection logic control module as input, and quickly performs a forward inference to output a high-fidelity pre-distorted target image.
[0126] When the system is working, these modules closely cooperate according to the process described in embodiment one. The environmental feature encoding module continuously observes the subject, the vision detection logic control module intelligently decides the test content, the synthetic target GAN module creatively generates test stimuli, and the processing unit presents the stimuli to the subject through the display unit. The whole system constitutes an intelligent detection instrument that can dynamically adapt to the environment and perform real-time self-calibration.
[0127] Embodiment three This embodiment will further illustrate the technical solutions, processes and effects of the present application by combining a specific application scenario, i.e. using the vision detection method and system described in the present application to perform rapid vision screening on a child in a professional ophthalmic clinic.
[0128] In this scenario, the technical requirement for the equipment is to still complete the detection quickly and accurately under the condition that the child's cooperation degree is limited and the head position may frequently move slightly.
[0129] The hardware selection of the equipment is as described in embodiment one, using NVIDIA Jetson AGX Orin 32GB module, 15.6-inch 4K medical-grade display and Basler global shutter industrial camera. The software system completely deploys the functional modules described in embodiment two.
[0130] Before the vision adaptive detection process starts, a one-time offline model training is needed for the system. This process is completed on a high-performance server. The embodiment of the present application constructs a retinal imaging simulator as shown in Figure 4 The parameter space of this simulator covers eye-screen distance D from 0.5 meters to 6 meters, pupil diameter from 2 millimeters to 8 millimeters, and statistical distribution of low-order and high-order aberrations based on Zernike polynomial representation. Millions of training samples are generated, each containing {a standard optotype image I_standard, a set of randomly sampled environmental parameters c, and the simulated retinal imaging S(I_standard) of I_standard under the condition of c}. Specifically, the generation process of the training sample is as follows: first, select standard optotype images of each level from the standard LogMAR visual acuity chart as I_standard; second, randomly sample within the preset parameter space (such as distance D in 0.5-6 meters, pupil diameter in 2-8 millimeters, and each order Zernike coefficient within the statistical range) to generate a set of environmental parameters c; finally, input I_standard and c into the retinal imaging simulator to calculate the corresponding simulated retinal imaging S(I_standard). By repeating this process, a training dataset containing millions of {I_standard, c, S(I_standard)} triplets is constructed.
[0131] Then, according to the process shown in Figure 3 , a cGAN model based on U-Net and PatchGAN structure is trained adversarially for several weeks until the model converges, and the generator has strong reverse optical solving ability. The trained generator model is optimized by TensorRT and deployed to the processing unit of the integrated screening instrument.
[0132] Now, enter the actual detection process.
[0133] A 7-year-old child subject sits in front of the device. The tester adjusts the chin rest to a comfortable position through the electric button on the device according to his height, and guides him to gently rest his forehead. Although the child is required to remain still, his head will still move slightly forward and backward, left and right when he feels curious or nervous.
[0134] The tester clicks the "start" button on the screen, and the detection process officially starts, which is fully automatic for the tester and the subject.
[0135] 1. Initialization and Perception (about 1-2 seconds): After the device is started, the child's eye images are immediately captured by the industrial camera at a frame rate of 120 fps. The environmental feature encoding module starts working, which quickly locks and stably tracks the child's facial key points in the first tens of frames of images. Assuming the initial moment, the module calculates that the child's eye-screen distance is 3.02 meters, and the head posture is basically central. This information is encoded as c_0.
[0136] 2. First round of testing: The QUEST algorithm of the logic control module determines the first target visual acuity level as LogMAR 0.2, taking the average visual acuity of the population as the prior. The instruction v_0 = 0.2 and the environmental vector c_0 are simultaneously sent to the target synthesis GAN module. Within about 30 milliseconds, the GAN module generates a 0.2 target image I_0 for a distance of 3.02 meters and displays it. The child recognizes the direction and presses the response device.
[0137] 3. Subsequent tests with dynamic adaptation: During the process of the child recognizing I_0, his head unconsciously moves forward by about 5 cm. After the system receives its response to I_0, it is ready to start the second round of testing. The environmental feature encoding module immediately captures this change and outputs a new environmental vector c_1, in which the eye-screen distance is updated to 2.97 meters. According to the correct response of the first time, the logic control module determines that the next test level is more difficult, LogMAR 0.1. At this time, the instruction v_1 = 0.1 and the latest environmental vector c_1 are sent to the GAN module. The GAN module generates a new 0.1 target image I_1 for a distance of 2.97 meters. The physical pixel size of this I_1 will be slightly smaller than the 0.1 target image generated at a distance of 3.02 meters, to ensure that the visual angles projected onto the child's retina at their respective distances are completely equivalent.
[0138] 4. Looping and Convergence (about 35 seconds): The above process proceeds quickly and seamlessly. No matter how small the child's head drifts, the target generated by each round of testing is tailored to the precise state of that moment. After about 18 iterations, the standard deviation of the visual acuity threshold posterior probability distribution maintained inside the QUEST algorithm is less than the preset threshold, and the algorithm converges.
[0139] 5. Result output: The detection is completed, and the final result is displayed on the screen: "Left eye visual acuity: LogMAR 0.0 (equivalent to 1.0)".
[0140] In order to more intuitively show the technical effect of the present application, please refer to Figure 7 . This figure is obtained by testing a simulated eye with standard visual acuity of 1.0 using the device of the present application and a device using traditional static calibration technology in a controlled experiment, respectively.
[0141] As Figure 7 shown, the present application has significant technical advantages and robustness compared with the prior art. Figure 7 (a) shows the comparison of the influence of head Z-axis position deviation on detection accuracy. The X-axis represents the Z-axis deviation distance (unit: mm), ranging from -10 mm to +10 mm, and the Y-axis represents the detection accuracy (normalized value 0 to 1.0). The prior art uses a static calibration scheme, and when the head position deviates, the detection accuracy drops sharply. The dashed curve shows that when the deviation distance reaches ±10 mm, the accuracy drops to 25-30%, which is difficult to meet the actual application requirements. While the present application uses a dynamic self-calibration mechanism based on a generative adversarial network, the solid curve shows that within a deviation range of ±10 mm, the detection accuracy remains above 98%, almost horizontal, showing strong robustness to distance disturbance.
[0142] Figure 7 (b) shows the comparison of the influence of head attitude angle change on detection accuracy, verifying the all-around robustness of the present application from multiple dimensions. The X-axis represents the attitude angle deviation (unit: degree), ranging from -15 degrees to +15 degrees, and the Y-axis represents the detection accuracy (normalized value 0 to 1.0). The prior art is highly sensitive to head attitude changes, and the three dashed lines represent the influence of pitch angle, yaw angle, and roll angle, respectively. As the attitude angle deviation increases, the accuracy decreases, and at a deviation of ±15 degrees, it drops to 40-50%. The influence of different attitude angles varies slightly, but overall, it shows high sensitivity to attitude disturbance. In contrast, the present application uses an adaptive attitude compensation mechanism to calculate the three-dimensional attitude of the head in real time and use it as one of the conditions for generating the target. The thick solid line represents the performance of all attitude angles, and the influence curves of the three attitude angles almost completely coincide, maintaining a detection accuracy of more than 95% within a ±15-degree attitude deviation range, fully verifying the all-around robustness and practical value of the present application to attitude disturbance.
[0143] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer programs. When the computer programs are loaded on a computer and executed, all or part of the processes or functions described in the embodiments of the present disclosure are generated. The computer can be a general purpose computer, a special purpose computer, a computer network, or other programmable devices. The computer programs can be stored in a computer readable storage medium or transferred from one computer readable storage medium to another computer readable storage medium, for example, the computer programs can be transferred from one website, computer, server or data center to another website, computer, server or data center through wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) mode. The computer readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. integrated with one or more available media. The available media can be magnetic media (such as floppy disk, hard disk, magnetic tape), optical media (such as high-density digital video disc (digital video disc, DVD)), or semiconductor media (such as solid state disk (solid state disk, SSD)) and the like.
[0144] Those skilled in the art can understand that the first, second, and the like various numerical numbers involved in the present disclosure are only for the convenience of description, and do not limit the scope of the embodiments of the present disclosure, nor represent the order of precedence.
[0145] At least one of the present disclosure can also be described as one or more, and the plurality can be two, three, four or more, which is not limited by the present disclosure. In the embodiments of the present disclosure, for a technical feature, the technical features in the technical feature are distinguished by "first", "second", "third", "A", "B", "C" and "D". There is no order or size order between the technical features described by "first", "second", "third", "A", "B", "C" and "D".
[0146] The correspondence relationship shown in each table in the present disclosure can be configured or predefined. The values of the information in each table are merely examples, and other values can be configured, and the present disclosure is not limited. When configuring the correspondence relationship of the information and each parameter, it is not necessarily required to configure all the correspondence relationships shown in each table. For example, the correspondence relationship shown in some rows in the table in the present disclosure can also not be configured. For another example, the above tables can be appropriately deformed, for example, split, merged, and the like. The names of the parameters shown in the titles of the above tables can also use other names understandable by the communication device, and the values or representations of the parameters can also use other values or representations understandable by the communication device. The above tables can also use other data structures when implemented, for example, arrays, queues, containers, stacks, linear tables, pointers, linked lists, trees, graphs, structures, classes, heaps, hash tables, or the like.
[0147] The predefinition in the present disclosure can be understood as definition, predefinition, storage, prestorage, prenegotiation, preconfiguration, solidification, or pre-burning. Those skilled in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present disclosure.
[0148] Those skilled in the art can clearly understand that, for the convenience and brevity of the description, the specific working processes of the above-described system, device and unit can refer to the corresponding processes in the foregoing method embodiments, which will not be described here.
[0149] The above is merely a specific implementation of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed by the present disclosure, which should be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.
Claims
1. A vision adaptive detection method based on generative adversarial networks, characterized in that, Includes the following steps: The detection environment of the test subject is perceived in real time to obtain a real-time environmental feature vector describing the detection environment; Based on the preset adaptive vision detection algorithm, determine the target vision level to be tested; The real-time environmental feature vector and the target visual acuity level are input into a pre-trained conditional generative adversarial network model to drive the model to generate a pre-distorted optotype image that can form an optically equivalent image to the standard optotype on the retina of the subject under the detection environment. Display the pre-distorted target image on a display device; The algorithm acquires the subject's response to the pre-distorted optotype image and iteratively executes the steps of determining the target visual acuity level, generating and displaying the pre-distorted optotype image based on the response, until the adaptive visual acuity detection algorithm converges to obtain the final visual acuity detection result.
2. The method according to claim 1, characterized in that, The step of real-time sensing the detection environment of the test subject to obtain a real-time environmental feature vector describing the detection environment includes: The image acquisition device is used to acquire a sequence of images containing the subject's face and eyes in real time. The image sequence is processed to calculate at least one of the following: the three-dimensional spatial position of the subject's eyes relative to the display device, the head posture angle, and the physical diameter of the pupil; The calculated at least one parameter is encoded into the real-time environment feature vector.
3. The method according to claim 1, characterized in that, The conditional generative adversarial network model includes a generator and a discriminator. The conditional generative adversarial network model is trained through an offline training process that includes a retinal imaging simulator, which is used to simulate the projection of a target image onto the retina after passing through the human eye's optical system. The goal of the offline training process is to ensure that the output of the pre-distorted target image generated by the generator, after being processed by the retinal imaging simulator, is consistent with the output of the standard target image after being processed by the retinal imaging simulator, under the dual constraints of adversarial loss and imaging consistency loss.
4. The method according to claim 3, characterized in that, The retinal imaging simulator characterizes the overall impact of the human eye's optical system by constructing a point spread function that can dynamically change with conditions. The point spread function is at least a function of the three-dimensional spatial position of the subject's eye relative to the display device and the physical diameter of the pupil.
5. The method according to claim 1, characterized in that, The adaptive vision detection algorithm is either the QUEST algorithm or the Best PEST algorithm.
6. A vision adaptive detection system based on generative adversarial networks, used to implement the method as described in any one of claims 1-5, characterized in that, include: The image acquisition unit is configured to perceive the detection environment of the subject in real time. The display unit is configured to display target images; The processing unit is connected to the image acquisition unit and the display unit; The processing unit is configured as follows: An environment feature encoding module is used to process the information acquired by the image acquisition unit to obtain a real-time environment feature vector describing the detection environment. Run the vision detection logic control module to determine the target vision level to be tested based on the preset adaptive vision detection algorithm; Run the optotype synthesis generative adversarial network module, which integrates a pre-trained conditional generative adversarial network model. The optotype synthesis generative adversarial network module is used to receive the real-time environmental feature vector and the target visual acuity level, and generate a pre-distorted optotype image that can form an optical image equivalent to the standard optotype on the retina of the subject under the detection environment. The display unit is controlled to display the pre-distorted target image; Based on the subject's response to the pre-distorted optotype image, the vision detection logic control module and the optotype synthesis generative adversarial network module are iteratively run until the adaptive vision detection algorithm converges to obtain the final vision detection result.
7. The system according to claim 6, characterized in that, The environmental feature encoding module is configured as follows: The image acquisition unit acquires a real-time image sequence containing the subject's face and eyes; The image sequence is processed by computer vision algorithms to calculate at least one of the following: the three-dimensional spatial position of the subject's eyes relative to the display device, the head posture angle, and the physical diameter of the pupil. The calculated parameters are then encoded into the real-time environmental feature vector.
8. The system according to claim 6, characterized in that, The conditional generative adversarial network model includes a generator and a discriminator, and is trained through an offline training process that includes a retinal imaging simulator. The retinal imaging simulator is constructed as a differentiable mathematical model to simulate the projection of a target image onto the retina after passing through a model of the human eye's optical system. The offline training process jointly optimizes the generator through adversarial loss and imaging consistency loss, enabling it to generate the pre-distorted target image.
9. The system according to claim 8, characterized in that, The core of the human eye optical system model is the point spread function model, which is configured to dynamically generate the corresponding point spread function based on the parameters contained in the real-time environment feature vector.
10. The system according to claim 6, characterized in that, The system is integrated into a single professional vision testing device, and the optical axis of the image acquisition unit is coaxially calibrated with the geometric center of the display unit.