Multimodal adaptive ophthalmic vr visual field detection method, system, device, and medium
Patent Information
- Application Number
- CN202611005008.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-07
- Publication Date
- 2026-08-28
AI Technical Summary
[0003]在相关技术中,基于虚拟现实的视野检测数据处理通常采用固定的检测场景与固定轨迹的刺激点呈现方式,刺激点多以独立闪烁点的形式出现;固视监测往往依赖单一参数,例如仅依据注视角度判断被测者是否保持固视;交互方式较为单一,多依赖手柄按键或单一语音指令;检测过程中产生的数据多依赖向远端服务器传输后再行处理与存储
通过对注视偏差数据与瞳孔波动数据的双模态融合判定,并在固视无效时动态平移固视目标以补偿头部微动,使检测流程在头部微动时不中断,提升所采集检测数据的有效性、减少重复检测;
Smart Images

Figure CN122642819A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of computer technology, specifically a multimodal adaptive ophthalmic VR field of vision detection method, system, device and medium. Background Technology
[0002] Visual field testing is an important means of assessing the extent and degree of visual field defects, and is often used for screening and monitoring the course of diseases such as glaucoma and retinal diseases. With the development of virtual reality display technology, data processing methods for visual field testing based on head-mounted display devices have gradually been applied. The basic process is as follows: stimuli are presented at different points within the display visual field, the subject's response to the stimuli is collected, and the visual field sensitivity at each point is assessed accordingly.
[0003] In related technologies, virtual reality-based visual field detection data processing typically employs a fixed detection scene and a fixed trajectory stimulus point presentation method, with the stimulus points often appearing as independent flashing points; fixation monitoring often relies on a single parameter, such as judging whether the subject is maintaining fixation based solely on the gaze angle; the interaction method is relatively simple, often relying on controller buttons or single voice commands; and the data generated during the detection process often depends on being transmitted to a remote server for processing and storage.
[0004] The aforementioned technologies have the following technical problems in practical applications: First, fixation monitoring relies on a single parameter, which can easily lead to misjudgment as fixation ineffectiveness when the subject makes slight head movements, thus interrupting the detection process, resulting in invalid collected data, the need for repeated testing, and low data validity. Second, the fixed scene and stimulus trajectory cannot be adjusted based on the subject's visual acuity level and previous test results, making it difficult to fully sample in areas with suspected defects, resulting in insufficient accuracy in identifying defect areas. Third, the single interaction method can easily lead to missed responses when the subject has difficulty speaking or operating. Fourth, reliance on remote transmission and processing results in high detection response delays, and there is a risk of insufficient data security during transmission.
[0005] Therefore, it is necessary to propose an ophthalmic VR visual field detection processing scheme to achieve low latency, uninterrupted operation, high defect recognition accuracy, and data security in visual field detection data processing locally. Summary of the Invention
[0006] To address the above problems, this application provides a multimodal adaptive ophthalmic VR field of view detection method, system, device, and medium. To achieve the above objectives, the technical solution adopted in this application is as follows: According to a first aspect of this application, a multimodal adaptive ophthalmic VR field of view detection method is provided, comprising the following steps: S1. Obtain the pre-detection data of the test subject, the pre-detection data including the vision level obtained by screening with a virtual vision chart, match and adaptively generate a detection scene in a preset scene-vision correspondence based on the vision level, and configure the stimulus points to be presented in the detection scene as the native elements of the detection scene; S2. In the detection scenario, a dynamically scaled fixation target is presented. The fixation deviation data and pupil fluctuation data of the subject are collected simultaneously. The fixation deviation data and pupil fluctuation data are fused in a dual-modal manner to determine the fixation validity. When the fixation is determined to be invalid, the fixation target is dynamically translated according to the fixation deviation data to compensate for head micro-movements until the fixation is determined to be valid and the fixation benchmark is output. S3. Based on the effective fixation, a multi-parameter dynamic stimulus point sequence is generated within the detection field of view according to the fixation benchmark, and the presentation trajectory of subsequent stimulus points is dynamically planned based on the sensitivity results of the preceding detection sites; S4. For each stimulus point presented in the stimulus point sequence, collect speech response signals, gaze timeout signals, and blink response signals. Perform trimodal fusion judgment on the speech response signals, gaze timeout signals, and blink response signals to determine the subject's perceptual response to the stimulus point. Based on the perceptual response, determine the sensitivity threshold of each detection site by step-wise threshold search to obtain a sensitivity threshold matrix. S5. Based on the sensitivity threshold matrix, a three-dimensional visual field sensitivity distribution is generated by interpolation. Defect areas in the three-dimensional visual field sensitivity distribution with sensitivity thresholds lower than the defect determination threshold are identified. The defect area, average sensitivity, and disease progression trend of the defect area are calculated to generate detection results. The detection results are then encrypted and stored.
[0007] Preferably, step S1 includes: The system receives the identity information entered by the test subject, triggers automatic pupillary distance measurement, and synchronously performs refractive adjustment based on the measured pupillary distance value. The visual acuity level is then obtained through screening using the virtual visual acuity chart. Based on the visual acuity level, the scene category of the detection scene is determined by querying the preset scene-visual acuity correspondence, and the scene element corresponding to the scene category is loaded; The overall brightness of the detection scene is dynamically reduced according to a preset decreasing pattern as the detection time progresses, and the stimulation point is presented in the form of the original elements of the detection scene.
[0008] Preferably, in step S2, the step of performing bimodal fusion determination of the fixation deviation data and the pupillary fluctuation data to determine fixation effectiveness includes: When the fixation deviation data is less than or equal to a preset first threshold and the pupil fluctuation data is less than or equal to a preset second threshold, fixation is determined to be effective; otherwise, fixation is determined to be ineffective. The method of dynamically translating the fixation target based on the gaze deviation data to compensate for head micro-movements includes: calculating the translation amount of the fixation target according to a preset translation gain based on the deviation direction and deviation amplitude represented by the gaze deviation data, translating the fixation target according to the translation amount, and keeping the detection process uninterrupted during the translation.
[0009] Preferably, in step S3, generating a multi-parameter dynamic stimulus point sequence within the detection visual field based on the fixation reference includes: The brightness of the stimulation point is adjusted within a preset brightness range according to a preset brightness step, the movement speed of the stimulation point is adaptively adjusted within a preset speed range, and the shape of the stimulation point is randomly switched from a preset shape set to generate a multi-parameter dynamic stimulation point sequence. The detection field of view includes both the horizontal and vertical ranges.
[0010] Preferably, in step S3, the dynamic planning of the presentation trajectory of subsequent stimulus points includes: For normal visual areas where the sensitivity of the preceding detection site is higher than a preset threshold, the stimulation points are sequentially presented using a polygonal trajectory. For suspected defect areas where the sensitivity result of the preceding detection site is not higher than the preset region threshold, a dwell-type stimulation point is used, wherein the dwell-type stimulation point stays at the corresponding detection site for a preset dwell time.
[0011] Preferably, in step S4, the trimodal fusion determination to determine the subject's perceptual response to the stimulus includes: Within a preset response window after the stimulus point is presented, if the voice response signal, or the gaze duration represented by the gaze timeout signal exceeds a preset gaze threshold, or any one of two consecutive blinks represented by the blink response signal is collected, it is determined that the subject has a perceptual response to the stimulus point; The stepwise threshold search for determining the sensitivity threshold of each detection site includes: reducing the stimulus intensity when a sensory response is generated, increasing the stimulus intensity when no sensory response is generated, recording multiple reversal points of the stimulus intensity, and taking the average brightness of the multiple reversal points as the sensitivity threshold of the detection site.
[0012] Preferably, step S5 includes: Using the sensitivity threshold of each detection site in the sensitivity threshold matrix as the sample value, weighted interpolation is performed on the non-sampled points within the detection field of view to generate the three-dimensional field of view sensitivity distribution; The detection sites in the three-dimensional visual field sensitivity distribution with sensitivity thresholds lower than the defect determination threshold are clustered into the defect regions, and the defect area, the average sensitivity, and the disease progression trend between the detection results and historical detection results are calculated for each defect region. The detection results are double-encrypted using a first encryption algorithm and a second encryption algorithm before being stored locally, and can be exported via an offline interface.
[0013] According to a second aspect of this application, a multimodal adaptive ophthalmic VR field of view detection system employing the above-described multimodal adaptive ophthalmic VR field of view detection method is provided, comprising: The scene generation module is configured to acquire the test subject's pre-detection data, which includes the visual acuity level obtained through screening with a virtual visual acuity chart. Based on the visual acuity level, the module matches and adaptively generates a test scene in a preset scene-visual acuity correspondence, and configures the stimulus points to be presented in the test scene as native elements of the test scene. The fixation calibration module is configured to present a dynamically scaled fixation target in the detection scenario, simultaneously collect the subject's fixation deviation data and pupil fluctuation data, perform dual-modal fusion judgment on the fixation deviation data and pupil fluctuation data to determine fixation validity, and when fixation is determined to be invalid, dynamically translate the fixation target according to the fixation deviation data to compensate for head micro-movements until fixation is determined to be valid and a fixation benchmark is output. The stimulus planning module is configured to generate a multi-parameter dynamic stimulus point sequence within the detection field of view based on the fixation benchmark, provided that fixation is effective, and to dynamically plan the presentation trajectory of subsequent stimulus points based on the sensitivity results of previous detection sites; The response determination module is configured to collect speech response signals, gaze timeout signals, and blink response signals for each presented stimulus point, perform trimodal fusion determination to determine the perceptual response, and determine the sensitivity threshold of each detection point based on the perceptual response using a step-wise threshold search to obtain a sensitivity threshold matrix; The results analysis module is configured to generate a three-dimensional visual field sensitivity distribution based on the sensitivity threshold matrix through interpolation, identify defect areas where the sensitivity threshold is lower than the defect determination threshold, calculate the defect area, average sensitivity and disease progression trend to generate detection results, and encrypt and store the detection results.
[0014] According to a third aspect of this application, an electronic device is provided, characterized in that the device comprises: One or more processors; and A storage device for storing one or more programs, which, when executed by one or more processors, enable the one or more processors to implement the above-described multimodal adaptive ophthalmic VR field of view detection method.
[0015] According to a fourth aspect of this application, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the above-described multimodal adaptive ophthalmic VR field of view detection method.
[0016] The beneficial effects of this application include: By fusing fixation deviation data and pupil fluctuation data in a dual-modal manner, and dynamically shifting the fixation target to compensate for head micro-movements when fixation is ineffective, the detection process is not interrupted when head micro-movements occur, thereby improving the effectiveness of the collected detection data and reducing repeated detections. By adaptively generating detection scenarios based on visual acuity levels and dynamically planning presentation trajectories based on the sensitivity results of previous detection sites, and by using dwell-type stimulation points for suspected defect areas, the sufficiency of sampling and the accuracy of identification of defect areas are improved; By fusing three modalities—voice, gaze timeout, and blink—to determine the perception response, the false negative rate is reduced. By performing data processing and dual-encrypted storage on the local processing unit using edge computing, detection response latency is reduced and detection data is processed locally throughout the entire process, thereby improving data security. Attached Figure Description
[0017] Figure 1 This is a flowchart of the multimodal adaptive ophthalmic VR field of view detection method according to an embodiment of this application; Figure 2 This is a structural block diagram of the multimodal adaptive ophthalmic VR field of view detection system according to an embodiment of this application; Figure 3 This is a structural block diagram of a computer device according to an embodiment of this application; Figure 4 This is a structural block diagram of a computer storage medium according to an embodiment of this application. Detailed Implementation
[0018] To enable those skilled in the art to better understand the technical solution, the present application will be described in detail below with reference to the embodiments. The description in this section is only exemplary and explanatory, and should not be used to limit the scope of protection of the present application in any way.
[0019] The example environment for implementing the multimodal adaptive ophthalmic VR field of view detection method in this application includes a user device, which comprises a data acquisition unit, a local processing unit, and an encrypted storage unit. The data acquisition unit collects signals such as eye movement, pupil size, speech, and blinking from the subject and inputs them into the local processing unit. The local processing unit runs the data processing method of this application, processes the collected signals, and generates detection results. The encrypted storage unit is used to encrypt and store the detection results. The method of this application is mainly executed by the local processing unit, and the data is processed entirely locally on the user device without transmission to a remote server, thereby reducing detection response latency and improving data security.
[0020] Example 1
[0021] Figure 1 A flowchart of a multimodal adaptive ophthalmic VR visual field detection method according to an embodiment of this application is shown. This embodiment uses glaucoma visual field defect screening as an application scenario to describe in detail the steps of the method of this application. The method is applied to a local processing unit and includes steps S1 to S5.
[0022] Step S1: Obtain the pre-detection data of the subject, adaptively generate the detection scene based on the visual acuity level, and configure the stimulus points as native elements of the detection scene.
[0023] Specifically, the local processing unit receives the identity information of the test subject via voice input to trigger pre-detection; then it triggers automatic pupillary distance measurement to obtain the pupillary distance value, and simultaneously performs refractive adjustment based on this pupillary distance value to ensure the clarity of the subsequently presented image; afterwards, a visual acuity level is obtained through screening with a virtual visual acuity chart. In this embodiment, the visual acuity level determination method of the virtual visual acuity chart screening is as follows: several levels of optotypes from large to small are presented sequentially, and the level corresponding to the smallest optotype that the test subject can correctly distinguish is recorded as the visual acuity level. The visual acuity level ranges from level 1 to level 5, with a smaller level value indicating lower visual acuity. The correspondence between the visual acuity level and the inter-level resolution can be determined by those skilled in the art based on the optotype size sequence used, which is a conventional design.
[0024] After obtaining the visual acuity level, the local processing unit queries a preset scene-visual acuity correspondence to determine the scene category of the detection scene. The preset scene-visual acuity correspondence is a mapping table from visual acuity level to scene category. Its tuning method is as follows: In clinical pre-experiments, candidate scenes with different contrasts and element densities are presented to subjects with different visual acuity levels. The fixation stability and stimulus recognition accuracy of the subjects under each candidate scene are statistically analyzed. The scene category with the optimal combination of fixation stability and stimulus recognition accuracy is tuned as the scene category corresponding to that visual acuity level. The selection criteria are: for lower visual acuity levels, scenes with sparse elements, high contrast, and soft brightness (e.g., forest or flower scenes) are configured to reduce visual load; for higher visual acuity levels, scenes with relatively abundant elements and moderate contrast (e.g., starry sky scenes) are configured to balance comfort and stimulus recognition. An exemplary scene-visual acuity correspondence is shown in Table 1 below.
[0025] Level 1 Forest scene higher sparse Level 2 Flower Scene higher sparse Level 3 Forest scene Moderate Moderate Level 4 Starry sky scene Moderate Moderate Level 5 Starry sky scene Moderate denser After determining the scene category, the local processing unit loads the scene elements corresponding to that scene category, adaptively generates the detection scene, and configures the stimulus points to be presented in the detection scene as the native elements of that detection scene. For example, in a forest scene, the stimulus points are configured as flying birds; in a starry sky scene, the stimulus points are configured as starlight; and in a flower scene, the stimulus points are configured as petals, rather than independent flashing points, thereby reducing the visual fatigue and tension of the test subject.
[0026] Furthermore, the local processing unit dynamically reduces the overall brightness of the detection scene according to a preset decreasing rule as the detection duration increases, thereby alleviating visual fatigue caused by prolonged detection. In this embodiment, the preset decreasing rule is a linear decreasing rule, and the overall brightness of the detection scene at detection duration t is calculated as follows: in, For detection duration Overall brightness at any given moment Initial brightness, The preset detection time limit, The decreasing coefficient ranges from 0.1 to 0.4. The reason for this value is that if the decreasing coefficient is too small, the decrease in brightness will not be obvious and the effect of relieving fatigue will be limited. If the decreasing coefficient is too large, the brightness will be too low in the later stage and the stimulus recognition will be affected. Therefore, the above range is taken. In this embodiment, 0.25 is taken. When the brightness calculated by the above formula is lower than the preset brightness lower limit, the preset brightness lower limit is taken to ensure that the stimulus can be recognized.
[0027] Step S2: Present a dynamically scaled fixation target in the detection scene, perform dual-modal fusion judgment on fixation deviation data and pupil fluctuation data to determine fixation effectiveness, and dynamically translate the fixation target to compensate for head micro-movements and output fixation benchmark when fixation is invalid.
[0028] Specifically, the local processing unit presents a dynamically scaling ring as a fixation target in the central area of the detection scene. Compared to a static fixation point, the dynamically scaling ring is more likely to attract and maintain the subject's gaze. Simultaneously with the fixation target, the local processing unit collects fixation deviation data and pupil fluctuation data via the data acquisition unit. The fixation deviation data is collected by an infrared eye-tracking camera and characterizes the deviation angle between the subject's actual gaze direction and the fixation target direction. The pupil fluctuation data is collected by a pupil sensor and characterizes the relative fluctuation amplitude of the pupil diameter per unit time.
[0029] The local processing unit performs dual-modal fusion judgment on fixation deviation data and pupil fluctuation data: when the fixation deviation data is less than or equal to a preset first threshold and the pupil fluctuation data is less than or equal to a preset second threshold, fixation is determined to be effective; otherwise, fixation is determined to be ineffective. In this embodiment, the first threshold is 0.3° and the second threshold is 5%. The first threshold is determined based on the following: when the fixation deviation is within 0.3°, the presentation error of the visual field point is within the acceptable range of visual field detection. If it is too large, the point positioning error will increase; if it is too small, it will be difficult to reach the target due to normal eye micromovement, resulting in excessive calibration time. Therefore, 0.3° is selected. The second threshold is determined based on the following: when the relative fluctuation of the pupil diameter is within 5%, it can be considered that the subject is in a stable fixation and stable illumination adaptation state. If it is too large, it may be caused by attention shift or changes in ambient light. Therefore, 5% is selected. By using simultaneous judgment of fixation deviation and pupil fluctuation as dual parameters, compared with judgment based on only the single parameter of fixation angle, it can avoid misjudging the subject as having effective fixation when the subject's fixation direction is correct but their attention has shifted, thereby improving the reliability of fixation judgment.
[0030] When fixation is determined to be invalid, the local processing unit further determines whether the fixation invalidity is caused by head micro-movement: if the deviation represented by the fixation deviation data shows a slow drift in the same direction over several consecutive frames, it is determined to be head micro-movement. In this case, the local processing unit does not interrupt the detection process, but instead dynamically translates the fixation target according to the fixation deviation data to compensate for the head micro-movement. The translation amount of the fixation target is calculated as follows: in, The rendering position of the fixed target in the current frame. To determine the position of the fixed target after translation. The deviation vector represents the gaze deviation data, with its direction being the deviation direction and its magnitude being the deviation amplitude. The preset translation gain has a value range greater than 0 and less than or equal to 1. The reason for this value is that if the translation gain is too small, the compensation will be insufficient and the fixation may still be interrupted; if the translation gain is too large, the fixation target will move too quickly, affecting the subject's gaze. Therefore, the above range is used, and in this embodiment, 0.6 is chosen. The detection process remains uninterrupted during the translation of the fixation target, thus avoiding the loss of collected data due to slight head movements. When the fixation validity determination condition is met again, the local processing unit outputs the current gaze direction as the fixation reference for use in subsequent step S3.
[0031] Step S3: Based on the effective fixation, generate a multi-parameter dynamic stimulus point sequence within the detection field of view according to the fixation benchmark, and dynamically plan the presentation trajectory of subsequent stimulus points based on the sensitivity results of the preceding detection sites.
[0032] Specifically, the local processing unit uses the fixation reference output in step S2 as the center of the visual field and determines several detection sites within the detection visual field. In this embodiment, the detection visual field includes a horizontal range and a vertical range. The horizontal range is ±60° and the vertical range is ±50°. The values are determined based on the fact that this range covers the central and peripheral visual field areas commonly used in clinical visual field testing. The local processing unit generates multi-parameter dynamic stimulation points at each detection site. The multi-parameters include brightness, movement speed, and shape. The brightness of the stimulation point is adjusted within a preset brightness range according to a preset brightness step. In this embodiment, the preset brightness range is 2dB to 30dB, and the preset brightness step is 2dB. The values are chosen based on the fact that this brightness range covers the commonly used stimulation intensity range for visual field detection, and the 2dB step matches the accuracy of the step-type threshold search. The movement speed of the stimulation point is adaptively adjusted within a preset speed range. In this embodiment, the preset speed range is 0.5° / s to 2° / s. The values are chosen based on the fact that this speed range is easy for the subject to follow without being too fast and causing omissions. The shape of the stimulation point is randomly switched from a preset shape set. In this embodiment, the preset shape set includes round dots, square dots, and triangular dots. Randomly switching shapes can reduce the subject's expectation of the regularity of the stimulus occurrence and reduce predictive false triggers.
[0033] Furthermore, the local processing unit dynamically plans the presentation trajectory of subsequent stimulus points based on the sensitivity results of preceding detection sites. Specifically, for normal visual areas where the sensitivity results of preceding detection sites are higher than a preset region threshold, stimulus points are presented sequentially using a broken line trajectory to improve detection efficiency; for suspected defect areas where the sensitivity results of preceding detection sites are not higher than the preset region threshold, a dwell-type stimulus point is used. The dwell-type stimulus point stays at the corresponding detection site for a preset dwell time to fully sample the suspected defect area and improve the accuracy of defect area identification. In this embodiment, the preset region threshold is 15dB, and the preset dwell time is 0.3s. The basis for these values is that a sensitivity below 15dB usually indicates a suspected defect in the area, requiring more thorough sampling, and a dwell time of 0.3s is sufficient to complete a complete stimulus presentation and response acquisition without excessively prolonging the detection time.
[0034] Step S4: Collect speech response signal, gaze timeout signal and blink response signal for each stimulus point, perform trimodal fusion judgment to determine the perceptual response, and determine the sensitivity threshold of each detection point by step threshold search to obtain the sensitivity threshold matrix.
[0035] Specifically, after each stimulus is presented, the local processing unit collects the speech response signal, gaze timeout signal, and blink response signal within a preset response window via the data acquisition unit, and performs trimodal fusion determination: within the preset response window, when the gaze duration represented by the collected speech response signal or gaze timeout signal exceeds a preset gaze threshold, or when either of the two consecutive blinks represented by the blink response signal is exceeded, it is determined that the subject has a perceptual response to the stimulus. In this embodiment, the preset response window is 1.5s, and the preset gaze threshold is 500ms; the preset response window is set based on the fact that the window covers the normal perceptual response, and the preset gaze threshold is set based on the fact that 500ms is sufficient to distinguish between intentional gaze triggering and unintentional brief fixation. The speech response signal is obtained by the local offline speech recognition model recognizing the audio collected by the microphone. The local offline speech recognition model can complete the recognition under conditions without network dependence. The three-modal fusion judgment uses a logical OR method. If any modality is satisfied, it is determined that a perceptual response has been generated. This allows the response to be triggered by other modalities even when the subject has difficulty using a certain modality (e.g., difficulty speaking), thus reducing the false negative rate.
[0036] After determining the perceived response at each stimulus point, the local processing unit determines the sensitivity threshold for each detection site using a stepped threshold search. Specifically, for each detection site, the local processing unit decreases the stimulus intensity in preset brightness steps when a perceived response is generated, and increases the stimulus intensity in preset brightness steps when no perceived response is generated, repeating this process. When the stimulus intensity changes from perceived to unperceived or from unperceived to perceived, a reversal point is recorded. The search continues until a preset number of reversal points are recorded, and the average brightness of the preset number of reversal points is taken as the sensitivity threshold for that detection site. The sensitivity threshold for a detection site is calculated as follows: in, For the first Sensitivity threshold for each detection site, For the first The brightness of the stimulation point at each flip point, The preset number of flip points ranges from 2 to 6. The reason for this is that too few flip points result in insufficient stability of the threshold estimation, while too many flip points result in excessively long detection times. In this embodiment, 4 is used. The sensitivity thresholds of each detection site are organized into a sensitivity threshold matrix according to their position within the detection field of view, for use in step S5.
[0037] Step S5: Based on the sensitivity threshold matrix, a three-dimensional visual field sensitivity distribution is generated through interpolation. The defect area is identified, and the defect area, average sensitivity, and disease progression trend are calculated to generate the detection results. The detection results are then encrypted and stored.
[0038] Specifically, the local processing unit uses the sensitivity threshold of each detection site in the sensitivity threshold matrix as the sampling value, and performs weighted interpolation on the non-sampling points within the detection field of view to generate a three-dimensional field of view sensitivity distribution. This three-dimensional field of view sensitivity distribution can be presented as a three-dimensional field of view defect heatmap with the horizontal and vertical angles of the field of view as planar coordinates and sensitivity as the height. In this embodiment, the weighted interpolation uses inverse distance weighted interpolation, and the sensitivity at non-sampling points is calculated as follows: in, Non-sampling points Sensitivity at the point, For the first Sensitivity threshold for each sampling point (detection site), The number of sampling points participating in the interpolation. For the first The weights of each sampling point are calculated as follows: in, Non-sampling points With the The distance between each sampling point The power exponent is the distance, and its value ranges from 1 to 3. The reason for choosing the value is that the larger the power exponent, the more the interpolation result is dominated by the nearest sampling points and the stronger the locality. The smaller the power exponent, the smoother the interpolation result. In this embodiment, it is 2.
[0039] After generating the three-dimensional visual field sensitivity distribution, the local processing unit clusters detection sites with sensitivity thresholds lower than the defect determination threshold into defect regions. In this embodiment, the defect determination threshold is set to 15 dB, based on the clinically accepted assumption that a sensitivity below 15 dB generally indicates a significant visual field defect. The local processing unit further calculates the defect area, average sensitivity, and disease progression trend of the defect regions. The defect area is characterized by the proportion of detection sites within the defect region to the total number of detection sites within the detection field of view, calculated as follows: in, The percentage of the defective area. The number of detection sites within the defect area. This measures the total number of detection sites within the detection field of view. The average sensitivity is the arithmetic mean of the sensitivity thresholds for all detection sites, calculated as follows: in, For average sensitivity, For the first The sensitivity threshold for each detection site. The disease progression trend refers to the change trend of the current test results and historical test results over time, characterized by the slope of a linear regression of the average sensitivity against the detection time, calculated as follows: in, The slope of the disease course trend. For the first The time of the second test For the first Average sensitivity of the tests, and These represent the mean values of detection time and average sensitivity, respectively. The number of tests involved in the regression. A negative slope indicates that the average sensitivity decreases over time, suggesting disease progression; a positive slope or close to zero indicates a stable disease course.
[0040] The local processing unit organizes the defect area, average sensitivity, and disease progression trend into detection results that include the aforementioned quantitative indicators, and then encrypts and stores the detection results. In this embodiment, the detection results are stored in a local encrypted storage unit after being doubly encrypted using a first encryption algorithm and a second encryption algorithm. The first encryption algorithm is the AES-256 algorithm, and the second encryption algorithm is the Chinese national standard SM4 algorithm. Using two encryption algorithms for double encryption ensures data confidentiality even if either algorithm is compromised, thus improving data security. The encrypted detection results can be exported via an offline interface (e.g., USB-C interface) and compared with historical detection results.
[0041] In another embodiment of this application, visual field monitoring of patients with retinal diseases is used as the application scenario. The visual acuity level obtained by the test subject through virtual visual acuity chart screening is level 5. Therefore, in step S1, a starry sky scene is matched according to the scene-visual acuity correspondence shown in Table 1, and the stimulus point is configured as starlight. In the brightness reduction in step S1, the reduction coefficient is taken near the upper end of the value range described in this application, that is, 0.4, to adapt to the characteristic that the test subject is more sensitive to long-term detection. In step S3, the brightness of the stimulus point is presented near the lower end of the preset brightness range to improve the detection capability of early mild defects. The remaining steps of this embodiment are the same as those of embodiment one, and will not be repeated here.
[0042] As can be seen from this embodiment, this application can adaptively generate detection scenarios and adaptively adjust parameters based on visual acuity levels, which can adapt to test subjects with different visual acuity levels and different application scenarios, without the need to use fixed scenarios and fixed parameters.
[0043] In another embodiment of this application, the brightness reduction coefficient is taken at the lower end of the value range described in this application, i.e., 0.1; the translation gain is taken at 0.2, which is close to the lower end of its value range; the interpolation distance exponent is taken at the upper end of the value range described in this application, i.e., 3; and the number of flip points is taken near the upper end of the value range described in this application, i.e., 6. Under the above parameter values, the method of this application can still normally execute steps S1 to S5 and generate detection results, indicating that it can be implemented near both ends of the parameter range described in this application. The remaining steps of this embodiment are the same as those of Embodiment 1.
[0044] Example 2
[0045] like Figure 2 As shown, this embodiment provides a multimodal adaptive ophthalmic VR field of view detection system, including a scene generation module, a fixation calibration module, a stimulus planning module, a response determination module, and a result analysis module. Each module corresponds one-to-one with steps S1 to S5 of the above method embodiment. Specifically: The scene generation module is configured to acquire the pre-detection data of the test subject, match and adaptively generate a detection scene in the preset scene-visual correspondence based on the visual acuity level, and configure the stimulus points to be presented in the detection scene as the native elements of the detection scene, i.e., to execute the above step S1.
[0046] The fixation calibration module is configured to present a dynamically scaled fixation target in the detection scene, perform dual-modal fusion judgment on fixation deviation data and pupil fluctuation data to determine fixation effectiveness, and dynamically translate the fixation target to compensate for head micro-movement and output fixation benchmark when fixation is invalid, i.e., execute the above step S2.
[0047] The stimulus planning module is configured to generate a multi-parameter dynamic stimulus point sequence based on the fixation benchmark and dynamically plan the presentation trajectory, i.e., to execute the above step S3.
[0048] The response determination module is configured to perform trimodal fusion determination on each stimulus point to determine the perceived response and to determine the sensitivity threshold of each detection site by step threshold search to obtain the sensitivity threshold matrix, i.e., to perform the above step S4.
[0049] The results analysis module is configured to generate a three-dimensional visual field sensitivity distribution based on the sensitivity threshold matrix through interpolation, identify the defect area and calculate the defect area, average sensitivity and disease trend to generate detection results and store them in encrypted form, i.e., to execute the above step S5.
[0050] The above modules can be implemented by software, hardware, or a combination of both. The division of modules is only a logical functional division, and other division methods may be used in actual implementation.
[0051] Example 3
[0052] Figure 3 A block diagram of an electronic device that can be used to implement embodiments of this application is shown. It includes a central processing unit (CPU), read-only memory (ROM), and random access memory (RAM), which are interconnected via a bus. Input / output interfaces are also connected to the bus. Multiple components in the computing device are connected to the input / output interfaces, including: input units such as a keyboard, microphone, eye-tracking camera, pupil sensor, etc.; output units such as a display, speaker, etc.; storage units such as a disk, optical disk, etc.; and communication units such as a network interface card (NIC), modem, etc.
[0053] The method steps described above can be executed by a central processing unit. For example, in some embodiments, the method can be implemented as a computer software program tangibly contained in a machine-readable medium, such as a storage unit. In some embodiments, part or all of the computer program can be loaded and / or installed on a computing device via read-only memory and / or a communication unit. When the computer program is loaded into random access memory and executed by the central processing unit, one or more steps of the method described above can be performed. The electronic device described in this application can be implemented by a computing device.
[0054] This application may be a method, apparatus, device, and / or computer-readable storage medium. The computer-readable storage medium carries computer program instructions for causing a processor to implement various aspects of this application. The computer-readable storage medium may be a tangible device capable of holding and storing instructions used by an instruction execution device, such as, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, or any suitable combination thereof. The method steps of this application can be implemented by a processor executing computer program instructions stored in the computer-readable storage medium, thereby realizing the function of this application.
[0055] It should be noted that, in this document, the terms "comprising," "including," and any other variations are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Specific examples have been used in this document to illustrate the principles and implementation methods of the technical solutions of this application. The above examples are only for the purpose of helping to understand the methods and core ideas of this application. The above descriptions are merely preferred embodiments of this application. It should be pointed out that, due to the limitations of written expression and the objective existence of infinite specific structures, those skilled in the art can make several improvements, modifications, or changes without departing from the principles of this application, and can also combine the above technical features in an appropriate manner; these improvements, modifications, changes, or combinations, or the direct application of the concept and technical solutions of this application to other situations without modification, should all be considered within the scope of protection of this application.
Claims
1. A multimodal adaptive ophthalmic VR field of view detection method, characterized in that, Includes the following steps: S1. Obtain the pre-detection data of the test subject, the pre-detection data including the vision level obtained by screening with a virtual vision chart, match and adaptively generate a detection scene in a preset scene-vision correspondence based on the vision level, and configure the stimulus points to be presented in the detection scene as the native elements of the detection scene; S2. In the detection scenario, a dynamically scaled fixation target is presented. The fixation deviation data and pupil fluctuation data of the subject are collected simultaneously. The fixation deviation data and pupil fluctuation data are fused in a dual-modal manner to determine the fixation validity. When the fixation is determined to be invalid, the fixation target is dynamically translated according to the fixation deviation data to compensate for head micro-movements until the fixation is determined to be valid and the fixation benchmark is output. S3. Based on the effective fixation, a multi-parameter dynamic stimulus point sequence is generated within the detection field of view according to the fixation benchmark, and the presentation trajectory of subsequent stimulus points is dynamically planned based on the sensitivity results of the preceding detection sites; S4. For each stimulus point presented in the stimulus point sequence, collect speech response signals, gaze timeout signals, and blink response signals. Perform trimodal fusion judgment on the speech response signals, gaze timeout signals, and blink response signals to determine the subject's perceptual response to the stimulus point. Based on the perceptual response, determine the sensitivity threshold of each detection site by step-wise threshold search to obtain a sensitivity threshold matrix. S5. Based on the sensitivity threshold matrix, a three-dimensional visual field sensitivity distribution is generated by interpolation. Defect areas in the three-dimensional visual field sensitivity distribution with sensitivity thresholds lower than the defect determination threshold are identified. The defect area, average sensitivity, and disease progression trend of the defect area are calculated to generate detection results. The detection results are then encrypted and stored.
2. The multimodal adaptive ophthalmic VR field of view detection method according to claim 1, characterized in that, Step S1 includes: The system receives the identity information entered by the test subject, triggers automatic pupillary distance measurement, and synchronously performs refractive adjustment based on the measured pupillary distance value. The visual acuity level is then obtained through screening using the virtual visual acuity chart. Based on the visual acuity level, the scene category of the detection scene is determined by querying the preset scene-visual acuity correspondence, and the scene element corresponding to the scene category is loaded; The overall brightness of the detection scene is dynamically reduced according to a preset decreasing pattern as the detection time progresses, and the stimulation point is presented in the form of the original elements of the detection scene.
3. The multimodal adaptive ophthalmic VR field of view detection method according to claim 1, characterized in that, In step S2, the process of performing bimodal fusion determination of the fixation deviation data and the pupillary fluctuation data to determine fixation effectiveness includes: When the fixation deviation data is less than or equal to a preset first threshold and the pupil fluctuation data is less than or equal to a preset second threshold, fixation is determined to be effective; otherwise, fixation is determined to be ineffective. The method of dynamically translating the fixation target based on the gaze deviation data to compensate for head micro-movements includes: calculating the translation amount of the fixation target according to a preset translation gain based on the deviation direction and deviation amplitude represented by the gaze deviation data, translating the fixation target according to the translation amount, and keeping the detection process uninterrupted during the translation.
4. The multimodal adaptive ophthalmic VR field of view detection method according to claim 1, characterized in that, In step S3, generating a multi-parameter dynamic stimulus point sequence within the detection field of view based on the fixation reference includes: The brightness of the stimulation point is adjusted within a preset brightness range according to a preset brightness step, the movement speed of the stimulation point is adaptively adjusted within a preset speed range, and the shape of the stimulation point is randomly switched from a preset shape set to generate a multi-parameter dynamic stimulation point sequence. The detection field of view includes both the horizontal and vertical ranges.
5. The multimodal adaptive ophthalmic VR field of view detection method according to claim 1, characterized in that, In step S3, the dynamic programming of the presentation trajectory of subsequent stimulus points includes: For normal visual areas where the sensitivity of the preceding detection site is higher than a preset threshold, the stimulation points are sequentially presented using a polygonal trajectory. For suspected defect areas where the sensitivity result of the preceding detection site is not higher than the preset region threshold, a dwell-type stimulation point is used, wherein the dwell-type stimulation point stays at the corresponding detection site for a preset dwell time.
6. The multimodal adaptive ophthalmic VR field of view detection method according to claim 1, characterized in that, In step S4, the trimodal fusion determination to determine the subject's perceptual response to the stimulus includes: Within a preset response window after the stimulus point is presented, if the voice response signal, or the gaze duration represented by the gaze timeout signal exceeds a preset gaze threshold, or any one of two consecutive blinks represented by the blink response signal is collected, it is determined that the subject has a perceptual response to the stimulus point; The stepwise threshold search for determining the sensitivity threshold of each detection site includes: reducing the stimulus intensity when a sensory response is generated, increasing the stimulus intensity when no sensory response is generated, recording multiple reversal points of the stimulus intensity, and taking the average brightness of the multiple reversal points as the sensitivity threshold of the detection site.
7. The multimodal adaptive ophthalmic VR field of view detection method according to claim 1, characterized in that, Step S5 includes: Using the sensitivity threshold of each detection site in the sensitivity threshold matrix as the sample value, weighted interpolation is performed on the non-sampled points within the detection field of view to generate the three-dimensional field of view sensitivity distribution; The detection sites in the three-dimensional visual field sensitivity distribution with sensitivity thresholds lower than the defect determination threshold are clustered into the defect regions, and the defect area, the average sensitivity, and the disease progression trend between the detection results and historical detection results are calculated for each defect region. The detection results are double-encrypted using a first encryption algorithm and a second encryption algorithm before being stored locally, and can be exported via an offline interface.
8. A multimodal adaptive ophthalmic VR visual field detection system employing the multimodal adaptive ophthalmic VR visual field detection method according to any one of claims 1 to 7, characterized in that, include: The scene generation module is configured to acquire the test subject's pre-detection data, which includes the visual acuity level obtained through screening with a virtual visual acuity chart. Based on the visual acuity level, the module matches and adaptively generates a test scene in a preset scene-visual acuity correspondence, and configures the stimulus points to be presented in the test scene as native elements of the test scene. The fixation calibration module is configured to present a dynamically scaled fixation target in the detection scenario, simultaneously collect the subject's fixation deviation data and pupil fluctuation data, perform dual-modal fusion judgment on the fixation deviation data and pupil fluctuation data to determine fixation validity, and when fixation is determined to be invalid, dynamically translate the fixation target according to the fixation deviation data to compensate for head micro-movements until fixation is determined to be valid and a fixation benchmark is output. The stimulus planning module is configured to generate a multi-parameter dynamic stimulus point sequence within the detection field of view based on the fixation benchmark, provided that fixation is effective, and to dynamically plan the presentation trajectory of subsequent stimulus points based on the sensitivity results of previous detection sites; The response determination module is configured to collect speech response signals, gaze timeout signals, and blink response signals for each presented stimulus point, perform trimodal fusion determination to determine the perceptual response, and determine the sensitivity threshold of each detection point based on the perceptual response using a step-wise threshold search to obtain a sensitivity threshold matrix; The results analysis module is configured to generate a three-dimensional visual field sensitivity distribution based on the sensitivity threshold matrix through interpolation, identify defect areas where the sensitivity threshold is lower than the defect determination threshold, calculate the defect area, average sensitivity and disease progression trend to generate detection results, and encrypt and store the detection results.
9. An electronic device, characterized in that... The device includes: One or more processors; and A storage device for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the multimodal adaptive ophthalmic VR field of view detection method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that... When the program is executed by the processor, it implements the multimodal adaptive ophthalmic VR field of vision detection method according to any one of claims 1 to 7.