A first-person view video construction method for people with visual field defects
By constructing first-person perspective videos and combining eye tracking and head posture perception, the shortcomings of existing visual field defect simulation methods in dynamic visual field changes have been addressed. This has enabled the reconstruction of realistic visual perception for patients with visual field defects, improving their rehabilitation training and independence in daily life.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- DALIAN MARITIME UNIVERSITY
- Filing Date
- 2026-03-27
- Publication Date
- 2026-07-14
AI Technical Summary
Existing methods for simulating visual field defects are mainly based on static images, which cannot truly reflect the dynamic changes in the patient's visual field. Furthermore, they lack first-person perspective alignment technology, resulting in a significant deviation from the patient's actual visual perception and making it difficult to meet the dynamic needs of rehabilitation training and daily activities.
Using an interdisciplinary research approach, we constructed a first-person perspective video for people with visual field defects. By building a full visual field model for normal people and a model for people with visual field defects, and combining eye-tracking and head posture perception data, we adjusted the visual field content in real time to generate dynamic videos, ensuring that the completed area is synchronized with the fixation point, so as to achieve consistency between visual field reconstruction and the patient's visual perception.
The generated dynamic video is highly consistent with the patient's visual perception, improving the independence of patients with visual field defects and the effectiveness of rehabilitation training. It is applicable to various types of visual field defects, improves orientation and obstacle avoidance abilities, and is suitable for visual rehabilitation training and efficacy evaluation.
Smart Images

Figure CN122391354A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision technology, and more particularly to a method for constructing first-person perspective videos for people with visual impairments. Background Technology
[0002] Visual field defect simulation can aid in clinical education by highlighting the signs and symptoms of visual disorders. Currently, most visual simulations are based on the "patch occlusion hypothesis," which uses black or gray patches to obscure the visual field, depicting visual field defects. Clinical studies have found these simulations to be flawed. Subjective feedback from patients with different types of visual field defects falls into two main categories: either the defective area disappears from consciousness, or localized visual distortion appears around the defective area. These two types of feedback sometimes occur individually or simultaneously. Existing research has broadly summarized the subjective visual feedback results of a small sample of patients with different types of visual field defects, primarily based on the flawed "patch occlusion hypothesis" or Amsler grid peripheral visual field examination results. This research fails to characterize visual field information in natural scenes and lacks in-depth research into the physiological characteristics, pathological mechanisms, and visual cognition of visual field defects, thus failing to provide sufficient theoretical support for the presentation of visual field information and visual computation in visual field defects.
[0003] Current visual field simulations for individuals with visual field defects remain static, and the results often differ significantly from actual patient observations. The simulation and reconstruction of patients' visual fields primarily use static images with fixed fixation points, failing to consider the dynamic changes in visual characteristics of the fixation point, resulting in a disconnect from real-world scenarios and limited auxiliary effects. Existing technologies lack specific methods for constructing dynamic videos for individuals with visual field defects and do not incorporate first-person perspective alignment techniques. This leads to significant discrepancies between the visual field completion area and the patient's actual visual perception, making it difficult to meet the dynamic needs of rehabilitation training and daily activities. Summary of the Invention
[0004] In response to the technical problems mentioned in the background section, this invention provides a method for constructing first-person perspective videos for individuals with visual impairments. This invention employs interdisciplinary research methods and approaches to construct a first-person perspective video method for individuals with visual impairments, objectively depicting the realistic visual dynamics of these individuals.
[0005] The technical means employed in this invention are as follows:
[0006] A method for constructing first-person perspective videos for people with visual impairments includes the following steps: Step 1: Construct a full visual field model for normal individuals; Based on ophthalmological clinical standards and publicly available physiological visual datasets, collect visual field data from people with normal vision, and combine this data with the physiological structure of the retina to construct a full visual field model for normal individuals; Step 2: Construct a first-person visual field model for individuals with visual field defects; Step 3: Adjust the content focus range of the reconstructed visual field based on the fixed fixation point to generate a standardized fixed fixation point visual field map; Step 4: Capture the user's gaze coordinates, saccade trajectory, and gaze duration in real time using an eye-tracking device, and construct a dynamic sequence of gaze changes by combining it with head posture perception data; Step 5: Based on the dynamic change sequence of the gaze point, process each frame of the video in real time; integrate the real-time adjusted frame sequence along the time axis and combine it with the scene dynamic features to generate a continuous video.
[0007] Furthermore, the visual field data of the normal vision population includes: central visual field, near-central visual field, intermediate peripheral visual field, peripheral visual field, and physiological blind spot; The central visual field is the spatial angle range of 0° to 5° from the center of the retina; the near-central visual field is the spatial angle range of 5° to 20° from the center of the retina; the intermediate peripheral visual field is the spatial angle range of 20° to 60° from the center of the retina, corresponding to the intermediate peripheral region of the retina; the peripheral visual field is the spatial angle range from 60° from the center of the retina to the physiological visual field limit; the physiological blind spot is the monocular local blind area corresponding to the optic nerve head.
[0008] Furthermore, step 2 includes the following steps: Step 21: Obtain multi-source heterogeneous data; obtain relevant data of patients with known visual field defects, and construct a multi-source heterogeneous dataset of patients with visual field defects based on the relevant data of patients with known visual field defects; the relevant data includes: basic patient information and medical history information, perimeter quantitative detection data, eye examination results, auxiliary examination results and neuroimaging examination results; wherein, the patient is a patient with visual field defects such as peripheral visual field constriction or hemianopsia, and the total visual field is not less than 20 degrees and the central visual acuity is not less than 0.5; Step 22: Based on the known multi-level photoreceptive depth model and full-field visual information calculation model, combined with the patient's baseline eye movement-head posture parameters, the visual field suturing strategy of the full-field visual information calculation model is used to suture the contents of the corresponding visual field defect area, and the multi-level photoreceptive depth model is used to adjust the spatial distribution of the surrounding normal visual field content. Step 23: Based on the patient's baseline eye movement and head posture data, and according to the known first-person visual field model, calibrate the presentation angle and content distribution of the reconstructed visual field to generate a personalized visual field reconstruction map that is consistent with the patient's visual perception.
[0009] Furthermore, in step 3, based on the fixed fixation point, the reconstructed visual field is divided into a known central clear area, a near-central transition area, a mid-peripheral area, and a peripheral area. The content focusing range of the reconstructed visual field is adjusted so that the central clear area maintains the highest content clarity, the near-central transition area performs a gradual content transition, and the mid-peripheral area and the peripheral area perform a decreasing resolution allocation, so that the fixation point area is clearly presented and the missing and filled area is naturally connected with the normal visual field area. By using known visual perception consistency verification models, we collect patients' subjective evaluation feedback on fixed fixation point visual field maps, optimize the clarity, contrast, and content distribution of visual field maps, and generate standardized fixed fixation point visual field maps. The subjective assessment includes: the degree of invisibility of the defect area, the degree of local distortion, the degree of blurring, the degree of boundary perceptibility, the naturalness of the scene, the comfort level, and the consistency with the patient's daily visual perception.
[0010] Furthermore, the head posture perception data includes: rotation angle and displacement.
[0011] Furthermore, in step 5, a known multi-layered optical perception depth model is used to adjust the visual field presentation range according to the current gaze point position, prioritizing the clear presentation of the area surrounding the gaze point; based on a known full-view visual information calculation model, the completion content of the missing area is dynamically updated to ensure that the completion area moves synchronously with the gaze point; visual coherence is maintained through a temporal stabilization and delay compensation mechanism; the temporal stabilization and delay compensation mechanism includes: aligning the eye movement sequence, head pose sequence, and video frame sequence according to a unified timestamp; predicting the target gaze point and target pose at the rendering time based on the gaze point and head pose change trends of the previous few frames; using the prediction results to perform pre-compensation transformation on the current frame; and after the actual observation arrives, using temporal filtering and error write-back methods to correct the current frame.
[0012] Further, in step 5, the real-time adjusted frame sequence is sorted by timestamp, and high-frequency eye movement and head posture data are interpolated to the corresponding video frame times to form a sequence of field of view transformation parameters corresponding to each frame; scene dynamic features between adjacent frames are extracted, including: optical flow field, camera self-motion, target motion trajectory, depth change and occlusion boundary; based on the scene dynamic features, motion compensation is performed on the result of the previous frame and weighted and fused with the result of the current frame, wherein the static background area adopts a higher temporal smoothing weight, and the moving target boundary and occlusion change area adopt a higher current frame weight; when a sudden scene change or large head movement is detected, the inherited weight of the previous frame is reduced and fast realignment is triggered, finally generating a first-person perspective video that is consistent with the dynamic changes of the gaze point, is temporally continuous and spatially smooth.
[0013] Compared with the prior art, the present invention has the following advantages: Existing visual field defect simulations often deviate significantly from patients' actual visual perception, and typically reconstruct static images, lacking the characteristics of gaze point changes in real-world scenarios. Therefore, this invention addresses this issue by employing interdisciplinary research methods and pathways to generate dynamic videos from the patient's first-person perspective. It adaptively selects a visual computation model based on the defect type, closely matching the patient's actual visual perception (no black patch occlusion, only removal of the defect content and minor surrounding distortions), thus resolving the visual perception bias problem of existing technologies. Furthermore, this invention customizes model parameters for visual field defects caused by different etiologies (visual pathway / photoreceptor damage), adapting to various visual field defect types such as glaucoma, macular degeneration, and hemianopsia, expanding the applicable population. By combining eye movement and posture data to adjust the visual field content in real time, it seamlessly connects the completed area with the gaze point and the normal visual field. Combined with real-time eye tracking, it perfectly recreates the dynamic characteristics of gaze point movement in human vision, synchronizing the video with the patient's visual movement, improving safety in mobile scenarios, providing patients with visual field defects with an auxiliary visual field that matches their actual visual perception, improving orientation and obstacle avoidance abilities, enhancing independence, and also serving as a tool for rehabilitation training and efficacy evaluation, thus promoting the development of visual rehabilitation technology. Attached Figure Description
[0014] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0015] Figure 1 This is a schematic diagram of the overall process of the present invention.
[0016] Figure 2 This is a schematic diagram of the field of view division in this invention. Detailed Implementation
[0017] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0018] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0019] like Figure 1-2 As shown, this invention provides a method for constructing first-person perspective videos for people with visual impairments, including the following steps: Step 1: Construct a full visual field model for normal individuals; Based on ophthalmological clinical standards and publicly available physiological visual datasets, collect visual field data from people with normal vision, and combine this data with the physiological structure of the retina to construct a full visual field model for normal individuals; Specifically, the first step is to collect and quantify data: based on ophthalmological clinical standards and publicly available physiological visual datasets, visual field data (including central visual field, peripheral visual field, range of physiological blind spots, and sensitivity distribution) of people with normal vision are collected. Combined with retinal physiological structures (such as differences in visual function in the macula and peripheral retinal regions), a full visual field model for normal people is quantified and defined.
[0020] In this application, the central visual field is defined as the spatial angle range of 0° to 5° from the center of the retina, corresponding to the fovea region of the macula. It is the core area with the highest visual sensitivity and is mainly responsible for fine visual discrimination, text reading, and color perception. In the model, it is configured with a parameter combination of high resolution, high contrast, and high color channel weights to match its high-precision visual characteristics.
[0021] Near-central visual field: defined as a spatial angle range of 5° to 20°, corresponding to the periretinal fovea region, mainly responsible for gaze assistance, rapid eye saccades and local moving target perception; configured in the model with medium to high resolution, taking into account the parameter combination of static visual resolution and dynamic response capability, so as to achieve rapid visual processing around the gaze point.
[0022] Mid-peripheral field of view: defined as a spatial angle range of 20° to 60°, corresponding to the mid-peripheral region of the retina, mainly responsible for dynamic target detection, spatial localization and overall environmental perception; configured in the model with a parameter combination of medium resolution, dynamic response priority and motion detection enhancement to improve the efficiency of capturing and locating mid-range moving targets.
[0023] Peripheral visual field: defined as the spatial angular range from 60° to the physiological visual field limit (approximately 100° to 110°), corresponding to the peripheral retina region, which is the edge vision region. It is mainly responsible for visual field boundary perception, large-scale motion detection, and alert visual warning. In the model, it is configured with a parameter combination of low resolution, wide coverage, high motion sensitivity, and low color weight to adapt to the low resolution and high motion sensitivity characteristics of peripheral vision.
[0024] Physiological blind spot: This refers to the localized blind area of one eye corresponding to the optic disc, located approximately 15° temporally with a diameter of 10°–15°. It is mainly distributed in the transition area between the near-central visual field (5°–20°) and the mid-peripheral visual field (20°–60°), and does not overlap with the central visual field (0°–5°) or the peripheral visual field (60°–110°). The visual processing model of this invention employs a visual field suturing mechanism within the near-central and mid-peripheral visual field divisions to address the physiological blind spot region, thereby eliminating the impact of the localized blind spot on visual information resolution.
[0025] Step 2: Construct a first-person visual field model for individuals with visual field defects. Specifically, this includes the following steps: Step 21: Acquire multi-source heterogeneous data. First, recruit eligible patients with visual field defects and collect multi-source subjective and objective data, including medical history, ophthalmological examinations, and neuroimaging, to clarify the location, size, and morphology of the defects. Combine this with publicly available datasets and literature to ensure sample representativeness and data comprehensiveness. Perform preprocessing on the multi-source heterogeneous data, including standardization, cleaning, and coordinate normalization, to lay the foundation for subsequent analysis and modeling. Finally, construct a cross-modal associated feature set through feature fusion, coupling defect morphology, physiological structure, functional indicators, and subjective experience. Combine this with prior information such as etiology and defect type to provide initialization and parameter constraints for the model. Based on the preprocessing and fusion results, construct a dynamically updated multi-source heterogeneous visual field information prior knowledge base.
[0026] Based on a prior knowledge base of multi-source heterogeneous visual field information (using the knowledge base framework and supplementing it with patient visual feedback data in dynamic scenarios), we collect patient visual field defect types (central defect, arcuate defect, hemianopsia, etc.), defect range, eye movement characteristics (basic fixation point distribution, saccade habits) and head posture baseline data. Step 22: Based on the multi-level light perception depth model and the full field of vision information calculation model, combined with the patient's personalized data, perform visual field suturing on the corresponding visual field defect area and adjust the spatial distribution of the surrounding normal visual field content. Step 23: Integrate a first-person perspective adaptation mechanism. Based on the patient's baseline eye movement and head posture data, calibrate the presentation angle and content distribution of the reconstructed visual field to generate a personalized visual field reconstruction map that is consistent with the patient's visual perception.
[0027] Step 3: Adjust the content focus range of the reconstructed visual field based on the fixed fixation point to generate a standardized fixed fixation point visual field map; Regarding the selection of fixation points: select fixed fixation points commonly used in clinical rehabilitation training (such as the center of the visual field, key points at the edge of the defect area, etc.), or set personalized fixed fixation points according to the patient's daily visual habits. Visual field optimization: Based on a fixed fixation point, the reconstructed visual field is divided into a central clear area, a near-central transition area, a mid-peripheral area, and a peripheral area. Resolution weights, contrast weights, and spatial distortion limits are assigned to each of these areas to ensure the central clear area maintains the highest content clarity, the near-central transition area implements a gradual content transition, and the mid-peripheral and peripheral areas implement a decreasing resolution allocation. Luminosity continuity, gradient continuity, and geometric continuity constraints are applied to the boundary between the missing area and the normal visual field to ensure a natural transition, thereby adjusting the content focus range of the reconstructed visual field and ensuring clear presentation of the fixation point area and a natural transition between the missing area and the normal visual field. A visual perception consistency verification model (a first-person perspective visual field alignment method for individuals with visual field defects) is used to collect patients' subjective evaluation feedback on the fixed fixation point visual field map, optimizing the clarity, contrast, and content distribution of the visual field map to generate a standardized fixed fixation point visual field map. The subjective evaluation includes at least the degree of invisibility of the missing area, the degree of local distortion, the degree of blurring, the degree of boundary perceptibility, scene naturalness, comfort, and consistency with the patient's daily visual perception.
[0028] Step 4: Capture the user's gaze coordinates, saccade trajectory, and gaze duration in real time using an eye-tracking device, and construct a dynamic sequence of gaze changes by combining it with head posture perception data. Real-time eye-tracking data acquisition: The eye-tracking device captures dynamic data such as the patient's gaze coordinates, saccade trajectory, and gaze duration in real time, and combines them with head posture perception data (rotation angle, displacement) to construct a dynamic sequence of gaze changes; Real-time video frame adjustment: Based on the dynamic change sequence of the gaze point, each frame of the video is processed in real time: using a multi-layered optical perception depth model, the field of view is adjusted according to the current gaze point position, prioritizing the clear presentation of the area surrounding the gaze point; based on the full field of view visual information calculation model, the content of the missing area is dynamically updated to ensure that the completed area is synchronized with the movement of the gaze point; through a temporal stabilization and delay compensation mechanism, video screen jumps or blurring are avoided, maintaining visual continuity; the temporal stabilization and delay compensation mechanism includes: aligning the eye movement sequence, head posture sequence and video frame sequence according to a unified timestamp; predicting the target gaze point and target posture at the rendering time based on the gaze point and head posture change trends of the previous few frames; using the prediction results to perform pre-compensation transformation on the current frame; after the actual observation arrives, temporal filtering and error write-back are used to correct the current frame to avoid video screen jumps, tearing or blurring, maintaining visual continuity.
[0029] Continuous video generation: The real-time adjusted frame sequence is integrated along the timeline, and the inter-frame transition effect is optimized by combining scene dynamic features (such as the speed of moving targets and changes in scene lighting) to generate a continuous video with dynamic changes in the gaze point. The specific process is as follows: The frame sequence is sorted based on the timestamp of each processed frame. The high-frequency eye movement and head posture sampling results are interpolated to the corresponding video frame time to form a sequence of field of view transformation parameters corresponding to each video frame. Scene dynamic features between adjacent frames are extracted (including at least optical flow, camera self-motion, target motion trajectory, depth changes, and occlusion boundaries). Motion compensation mapping is performed on the processing result of the previous frame according to the scene dynamic features, and weighted fusion is performed with the processing result of the current frame. Static background areas are given a higher temporal smoothing weight, while moving target boundaries and occlusion change areas are given a higher weight in the current frame. When a sudden scene change or large head movement is detected, the inherited weight of the previous frame is reduced and fast realignment is triggered to avoid ghosting and blurring. Finally, a first-person perspective video that is continuous in time, smooth in space, and consistent with the dynamic changes in the gaze point is output.
[0030] Step 5: Based on the dynamic change sequence of the gaze point, process each frame of the video in real time; integrate the real-time adjusted frame sequence along the time axis and combine it with the scene dynamic features to generate a continuous video.
[0031] Objective verification can be employed by collecting optodynamic interaction data (fixation stability, saccade response time, and synchronization between eye movements and video content) from patients while watching videos to verify the compatibility between the video and the patient's visual movements. Simultaneously, subjective evaluation can be conducted by combining patient subjective ratings (clarity, naturalness, and comfort) with assessments from clinical medical staff to optimize the dynamic adjustment parameters of the video (fixation response speed and defect completion smoothness). Finally, iterative optimization can be achieved by continuously adjusting the model parameters based on the verification results to ensure that the video construction effect conforms to individual patient differences and improves the dynamic visual experience.
[0032] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments. In the above embodiments of the present invention, the descriptions of each embodiment have their own emphasis; parts not described in detail in a certain embodiment can be referred to in the relevant descriptions of other embodiments. It should be understood that the disclosed technical content in the several embodiments provided in this application can be implemented in other ways.
[0033] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for constructing first-person perspective videos for people with visual impairments, characterized in that, Includes the following steps: Step 1: Construct a full visual field model for normal individuals; Based on ophthalmological clinical standards and publicly available physiological visual datasets, collect visual field data from people with normal vision, and combine this data with the physiological structure of the retina to construct a full visual field model for normal individuals; Step 2: Construct a first-person visual field model for individuals with visual field defects; Step 3: Adjust the content focus range of the reconstructed visual field based on the fixed fixation point to generate a standardized fixed fixation point visual field map; Step 4: Capture the user's gaze coordinates, saccade trajectory, and gaze duration in real time using an eye-tracking device, and construct a dynamic sequence of gaze changes by combining it with head posture perception data. Step 5: Based on the dynamic change sequence of the gaze point, process each frame of the video in real time; integrate the real-time adjusted frame sequence along the time axis and combine it with the scene dynamic features to generate a continuous video.
2. The method for constructing first-person perspective videos for people with visual impairments according to claim 1, characterized in that, The visual field data of the normal vision population includes: central visual field, near-central visual field, intermediate visual field, peripheral visual field, and physiological blind spot; The central visual field is the spatial angle range of 0° to 5° from the center of the retina; the near-central visual field is the spatial angle range of 5° to 20° from the center of the retina; the intermediate peripheral visual field is the spatial angle range of 20° to 60° from the center of the retina, corresponding to the intermediate peripheral region of the retina; the peripheral visual field is the spatial angle range from 60° from the center of the retina to the physiological visual field limit; the physiological blind spot is the monocular local blind area corresponding to the optic nerve head.
3. The method for constructing first-person perspective videos for people with visual impairments according to claim 1, characterized in that, Step 2 includes the following steps: Step 21: Obtain multi-source heterogeneous data; obtain relevant data of patients with known visual field defects, and construct a multi-source heterogeneous dataset of patients with visual field defects based on the relevant data of patients with known visual field defects; the relevant data includes: basic patient information and medical history information, perimeter quantitative detection data, eye examination results, auxiliary examination results and neuroimaging examination results; wherein, the patient is a patient with visual field defects such as peripheral visual field constriction or hemianopsia, and the total visual field is not less than 20 degrees and the central visual acuity is not less than 0.5; Step 22: Based on the known multi-level photoreceptive depth model and full-field visual information calculation model, combined with the patient's baseline eye movement-head posture parameters, the visual field suturing strategy of the full-field visual information calculation model is used to suture the contents of the corresponding visual field defect area, and the multi-level photoreceptive depth model is used to adjust the spatial distribution of the surrounding normal visual field content. Step 23: Based on the patient's baseline eye movement and head posture data, and according to the known first-person visual field model, calibrate the presentation angle and content distribution of the reconstructed visual field to generate a personalized visual field reconstruction map that is consistent with the patient's visual perception.
4. The method for constructing first-person perspective videos for people with visual impairments according to claim 1, characterized in that, In step 3, based on the fixed fixation point, the reconstructed visual field is divided into the known central clear area, near-central transition area, mid-peripheral area and peripheral area. The content focusing range of the reconstructed visual field is adjusted so that the central clear area maintains the highest content clarity, the near-central transition area performs a gradual content transition, and the mid-peripheral area and peripheral area perform a decreasing resolution allocation, so that the fixation point area is clearly presented and the missing and filled area is naturally connected with the normal visual field area. By using known visual perception consistency verification models, we collect patients' subjective evaluation feedback on fixed fixation point visual field maps, optimize the clarity, contrast, and content distribution of visual field maps, and generate standardized fixed fixation point visual field maps. The subjective assessment includes: the degree of invisibility of the defect area, the degree of local distortion, the degree of blurring, the degree of boundary perceptibility, the naturalness of the scene, the comfort level, and the consistency with the patient's daily visual perception.
5. The method for constructing first-person perspective video for people with visual impairments according to claim 1, characterized in that, The head posture perception data includes: rotation angle and displacement.
6. The method for constructing first-person perspective video for people with visual impairments according to claim 1, characterized in that, In step 5, the known multi-level light perception depth model is used to adjust the visual field presentation range according to the current gaze point position, prioritizing the clear presentation of the area surrounding the gaze point; based on the known full-field visual information calculation model, the content of the missing area is dynamically updated to ensure that the completed area moves synchronously with the gaze point. Visual coherence is maintained through temporal stability and delay compensation mechanisms; The temporal stabilization and delay compensation mechanism includes: aligning the eye-tracking sequence, head pose sequence, and video frame sequence according to a unified timestamp; predicting the target gaze point and target pose at the rendering time based on the gaze point and head pose change trend of the previous few frames; using the prediction results to perform pre-compensation transformation on the current frame; and correcting the current frame by using temporal filtering and error write-back after the actual observation arrives.
7. A method for constructing first-person perspective video for people with visual impairments according to claim 1, characterized in that, In step 5, the real-time adjusted frame sequence is sorted by timestamp, and high-frequency eye movement and head posture data are interpolated to the corresponding video frame times to form a sequence of field of view transformation parameters corresponding to each frame. Scene dynamic features between adjacent frames are extracted. These scene dynamic features include: optical flow field, camera self-motion, target motion trajectory, depth change, and occlusion boundary. Based on the scene dynamic features, motion compensation is performed on the previous frame result and weighted fusion is performed with the current frame result. The static background area is given a higher temporal smoothing weight, while the moving target boundary and occlusion change area are given a higher current frame weight. When a sudden scene change or large head movement is detected, the inherited weight of the previous frame is reduced and fast realignment is triggered. Finally, a first-person perspective video that is consistent with the dynamic changes of the gaze point, is temporally continuous, and spatially smooth is generated.