Dynamic environment three-dimensional reconstruction system and method for AD patient space cognitive training

By combining augmented reality technology with incremental neural point cloud coding, the spatial cognitive training of AD patients can be evaluated and dynamically adjusted in real time, solving the problem of insufficient integration of real environment and virtual guidance in existing systems, and achieving efficient and personalized spatial cognitive training effects.

CN120673977AInactive Publication Date: 2025-09-19HUNAN PROVINCIAL HOSPITAL OF INTEGRATED TRADITIONAL CHINESE & WESTERN MEDICINE (AFFILIATED HOSPITAL OF HUNAN PROVINCIAL RES INST OF TRADITIONAL CHINESE MEDICINE HUNAN PROVINCIAL RES INST OF TRADITIONAL CHINESE MEDICINE CLINICAL RES INST HUNAN PROVINCIAL RES INST OF TRADITIONAL CHINESE MEDICINE ONCOLOGY RES INST)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510777526.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-09-19
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing spatial cognitive training system for AD patients cannot effectively combine the real environment with virtual guidance, lacks a dynamic adjustment mechanism, making it difficult to transfer the training effects to daily life, and lacks quantitative evaluation and real-time feedback, resulting in low training efficiency.

Method used

It uses augmented reality technology and incremental neural point cloud coding technology, combined with eye tracking and multimodal perception fusion, to evaluate the patient's cognitive status in real time, dynamically adjust the training difficulty and content, and provide personalized training guidance through the augmented reality display module.

Benefits of technology

It improves the transferability of training effects, enhances the accuracy and stability of environmental understanding, avoids cognitive overload or insufficient challenges, and builds a closed-loop, efficient training system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673977A_ABST
    Figure CN120673977A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medical rehabilitation, in particular to a dynamic environment three-dimensional reconstruction system and method for AD patient spatial cognitive training, and the system captures a training environment in real time through a binocular imaging module, collects motion parameters of a patient in combination with an inertial measurement unit, extracts scene features through an incremental neural point cloud coding module, and generates semantic information. The augmented reality display module projects a virtual landmark and a training scene to the AR glasses, the eye movement tracking module monitors eye movement data of a patient, the cognitive state evaluation module analyzes the spatial cognitive state of the patient, and the training guide module provides adaptive training. And the data management module stores data and generates an evaluation report. The system fuses virtuality and reality through the AR technology, improves the training effect, helps AD patients to improve the spatial cognitive ability, overcomes the cognitive gap between traditional virtual reality training and the real environment, and enhances the ability of the training effect to migrate to daily life.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical rehabilitation technology, and in particular to a dynamic environment three-dimensional reconstruction system and method for spatial cognitive training of AD patients, which is suitable for the spatial cognitive ability assessment and training of AD patients. Background Art

[0002] Alzheimer's disease is a progressive neurodegenerative disorder. Patients often exhibit spatial cognitive impairments, such as difficulty with spatial orientation, decreased path planning ability, and environmental cognition impairment. These impairments not only severely impact patients' daily living abilities but also increase their risk of getting lost and experiencing accidental injury.

[0003] Currently, spatial cognition training for patients with Alzheimer's disease (AD) primarily involves traditional paper-and-pencil tests, virtual reality training, and on-site navigation training. However, these methods present several challenges: First, traditional paper-and-pencil tests lack real-world interaction, making them ineffective in reflecting patients' spatial cognition in real environments. Second, while virtual reality training offers an immersive experience, there's a cognitive gap between it and the real world, making it difficult to transfer training results to daily life. Furthermore, while on-site navigation training offers a high degree of realism, it lacks quantitative assessment methods, making it difficult to accurately assess patients' progress. Furthermore, the training process lacks real-time guidance and adjustment mechanisms, resulting in low training efficiency.

[0004] Furthermore, existing spatial cognition training systems often use fixed training protocols and are unable to dynamically adjust to the patient's real-time cognitive state. Furthermore, they lack consideration for the patient's cognitive load, which can easily lead to cognitive overload or insufficient cognitive challenge during training, compromising training effectiveness.

[0005] Therefore, how to develop a spatial cognitive training system that can effectively combine the real environment with virtual guidance and dynamically adjust according to the patient's cognitive state has become a technical problem that needs to be solved urgently. Summary of the Invention

[0006] The purpose of the present invention is to provide a dynamic environment three-dimensional reconstruction system and method for spatial cognitive training of AD patients. The system seamlessly integrates virtual guidance with the real environment through augmented reality technology, combines incremental neural point cloud coding technology to achieve dynamic three-dimensional reconstruction of the environment, and evaluates the patient's cognitive status in real time based on eye tracking and multimodal perception fusion technology, dynamically adjusts the training difficulty and content, thereby improving the pertinence and effectiveness of spatial cognitive training.

[0007] The present invention proposes a dynamic environment 3D reconstruction system for spatial cognition training of AD patients, comprising:

[0008] Binocular imaging module, used to shoot the training environment in real time and obtain scene image data;

[0009] Inertial measurement unit, used to collect motion parameters of AD patients during training;

[0010] an incremental neural point cloud encoding module, electrically connected to the binocular imaging module, configured to receive the scene image data, extract feature points based on a cognitive load adaptive feature sampling mechanism, generate scene semantic information through multi-scale fused semantic perception feature description, and implement dynamic map updates using a spatiotemporal constraint incremental learning architecture;

[0011] an augmented reality display module, communicatively connected to the incremental neural point cloud encoding module, configured to receive the scene semantic information, generate virtual landmarks and training scenes, and project the virtual landmarks and training scenes onto an AR glasses display interface;

[0012] An eye tracking module, installed in the augmented reality display module, for collecting eye movement data of AD patients;

[0013] a cognitive state assessment module, electrically connected to the eye tracking module and the inertial measurement unit, configured to receive the eye movement data and the motion parameters and analyze the spatial cognitive state of the AD patient based on an attention-driven feature fusion strategy;

[0014] a training guidance module, communicatively connected to the cognitive state assessment module and the augmented reality display module, configured to receive the spatial cognitive state, generate adaptive training guidance information, and display the information through the augmented reality display module;

[0015] The data management module is connected to the incremental neural point cloud encoding module, the cognitive state assessment module and the training guidance module respectively, and is used to store training process data and patient behavior data and generate a training assessment report.

[0016] Preferably, the cognitive load adaptive feature sampling mechanism includes:

[0017] A field of view region division unit is used to divide the field of view into the fovea centralis, the parafovea centralis, and the peripheral region, and to set different feature point sampling densities according to the importance of the regions;

[0018] A semantic importance determination unit, configured to identify key navigation elements in a scene and assign a higher sampling weight to the key navigation elements;

[0019] A cognitive state response unit is used to estimate the current cognitive load level based on the eye movement data, reduce sampling complexity in a high cognitive load state, and increase sampling complexity in a low cognitive load state.

[0020] Preferably, the multi-scale fused semantic perception feature description includes:

[0021] A multi-scale representation unit for extracting feature descriptions at three different scales: a small scale for capturing microscopic features, a medium scale for capturing object and regional features, and a large scale for capturing spatial layout relationships;

[0022] Semantic enhancement unit, which is used to integrate semantic label information into the feature description process and apply different description parameters to different semantic categories;

[0023] The feature matching unit is used to match feature points using a hierarchical matching strategy, calculate similarity using normalized cross-correlation, and filter false matches based on geometric consistency constraints.

[0024] Preferably, the spatiotemporal constrained incremental learning architecture includes:

[0025] Temporal continuity guarantee unit, used to limit the feature change rate between adjacent frames, establish a temporal sliding window, and maintain historical coherence;

[0026] The spatial structure preserving unit is used to maintain the topological relationship graph between feature points, calculate the deformation degree of the topological structure, and perform feature association based on topological invariance constraints;

[0027] The incremental parameter update unit is used to adopt a hierarchical update strategy for the network parameter Kp and dynamically adjust the learning rate according to the degree of scene change;

[0028] The resource efficiency optimization unit is used to implement a selective update strategy, updating only the parameters in the changed area, and adopting a floating-point precision adaptive mechanism to balance accuracy and computational complexity.

[0029] Preferably, the attention-driven feature fusion strategy includes:

[0030] The spatial attention unit is used to generate a spatial attention map based on the semantic importance of the scene, allocating higher computing resources and representation accuracy to important areas;

[0031] a modality fusion unit for integrating visual features, the motion parameters, and depth information, designing an adaptive weight allocation strategy, and adjusting weights according to modality reliability;

[0032] Semantic enhancement processing unit, used to enhance the features of identified key navigation elements. Enhancement processing includes contrast improvement, edge sharpening, and color saliency adjustment.

[0033] The dynamic adaptation unit is used to monitor the response degree of AD patients to different features and adjust the feature fusion weight according to the response strength.

[0034] Preferably, the cognitive status assessment module includes:

[0035] an eye movement analysis unit, configured to calculate fixation duration, saccade pattern, and pupil diameter change based on the eye movement data;

[0036] Spatial Comprehension Assessment Unit, used to analyze spatial comprehension abilities based on feature gaze preferences;

[0037] Navigational ability assessment unit, used to evaluate navigational ability through landmark recognition accuracy;

[0038] a depth perception evaluation unit for evaluating spatial perception capabilities based on depth perception accuracy;

[0039] A comprehensive scoring unit is used to generate a comprehensive cognitive status score based on the output results of the eye movement analysis unit, the spatial understanding evaluation unit, the navigation ability evaluation unit and the depth perception evaluation unit.

[0040] Preferably, the training guidance module includes:

[0041] A difficulty control unit is used to determine the training difficulty level based on the spatial cognition state, and define five levels of training difficulty, from simple recognition to complex navigation;

[0042] The path planning unit is used to optimize the posture using the Bayesian minimum mean square error method, predict the movement trajectory and posture of AD patients through the path integral method, and construct the walking path based on the updated map;

[0043] A virtual guidance unit, which generates virtual navigation markers and visual cues to guide AD patients to match virtual and real landmarks through AR projection;

[0044] The feedback generation unit is used to provide real-time feedback information based on the behavioral performance of AD patients, including visual and audio prompts.

[0045] Preferably, the augmented reality display module includes:

[0046] a virtual content generation unit, configured to generate virtual landmarks and training scenes based on the scene semantic information;

[0047] a three-dimensional registration unit, configured to achieve spatial alignment between the virtual landmark and the real landmark;

[0048] A rendering optimization unit, used to adjust rendering parameters according to ambient lighting conditions and the visual characteristics of AD patients;

[0049] The display control unit is used to project the optimized virtual content onto the AR glasses display interface.

[0050] Preferably, the data management module includes:

[0051] Training data storage unit, used to record the walking trajectory, posture changes, eye movement data and task completion status of AD patients;

[0052] A data analysis unit, used to compare the deviation between the planned path and the actual path and analyze the accuracy and speed of virtual landmark recognition;

[0053] Progress tracking unit, used to establish long-term training records and track changes in spatial cognitive abilities;

[0054] Report generation unit, used to generate personalized training reports and progress indicators to support medical staff in evaluating treatment effects.

[0055] A dynamic environment 3D reconstruction method for spatial cognition training of AD patients includes the following steps:

[0056] Obtaining scene image data of the training environment and motion parameters of AD patients;

[0057] Extract feature points from scene image data based on cognitive load adaptive feature sampling mechanism;

[0058] Generate scene semantic information through multi-scale fused semantic-aware feature description;

[0059] Adopting a spatiotemporal-constrained incremental learning architecture to achieve dynamic map updates;

[0060] Generate virtual landmarks and training scenes based on scene semantic information, and project the virtual landmarks and training scenes onto the AR glasses display interface;

[0061] Collect eye movement data from AD patients;

[0062] Analyze the spatial cognitive status of AD patients based on attention-driven feature fusion strategy;

[0063] Generate adaptive training guidance information based on spatial cognitive status and display it through the AR glasses display interface;

[0064] Store training process data and patient behavior data, and generate training evaluation reports.

[0065] The present invention has the following beneficial effects:

[0066] 1. By integrating virtual training content with the real environment through augmented reality technology, the cognitive gap between traditional virtual reality training and the real environment is overcome, and the ability to transfer training effects to daily life is improved;

[0067] 2. Incremental neural point cloud coding technology is used to achieve dynamic 3D reconstruction of the environment, reducing computing resource requirements, making it suitable for real-time operation on mobile devices, while improving the accuracy and stability of environmental understanding;

[0068] 3. Based on a cognitive load adaptive feature sampling mechanism, the system can dynamically adjust the density and complexity of feature extraction according to the patient's cognitive state, avoiding cognitive overload or insufficient challenge.

[0069] 4. Through eye tracking and multimodal perception fusion technology, the patient's spatial cognitive status is assessed in real time, providing a scientific basis for dynamic adjustment of training difficulty and achieving personalized training;

[0070] 5. A complete evaluation-adjustment-training-feedback closed-loop system has been built to make the training process more accurate and efficient, and improve the compliance and effectiveness of training. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] Figure 1 Schematic diagram of the overall architecture of the system of the present invention;

[0072] Figure 2 Schematic diagram of the structure of the incremental neural point cloud encoding module of the present invention;

[0073] Figure 3 This is a flow chart of the cognitive load adaptive feature sampling mechanism of the present invention;

[0074] Figure 4 Schematic diagram of the semantic perception feature description of multi-scale fusion of the present invention;

[0075] Figure 5 This is a flow chart for implementing the incremental learning architecture with spatiotemporal constraints of the present invention;

[0076] Figure 6 Schematic diagram of the attention-driven feature fusion strategy of the present invention;

[0077] Figure 7 Schematic diagram of the structure of the cognitive state assessment module of the present invention;

[0078] Figure 8 This is a workflow diagram of the training guidance module of the present invention;

[0079] Figure 9 This is a schematic structural diagram of the augmented reality display module of the present invention;

[0080] Figure 10 This is a functional diagram of the data management module of the present invention;

[0081] Figure 11 This is a flow chart of the dynamic environment three-dimensional reconstruction method for spatial cognition training of AD patients according to the present invention. DETAILED DESCRIPTION

[0082] Please refer to the attached Figure 1-11 The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0083] like Figure 1 As shown, the dynamic environment three-dimensional reconstruction system for spatial cognitive training of AD patients provided by the present invention includes a binocular imaging module 1, an inertial measurement unit 2, an incremental neural point cloud encoding module 3, an augmented reality display module 4, an eye tracking module 5, a cognitive state assessment module 6, a training guidance module 7 and a data management module 8.

[0084] In one embodiment of the present invention, the binocular imaging module 1 is used to shoot the training environment in real time and obtain scene image data. The binocular imaging module 1 is preferably composed of two independent and oppositely placed image sensors in conjunction with a fisheye lens, with an acquisition frequency of 30 frames per second and a resolution of 1920×1080 to ensure the acquisition of high-quality stereo images. For example, in a real training environment such as a hospital corridor or a rehabilitation center, the binocular imaging module 1 can clearly capture key visual clues in the daily navigation of AD patients, such as corridor corners, door and window signs. The binocular imaging module 1 is connected to the system main controller via a high-speed interface such as USB3.0 or MIPI to achieve high-speed transmission of image data.

[0085] In another embodiment, the inertial measurement unit 2 is used to collect motion parameters of AD patients during training. The inertial measurement unit 2 includes a three-axis accelerometer and a three-axis gyroscope, and the sampling rate is set to 200Hz, which can accurately capture the patient's position changes and posture changes. In actual applications, when the AD patient walks in the training scene, the inertial measurement unit 2 will record his walking speed, steering angle and acceleration changes. These data are crucial for evaluating the patient's motor ability and spatial positioning ability. The data of the inertial measurement unit 2 is transmitted to the system main controller through the I2C or SPI interface, and is time-synchronized with the image data to provide auxiliary information for posture estimation.

[0086] In the core embodiment of the present invention, Figure 2 As shown, the incremental neural point cloud coding module 3 is electrically connected to the binocular imaging module 1 and is used to receive scene image data, extract feature points, and generate scene semantic information. The incremental neural point cloud coding module 3 includes multiple functional units, implementing the entire process from feature extraction to 3D map construction.

[0087] like Figure 3 As shown, the incremental neural point cloud coding module 3 first extracts feature points based on a cognitive load adaptive feature sampling mechanism. The mechanism includes a visual field region division unit 31, a semantic importance determination unit 32, and a cognitive state response unit 33.

[0088] The field of view area division unit 31 divides the field of view into three areas: the fovea area (within 10° of visual angle), the parafovea area (10°-30° of visual angle) and the peripheral area (above 30° of visual angle). In the fovea area, the feature point sampling density is set to 100% of the baseline value; in the parafovea area, the sampling density is 60% to 80% of the baseline value; in the peripheral area, the sampling density is 30% to 50% of the baseline value. This sampling strategy is consistent with the visual perception characteristics of AD patients, because studies have shown that early AD patients usually have relatively retained information processing capabilities in the central visual field, while their peripheral visual field information processing capabilities are damaged earlier. For example, when a patient looks at the exit sign at the end of a hospital corridor, the system will perform high-density sampling on the sign and its surrounding areas, and perform lower-density sampling on surrounding areas such as the walls on both sides of the corridor, thereby optimizing the allocation of computing resources.

[0089] The semantic importance determination unit 32 is responsible for identifying key navigation elements in the scene, such as doors, windows, corridors, and landmarks, and assigning higher sampling weights to these elements. In the specific implementation, the sampling weight coefficient of key navigation elements is set to 1.5-2.5, and the sampling weight coefficient of general areas is set to 1.0. For example, in a nursing home environment, the system will identify key navigation elements such as room numbers, direction signs, and rest area signs, and perform high-density feature point sampling on these elements to ensure that AD patients can obtain sufficient navigation clues.

[0090] The cognitive state response unit 33 estimates the current cognitive load level based on eye movement data and dynamically adjusts the feature sampling strategy. When a high cognitive load state is detected (pupil diameter expands by more than 15% or the gaze point jumps frequently, with more than 3 jumps per second), the system will reduce the sampling complexity and reduce the number of feature points by 15% to 30%; when a low cognitive load state is detected (pupil diameter is stable or gaze behavior is concentrated, and the single-point gaze time exceeds 500ms), the system will appropriately increase the sampling complexity and increase the number of feature points by 5% to 15%. For example, when the system detects that the patient is showing a high cognitive load state in a complex intersection environment, it will reduce the feature sampling density, highlight the main navigation markers, and reduce the patient's cognitive stress.

[0091] like Figure 4 As shown, the incremental neural point cloud coding module 3 generates scene semantic information through multi-scale fusion semantic perception feature description. The description method includes a multi-scale representation unit 34, a semantic enhancement unit 35 and a feature matching unit 36.

[0092] The multi-scale representation unit 34 extracts feature descriptions at three different scales: small scale (3×3 window) captures micro features such as texture and edges; medium scale (9×9 window) captures object and regional features; large scale (27×27 window) captures spatial layout relationships. This multi-scale representation strategy is suitable for the cognitive characteristics of AD patients at different stages. For example, mild AD patients can usually recognize large-scale spatial layouts, but their ability to recognize small-scale details may decrease; while moderate AD patients may rely more on medium-scale features for environmental recognition. In actual applications, when patients are trained in a community park environment, the system will simultaneously extract small-scale road texture features, medium-scale object features such as seats and lampposts, and large-scale path layout features to fully support the patient's spatial cognitive process.

[0093] The semantic enhancement unit 35 incorporates semantic tag information into the feature description process, applying different description parameters to different semantic categories. For example, for key navigation elements such as doors and windows, the feature description increases the weight of edge and geometric shape information; for flat areas such as walls, the weight of texture information is increased. The semantic enhancement unit 35 uses a nonlinear conversion function to enhance semantic differences and improve differentiation. This nonlinear conversion function can be expressed as:

[0094]

[0095] Where x is the original eigenvalue, representing the raw response strength of the feature point; α is the enhancement factor for key navigation elements, typically set to 1.5-2.0 to increase their salience; β is the enhancement factor for general environmental elements, typically set to 0.8-1.0 to appropriately reduce the salience of non-critical areas. In actual training scenarios, such as hospital corridors, this function enhances the eigenvalues ​​of navigation elements like exit signs and directional signs, while maintaining or slightly reducing the eigenvalues ​​of general environmental elements like walls and ceilings, helping AD patients more easily identify key navigation cues.

[0096] The feature matching unit 36 ​​adopts a hierarchical matching strategy to match feature points, first performing coarse matching and then fine matching. The coarse matching stage uses the Hamming distance to quickly screen candidate matching points, and the fine matching stage uses normalized cross-correlation to calculate the similarity. The similarity threshold is dynamically adjusted according to the complexity of the scene and is usually set between 0.75-0.85. In addition, the feature matching unit 36 ​​also filters out false matches based on geometric consistency constraints to improve the matching accuracy. For example, when an AD patient moves from one room to another, the system needs to quickly match the common feature points in the two scenes to achieve continuous reconstruction of the environment. In experimental verification, the matching accuracy of this matching strategy in complex indoor environments can reach more than 85%, which is sufficient to support stable environmental reconstruction and positioning.

[0097] like Figure 5 As shown, the incremental neural point cloud coding module 3 uses a spatiotemporal constraint incremental learning architecture to achieve dynamic map updates. This architecture includes a temporal continuity guarantee unit 37, a spatial structure maintenance unit 38, an incremental parameter update unit 39, and a resource efficiency optimization unit 40.

[0098] The time continuity guarantee unit 37 limits the feature change rate between adjacent frames to no more than 15%, establishes a time sliding window containing 30-50 frames, and maintains historical consistency. At the same time, the unit also implements the abnormal frame detection and processing function. When a mutation frame (feature change rate exceeds 30%) is detected, the abnormal processing mechanism will be triggered to avoid mutation interference. During the training process of AD patients, interference factors such as sudden changes in illumination and rapid movement of people may occur in the environment. The unit can effectively filter out these interferences and maintain the stability of environmental reconstruction. For example, when a patient passes through a window area with direct sunlight, the change in illumination may cause a mutation in image features. The system will identify and filter such mutations to ensure the continuity of spatial cognitive training.

[0099] The spatial structure maintenance unit 38 maintains the topological relationship graph between feature points, calculates the deformation degree of the topological structure, and triggers correction when the deformation degree exceeds a set threshold (usually 0.1). This unit performs feature association based on topological invariance constraints to ensure the consistency and stability of the spatial structure. In the spatial cognition training of AD patients, it is crucial to maintain the consistency of the spatial structure of the environment, because sudden changes in the spatial structure may cause confusion in the patient's spatial cognition. For example, when the system reconstructs the public area of ​​a nursing home, even if local objects move (such as changes in the position of a chair), the overall spatial layout remains stable, providing patients with a reliable spatial reference.

[0100] The incremental parameter update unit 39 adopts a hierarchical update strategy for the network parameter Kp. The update rate of the bottom network parameters (responsible for basic feature extraction) is relatively low, set to 0.01-0.02; the update rate of the high-level network parameters (responsible for semantic understanding) is relatively high, set to 0.03-0.05. The update formula can be expressed as:

[0101]

[0102] Among them, Kp new is the updated network parameter, representing the network weights of each layer of the neural point cloud encoder; Kp oldis the original network parameter, indicating the network weight before the update; η is the learning rate, ranging from 0.01 to 0.05, which controls the step size of the parameter update; λ is the adaptation coefficient, ranging from 0.8 to 1.2, which adjusts the degree of influence of the error on the update; |error| is the absolute value of the current error, indicating the absolute size of the difference between the predicted result and the actual observation; tanh is the hyperbolic tangent function, which maps the error to the interval (-1, 1) to ensure that the update amplitude increases smoothly as the error increases; is the loss gradient, indicating the direction of parameter update.

[0103] This modified update strategy has the following advantages: (1) Using the absolute value of the error |error| avoids the update direction problem caused by positive and negative errors; (2) The update direction is kept constant by the gradient. In practice, when the system migrates from a hospital corridor to a ward, this update strategy stabilizes the underlying feature extraction parameters, while accelerating the update of high-level semantic understanding parameters based on the degree of environmental change, achieving environmentally adaptive learning.

[0104] The resource efficiency optimization unit 40 implements a selective update strategy and only updates the parameters of the changed areas. The unit sets the change threshold at 10% to 15%. Only areas where the change exceeds the threshold will trigger parameter updates, which greatly reduces the amount of calculation. In addition, the unit also adopts a floating-point precision adaptive mechanism to dynamically adjust the accuracy according to the importance of the parameters, using 32-bit floating points for key parameters and 16-bit floating points for non-key parameters, optimizing memory usage while ensuring accuracy. This resource optimization strategy enables the system to run in real time on portable devices and provide mobile training support for AD patients. For example, when assisting in home training, the system only needs to update the environmental map for changing areas such as furniture movement, while keeping the maps of stable structures such as walls and doors unchanged, significantly improving the update efficiency.

[0105] like Figure 6 As shown, the cognitive state assessment module 6 analyzes the spatial cognitive state of AD patients based on an attention-driven feature fusion strategy, which includes a spatial attention unit 61 , a modality fusion unit 62 , a semantic enhancement processing unit 63 , and a dynamic adaptation unit 64 .

[0106] The spatial attention unit 61 generates a spatial attention map based on the semantic importance of the scene, allocating higher computing resources and representation accuracy to important areas. The spatial attention calculation formula is:

[0107] A(x,y)=softmax(W·F(x,y)),

[0108] Here, A(x, y) is the attention value at position (x, y), indicating its importance; F(x, y) is the feature map for that position, containing multi-channel feature information; W is a learnable weight matrix used to map features to the attention space; and softmax is a normalization function that ensures that the sum of attention values ​​is 1. In training AD patients, this attention mechanism can highlight key navigation points in the environment. For example, when a patient is walking in a community park, the system will increase the attention value of key navigation points such as road signs and forks in the road, guiding the patient's attention to these key areas and improving spatial cognition efficiency.

[0109] The modal fusion unit 62 integrates visual features, motion parameters, and depth information, designs an adaptive weight allocation strategy, and adjusts the weights according to the modal reliability. The modal fusion formula is:

[0110] F fused =α·F visual +β·F spatial +γ·F temporal ,

[0111] Among them, F fused is the fused feature, indicating the result of multimodal information fusion; F visual is the visual feature, the image information from the binocular imaging module; F spatial is a spatial feature, including depth and position information; F temporal α is a temporal feature containing information about motion and change; α, β, and γ are corresponding weights, satisfying α + β + γ = 1. The weight values ​​are dynamically adjusted based on the complexity of the scene. In simple scenes with good lighting, α is set higher (0.5-0.7); in complex scenes with large lighting changes, β and γ are set higher (0.3-0.4 each). This multimodal fusion mechanism can adapt to training needs in different environmental conditions. For example, in indoor environments with insufficient lighting, the system will reduce the weight of visual features and increase the weight of spatial and temporal features to ensure accurate positioning and spatial perception.

[0112] The semantic enhancement processing unit 63 performs feature enhancement on the identified key navigation elements. The enhancement processing includes contrast enhancement (increase by 20% to 30%), edge sharpening (sharpening factor 1.2-1.5) and color saliency adjustment (saturation increased by 15% to 25%). These parameters are specially designed according to the visual perception characteristics of AD patients to improve the visual saliency of key navigation elements. Studies have shown that AD patients respond better to visual stimuli with high contrast and clear edges, so these enhancement processes can effectively improve patients' environmental perception ability. For example, when the system identifies the safety exit sign at the end of the corridor, it will enhance its contrast and edge clarity to help patients more easily identify this key navigation sign.

[0113] The dynamic adaptation unit 64 monitors the response degree of AD patients to different features (through eye movement data analysis) and adjusts the feature fusion weight according to the response strength. For example, when the patient's gaze time on the edge feature is significantly longer than the texture feature (the gaze time ratio exceeds 2:1), the system will increase the weight of the edge feature in the fusion (increase by 20% to 30%). This dynamic adaptation mechanism can optimize the feature representation according to the individual differences of the patient and improve the training effect. In actual application, the system will learn and record the feature preference pattern of each patient, and gradually establish a personalized feature representation strategy to make the training process more in line with the patient's cognitive characteristics.

[0114] like Figure 7 As shown, the cognitive state assessment module 6 includes an eye movement analysis unit 65 , a spatial understanding assessment unit 66 , a navigation ability assessment unit 67 , a depth perception assessment unit 68 and a comprehensive scoring unit 69 .

[0115] The eye movement analysis unit 65 calculates the gaze duration (usually the judgment threshold is 100-200ms), the scanning pattern (the effective scanning speed is 100-500° / s) and the pupil diameter change (the normal fluctuation range is ±15%) based on the eye movement data. These eye movement parameters are important indicators for evaluating cognitive state, reflecting the patient's attention allocation and cognitive load level. For example, when AD patients face a complex intersection, frequent short-term gaze point switching (gaze duration <100ms) and enlarged pupil diameter (enlargement >15%) usually indicate that the patient is in a high cognitive load state and may need more navigation support.

[0116] The spatial understanding assessment unit 66 analyzes spatial understanding ability based on feature gaze preferences. This unit records the frequency and duration of the patient's gaze on different environmental features, forming a gaze heat map. The uniformity of gaze distribution and the effectiveness of gaze switching are key indicators for evaluating spatial understanding ability. For example, patients who can effectively switch their gaze between global environmental features and local navigation cues generally exhibit good spatial understanding ability; while patients whose gaze points are overly concentrated or randomly jump around may have spatial understanding impairments. In actual training, the system will evaluate the patient's gaze patterns in scenes such as community parks or hospital corridors to determine their level of spatial cognition ability.

[0117] The navigation ability assessment unit 67 evaluates navigation ability through landmark recognition accuracy. This unit sets landmark recognition tasks of varying difficulty and records the patient's recognition accuracy and reaction time. The accuracy threshold is typically set at 70%, and the reaction time threshold is dynamically adjusted based on task complexity, ranging from 1 to 2 seconds for simple tasks to 3 to 5 seconds for complex tasks. For example, the system may ask the patient to identify navigation tasks such as the exit sign at the end of the corridor or turn left to the third room, and their navigation ability is assessed based on their completion. This assessment method directly reflects the patient's navigation performance in a real-world environment, providing targeted guidance for training.

[0118] The depth perception assessment unit 68 assesses spatial perception based on depth perception accuracy. This unit assesses the patient's accuracy in perceiving distance and depth by setting up interactive tasks involving virtual objects and the real environment. The error threshold is typically set to ±15% of the actual distance. For example, the system may ask the patient to determine the approximate distance of a door ahead or the distance relationship between two landmarks, and depth perception is assessed based on the accuracy of the answer. Depth perception impairment is a common spatial cognition problem in AD patients, and this assessment unit can promptly identify this problem and provide targeted training.

[0119] The comprehensive scoring unit 69 generates a comprehensive cognitive status score based on the output results of each evaluation unit. The score adopts a weighted average method, and the weight of each dimension is dynamically adjusted according to the training goal. The scoring formula is:

[0120] Score=w1·S eye +w2·S spatial +w3·S navigation +w4·S depth ,

[0121] Among them, Score is a comprehensive score, which indicates the patient's overall spatial cognition status; S eye Score for eye movement analysis, reflecting attention status; S spatial Score for spatial understanding, reflecting the ability to understand spatial layout; S navigation Score navigation ability, reflecting the ability to plan and execute paths; S depth Depth perception is a score reflecting distance judgment ability; w1, w2, w3, and w4 are corresponding weights, satisfying w1 + w2 + w3 + w4 = 1. In practice, for patients with mild AD, the weights of the four dimensions may be relatively balanced (approximately 0.25 each). For patients with moderate AD, the weights of navigation and spatial understanding may be increased (approximately 0.3-0.35 each), while the weights of other dimensions may be reduced accordingly to focus on core skills training.

[0122] like Figure 8As shown, the training guidance module 7 includes a difficulty adjustment unit 71 , a path planning unit 72 , a virtual guidance unit 73 and a feedback generation unit 74 .

[0123] The difficulty control unit 71 determines the training difficulty level based on the spatial cognitive state and defines five levels of training difficulty, from simple recognition to complex navigation. Difficulty level 1 is simple landmark recognition, such as recognizing the exit sign at the end of the corridor; Difficulty level 2 is dual landmark association, such as walking from the nurse station to the elevator entrance; Difficulty level 3 is simple path planning, such as walking from the room to the restaurant; Difficulty level 4 is multi-landmark navigation, such as passing the pharmacy and rest area in sequence to reach the doctor's office; Difficulty level 5 is complex environment navigation, such as finding a designated destination in an unfamiliar community environment. The difficulty adjustment step is 0.1-0.3, which is dynamically adjusted according to the patient's performance. For example, when a patient successfully completes a task of difficulty 2.5 three times in a row, the system will increase the difficulty to 2.7; and when the patient fails twice in a row, the system will reduce the difficulty to 2.3 to keep the training in the optimal challenge range.

[0124] The path planning unit 72 uses the Bayesian minimum mean square error method to optimize the posture and predict the movement trajectory and posture of the AD patient through the path integral method. The optimization formula of the Bayesian minimum mean square error method is:

[0125]

[0126] in, is the final optimal pose estimation value, including position and pose information; is the possible pose estimate, which is a variable in the optimization process; θ is the true pose parameter, which represents the actual position and posture of the patient (it is a random variable); y is the observation data, including image and IMU data; E is the expectation operator, which calculates the conditional expectation value; It means finding the parameter value that minimizes the objective function.

[0127] The mathematical analytical solution of the above formula is: This is the conditional expectation of the true pose θ given the observation data y. This Bayesian estimation method combines historical information with current observations, providing stable pose estimation in noisy environments. In practice, the system uses a particle filter algorithm to approximate this expectation by maintaining multiple pose hypotheses (particles) and updating their weights based on the observation data.

[0128] This optimization method, combined with a spatiotemporal-constrained incremental learning architecture, provides a stable and reliable pose basis for precise projection of virtual guidance, ensuring accurate alignment and smooth presentation of AR content, greatly improving training effectiveness and patient experience.

[0129] The path integral prediction formula is:

[0130]

[0131] Among them, X(t) is the position at time t, which represents the three-dimensional space coordinates; X(0) is the initial position, which represents the coordinates of the starting point; v(τ) is the velocity function, which represents the movement speed at different times; a(s) is the acceleration function, which represents the change in movement acceleration; represents the integral from time 0 to t; Denotes double integral. This formula uses integral calculations to predict future positions, providing a basis for path planning. In practical applications, for example, when an AD patient is training in a nursing home, the system predicts the patient's movement trajectory for the next 5 to 10 seconds based on their current walking speed and direction. This prediction is used to plan guidance information in advance, ensuring timely and smooth guidance.

[0132] The virtual guidance unit 73 generates virtual navigation markers and visual cues, and guides AD patients to match virtual and real landmarks through AR projection. Navigation markers include directional arrows, path indicator lines, and target signs. The visual salience is dynamically adjusted according to the complexity of the environment to ensure clear identification under different lighting conditions. For example, in a well-lit outdoor environment, the transparency of the virtual navigation marker is set to 40% to 60%, and the color contrast is moderate; in an indoor environment with insufficient light, the transparency of the marker is reduced to 20% to 30%, while the color contrast and brightness are increased to ensure the visibility of the marker. This environmentally adaptive virtual guidance method can provide a consistent visual experience in various training scenarios.

[0133] The feedback generation unit 74 provides real-time feedback information based on the behavioral performance of the AD patient, including visual and audio prompts. When the patient is moving correctly, the system will provide positive reinforcement (such as displaying a green check mark and playing a short success prompt sound); when the patient deviates from the path, the system will provide correction prompts (such as displaying a red arrow pointing in the correct direction and playing a slight reminder sound). The feedback intensity is dynamically adjusted according to the patient's cognitive state to avoid feedback that is too strong or too weak. Studies have shown that moderate positive reinforcement can significantly improve the training motivation and concentration of AD patients, while timely correction prompts can reduce training frustration. For example, for mild AD patients, the system may use more obscure visual prompts; for moderate patients, it may use more obvious visual and audio combination prompts to ensure that the patient can understand the feedback information.

[0134] like Figure 9 As shown, the augmented reality display module 4 includes a virtual content generation unit 41 , a three-dimensional registration unit 42 , a rendering optimization unit 43 and a display control unit 44 .

[0135] The virtual content generation unit 41 generates virtual landmarks and training scenes based on scene semantic information. These virtual landmarks include directional indicators, target markers, and path prompts, with their complexity and salience dynamically adjusted based on the training difficulty and the patient's cognitive state. For example, in relatively low-difficulty training, the system might generate continuous path lines to directly guide the patient to the target; whereas in more difficult training, the system might only generate directional indicators at key turning points, requiring the patient to independently plan their path. This progressively difficult virtual content design can effectively support the gradual improvement of patients' spatial cognitive abilities.

[0136] The 3D registration unit 42 spatially aligns virtual landmarks with real-world landmarks. This unit utilizes feature point matching and pose optimization algorithms to ensure precise alignment of virtual content with the real environment, with alignment accuracy controlled within ±1 cm. Accurate 3D registration is crucial for augmented reality training, directly impacting both training effectiveness and patient experience. For example, when the system generates a virtual arrow pointing to a real doorway, if the arrow accurately points to the actual doorway, the patient can clearly establish a virtual-real correspondence. However, inaccurate alignment can lead to patient confusion and training frustration.

[0137] The rendering optimization unit 43 adjusts the rendering parameters according to the ambient lighting conditions and the visual characteristics of AD patients. For example, in a strong light environment (ambient light intensity>1000lux), the system will increase the contrast of the virtual content (increase by 20% to 30%) and saturation (increase by 15% to 25%); in a low light environment (ambient light intensity<100lux), the system will adjust the brightness (increase by 30% to 50%) and color temperature (reduce by 200 to 300K, biased towards warmer tones) to ensure that the virtual content is clearly visible in all environments. This environment-adaptive rendering technology solves the application problem of augmented reality systems in complex lighting environments and expands the scope of application of training scenarios.

[0138] The display control unit 44 projects the optimized virtual content onto the AR glasses' display interface. The display refresh rate is maintained above 60Hz, ensuring a smooth visual experience; the field of view covers an angle of view exceeding 90°, providing an immersive training environment. Display parameters suitable for AD patients are crucial to training effectiveness. Research has shown that a stable high refresh rate and a sufficiently wide field of view can reduce dizziness and discomfort in patients, improving training comfort and compliance.

[0139] like Figure 10 As shown, the data management module 8 includes a training data storage unit 81 , a data analysis unit 82 , a progress tracking unit 83 and a report generation unit 84 .

[0140] Training data storage unit 81 records the AD patient's walking trajectory, posture changes, eye movement data, and task completion status. The data sampling frequency varies depending on the data type: the position data sampling frequency is 10Hz, sufficient to capture natural human movement; the eye movement data sampling frequency is 120Hz, capable of accurately recording rapid eye movements; and task completion status is recorded in real time. This multimodal data provides a scientific basis for evaluating patient training effectiveness and developing personalized training plans. For example, by analyzing changes in a patient's walking trajectory in a hospital corridor environment, it is possible to assess improvements in their spatial navigation ability; and through eye movement data analysis, changes in their environmental perception and attention allocation abilities can be evaluated.

[0141] The data analysis unit 82 compares the deviation between the planned path and the actual path, and analyzes the accuracy and speed of virtual landmark recognition. The path deviation evaluation uses two indicators: average deviation distance and maximum deviation distance, and the landmark recognition evaluation uses two indicators: recognition accuracy and average recognition time. These quantitative evaluation indicators can scientifically measure the training effect and provide a basis for adjusting the training program. For example, if the patient's path deviation in a certain type of environment is continuously large (average deviation distance>2m), the system will increase the training frequency of this type of environment; if the patient's recognition accuracy of a certain type of landmark is significantly lower than the average level (less than 60%), the system will increase the recognition training of this type of landmark.

[0142] The progress tracking unit 83 establishes long-term training records and tracks the changing trends in spatial cognitive abilities. This unit uses a sliding window analysis method, typically selecting the most recent 5 to 10 training sessions for trend analysis and generating a progress curve. Long-term progress tracking is crucial for evaluating training effectiveness and maintaining patient motivation. For example, the system generates a curve showing the average navigation accuracy of the most recent 10 training sessions, visually demonstrating the patient's progress. It also identifies training bottlenecks, such as slow progress in corner recognition, and provides targeted guidance for subsequent training.

[0143] The report generation unit 84 generates personalized training reports and progress indicators to support medical staff in evaluating treatment effectiveness. The training report includes quantitative indicators such as the number of training sessions, total duration, number of completed tasks, and accuracy trend, as well as qualitative assessments and suggestions. These scientific and data-based training reports provide medical staff with comprehensive patient assessment information, making up for the shortcomings of traditional subjective assessments. For example, the spatial memory retention time indicator in the report can objectively reflect changes in the patient's spatial memory ability, and the environmental adaptation speed indicator can reflect the patient's ability to adapt to the new environment. These indicators provide an accurate reference basis for clinical intervention.

[0144] like Figure 11 As shown, the present invention also provides a dynamic environment three-dimensional reconstruction method for AD patients' spatial cognition training, which includes the following steps:

[0145] Step 1: Obtain scene image data of the training environment and motion parameters of the AD patient. The binocular imaging module captures environmental images at a frequency of 30 frames per second; the inertial measurement unit collects patient motion parameters at a sampling rate of 200 Hz. Both data types are synchronized using timestamps to ensure data consistency. For example, when a patient is training in a public area of ​​a nursing home, the system simultaneously captures stereo images of the environment and motion data such as the patient's walking speed and steering angle, providing basic data for subsequent processing.

[0146] Step 2: Extract feature points from the scene image data using a cognitive load-adaptive feature sampling mechanism. First, the field of view is divided into the foveal, parafoveal, and peripheral regions, with different sampling densities set for each. Next, key navigation elements are identified and assigned higher sampling weights. Finally, the sampling complexity is dynamically adjusted based on the patient's cognitive load. This adaptive sampling mechanism ensures sufficient feature extraction in key regions while optimizing computational resource allocation, adapting to the cognitive characteristics of AD patients.

[0147] Step 3: Generate scene semantic information through multi-scale fusion of semantically aware feature descriptions. Feature descriptions are extracted at small, medium, and large scales to capture environmental features at different levels. Semantic label information is integrated into the feature description process to enhance feature differentiation. A hierarchical matching strategy is used for feature point matching to improve matching accuracy. For example, in a community park environment, the system simultaneously extracts multi-scale features such as pavement texture, seat shape, and overall path layout, assigning different weights based on semantic importance to generate rich scene semantic information.

[0148] Step 4: Dynamically update the map using a spatiotemporal-constrained incremental learning architecture. This architecture limits the rate of feature change between adjacent frames, establishes a temporal sliding window, and maintains historical coherence. It also maintains a topological relationship graph between feature points to ensure consistency in the spatial structure. A hierarchical update strategy is employed for the network parameter Kp, balancing stability and adaptability. Finally, a selective update strategy is implemented, updating only parameters in changed areas to optimize computational efficiency. This incremental learning architecture enables the system to adapt to environmental changes, such as furniture movement or lighting changes, while maintaining the continuity and stability of the environmental representation.

[0149] Step 5: Generate virtual landmarks and training scenes based on scene semantics and project them onto the AR glasses' display interface. Virtual landmarks include directional indicators, target markers, and path prompts. 3D registration technology ensures precise alignment of virtual content with the real environment. Rendering parameters are optimized based on ambient lighting conditions and the patient's visual characteristics. The optimized virtual content is projected onto the AR glasses' display interface. For example, in a hospital corridor, the system generates virtual arrows indicating direction at key turning points and virtual markers at target locations to assist patients in completing navigation tasks.

[0150] Step 6: Collect eye movement data from the AD patient. Using the eye tracking module built into the AR glasses, eye movement data, including gaze point, gaze duration, saccade trajectory, and pupil diameter, is collected at a sampling rate of 120Hz. This eye movement data is crucial for assessing the patient's cognitive status and can reflect information such as attention allocation, cognitive load, and visual preferences.

[0151] Step 7: Analyze the spatial cognitive status of AD patients based on an attention-driven feature fusion strategy. Generate a spatial attention map based on the semantic importance of the scene; integrate visual features, motion parameters, and depth information, and adopt an adaptive weight allocation strategy; enhance the features of key navigation elements to increase visual salience; monitor the patient's response to different features and dynamically adjust the feature fusion weights; and generate a comprehensive cognitive status score based on eye movement data, spatial understanding ability, navigation ability, and depth perception ability. For example, the system will analyze the patient's gaze pattern, landmark recognition ability, and spatial understanding ability during navigation to form a comprehensive cognitive status assessment, providing a basis for adjusting training difficulty.

[0152] Step 8: Generate adaptive training guidance information based on spatial cognitive status and display it through the AR glasses display interface. Determine the training difficulty level based on the cognitive status score; use the Bayesian minimum mean square error method to optimize the posture and predict the patient's movement trajectory; generate virtual navigation markers and visual cues to guide the patient through the training task; and provide real-time feedback information based on the patient's behavioral performance, including visual and audio cues. For example, when the system detects that the patient is exhibiting a high cognitive load in a complex environment, it will reduce the training difficulty and increase the frequency and visibility of guidance cues to help the patient gradually build spatial cognitive ability.

[0153] Step 9: Store training process data and patient behavior data and generate a training evaluation report. Record the patient's walking trajectory, posture changes, eye movement data, and task completion status; analyze deviations between the planned and actual paths to evaluate training effectiveness; establish long-term training records to track trends in spatial cognitive abilities; and generate personalized training reports and progress indicators to support medical staff in evaluating treatment effectiveness. These data and reports provide a scientific basis for the patient's long-term recovery and inform medical staff's intervention decisions.

[0154] Through the above steps, the present invention achieves dynamic environmental 3D reconstruction and spatial cognition training for AD patients. It can dynamically adjust the training content and difficulty based on the patient's real-time cognitive state, improving the targetedness and effectiveness of the training. Furthermore, the system seamlessly integrates virtual guidance with the real environment through augmented reality technology, overcoming the limitations of traditional training methods and providing a new technical approach for spatial cognitive rehabilitation in AD patients.

[0155] The foregoing is merely a preferred embodiment of the present invention and is not intended to limit the present invention. It will be apparent to those skilled in the art that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.

Claims

1. A dynamic environment 3D reconstruction system for spatial cognition training of AD patients, characterized by: include: Binocular imaging module, used to shoot the training environment in real time and obtain scene image data; Inertial measurement unit, used to collect motion parameters of AD patients during training; an incremental neural point cloud encoding module, electrically connected to the binocular imaging module, configured to receive the scene image data, extract feature points based on a cognitive load adaptive feature sampling mechanism, generate scene semantic information through multi-scale fused semantic perception feature description, and implement dynamic map updates using a spatiotemporal constraint incremental learning architecture; an augmented reality display module, communicatively connected to the incremental neural point cloud encoding module, configured to receive the scene semantic information, generate virtual landmarks and training scenes, and project the virtual landmarks and training scenes onto an AR glasses display interface; An eye tracking module, installed in the augmented reality display module, for collecting eye movement data of AD patients; a cognitive state assessment module, electrically connected to the eye tracking module and the inertial measurement unit, configured to receive the eye movement data and the motion parameters and analyze the spatial cognitive state of the AD patient based on an attention-driven feature fusion strategy; a training guidance module, communicatively connected to the cognitive state assessment module and the augmented reality display module, configured to receive the spatial cognitive state, generate adaptive training guidance information, and display the information through the augmented reality display module; The data management module is connected to the incremental neural point cloud encoding module, the cognitive state assessment module and the training guidance module respectively, and is used to store training process data and patient behavior data and generate a training assessment report.

2. The system according to claim 1, wherein: The cognitive load adaptive feature sampling mechanism includes: A field of view region division unit is used to divide the field of view into the fovea centralis, the parafovea centralis, and the peripheral region, and to set different feature point sampling densities according to the importance of the regions; A semantic importance determination unit, configured to identify key navigation elements in a scene and assign a higher sampling weight to the key navigation elements; A cognitive state response unit is used to estimate the current cognitive load level based on the eye movement data, reduce sampling complexity in a high cognitive load state, and increase sampling complexity in a low cognitive load state.

3. The system according to claim 1, wherein: The multi-scale fusion semantic perception feature description includes: A multi-scale representation unit for extracting feature descriptions at three different scales: a small scale for capturing microscopic features, a medium scale for capturing object and regional features, and a large scale for capturing spatial layout relationships; Semantic enhancement unit, which is used to integrate semantic label information into the feature description process and apply different description parameters to different semantic categories; The feature matching unit is used to match feature points using a hierarchical matching strategy, calculate similarity using normalized cross-correlation, and filter false matches based on geometric consistency constraints.

4. The system according to claim 1, wherein: The spatiotemporal constrained incremental learning architecture includes: Temporal continuity guarantee unit, used to limit the feature change rate between adjacent frames, establish a temporal sliding window, and maintain historical coherence; The spatial structure preserving unit is used to maintain the topological relationship graph between feature points, calculate the deformation degree of the topological structure, and perform feature association based on topological invariance constraints; The incremental parameter update unit is used to adopt a hierarchical update strategy for the network parameter Kp and dynamically adjust the learning rate according to the degree of scene change; The resource efficiency optimization unit is used to implement a selective update strategy, updating only the parameters in the changed area, and adopting a floating-point precision adaptive mechanism to balance accuracy and computational complexity.

5. The system according to claim 1, wherein: The attention-driven feature fusion strategy includes: The spatial attention unit is used to generate a spatial attention map based on the semantic importance of the scene, allocating higher computing resources and representation accuracy to important areas; a modality fusion unit for integrating visual features, the motion parameters, and depth information, designing an adaptive weight allocation strategy, and adjusting weights according to modality reliability; Semantic enhancement processing unit, used to enhance the features of identified key navigation elements. Enhancement processing includes contrast improvement, edge sharpening, and color saliency adjustment. The dynamic adaptation unit is used to monitor the response degree of AD patients to different features and adjust the feature fusion weight according to the response strength.

6. The system according to claim 1, wherein: The cognitive status assessment module includes: an eye movement analysis unit, configured to calculate fixation duration, saccade pattern, and pupil diameter change based on the eye movement data; Spatial Comprehension Assessment Unit, used to analyze spatial comprehension abilities based on feature gaze preferences; Navigational ability assessment unit, used to evaluate navigational ability through landmark recognition accuracy; a depth perception evaluation unit for evaluating spatial perception capabilities based on depth perception accuracy; A comprehensive scoring unit is used to generate a comprehensive cognitive status score based on the output results of the eye movement analysis unit, the spatial understanding evaluation unit, the navigation ability evaluation unit and the depth perception evaluation unit.

7. The system according to claim 1, wherein: The training guidance module includes: A difficulty control unit is used to determine the training difficulty level based on the spatial cognition state, and define five levels of training difficulty, from simple recognition to complex navigation; The path planning unit is used to optimize the posture using the Bayesian minimum mean square error method, predict the movement trajectory and posture of AD patients through the path integral method, and construct the walking path based on the updated map; A virtual guidance unit, which generates virtual navigation markers and visual cues to guide AD patients to match virtual and real landmarks through AR projection; The feedback generation unit is used to provide real-time feedback information based on the behavioral performance of AD patients, including visual and audio prompts.

8. The system according to claim 1, wherein: The augmented reality display module includes: a virtual content generation unit, configured to generate virtual landmarks and training scenes based on the scene semantic information; a three-dimensional registration unit, configured to achieve spatial alignment between the virtual landmark and the real landmark; A rendering optimization unit, used to adjust rendering parameters according to ambient lighting conditions and the visual characteristics of AD patients; The display control unit is used to project the optimized virtual content onto the AR glasses display interface.

9. The system according to claim 1, wherein: The data management module includes: Training data storage unit, used to record the walking trajectory, posture changes, eye movement data and task completion status of AD patients; A data analysis unit, used to compare the deviation between the planned path and the actual path and analyze the accuracy and speed of virtual landmark recognition; Progress tracking unit, used to establish long-term training records and track changes in spatial cognitive abilities; Report generation unit, used to generate personalized training reports and progress indicators to support medical staff in evaluating treatment effects.

10. A dynamic environment three-dimensional reconstruction method for spatial cognition training of AD patients, using the system according to any one of claims 1 to 9, characterized in that: The following steps are involved: Obtaining scene image data of the training environment and motion parameters of AD patients; Extract feature points from scene image data based on cognitive load adaptive feature sampling mechanism; Generate scene semantic information through multi-scale fused semantic-aware feature description; Adopting a spatiotemporal-constrained incremental learning architecture to achieve dynamic map updates; Generate virtual landmarks and training scenes based on scene semantic information, and project the virtual landmarks and training scenes onto the AR glasses display interface; Collect eye movement data from AD patients; Analyze the spatial cognitive status of AD patients based on attention-driven feature fusion strategy; Generate adaptive training guidance information based on spatial cognitive status and display it through the AR glasses display interface; Store training process data and patient behavior data, and generate training evaluation reports.