Myopia prevention method based on neural network model

By acquiring real-time visual behavior data in immersive experience devices and using neural network models to evaluate the state of visual accommodation compensation, and combining visual behavior entropy and content semantic focus change rate, display parameters are dynamically adjusted, solving the problem that existing technologies cannot perceive individualized visual physiological states in real time, and achieving personalized myopia prevention effects.

CN121662286APending Publication Date: 2026-03-13SHANGHAI YUANHE VISION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-05
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Existing technologies cannot perceive and respond to users' individual dynamic visual physiological states in real time in immersive experience devices, resulting in a lack of personalization and real-time nature in myopia risk warning and prevention measures.

Method used

By acquiring real-time visual behavior data of users in immersive experience devices, a neural network model is used to evaluate the state of visual accommodation compensation. By combining visual behavior entropy and content semantic focus change rate, the dominant causes of abnormal visual accommodation compensation are analyzed, and display parameters are dynamically adjusted to prevent myopia.

Benefits of technology

It enables personalized, real-time monitoring and prevention of users' visual health in an immersive experience, ensuring that display parameters are adjusted closely to the root causes of visual load, thus improving the accuracy and effectiveness of myopia prevention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121662286A_ABST
    Figure CN121662286A_ABST
Patent Text Reader

Abstract

The invention discloses a myopia prevention method based on a neural network model, particularly relates to the technical field of visual health monitoring, and is used for solving the problems that the real-time dynamic visual physiological load of a user cannot be perceived and responded due to dependence on a static rule and personalized myopia risk early warning and prevention are difficult to realize in the prior art. Real-time visual behavior data of a user in immersive experience equipment is acquired, a visual adjustment compensation state of the user is evaluated based on a neural network model, and when the state exceeds a preset threshold value, a visual behavior entropy is further calculated and a content semantic focus change rate of a virtual scene is synchronously analyzed so as to analyze a dominant inducement type of state abnormality; matching the fixation point distribution with the scene visual saliency topological graph according to the inducement type to identify specific dominant factors, and dynamically adjusting display parameters of the immersive experience equipment according to the specific dominant factors; accurate evaluation and inducement analysis of the real-time load of the user visual system are achieved, and the myopia prevention effect is achieved while the immersion experience is maintained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of visual health monitoring technology, and more specifically, to a method for myopia prevention based on a neural network model. Background Technology

[0002] The visual health of teenagers using immersive experience devices has become an increasingly important concern. To prevent the onset or worsening of myopia, existing technologies typically manage eye strain by monitoring device usage time, setting fixed rest reminders, or pre-setting uniform display brightness and contrast parameters based on the user's age. In addition, some methods attempt to analyze basic user behavioral data, such as average viewing distance or simple blinking frequency, and provide standardized eye care recommendations based on general medical research findings. These practices constitute the main technical means for health risk intervention in immersive environments.

[0003] The drawback of the aforementioned existing technical solutions lies in the fact that the intervention rules and parameters they rely on are essentially static and universal, failing to perceive and respond to the highly personalized dynamic visual physiological state generated by users in real time during immersive experiences. Specifically, because different users exhibit significant individual differences in the real-time adjustment and compensatory responses of their visual systems when facing the same virtual visual environment, existing methods based on fixed rules or simple statistics cannot capture and assess inherent, non-obvious differences in physiological load, making it difficult to achieve truly personalized myopia risk warning and preventive control that matches the user's real-time physiological state in immersive experience scenarios. Summary of the Invention

[0004] In order to overcome the above-mentioned defects of the prior art, the present invention provides a myopia prevention method based on a neural network model to solve the problems mentioned in the background art.

[0005] To achieve the above objectives, the present invention provides the following technical solution: Myopia prevention methods based on neural network models include: S1. Acquire real-time visual behavior data of users in immersive experience devices; S2. Based on real-time visual behavior data, the user's visual accommodation compensation status is evaluated through a neural network model. S3. When the visual adjustment compensation state exceeds the preset threshold, calculate the visual behavior entropy corresponding to the real-time visual behavior data, and simultaneously analyze the content semantic focus change rate of the virtual scene currently displayed by the immersive experience device. S4. Based on the coupling relationship between visual behavioral entropy and the rate of change of content semantic focus, analyze the dominant cause type of abnormal visual accommodation compensation state. S5. Match the gaze point distribution in real-time visual behavior data with the visual saliency topology map of the virtual scene according to the dominant cause type, and analyze the composition characteristics of the topology mapping deviation to identify at least one specific dominant factor that causes the visual accommodation compensation state to exceed the preset threshold. S6. Based on specific dominant factors, dynamically adjust the display parameters of immersive experience devices to prevent myopia.

[0006] Furthermore, real-time visual behavior data of users in immersive experience devices is acquired, including: The eye-tracking unit integrated into the immersive experience device acquires the user's eye movement trajectory data and pupil diameter change data. The user's head displacement and posture change data are acquired through the inertial measurement unit integrated into the immersive experience device; Eye movement data, pupil diameter change data, head displacement data, and posture change data are synchronized and time-stamped to generate real-time visual behavior data.

[0007] Furthermore, based on real-time visual behavior data, a neural network model is used to assess the user's visual accommodation compensation status, including: Eye movement data, pupil diameter change data, head displacement data, and posture change data from real-time visual behavior data are input into a trained neural network model. The neural network model analyzes saccade and fixation patterns in eye movement trajectory data, performs correlation analysis between pupillary light reflection and accommodation in pupil diameter change data, and calculates vestibular-visual interaction compensation in head displacement and posture change data. The neural network model performs multi-source feature fusion based on the results of saccade and fixation pattern analysis, the results of pupillary light reflection and accommodation correlation analysis, and the results of vestibular-visual interaction compensation calculation, and outputs a visual accommodation compensation state assessment value that characterizes the real-time load of the user's visual system.

[0008] Furthermore, the trained neural network model is obtained by acquiring historical visual behavior data and synchronously recorded refraction parameter change data generated by multiple groups of users in the pre-experimental immersive experience; using the historical visual behavior data as input and the visual fatigue level label corresponding to the refraction parameter change data as a supervision signal, the neural network model is trained until it can output a visual accommodation compensation state evaluation value consistent with the visual fatigue level label based on the input data.

[0009] Furthermore, when the visual accommodation compensation state exceeds a preset threshold, the visual behavior entropy corresponding to the real-time visual behavior data is calculated, and the semantic focus change rate of the content of the virtual scene currently displayed on the immersive experience device is analyzed simultaneously, including: When the visual accommodation compensation state assessment value exceeds the preset threshold, the visual behavior entropy, which represents the randomness and disorder of the user's visual exploration behavior, is calculated by the information entropy algorithm based on the eye movement trajectory data in the real-time visual behavior data. The virtual scene depth map and object identification information of the current frame and historical frames are synchronously output based on the rendering pipeline of the immersive experience device. By analyzing the displacement rate of the semantic center of gravity of the scene in three-dimensional space between consecutive frames, the content semantic focus change rate is obtained.

[0010] Furthermore, the visual behavior entropy is calculated using the information entropy algorithm in the following way: the eye movement trajectory data in the real-time visual behavior data is discretized, the gaze regions are divided, and the number of gaze points in each region is counted; the probability distribution of gaze in each gaze region is calculated based on the number of gaze points; and the visual behavior entropy value, which characterizes the randomness and disorder of the user's visual exploration behavior, is calculated using the information entropy formula according to the probability distribution.

[0011] Furthermore, the semantic focus change rate of the content is analyzed by analyzing the displacement rate of the semantic center of gravity of the scene between consecutive frames in the following way: extract the three-dimensional coordinates of the semantic center of gravity of each frame from the continuous multi-frame virtual scene output by the rendering pipeline of the immersive experience device; calculate the Euclidean distance between the three-dimensional coordinates of the semantic center of gravity of the current frame and the previous frame; and calculate the semantic focus change rate of the content based on the Euclidean distance and the rendering frame interval of the continuous multi-frame virtual scene.

[0012] Furthermore, based on the coupling relationship between visual behavioral entropy and the rate of change of content semantic focus, the dominant triggering factors for abnormal visual accommodation compensatory states are analyzed, including: The visual behavior entropy is compared with a preset entropy threshold, and the content semantic focus change rate is compared with a preset change rate threshold. When the visual behavior entropy is higher than the entropy threshold and the change rate of the semantic focus of the content is also higher than the change rate threshold, the dominant cause type is determined to be exogenous content-driven. When the visual behavioral entropy is higher than the entropy threshold and the change rate of the semantic focus of the content is lower than or equal to the change rate threshold, the dominant cause type is determined to be an abnormal endogenous physiological state.

[0013] Furthermore, based on the dominant trigger type, the gaze point distribution in real-time visual behavior data is matched with the visual saliency topology map of the virtual scene. The constituent features of the topology mapping deviation are analyzed to identify at least one specific dominant factor causing the visual accommodation compensation state to exceed a preset threshold, including: A visual saliency topology map is generated based on the depth and color contrast information of the current virtual scene. When the dominant cause type is exogenous content-driven, the matching degree between the gaze point distribution and the highly salience region in the visual salience topology map is analyzed, and the scene content elements corresponding to the highly salience region with a matching degree lower than the preset matching degree threshold are identified as specific dominant factors. When the dominant cause is an abnormal endogenous physiological state, the dispersion and drift frequency of the fixation point distribution relative to any region in the visual salience topology are analyzed, and fixation behavior patterns with a dispersion higher than a preset dispersion threshold or a drift frequency higher than a preset frequency threshold are identified as specific dominant factors.

[0014] Furthermore, based on specific dominant factors, the display parameters of immersive experience devices are dynamically adjusted to prevent myopia, including: When the specific dominant factor is the scene content element driven by external content, reduce the display brightness and contrast of the visual area related to the corresponding scene content element in the immersive experience device, and increase the color saturation of the area adjacent to the corresponding scene content element in the immersive experience device. When the specific dominant factor is the gaze behavior pattern corresponding to an abnormal endogenous physiological state, reduce the global display brightness and global color contrast of the immersive experience device, and activate the depth-of-field blur rendering effect of the immersive experience device to increase the out-of-focus area of ​​the virtual scene.

[0015] Compared with the prior art, the present invention has the following beneficial effects: 1. By acquiring multi-dimensional real-time visual behavior data and combining it with a neural network model to evaluate the visual accommodation compensation state, the system can capture and quantify the non-obvious physiological load changes generated by the user's visual system under the stimulation of specific virtual content in real time. This allows the system to go beyond simple behavioral statistics and delve into the dynamic perception of the user's real-time compensatory level of visual physiological function. This provides a truly personalized judgment benchmark based on real-time physiological state for subsequent precise intervention, thereby solving the problem of delayed or inaccurate early warning and regulation caused by the inability of existing technologies to perceive individualized dynamic physiological responses.

[0016] 2. By introducing coupled analysis of visual behavior entropy and content semantic focus change rate, the system can intelligently distinguish whether abnormal visual load is driven by external virtual content or dominated by fluctuations in the user's own physiological state. This makes the subsequent adjustment strategy of the system highly targeted. Whether it is to locate specific overload scene elements by matching gaze points and visual salience topology maps, or to identify endogenous disorder characteristics by analyzing gaze behavior patterns, it ultimately leads to differentiated and dynamic adjustments of the display parameters of immersive devices that are precisely corresponding to the causes. This ensures that every parameter adjustment is closely aligned with the specific root cause of the user's current visual load. Thus, while maintaining the continuity of the immersive experience, it achieves a proactive, adaptive, and highly personalized myopia prevention effect, effectively promoting the user's visual health in the virtual environment. Attached Figure Description

[0017] Figure 1 This is a flowchart of the myopia prevention method based on a neural network model according to the present invention. Detailed Implementation

[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] Example: Figure 1 This invention presents a myopia prevention method based on a neural network model, comprising: S1. Acquire real-time visual behavior data of users in immersive experience devices; S2. Based on real-time visual behavior data, the user's visual accommodation compensation status is evaluated through a neural network model. S3. When the visual adjustment compensation state exceeds the preset threshold, calculate the visual behavior entropy corresponding to the real-time visual behavior data, and simultaneously analyze the content semantic focus change rate of the virtual scene currently displayed by the immersive experience device. S4. Based on the coupling relationship between visual behavioral entropy and the rate of change of content semantic focus, analyze the dominant cause type of abnormal visual accommodation compensation state. S5. Match the gaze point distribution in real-time visual behavior data with the visual saliency topology map of the virtual scene according to the dominant cause type, and analyze the composition characteristics of the topology mapping deviation to identify at least one specific dominant factor that causes the visual accommodation compensation state to exceed the preset threshold. S6. Based on specific dominant factors, dynamically adjust the display parameters of immersive experience devices to prevent myopia.

[0020] S1. Obtain real-time visual behavior data of the user in the immersive experience device, specifically implemented as follows: First, the user's eye movement trajectory data and pupil diameter change data are acquired through an eye-tracking unit integrated into the immersive experience device. This eye-tracking unit utilizes hardware components based on near-infrared pupil-corneal reflection. Specifically, an array of near-infrared light-emitting diodes (LEDs) with a wavelength of, for example, 850 nanometers, is used to uniformly illuminate the user's eyeball, while a high-speed image sensor simultaneously acquires the reflected images from the eye's surface. Image processing is performed on each frame of the reflected image, including but not limited to grayscale conversion, binarization, and contour detection, to identify the elliptical contour of the pupil and the reflected light spot on the cornea generated by a fixed infrared light source. By calculating the relative positional offset between the center of the pupil contour and the center of the corneal reflected light spot in the image coordinate system, and based on a pre-established mapping relationship model between the offset and the gaze direction through individual calibration, the coordinates of the user's gaze point on the two-dimensional display plane of the virtual scene are calculated. This calculation process is continuously performed at a preset sampling frequency of, for example, 90 Hz, and the output sequence of continuous gaze point coordinates constitutes the eye movement trajectory data. The individual calibration mentioned here refers to guiding the user to gaze at multiple known coordinate calibration points that appear sequentially on the screen when they first use the device, while simultaneously recording the corresponding pupil and corneal reflection vectors to establish a user-specific mapping model. Simultaneously, by calculating the pixel region enclosed by the detected pupil contour in the image, the equivalent diameter of the pupil in each frame is obtained. Arranging the equivalent pupil diameters of multiple consecutive frames in chronological order yields the pupil diameter variation data. The preset sampling frequency is set based on the principle of accurately capturing rapid eye saccades without distortion of the movement trajectory due to insufficient sampling. Experiments have verified that sampling frequencies higher than 60 Hz can meet the analysis requirements for the continuity of eye movement trajectories in this scenario; therefore, 90 Hz, for example, is chosen as the implementation value.

[0021] Secondly, the user's head displacement and posture change data are acquired through the inertial measurement unit integrated into the immersive experience device. This inertial measurement unit includes a three-axis accelerometer and a three-axis gyroscope with a microelectromechanical system (MEMS) architecture. Specifically, the accelerometer continuously measures the linear acceleration components in the X, Y, and Z axes of the device's coordinate system at a sampling frequency of, for example, 200 Hz, denoted as Ax, Ay, and Az. To obtain the head displacement data, the linear acceleration after removing the gravitational acceleration component needs to be integrated twice. First, based on the device's initial stationary or known posture, the components of gravitational acceleration in each axis under the current posture are subtracted from the original acceleration measurement value to obtain the pure linear acceleration caused by the user's head movement. Then, within a short time window, this pure linear acceleration signal is integrated for the first time to obtain the change in linear velocity. A second time integration is then performed on the linear velocity, and combined with the initial position (zero point), the displacement vector sequence of the head relative to the initial point in three-dimensional space is finally calculated. This sequence is the head displacement data. Simultaneously, the gyroscope continuously measures the angular velocity components of the device's rotation around the X, Y, and Z axes at the same sampling frequency, for example, 200 Hz, denoted as ωx, ωy, and ωz. By performing a time integration on the angular velocity signal and combining it with the initial attitude angle (zero value), the Euler angle changes of the device relative to the initial attitude at each moment can be calculated, i.e., the sequence of pitch, yaw, and roll angle changes. This sequence is the attitude change data. The 200 Hz sampling frequency is set to accurately capture rapid rotations and jitters that the head may produce, avoiding signal aliasing caused by excessively low sampling frequencies; this frequency value is more than twice that of the main frequency components of head movement.

[0022] Finally, the eye-tracking data, pupil diameter change data, head displacement data, and posture change data are timestamped and aligned to generate real-time visual behavior data. Since the eye-tracking unit and the inertial measurement unit are independent hardware with different internal clock sources, the directly acquired data streams exhibit time deviations and drift. To achieve precise synchronization, a method combining hardware triggering and software timestamp interpolation is employed. Specifically, during system initialization, a hardware synchronization pulse signal is simultaneously sent to both units. Subsequently, each unit, when generating each data sample, not only timestamps itself using its own clock but also records the number of clock cycles since receiving the synchronization pulse. At the data processing end, a unified high-precision software time reference is maintained. When the data processing end receives any data sample, it first corrects the sample's local timestamp to an absolute timestamp referenced by the unified software time reference, based on the number of cycles since the synchronization pulse recorded by the unit to which the sample belongs and the pre-measured clock frequency of that unit. Data generated by the eye-tracking unit at 90 Hz and data generated by the inertial measurement unit at 200 Hz are not strictly aligned in time. When generating real-time visual behavior data, a fixed fusion time interval, such as 10 milliseconds, is set. Using the start time of each fusion time interval as an alignment point, for eye movement trajectory data and pupil diameter change data, the data sample with the smallest time difference between its absolute timestamp and the alignment point is selected as the representative value for that interval. If no new samples are received within that interval, the representative value from the previous interval is used. The same selection logic is applied to head displacement data and posture change data. Subsequently, the final selected eye movement trajectory coordinate values, pupil diameter change data values, head displacement data vectors, and posture change data values ​​within each 10-millisecond fusion interval are combined into a structured data packet. Combining all the structured data packets arranged in chronological order constitutes a multi-channel synchronous time series, which serves as the real-time visual behavior data for all subsequent analysis steps. The 10-millisecond fusion time interval is a balance between data timeliness and processing load; it is shorter than the duration of a typical human fixation, thus ensuring the sensitivity of subsequent analysis to dynamic changes in visual behavior.

[0023] S2. Based on real-time visual behavior data, the user's visual accommodation compensation state is evaluated through a neural network model, specifically as follows: First, eye-tracking data, pupil diameter change data, head displacement data, and posture change data from real-time visual behavior data are input into a trained neural network model. This input process requires normalization preprocessing of different dimensions of data in the real-time visual behavior data to ensure consistent data scale for easy model processing. For eye-tracking data, the coordinate values ​​are normalized to the range of 0 to 1 according to the resolution of the virtual screen, specifically by dividing the original pixel coordinates by the screen pixel width and height. For pupil diameter change data, the values ​​are standardized based on the resting pupil diameter measured by the individual during the pre-experiment calibration phase, converted into a percentage change relative to the resting diameter, calculated by subtracting the resting pupil diameter from the real-time pupil diameter and then dividing by the resting pupil diameter. For head displacement data, the linear displacement values ​​are normalized according to the physical activity space dimensions allowed by the immersive experience device, for example, by dividing the measured displacement by the length of the spatial diagonal. For posture change data, the Euler angle values ​​are normalized to the corresponding numerical range of -180 degrees to +180 degrees, specifically by dividing the original angle value by 180. After normalization, the four types of data are organized into data samples with a fixed time window length according to a strictly corresponding time series order. Each sample contains a data sequence with a continuous duration of 500 milliseconds, which serves as one input to the neural network model.

[0024] Next, the neural network model performs parallel analysis of the input data. The model analyzes saccade and fixation patterns in the eye-tracking data. Specifically, the model's internal temporal convolutional layer and long short-term memory layer collaborate to extract motion features from the eye-tracking coordinate sequence, including temporal variations in instantaneous velocity and acceleration. The extracted instantaneous angular velocity features are compared to a preset velocity threshold used to distinguish between saccades and fixations. When the instantaneous angular velocity exceeds the threshold, for example, 100 degrees / second, the corresponding motion segment is identified as a saccade event; when the instantaneous angular velocity is less than or equal to 100 degrees / second, the corresponding motion segment is identified as a fixation event. The model then calculates parameters such as the duration of the fixation event, the amplitude of the saccade event, and the peak velocity. These parameters collectively constitute the result of the saccade and fixation pattern analysis. The 100-degree / second velocity threshold is a typical value set based on the common range of minimum saccade velocity definitions found in numerous eye-tracking research papers. Simultaneously, the model performs a correlation analysis between pupillary light reflection and accommodation on pupil diameter change data. Specifically, the model analyzes the pupil diameter curve over time and combines it with the brightness information flow of the current virtual scene. The model uses a learnable filter to separate the rapid contraction and expansion components caused by sudden changes in overall scene brightness (i.e., light reflection), and the slow changes caused by switching visual focus between near and far virtual objects (i.e., accommodation). It then quantifies the fluctuation amplitude and response delay time of these two components, and these quantified indicators constitute the results of the pupillary light reflection and accommodation correlation analysis. Furthermore, the model performs vestibular-visual interaction compensation calculations on head displacement and posture change data. Specifically, based on the angular velocity and linear acceleration data of head movement, and using a built-in vestibular-ocular reflex mathematical model, the model predicts the compensatory motion signal that the eye should theoretically produce for a stable visual field. Then, the model compares the predicted eye movement compensation signal with the actual input eye movement trajectory data, calculates the matching error between the two, i.e., the root mean square error, and uses this error value and the characteristics of the predicted signal as the result of the vestibular-visual interaction compensation solution.

[0025] Subsequently, the neural network model performs multi-source feature fusion based on the results obtained from the above three analytical analyses. The fusion process takes place in a fully connected layer, which connects all feature parameters—such as fixation duration, saccade amplitude, pupillary light reflection and accommodation correlation analysis, and vestibular-visual interaction compensation calculation—into a high-dimensional feature vector. This fully connected layer assigns a trainable weight to each feature in the vector. These weights are not pre-set but are automatically learned numerical parameters based on the loss function through backpropagation during model training. Their role is to adjust the importance contribution of different features to the final evaluation result. Information integration is achieved through weighted summation and a nonlinear activation function. This integration process essentially learns an implicit expression from multi-angle physiological behavioral features that comprehensively represents the overall tension of the visual system. Finally, the model maps this implicit expression to a scalar value through an output layer—a neuron with a linear activation function. This value is the evaluation value of the visual accommodation compensation state. The evaluation value is designed as a continuous quantity. The higher the value, the heavier the real-time load on the user's visual system and the further the visual accommodation compensation state deviates from the relaxation benchmark.

[0026] The trained neural network model described above was obtained through the following method. First, historical visual behavior data and synchronously recorded refraction parameter changes from multiple groups of users during a pre-experimental immersive experience were acquired. The pre-experimental immersive experience refers to a controlled environment where a representative group of users is invited to wear immersive experience devices and experience a 30-minute program containing various visual load scenarios. Throughout the experience, historical visual behavior data of the users was synchronously collected and generated using the methods described in the steps. Simultaneously, before the experience began, at fixed time intervals of 5 minutes during the experience, and immediately after the experience, key refraction parameters such as accommodative hysteresis and accommodative micro-fluctuations were measured using refraction equipment. The refraction parameter values ​​measured at different time points during the experience for the same user were compared with the baseline values ​​before the experience to calculate the refraction parameter changes. Next, ophthalmologists, using a pre-defined set of criteria, determined significant visual fatigue when the increase in accommodative hysteresis exceeded 0.5 diopters and the increase in accommodative micro-fluctuations exceeded 0.1 diopters. Each historical visual behavior data sample was then assigned a visual fatigue level label, for example, 0 for no fatigue, 1 for mild fatigue, and 2 for significant fatigue. The 0.5 and 0.1 diopters were set based on clinically recognized thresholds for visual fatigue to induce significant changes in accommodative function. Then, the historical visual behavior data was used as input features, and its corresponding visual fatigue level labels were used as supervisory signals to train the neural network model. The training process specifically included using the backpropagation algorithm to calculate the difference between the model's predicted visual accommodative compensation state assessment value and the actual visual fatigue level label—the loss function value. An optimizer iteratively adjusted all trainable weights within the model to minimize this loss function value. Training continues until, on a separately reserved validation dataset, the correlation between the model's predicted evaluation values ​​and the labels reaches a preset convergence criterion: the correlation coefficient stabilizes above 0.85 and its fluctuation range is less than 0.01 over 10 consecutive training cycles. At this point, the model is considered to have output visual accommodation compensation state evaluation values ​​that are consistent with the visual fatigue level labels based on the input data, and training is considered complete. This 0.85 is set based on a commonly used standard for statistically high correlation.

[0027] S3. When the visual accommodation compensation state exceeds a preset threshold, calculate the visual behavior entropy corresponding to the real-time visual behavior data, and simultaneously analyze the semantic focus change rate of the content of the virtual scene currently displayed on the immersive experience device. Specifically, this is implemented as follows: When the visual accommodation compensation state assessment value exceeds a preset visual load threshold, the simultaneous analysis process of calculating visual behavior entropy and content semantic focus change rate is initiated. This visual load threshold is a pre-set scalar value, obtained based on statistical analysis of pre-experimental data. Specifically, it involves collecting a large number of user visual accommodation compensation state assessment values ​​under normal comfort and visual fatigue states, and setting the threshold as the critical value that can distinguish between the two states in more than 95% of cases; for example, this threshold could be set to 0.75. When the visual accommodation compensation state assessment value output from the neural network model in real time continuously exceeds this threshold, for example, if it is greater than 0.75 for three consecutive assessment cycles, the visual accommodation compensation state is determined to have exceeded the preset threshold, triggering subsequent calculations. If the assessment value does not exceed the threshold, monitoring continues without initiating calculations.

[0028] The calculation of visual behavior entropy is based on eye movement trajectory data from real-time visual behavior data and is implemented using the information entropy algorithm. The specific process is as follows: First, the input eye movement trajectory data is preprocessed and discretized. Eye movement trajectory data is a series of two-dimensional gaze point coordinates with timestamps. Discretization refers to dividing the display field of view of the virtual scene into multiple non-overlapping regular grid regions. For example, the entire screen is evenly divided into 10 parts in both the horizontal and vertical directions, resulting in 100 rectangular gaze regions of equal area. The division is based on ensuring that each region has a visually distinguishable spatial range, while the number of regions balances computational complexity and behavioral resolution granularity. For example, the side length of the region is usually not less than the number of screen pixels corresponding to 1 degree of visual angle. Next, the number of times the user's gaze point falls into each rectangular gaze region within a specified analysis time window, such as the past 5 seconds, is counted. During the count, each gaze point coordinate with timestamp is mapped to the corresponding rectangular gaze region according to its two-dimensional coordinate value. If a gaze point happens to fall on the boundary of a region, it is assigned according to predefined rules, such as uniformly belonging to the left or upper region. Then, based on the statistically obtained number of fixations in each region, the probability of each fixation region being fixated is calculated. The probability calculation method is to divide the number of fixations in a region by the total number of fixations within the analysis time window, thus obtaining the probability value of that region being fixated. The set of probability values ​​for all fixation regions constitutes the user's fixation probability distribution within that time window. Finally, the visual behavior entropy value is calculated using the information entropy formula. The information entropy formula states that visual behavior entropy equals a negative summation sign, with the summation iterating through all fixation regions, multiplying the probability value of each region by the base-2 logarithm of that probability value. Specifically, the calculation first calculates the product of the probability value of each region and the base-2 logarithm of that probability value, then adds the products of all regions, and finally takes the negative value of the sum. The resulting scalar is the visual behavior entropy value. This value directly characterizes the randomness and disorder of the distribution of user fixations in different spatial regions within the analysis time window.

[0029] The synchronous content semantic focus change rate analysis process is implemented independently of, but temporally aligned with, the visual behavior entropy calculation. This analysis is based on the virtual scene data output in real time by the rendering pipeline of the immersive experience device. Specifically, the rendering pipeline outputs a corresponding scene depth map and object identification information while generating each frame of the virtual scene image. The scene depth map is a matrix aligned with the image pixels, where the value of each element represents the distance from the surface of the virtual object corresponding to that pixel location to the camera. The object identification information is another matrix of the same size, where each element is an integer label that uniquely identifies the virtual scene object to which that pixel location belongs; for example, background is label 0, character is label 1, and prop is label 2. The analysis process first extracts the three-dimensional coordinates of the scene semantic centroid for each frame from continuous multi-frame data. For single-frame data, the calculation of the three-dimensional coordinates of the semantic centroid is divided into two steps. The first step is to calculate the two-dimensional image space centroid of each type of object, based on the fact that all pixels of that type of object have the same label in the object identification information matrix. The average of the screen coordinates of these pixels is used to obtain the two-dimensional centroid of that type of object. The second step involves using the two-dimensional centroid coordinates to index the corresponding depth value in the scene depth map matrix. Combined with the camera intrinsic parameter matrix, the coordinates of the centroid of this object type in the virtual world's three-dimensional coordinate system are calculated through back projection. Next, from all object categories, the three object categories with the largest screen projection area in the current frame are selected. Their three-dimensional centroid coordinates are then weighted and averaged, with the weight representing the proportion of each object's screen projection area. The resulting weighted average coordinates are the three-dimensional coordinates of the scene's semantic centroid for that frame. If there are fewer than three identifiable objects in the scene, the weighted average of all objects is used. Subsequently, the content semantic focus change rate is calculated. The three-dimensional coordinates of the scene's semantic centroid in the current frame and the immediately preceding frame are taken, and the Euclidean distance between them is calculated.

[0030] The Euclidean distance is calculated by squaring the differences between two coordinate points in the X, Y, and Z dimensions, summing the three squares, and then taking the square root of the sum. The result is the displacement distance measured in virtual world units. Simultaneously, the rendering frame interval between these two frames is obtained from the rendering pipeline. This interval is usually fixed; for example, at a rendering rate of 90 frames per second, the frame interval is approximately 11.1 milliseconds. The final calculation method for the content semantic focus change rate is to divide the calculated Euclidean distance by the corresponding rendering frame interval. The result is the displacement rate per unit time, which is the content semantic focus change rate. The entire analysis process uses the same starting trigger signal and the same analysis time window as the visual behavior entropy calculation, ensuring the synchronization and comparability of the two metrics in the time dimension. If a valid semantic centroid coordinate cannot be calculated within the time window, for example, if all objects are background, the content semantic focus change rate is set to zero.

[0031] S4. Based on the coupling relationship between visual behavioral entropy and the rate of change of content semantic focus, analyze the dominant cause types of abnormal visual accommodation compensation states, specifically as follows: First, obtain the visual behavior entropy value and the content semantic focus change rate value calculated and output by step S3. These two values ​​are synchronization indicators calculated for the exact same analysis time window. The visual behavior entropy value represents the spatial disorder of the user's visual exploration behavior within this time window, and the content semantic focus change rate value represents the degree of change of the core semantic content of the virtual scene over time within the same time window.

[0032] The analysis process relies on comparing the acquired two indicator values ​​with two pre-defined independent judgment thresholds. The first judgment threshold is the entropy threshold, and the second is the rate of change threshold. The specific method for setting the entropy threshold is as follows: In the pre-experiment during the system development phase, a group of representative users are organized to experience the system in a static virtual scene with low visual load, and their visual behavior entropy data in a relaxed state is continuously collected. The statistical distribution of all collected visual behavior entropy data is calculated, and the value corresponding to a higher percentile of this distribution is selected as the entropy threshold. For example, the value corresponding to the 95th percentile can be selected, which means that only five percent of the observations in a normal relaxed state will exceed this value. The entropy threshold is set as a statistical boundary value used to distinguish between ordered and disordered visual exploration behavior. For example, this threshold may be quantized to 3.2 bits. The specific method for setting the rate of change threshold is as follows: In the pre-experiment phase, collect multiple typical virtual experience content segments covering different scene types, such as game or video clips containing smooth scenes and action scenes; calculate the rate of change of semantic focus of the content corresponding to each frame in these content segments to form a rate of change dataset; calculate the statistical distribution of the dataset, and select a percentile value that can represent the general level of change of scene content as the rate of change threshold. For example, the value corresponding to the 80th percentile can be selected. The rate of change threshold is set as a reference benchmark for judging whether the change of scene content is drastic. For example, the threshold may be quantified as 1.5 meters per second.

[0033] After obtaining the visual behavior entropy value and the content semantic focus change rate value for the current analysis time window, two independent comparison operations are performed simultaneously. The first comparison operation is a numerical comparison, comparing the visual behavior entropy value with an entropy threshold to determine if the former is greater than the latter. The second comparison operation is also a numerical comparison, comparing the content semantic focus change rate value with a change rate threshold to determine if the former is greater than the latter.

[0034] Based on the results of the two comparison operations above—namely, the states of being higher than or lower than or equal to—the dominant cause type is determined according to a predetermined decision-making logic. The specific determination rules are as follows: If the result of the first comparison operation is that the visual behavior entropy value is higher than the entropy threshold, and the result of the second comparison operation is that the content semantic focus change rate is higher than the change rate threshold, then the dominant cause type of the current abnormal visual accommodation compensation state is determined to be exogenous content-driven. This determination is based on the fact that the simultaneous occurrence of high disorder in visual behavior and high variability in scene content indicates that rapidly changing external content is the main driving factor causing the user's visual system disorder. If the result of the first comparison operation is that the visual behavior entropy value is higher than the entropy threshold, but the result of the second comparison operation is that the content semantic focus change rate is lower than or equal to the change rate threshold, then the dominant cause type of the current abnormal visual accommodation compensation state is determined to be endogenous physiological state abnormality. This determination is based on the fact that even when the scene content itself changes gradually, the user's visual behavior still exhibits high disorder, indicating that the disorder mainly stems from abnormalities in the user's own visual accommodation function and other physiological states. If the result of the first comparison operation is that the visual behavior entropy value is lower than or equal to the entropy threshold, then regardless of the content semantic focus change rate value, it is determined that there is no significant anomaly dominated by the defined cause type. Therefore, no cause type determination is performed, the process returns and continues to monitor the indicators in subsequent time windows.

[0035] The entire analysis, comparison, and judgment process is executed cyclically at a fixed time period, such as once every 1 second, thereby achieving continuous and dynamic identification of the causes of visual fatigue and providing clear classification input for subsequent steps.

[0036] S5. Match the gaze point distribution in real-time visual behavior data with the visual saliency topology map of the virtual scene according to the dominant cause type, and analyze the compositional characteristics of the topology mapping deviation to identify at least one specific dominant factor that causes the visual accommodation compensation state to exceed a preset threshold. The specific implementation is as follows: First, a visual saliency topology map is generated based on the depth and color contrast information of the current virtual scene. The visual saliency topology map is a two-dimensional data matrix aligned with the pixel spatial positions of the current virtual scene frame. Each element in the matrix stores a value representing the visual saliency of the corresponding pixel area in the virtual scene, attracting the user's attention. The depth information required to generate the visual saliency topology map comes directly from the scene depth map output in real-time by the rendering pipeline of the immersive experience device. Each pixel value in this depth map records the distance from the surface of the corresponding object to the virtual camera. The color contrast information required to generate the visual saliency topology map is obtained as follows: the current frame color image output synchronously by the rendering pipeline is preprocessed, converting the image from the red-green-blue color space to the hue-saturation-lightness color space; the lightness component image is extracted; the standard deviation of the lightness values ​​of each pixel in the lightness component image and its surrounding fixed neighborhood pixels is calculated, and the calculated standard deviation is used as the color contrast value of that pixel. The core calculation process for generating a visual saliency topology map is as follows: For each pixel location in the color image, its corresponding depth map distance value and color contrast value are read simultaneously; the depth value is normalized, mapping the distance value to a depth saliency score between 0 and 1, with the mapping principle being that the closer the distance, the higher the score; the color contrast value is normalized, mapping it to a color saliency score between 0 and 1, with the mapping principle being that the higher the contrast, the higher the score; the depth saliency score and color saliency score of each pixel are linearly weighted and fused according to a set of predetermined fusion weight coefficients to obtain the comprehensive visual saliency score of the pixel, which is used as the element value of the visual saliency topology map matrix at the corresponding position. The predetermined fusion weight coefficients, such as a depth weight coefficient of 0.6 and a color weight coefficient of 0.4, were obtained through pre-experiment calibration. In the pre-experiment, a set of test scenes with clear visual appeal targets was prepared, and a large number of users' average gaze heatmaps were collected. The fusion weights of depth and color were adjusted to maximize the spatial correlation between the generated visual saliency topology map and the users' average gaze heatmaps. Finally, the weight coefficients that resulted in the highest correlation were determined as the predetermined fusion weight coefficients. After generating the visual saliency topology map, the comprehensive visual saliency scores of all its pixels were sorted, and the regions connected by the pixels with the top 20% scores were defined as high saliency regions.

[0037] When the dominant trigger type determined by step S4 is exogenous content-driven, targeted analysis is performed to identify specific dominant factors. The core of this analysis is calculating the matching degree between the gaze point distribution in real-time visual behavior data and the highly salient regions in the visual salience topology map. The gaze point distribution refers to the set of two-dimensional screen coordinates of all gaze points extracted from the eye movement trajectory data of real-time visual behavior data within the current analysis time window. The specific calculation method for the matching degree is as follows: First, traverse the coordinates of each gaze point in the gaze point distribution set and determine whether the coordinate falls within the pixel range of any highly salient region defined by the visual salience topology map; count the number of gaze points falling within the highly salient regions; divide this number by the total number of gaze points in the gaze point distribution set within the current analysis time window, and the resulting quotient is the matching degree, which is a dimensionless value between 0 and 1. Subsequently, the calculated matching degree is compared with a preset matching degree threshold. The method for obtaining the preset matching threshold is as follows: In the pre-experiment during the system development phase, users are organized to watch a series of visually guided virtual scenes in a visually comfortable state, and their gaze distribution is recorded simultaneously to generate a visual saliency topology map of the corresponding scene; the matching degree of each data segment is calculated to form a set of matching degree sample data; the distribution of this sample data is statistically analyzed, and the value corresponding to its 10th percentile is taken as the preset matching threshold, for example, this threshold may be set to 0.65. If the currently calculated matching degree is lower than the preset matching threshold, it indicates that the user's actual gaze behavior has not effectively followed the most salient content in the scene. At this time, in all highly saliency regions of the visual saliency topology map, the proportion of gaze points covering each region is calculated one by one, that is, the number of gaze points falling into the region divided by the total number of pixels in the region; one or more highly saliency regions with the lowest proportion of gaze point coverage are selected; the entity objects or visual elements in the virtual scene corresponding to these regions, such as a fast-moving character, a flashing special effect spot, or a high-contrast texture, are identified as the specific dominant factors causing the current abnormal visual accommodation compensation state.

[0038] When the dominant triggering factor determined by step S4 is an endogenous physiological abnormality, another set of analytical logic is executed to identify the specific dominant factor. This analysis focuses on the spatiotemporal characteristics of the fixation point distribution itself, calculating two key indicators: dispersion and drift frequency. The dispersion is calculated on the set of fixation point distributions within the current analysis time window. The calculation steps are as follows: First, calculate the arithmetic mean center point of all fixation point coordinates in the set of fixation point distributions in the two-dimensional space of the screen; then, calculate the Euclidean distance from each fixation point coordinate in the set to this mean center point; finally, calculate the standard deviation of all these distance values, which is the dispersion, measured in pixels. A larger value indicates that the fixation points are more spatially dispersed. The drift frequency is calculated on the fixation points arranged over time. The calculation steps are as follows: Based on the timestamp order of the gaze points, examine each pair of consecutive gaze points; determine whether these two consecutive gaze points fall within different and spatially non-adjacent salience regions in the visual salience topology map; if so, record it as a gaze drift event; count the total number of gaze drift events occurring within the current analysis time window; divide the total number by the duration of the analysis time window to obtain the drift frequency, measured in times per second. Then, compare the calculated dispersion with a preset dispersion threshold, and simultaneously compare the calculated drift frequency with a preset frequency threshold. The preset dispersion threshold is obtained by: collecting gaze point data from users viewing static or flat scenes in a relaxed state during a pre-experiment, calculating its dispersion to form a sample distribution, and taking the value corresponding to the 90th percentile of this distribution as the threshold, for example, this threshold might be 150 pixels. The preset frequency threshold is obtained similarly, taking the 90th percentile of the user gaze drift frequency sample distribution under the same conditions as the threshold, for example, this threshold might be 2.5 times per second. If the calculated dispersion is higher than a preset dispersion threshold, or the calculated drift frequency is higher than a preset frequency threshold, then the gaze behavior pattern exhibiting this numerical characteristic is identified as the specific dominant factor. This may correspond to a scattered state in which stable gaze cannot be maintained, or a rapid and aimless visual search state. The entire recognition process provides a clear target for targeted intervention in subsequent steps.

[0039] S6. Based on specific dominant factors, dynamically adjust the display parameters of the immersive experience device to prevent myopia. The specific implementation is as follows: When a specific dominant factor is identified as a scene content element driven by exogenous content, the goal of the adjustment strategy is to proactively weaken the visual stimulus intensity of that element and create conditions to guide the user's visual attention to shift to surrounding, gentler areas. The specific implementation process includes the following steps: First, determine the target visual region requiring parameter adjustment. This target visual region is the entire pixel range occupied by the specific scene content element(s) identified in step S5 within the current frame image of the virtual scene. The range of this region is precisely defined by calling the object identifier buffer provided by the immersive experience device rendering pipeline or real-time semantic segmentation results to ensure accurate target positioning. Next, perform a decay adjustment of local display parameters on the target visual region. The first adjustment is to reduce the display brightness within the target visual region. Specifically, during the post-processing stage of graphics rendering, obtain the brightness value of the current frame for the pixel range corresponding to the target visual region. The brightness value is typically obtained from the image's luminance channel. Then, multiply the original brightness value of each pixel by a pre-set brightness decay coefficient to obtain the reduced new brightness value. The brightness attenuation coefficient is a decimal between 0 and 1. Its specific value, such as 0.7, is determined through the system's pre-calibration process. The pre-calibration process involves designing a set of test scenarios containing highly attractive elements during the development phase. Test users are allowed to experience these scenarios with default parameters and different attenuation coefficients, and their gaze duration and blink frequency are monitored. The attenuation coefficient value that stably reduces gaze duration to within a safe threshold without causing significant discomfort or image distortion is ultimately selected as the global setting. The second adjustment is to reduce color contrast within the target visual area. The specific method involves first calculating the difference in saturation and brightness between the pixel colors within the target area and the average color of the entire scene. Then, a color mixing algorithm is used to moderately blend the colors of the target area towards the average color of the scene background. The blending strength of this algorithm is controlled by a parameter, for example, setting it to reduce the peak saturation of the colors within the area by 30%. This strength parameter is also determined through pre-calibration experiments, using the standard of effectively distracting attention while maintaining image naturalness. Then, compensatory enhancement adjustments are performed on the surrounding areas adjacent to the target visual area. The adjacent region is defined as a ring-shaped area directly adjacent to the outer boundary of the target region, with a preset width. This preset width is set according to a certain proportion of the total screen width, for example, five percent of the screen width. The color saturation of this adjacent region is increased. Specifically, the original saturation value of the pixels within this region is obtained and multiplied by a saturation enhancement factor greater than 1, such as 1.3, to obtain the enhanced saturation. The upper limit of this enhancement factor is set to ensure that the enhanced saturation does not exceed the generally accepted range for comfortable viewing by the human eye. Its purpose is to provide a visually relatively active but gentle transition zone after the core stimulus source is weakened, so as to smoothly guide the movement of the gaze.All the above parameter adjustments are performed gradually. For example, within a 0.5-second time window, the parameters are gradually changed from the current value to the target value through linear interpolation to avoid discomfort caused by sudden changes in the image. At the same time, the system monitors the visual accommodation compensation status evaluation value in real time. If this value begins to decline, the adjustment intensity can be dynamically and slightly reduced accordingly.

[0040] When the specific dominant factor is identified as a gaze behavior pattern corresponding to an abnormal endogenous physiological state, the goal of the adjustment strategy is to comprehensively reduce the overall visual load of the virtual environment and provide a relaxing viewing state for the user's visual accommodation system. The specific implementation process includes the following steps. First, perform global parameter adjustments covering the entire display screen. The first adjustment is to reduce the global display brightness of the immersive experience device. The specific method is to calculate the average brightness value of the current frame and refer to a preset global target brightness benchmark, for example, set to 40% of the device's maximum brightness capability. In each subsequent rendered frame, a smoothing filtering algorithm is used to gradually approach the overall output brightness of the screen towards this target brightness benchmark, for example, allowing a maximum brightness change of 1% of the original value per frame, until stability is achieved. This target brightness benchmark is set based on the comfortable brightness level range recommended in visual ergonomics research for prolonged viewing. The second adjustment is to reduce the global color contrast of the immersive experience device. The specific method is to apply a global tone mapping operator or contrast compression filter at the end of the graphics rendering pipeline. The core function of this operator is to compress the brightness difference between the brightest and darkest parts of the image. Specifically, this means moderately reducing the brightness of highlights and increasing the brightness of shadows without losing the details of the main subject. The degree of compression is quantified and controlled by a global contrast compression coefficient, for example, set to compress the dynamic range of the image to 80% of its original range. This coefficient is set with reference to the appropriate contrast range described in research literature on visual comfort. Then, the depth-of-field blur rendering effect of the immersive experience device is activated and configured. Depth-of-field blur is a post-processing technique that simulates the imaging characteristics of real optical lenses, causing objects outside the focal plane of the virtual camera to appear blurred in accordance with physical laws. The core of the adjustment operation is to significantly expand the area outside the sharp focus range in the virtual scene, i.e., the out-of-focus area. The specific implementation involves two key steps: the first step is to determine the position of the virtual focal plane. As a conservative and easy-to-implement strategy, the focal plane can be fixed at a moderate distance in virtual space, such as 2 meters from the virtual camera's observation point. The second step is to adjust the key rendering parameters of the depth-of-field effect to expand the blur range. This mainly includes reducing the aperture value of the virtual camera, for example, adjusting it from the default f / 16 to f / 4, to enhance the intensity of optical blur; at the same time, reducing the depth of focus of the lens, that is, shortening the depth range of sharp imaging.By combining and adjusting these parameters, the final rendered image shows only one object within a very narrow depth range centered on the focal plane in sharpness. The distant and near objects, occupying most of the image area, exhibit a natural, soft blur. The physiological basis for this is that a large area of ​​blurred background effectively reduces the accommodative fluctuations and tension experienced by the ciliary muscle in continuously focusing on details at different depths, thus providing a low-demand, relaxing environment for the fatigued visual accommodation system. All the aforementioned adjustments targeting the intrinsic state also follow the principle of smooth transition. Once the user's visual accommodation compensation state assessment value returns to below the preset visual load threshold, the system does not immediately restore the default settings. Instead, it gradually reverts to the adjustments at a slower rate, such as over 5 to 10 seconds, allowing the display parameters to smoothly return to normal values ​​to prevent repeated relapses.

Claims

1. A myopia prevention method based on a neural network model, characterized in that, include: S1. Acquire real-time visual behavior data of users in immersive experience devices; S2. Based on real-time visual behavior data, the user's visual accommodation compensation status is evaluated through a neural network model. S3. When the visual adjustment compensation state exceeds the preset threshold, calculate the visual behavior entropy corresponding to the real-time visual behavior data, and simultaneously analyze the content semantic focus change rate of the virtual scene currently displayed by the immersive experience device. S4. Based on the coupling relationship between visual behavioral entropy and the rate of change of content semantic focus, analyze the dominant cause type of abnormal visual accommodation compensation state. S5. Match the gaze point distribution in real-time visual behavior data with the visual saliency topology map of the virtual scene according to the dominant cause type, and analyze the composition characteristics of the topology mapping deviation to identify at least one specific dominant factor that causes the visual accommodation compensation state to exceed the preset threshold. S6. Based on specific dominant factors, dynamically adjust the display parameters of immersive experience devices to prevent myopia.

2. The myopia prevention method based on a neural network model according to claim 1, characterized in that, Acquire real-time visual behavior data of users in immersive experience devices, including: The eye-tracking unit integrated into the immersive experience device acquires the user's eye movement trajectory data and pupil diameter change data. The user's head displacement and posture change data are acquired through the inertial measurement unit integrated into the immersive experience device; Eye movement data, pupil diameter change data, head displacement data, and posture change data are synchronized and time-stamped to generate real-time visual behavior data.

3. The myopia prevention method based on a neural network model according to claim 1, characterized in that, Based on real-time visual behavior data, a neural network model is used to assess the user's visual accommodation compensation status, including: Eye movement data, pupil diameter change data, head displacement data, and posture change data from real-time visual behavior data are input into a trained neural network model. The neural network model analyzes saccade and fixation patterns in eye movement trajectory data, performs correlation analysis between pupillary light reflection and accommodation in pupil diameter change data, and calculates vestibular-visual interaction compensation in head displacement and posture change data. The neural network model performs multi-source feature fusion based on the results of saccade and fixation pattern analysis, the results of pupillary light reflection and accommodation correlation analysis, and the results of vestibular-visual interaction compensation calculation, and outputs a visual accommodation compensation state assessment value that characterizes the real-time load of the user's visual system.

4. The myopia prevention method based on a neural network model according to claim 3, characterized in that, The trained neural network model was obtained by acquiring historical visual behavior data and synchronously recorded refraction parameter change data generated by multiple groups of users in the pre-experimental immersive experience; using the historical visual behavior data as input and the visual fatigue level label corresponding to the refraction parameter change data as a supervision signal, the neural network model was trained until it could output a visual accommodation compensation state evaluation value consistent with the visual fatigue level label based on the input data.

5. The myopia prevention method based on a neural network model according to claim 1, characterized in that, When the visual accommodation compensation state exceeds a preset threshold, the visual behavior entropy corresponding to the real-time visual behavior data is calculated, and the semantic focus change rate of the content currently displayed virtual scene on the immersive experience device is analyzed simultaneously, including: When the visual accommodation compensation state assessment value exceeds the preset threshold, the visual behavior entropy, which represents the randomness and disorder of the user's visual exploration behavior, is calculated by the information entropy algorithm based on the eye movement trajectory data in the real-time visual behavior data. The virtual scene depth map and object identification information of the current frame and historical frames are synchronously output based on the rendering pipeline of the immersive experience device. By analyzing the displacement rate of the semantic center of gravity of the scene in three-dimensional space between consecutive frames, the content semantic focus change rate is obtained.

6. The myopia prevention method based on a neural network model according to claim 5, characterized in that, The visual behavior entropy is calculated using the information entropy algorithm as follows: the eye movement trajectory data in the real-time visual behavior data is discretized, the fixation regions are divided, and the number of fixation points in each region is counted; the probability distribution of each fixation region being fixated on is calculated based on the number of fixation points. Based on the probability distribution, the visual behavior entropy value, which characterizes the randomness and disorder of user visual exploration behavior, is calculated using the information entropy formula.

7. The myopia prevention method based on a neural network model according to claim 5, characterized in that, The semantic focus change rate of the content is analyzed by analyzing the displacement rate of the semantic center of gravity of the scene between consecutive frames. This is achieved by extracting the three-dimensional coordinates of the semantic center of gravity of each frame from the continuous multi-frame virtual scene output by the rendering pipeline of the immersive experience device; calculating the Euclidean distance between the three-dimensional coordinates of the semantic center of gravity of the current frame and the previous frame; and calculating the semantic focus change rate of the content based on the Euclidean distance and the rendering frame interval of the continuous multi-frame virtual scene.

8. The myopia prevention method based on a neural network model according to claim 1, characterized in that, Based on the coupling relationship between visual behavioral entropy and the rate of change of content semantic focus, the dominant triggering factors for abnormal visual accommodation compensatory states are analyzed, including: The visual behavior entropy is compared with a preset entropy threshold, and the content semantic focus change rate is compared with a preset change rate threshold. When the visual behavior entropy is higher than the entropy threshold and the change rate of the semantic focus of the content is also higher than the change rate threshold, the dominant cause type is determined to be exogenous content-driven. When the visual behavioral entropy is higher than the entropy threshold and the change rate of the semantic focus of the content is lower than or equal to the change rate threshold, the dominant cause type is determined to be an abnormal endogenous physiological state.

9. The myopia prevention method based on a neural network model according to claim 1, characterized in that, Based on the dominant trigger type, the gaze point distribution in real-time visual behavior data is matched with the visual saliency topology map of the virtual scene. The compositional characteristics of the topology mapping deviation are analyzed to identify at least one specific dominant factor that causes the visual accommodation compensation state to exceed a preset threshold, including: A visual saliency topology map is generated based on the depth and color contrast information of the current virtual scene. When the dominant cause type is exogenous content-driven, the matching degree between the gaze point distribution and the highly salience region in the visual salience topology map is analyzed, and the scene content elements corresponding to the highly salience region with a matching degree lower than the preset matching degree threshold are identified as specific dominant factors. When the dominant cause is an abnormal endogenous physiological state, the dispersion and drift frequency of the fixation point distribution relative to any region in the visual salience topology are analyzed, and fixation behavior patterns with a dispersion higher than a preset dispersion threshold or a drift frequency higher than a preset frequency threshold are identified as specific dominant factors.

10. The myopia prevention method based on a neural network model according to claim 1, characterized in that, Based on specific dominant factors, the display parameters of immersive experience devices are dynamically adjusted to prevent myopia, including: When the specific dominant factor is the scene content element driven by external content, reduce the display brightness and contrast of the visual area related to the corresponding scene content element in the immersive experience device, and increase the color saturation of the area adjacent to the corresponding scene content element in the immersive experience device. When the specific dominant factor is the gaze behavior pattern corresponding to an abnormal endogenous physiological state, reduce the global display brightness and global color contrast of the immersive experience device, and activate the depth-of-field blur rendering effect of the immersive experience device to increase the out-of-focus area of ​​the virtual scene.