Multimodal scene perception-based hyperopia method and system, electronic equipment and medium
By using multimodal data perception technology to identify learning scenarios and correct visual fatigue index, the shortcomings of traditional visual fatigue reminder methods are solved, enabling accurate prediction and personalized protection of visual fatigue, and improving the efficiency of visual health protection.
Patent Information
- Application Number
- CN202511035345.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-11-11
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In existing technologies, reminders at fixed time intervals cannot effectively prevent eye strain and cannot adapt to differences in visual load under different usage scenarios, resulting in reminders that are too frequent or untimely, thus reducing the efficiency of eye strain protection.
By acquiring ambient lighting data, eye movement data, and screen interaction data, the system identifies learning scenarios, calculates visual load quantification values, predicts visual fatigue indices, and uses lighting adaptation coefficients to determine whether to perform telephoto training operations, thus avoiding the shortcomings of traditional reminder methods.
It enables accurate prediction and timely protection against eye strain, improves the efficiency of eye strain protection, avoids problems of excessive disturbance or untimely reminders, and provides personalized visual health protection.
Smart Images

Figure CN120918924A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, specifically to a farsight method, system, electronic device, and medium based on multimodal scene perception. Background Technology
[0002] With the rapid development of digital education and the widespread use of smart learning devices, students and workers are spending significantly more time using electronic screens for learning and work. Prolonged close-range viewing of electronic screens easily leads to visual health problems such as eye strain, dry eye syndrome, and worsening myopia, which has become a significant factor affecting the visual health of modern people. Therefore, how to effectively prevent and alleviate eye strain during the use of digital devices has become an urgent technical challenge to be solved.
[0003] Currently, methods to protect against eye strain from digital devices typically employ fixed-time reminder mechanisms, such as reminding users to take a break every 20 or 30 minutes, or determining whether a break is needed based on simple cumulative usage time.
[0004] However, in practical applications, due to the differences in the scenarios in which users use their devices, using only fixed-interval reminders ignores the differences in the accumulation of visual fatigue under different usage scenarios. This can easily lead to reminders being too frequent and interfering with normal use, or reminders being untimely and unable to effectively prevent the occurrence of visual fatigue, thereby reducing the efficiency of assisting users in preventing visual fatigue. Summary of the Invention
[0005] This application provides a farsightedness method, system, electronic device, and medium based on multimodal scene perception, which can promptly remind users and improve the efficiency of assisting users in preventing eye fatigue.
[0006] Firstly, this application provides a far-seeing method based on multimodal scene perception, including: Acquire ambient light data, eye movement data, and screen interaction data during the user's use of the learning device; The user's learning scenario is identified based on the screen interaction data, and the corresponding visual load quantification value is calculated based on the identified learning scenario and the eye movement data. Based on the visual load quantification value and the user's continuous eye use duration, predict the user's visual fatigue index within a preset time window; The ambient light data is compared and analyzed with the screen brightness of the learning device to calculate the light adaptation coefficient, and the visual fatigue index is corrected based on the light adaptation coefficient. When the corrected visual fatigue index exceeds the fatigue index threshold, a preset telephoto training operation is performed on the user.
[0007] By adopting the above technical solution, multimodal data such as ambient light data, eye movement data, and screen interaction data during the user's use of the learning device can be acquired to comprehensively perceive the user's usage scenario. Furthermore, based on the screen interaction data, the user's specific learning scenario is identified, and the corresponding visual load quantification value is calculated in combination with the eye movement data, thereby achieving accurate quantification of visual load under different learning scenarios. On this basis, the visual fatigue index of the user within a preset time window is predicted based on the visual load quantification value and continuous eye use duration. The visual fatigue index is corrected by the light adaptation coefficient obtained by comparing ambient light data and screen brightness, making the visual fatigue prediction results more accurate and reliable. Finally, based on the corrected visual fatigue index, it is determined whether to perform far-focus training operations, thereby avoiding the problems of excessive disturbance or untimely reminders that may be caused by traditional fixed time interval reminder methods, and improving the efficiency of assisting users in preventing visual fatigue.
[0008] Optionally, the application type, page dwell time, scrolling frequency, and click density are extracted from the screen interaction data, and corresponding interaction feature vectors are constructed; the matching degree between the interaction feature vectors and the scene feature vectors corresponding to each learning scene in the preset learning scene template library is calculated; and the learning scene with the highest matching degree is selected as the user's learning scene.
[0009] Optionally, blink frequency, fixation duration, pupil diameter change rate, and eye movement speed are extracted from the eye movement data as eye fatigue feature parameters; the corresponding scene weight coefficients are obtained from a preset weight configuration table according to the identified learning scene; each of the eye fatigue feature parameters is weighted and calculated with the scene weight coefficients to obtain the weighted fatigue value corresponding to each eye fatigue feature parameter; the weighted fatigue values corresponding to each eye fatigue feature parameter are normalized and summed to obtain the user's visual load quantification value.
[0010] Optionally, the basic fatigue increment per unit time is calculated based on the visual load quantification value; a fatigue accumulation coefficient is determined based on the continuous eye use duration, wherein the fatigue accumulation coefficient is positively correlated with the continuous eye use duration; the basic fatigue increment is multiplied by the fatigue accumulation coefficient to obtain the actual fatigue increment; the actual fatigue increment is integrated over time within the preset time window to obtain the user's visual fatigue index within the preset time window.
[0011] Optionally, the intensity level of telephoto training is determined based on the corrected visual fatigue index, and a corresponding virtual distant scene is selected from a preset scene library based on the intensity level; the screen imaging distance of the learning device is adjusted to a preset distant focal range, and a dynamic guide marker is generated in the virtual distant scene, the movement trajectory of which is determined by the eye movement direction identified in the eye movement data; the user's eye movement response data during telephoto training is monitored in real time, and the movement parameters and guidance frequency of the dynamic guide marker are dynamically adjusted based on the eye movement response data; when the user's eye movement response data is detected to reach a preset recovery standard, the screen imaging distance is gradually adjusted back to the standard focal length according to a preset step size, and eye protection prompts are displayed on the screen of the learning device.
[0012] Optionally, based on the eye movement data, a target guidance direction for eye movement is determined; based on the target guidance direction, a motion trajectory template is constructed, the motion trajectory template including at least one of horizontal movement, vertical movement, diagonal movement, and circular movement; the initial movement speed and dwell time of the dynamic guidance mark are set according to the intensity level; a dynamic guidance mark is generated in a preset depth level of the virtual distant scene, and the dynamic guidance mark is controlled to perform directional movement according to the motion trajectory template; the following accuracy of the user's eye tracking of the dynamic guidance mark is calculated in real time, and when the following accuracy is less than the accuracy threshold, the movement speed of the dynamic guidance mark is reduced and the dwell time is extended, and when the following accuracy is greater than or equal to the accuracy threshold, the movement speed of the dynamic guidance mark is increased and the dwell time is shortened.
[0013] Optionally, the eye movement data is reconstructed to extract the amplitude parameters and smoothness parameters of eye movement in the horizontal, vertical, and diagonal directions; the amplitude parameters are compared with a preset standard amplitude range to identify directions with amplitudes lower than the standard range as target guidance directions for eye movement; the smoothness parameters are statistically analyzed to identify directions with a rate of change of movement speed higher than a rate of change threshold or a pause frequency higher than a frequency threshold as target guidance directions for eye movement.
[0014] A second aspect of this application provides a far-seeing system based on multimodal scene perception, the system comprising: The data acquisition module is used to acquire ambient light data, eye movement data, and screen interaction data during the user's use of the learning device; The visual fatigue index determination module is used to identify the user's learning scenario based on the screen interaction data, and calculate the corresponding visual load quantification value based on the identified learning scenario and the eye movement data; and predict the user's visual fatigue index within a preset time window based on the visual load quantification value and the user's continuous eye use duration. The index correction module is used to compare and analyze the ambient light data with the screen brightness of the learning device, calculate the light adaptation coefficient, and correct the visual fatigue index based on the light adaptation coefficient. The telephoto training module is used to perform preset telephoto training operations on the user when the corrected visual fatigue index exceeds the fatigue index threshold.
[0015] A third aspect of this application provides an electronic device including a memory, a processor, and a program stored in the memory and executable on the processor, the program being loaded and executed by the processor to implement a farsight method based on multimodal scene perception.
[0016] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to implement a farsight method based on multimodal scene perception.
[0017] In summary, one or more technical solutions provided in this application have at least the following technical effects or advantages: By adopting the above technical solution, multimodal data such as ambient light data, eye movement data, and screen interaction data during the user's use of the learning device can be acquired to comprehensively perceive the user's usage scenario. Furthermore, based on the screen interaction data, the user's specific learning scenario is identified, and the corresponding visual load quantification value is calculated in combination with the eye movement data, thereby achieving accurate quantification of visual load under different learning scenarios. On this basis, the visual fatigue index of the user within a preset time window is predicted based on the visual load quantification value and continuous eye use duration. The visual fatigue index is corrected by the light adaptation coefficient obtained by comparing ambient light data and screen brightness, making the visual fatigue prediction results more accurate and reliable. Finally, based on the corrected visual fatigue index, it is determined whether to perform far-focus training operations, thereby avoiding the problems of excessive disturbance or untimely reminders that may be caused by traditional fixed time interval reminder methods, and improving the efficiency of assisting users in preventing visual fatigue. Attached Figure Description
[0018] Figure 1 This is a flowchart illustrating a far-seeing method based on multimodal scene perception provided in an embodiment of this application; Figure 2This is a schematic diagram of the structure of a far-seeing system based on multimodal scene perception provided in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0019] Explanation of reference numerals in the attached drawings: 300, electronic device; 301, processor; 302, communication bus; 303, user interface; 304, network interface; 305, memory. Detailed Implementation
[0020] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.
[0021] In the description of the embodiments of this application, the words "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design that is described as "for example" or "for instance" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design options. Rather, the use of the words "for example" or "for instance" is intended to present the relevant concepts in a specific manner.
[0022] In the description of the embodiments of this application, the term "multiple" means two or more. For example, multiple systems means two or more systems, and multiple screen terminals means two or more screen terminals. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, a feature defined with "first" or "second" may explicitly or implicitly include one or more of that feature. The terms "comprising," "including," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.
[0023] This application provides a farsight method based on multimodal scene perception. In one embodiment, please refer to... Figure 1 , Figure 1 This is a flowchart illustrating a farsightedness method based on multimodal scene perception provided in an embodiment of this application. This method can be implemented using a computer program, which can be integrated into an application or run as a standalone utility application. The method can also be implemented using a microcontroller and can run on a farsightedness system based on multimodal scene perception and the von Neumann architecture. Specifically, the method may include the following steps: Step 101: Obtain ambient light data, eye movement data, and screen interaction data during the user's use of the learning device.
[0024] Ambient lighting data refers to the lighting parameters in the user's environment collected by the lighting sensors configured on the learning device. These parameters include illuminance, color temperature, and uniformity of light distribution. This lighting data reflects the visual comfort level of the user's environment and has a significant impact on assessing the occurrence and development of eye strain. It also serves as an important basis for the system to perform lighting adaptability analysis and correct the eye strain index.
[0025] Eye movement data refers to the user's eye activity features collected by the front-facing camera of the learning device and extracted by image processing algorithms. These features mainly include eye fatigue characteristic parameters such as blink frequency, fixation duration, pupil diameter change rate, and eye movement speed. These parameters can intuitively reflect the user's real-time visual fatigue state and are the core data source for the system to calculate the quantification of visual load and predict the visual fatigue index.
[0026] Screen interaction data refers to human-computer interaction data generated during user interaction with learning devices. This includes information such as the type of application used, the duration of time spent on specific pages, the frequency of screen content scrolling, and the density distribution of click operations. This interaction data helps the system identify the specific learning scenario the user is currently in, thus providing a basis for subsequent scenario weighting and visual load calculation.
[0027] Learning devices refer to electronic devices used by users for learning activities, such as tablets, laptops, and desktop computer monitors, which have screen display capabilities. These devices need to be equipped with data acquisition devices such as light sensors and front-facing cameras, as well as an operating system capable of supporting interactive data recording, thereby enabling the collection of ambient light data, eye movement data, and screen interaction data.
[0028] Specifically, the learning device first needs to collect ambient light data in real time using a light sensor. This sensor, which can be located on the front panel of the device, detects parameters such as light intensity and color temperature in the user's environment. These parameters directly affect the user's visual comfort. Simultaneously, the device's front-facing camera continuously captures images of the user's eyes. Image processing algorithms extract eye movement data, including blink frequency, fixation duration, pupil diameter change rate, and eye movement speed—parameters characteristic of eye fatigue. These parameters directly reflect the user's visual fatigue state. During the collection of eye movement data, the device also records the user's screen interaction data, including the type of application being used, page dwell time, screen content scrolling frequency, and click density distribution. This interaction data helps the system accurately identify the user's current learning scenario. By simultaneously collecting these three types of data, the system can comprehensively perceive the user's environment, visual state, and learning behavior, providing a complete data foundation for subsequent scene recognition and visual fatigue prediction. The ambient light data sampling frequency can be set to once per second to capture changes in ambient light in a timely manner; the eye movement data acquisition frequency needs to reach more than 30 frames per second to ensure accurate capture of rapid eye movements; and screen interaction data can be recorded according to actual operation events. This multimodal data collaborative acquisition method can not only comprehensively reflect the user's usage status, but also improve the accuracy of the system's judgment through data cross-validation, thus laying a solid data foundation for achieving precise eye fatigue protection. In this way, the system can effectively overcome the limitations of existing technologies that rely on only a single data source for judgment, and achieve comprehensive monitoring and accurate assessment of the user's visual health status.
[0029] Step 102: Identify the user's learning scenario based on screen interaction data, and calculate the corresponding visual load quantification value based on the identified learning scenario and eye movement data.
[0030] Learning scenarios refer to the context in which users engage in different types of learning activities using learning devices, such as watching video courses, reading e-books, answering online questions, and editing documents. Each learning scenario has its unique interactive characteristics and visual requirements, which are distinguished and identified through feature parameters in screen interaction data, such as application type, page dwell time, scrolling frequency, and click density. Accurate identification of learning scenarios is crucial for the system to assess the user's visual load and predict visual fatigue.
[0031] Visual load quantification is a numerical representation of the user's visual system load level, calculated by the system using a specific algorithm based on the identified learning scenario and eye movement data. This quantification comprehensively considers the user's eye fatigue characteristic parameters in a specific learning scenario (such as blink frequency, fixation duration, pupil diameter change rate, and eye movement speed) and the corresponding scenario's weighting coefficients, obtained through weighted calculation and normalization. Visual load quantification directly reflects the user's visual system load status in the current learning scenario and is a crucial foundational data for the system's visual fatigue prediction.
[0032] Specifically, after acquiring user usage data, the system first needs to identify the user's specific learning scenario based on screen interaction data. This is because different learning scenarios place varying degrees of load on the user's visual system. For example, the eye movement patterns and visual attention allocation differ significantly between watching video courses and reading e-books, thus requiring different evaluation criteria. By analyzing screen interaction data, the system can accurately determine the user's current learning scenario, providing a scenario-based basis for subsequent visual load assessment. After completing the learning scenario identification, the system combines the identified learning scenario features with real-time collected eye movement data to calculate the user's current visual load quantification value. This quantification value reflects the degree of visual system load on the user in a specific learning scenario and is an important indicator for assessing the risk of visual fatigue. Through this scenario-based visual load calculation method, the system can more accurately assess the user's visual fatigue state, avoiding the limitations of traditional methods that rely solely on single data points. This scenario-aware calculation method allows the system to adopt more targeted visual fatigue assessment criteria based on the characteristics of different learning scenarios, thereby improving the accuracy and timeliness of visual fatigue prediction and providing a reliable basis for subsequent protective measures.
[0033] Based on the above embodiments, as an optional embodiment, step 102: identifying the user's learning scenario based on screen interaction data, this step may further include the following steps: Step 201: Extract application type, page dwell time, scrolling frequency and click density from screen interaction data, and construct the corresponding interaction feature vector.
[0034] Specifically, the first step is to extract feature parameters from screen interaction data. These parameters include application type, page dwell time, scrolling frequency, and click density. The system determines the application type by calling the operating system interface to obtain the currently running application's identifier information; it obtains the page dwell time by recording the actual time the user spends on each page; it calculates the scrolling frequency by detecting the number of times the user scrolls up and down the screen content per unit time; and it obtains the click density by counting the number of clicks per unit area. After obtaining these raw feature parameters, the system organizes them into an interaction feature vector according to a preset dimensional order. For example, the application type can be encoded as a discrete value, and the page dwell time, scrolling frequency, and click density can be normalized to construct a four-dimensional feature vector. This method of constructing feature vectors can transform user interaction behavior data into a quantifiable and computable mathematical expression, providing standardized input data for subsequent scene matching.
[0035] Step 202: Calculate the matching degree between the interaction feature vector and the scene feature vector corresponding to each learning scene in the preset learning scene template library; select the learning scene with the highest matching degree as the user's learning scene.
[0036] Specifically, the system needs to match the constructed interaction feature vector with a pre-defined learning scenario template library. This library stores standard feature vectors for various typical learning scenarios, obtained through statistical analysis of extensive user data. The system uses cosine similarity to calculate the similarity score between the current interaction feature vector and each scenario feature vector in the template library; a higher similarity score indicates a higher match. Specifically, the interaction feature vector is multiplied by each scenario feature vector, and then divided by the product of the magnitudes of the two vectors to obtain a normalized similarity value. The system selects the scenario with the highest similarity score as the current user's learning scenario. This scenario recognition method based on feature vector matching accurately captures the user's current learning behavior pattern, providing a reliable scenario basis for subsequent visual load calculation, thereby achieving accurate assessment of the user's visual health status. Compared to simple rule-based judgment, this method has better adaptability and accuracy, and can handle complex changes in learning scenarios.
[0037] Based on the above embodiments, as an optional embodiment, step 102, which calculates the corresponding visual load quantification value based on the identified learning scene and eye movement data, may further include the following steps: Step 203: Extract blink frequency, fixation duration, pupil diameter change rate, and eye movement speed from eye movement data as characteristic parameters of eye fatigue.
[0038] Specifically, the system needs to extract eye fatigue characteristic parameters from the collected eye movement data. For blink frequency extraction, the system calculates the number of complete blinks per minute, where a complete blink is defined as the process of the upper and lower eyelids completely closing and reopening. For fixation duration extraction, the system tracks the user's gaze direction, recording a valid fixation when the gaze lingers within a circular area with a diameter of 2 cm for more than 200 milliseconds, accumulating the total effective fixation duration per minute. For pupil diameter change rate extraction, the system collects pupil diameter data every 100 milliseconds and calculates the percentage change between two adjacent samples to obtain the pupil diameter change rate. For eye movement velocity extraction, the system records the distance the eyeball moves every 100 milliseconds and divides it by the time interval to obtain the average eye movement velocity. These characteristic parameter extraction processes are performed in real-time to ensure timely reflection of changes in the user's visual fatigue state.
[0039] Step 204: Obtain the corresponding scene weight coefficient from the preset weight configuration table based on the identified learning scene; calculate the weighted fatigue value corresponding to each eye fatigue feature parameter by weighting each eye fatigue feature parameter with the scene weight coefficient.
[0040] Specifically, the system first needs to access a preset weight configuration table to obtain scene weight coefficients. The weight configuration table sets four weight coefficients for each learning scenario, corresponding to blink frequency weight coefficient W1, fixation duration weight coefficient W2, pupil diameter change rate weight coefficient W3, and eye movement speed weight coefficient W4. For example, in the e-reading scenario, W1=0.3, W2=0.4, W3=0.2, W4=0.1; in the video viewing scenario, W1=0.4, W2=0.2, W3=0.3, W4=0.1. Based on the specific learning scenario identified in step 202, the system searches for the corresponding weight coefficient group in the weight configuration table. Then, the system multiplies the extracted eye fatigue feature parameters by their corresponding weighting coefficients. The specific calculation formulas are as follows: weighted fatigue value V1 for blink frequency = blink frequency × W1, weighted fatigue value V2 for fixation duration = fixation duration × W2, weighted fatigue value V3 for pupil diameter change rate = pupil diameter change rate × W3, and weighted fatigue value V4 for eye movement velocity = eye movement velocity × W4.
[0041] Step 205: Normalize and sum the weighted fatigue values corresponding to each eye fatigue characteristic parameter to obtain the user's visual load quantification value.
[0042] Specifically, the system first normalizes the four weighted fatigue values. Normalization uses a maximum-minimum normalization method, calculated as: Normalized value = (Current value - Minimum value) / (Maximum value - Minimum value). The maximum and minimum values are predetermined reference ranges determined through statistical analysis of a large amount of user data. For example, for the weighted fatigue value V1 based on blink frequency, the normalization formula is: V1_norm = (V1 - 5) / (30 - 5), where 5 blinks / minute is the minimum reference value and 30 blinks / minute is the maximum reference value. Other weighted fatigue values are normalized in a similar manner, yielding V2_norm, V3_norm, and V4_norm. Finally, the system sums the four normalized weighted fatigue values to obtain the final quantified visual load value L = V1_norm + V2_norm + V3_norm + V4_norm. This quantified value ranges from [0, 4], with a larger value indicating more severe visual fatigue. This calculation method takes into account the importance of each feature parameter under different scenarios, ensures the comparability of data, and can accurately reflect the user's visual load status.
[0043] Step 103: Based on the visual load quantification value and the user's continuous eye use duration, predict the user's visual fatigue index within a preset time window.
[0044] Continuous eye use duration refers to the cumulative usage time from the start of using the learning device to the current moment. The system records the start time of user login or device unlocking, and subtracts the detected significant rest intervals (such as device screen lock, application closure, etc. exceeding the preset duration) from the actual usage time. Continuous eye use duration is an important time dimension parameter for assessing the risk of visual fatigue, reflecting the cumulative load process of the user's visual system.
[0045] The visual fatigue index is a numerical value calculated by the system based on the quantified visual load and continuous eye use duration using a preset prediction model, indicating the degree of visual fatigue experienced by a user within a predetermined time window. This index comprehensively reflects the fatigue state of the user's visual system and serves as a quantitative indicator of their visual health. A higher visual fatigue index indicates more severe visual fatigue. Based on the trend of this index's changes, the system can promptly warn of potential visual health risks, providing a basis for triggering protective measures.
[0046] Specifically, after obtaining the quantified value of visual load, the system needs to combine it with the user's continuous eye use duration to predict the user's visual fatigue index within a preset time window. This is because the occurrence and development of visual fatigue are not only related to the current visual load state, but also closely related to the user's continuous eye use time. By recording the user's continuous use time from the start of using the learning device to the current moment, combined with the real-time calculated quantified value of visual load, the system can predict the degree of visual fatigue the user may reach in the future. Specifically, the system uses a preset visual fatigue prediction model, taking the quantified value of visual load and the continuous eye use duration as input parameters, to calculate the visual fatigue index within the preset time window. This prediction method based on multi-dimensional data can identify potential visual fatigue risks in advance, providing a basis for timely protective measures. Compared with the traditional single threshold judgment method, this predictive assessment method has better foresight and prevention, helping users better protect their visual health.
[0047] Based on the above embodiments, as an optional embodiment, step 103: predicting the user's visual fatigue index within a preset time window based on the visual load quantification value and the user's continuous eye use duration, may further include the following steps: Step 301: Calculate the basic fatigue increment per unit time based on the visual load quantification value.
[0048] Specifically, the system needs to calculate the basic fatigue increment per unit time based on the currently calculated visual load quantification value L. The preset formula for calculating the basic fatigue increment is used: ΔF = k × L, where k is a proportionality coefficient with a value of 0.1, and L is the visual load quantification value. Since the visual load quantification value L ranges from [0, 4], the basic fatigue increment ΔF ranges from [0, 0.4]. This linear mapping relationship can convert the visual load state into the fatigue accumulation per unit time (e.g., per minute), providing basic data for subsequent fatigue prediction.
[0049] Step 302: Determine the fatigue accumulation coefficient based on the continuous eye use duration, where the fatigue accumulation coefficient is positively correlated with the continuous eye use duration; multiply the basic fatigue increment by the fatigue accumulation coefficient to obtain the actual fatigue increment.
[0050] Specifically, the system first needs to calculate the fatigue accumulation coefficient α based on the user's continuous eye usage duration T. The fatigue accumulation coefficient is calculated using a piecewise function: when T ≤ 30 minutes, α = 1 + 0.01T; when 30 minutes < T ≤ 60 minutes, α = 1 + 0.02T; when T > 60 minutes, α = 1 + 0.03T. This piecewise function design reflects the characteristic that visual fatigue accumulates more rapidly as the usage time extends, and embodies the fatigue weighting degree at different eye usage duration stages. For example, when the continuous eye usage duration is 45 minutes, the fatigue accumulation coefficient α = 1 + 0.02 × 45 = 1.9. Then, the system multiplies the basic fatigue increment ΔF by the calculated fatigue accumulation coefficient α to obtain the actual fatigue increment ΔF' = ΔF × α. This weighted calculation method can dynamically adjust the fatigue growth rate to make it more in line with the actual development law of human eye visual fatigue.
[0051] Step 303: Perform a time integration operation on the actual fatigue increment within a preset time window to obtain the visual fatigue index of the user within the preset time window.
[0052] Specifically, the system needs to perform a time integration operation on the actual fatigue increment ΔF' within a preset time window. The preset time window is usually set to 15 minutes, which is used to predict the development trend of visual fatigue in the next 15 minutes. The system adopts a discrete integration method, divides the time window into N sampling points (for example, if there is one sampling point per minute, then N = 15), calculates the actual fatigue increment at each sampling point, and accumulates them. The specific integration calculation formula is: F = F0 + ∑(ΔF' × Δt), where F0 is the visual fatigue reference value at the current moment, and Δt is the sampling interval (1 minute). For example, if the visual fatigue reference value F0 = 2.0 at the current moment and the actual fatigue increment ΔF' = 0.3, then the predicted value of the visual fatigue index after 15 minutes is F = 2.0 + 0.3 × 15 = 6.5. This prediction method based on time integration can simulate the cumulative process of visual fatigue over time and obtain a more accurate prediction result. The visual fatigue index calculated in this way not only considers the current visual load state but also reflects the cumulative effect of continuous eye usage, and can provide a more accurate visual fatigue warning for users.
[0053] Step 104: Compare and analyze the environmental light data with the screen brightness of the learning device, calculate the light adaptation coefficient, and correct the visual fatigue index based on the light adaptation coefficient.
[0054] Among them, the light adaptation coefficient is a correction parameter obtained by comparing and analyzing the matching degree between the environmental light data and the screen brightness of the learning device. This coefficient reflects the suitability of the light conditions in the user's current usage environment and is used to adjust the predicted value of the visual fatigue index.
[0055] Specifically, the system needs to correct the predicted visual fatigue index for ambient light conditions, as the mismatch between ambient light and device screen brightness can exacerbate user visual fatigue. Specifically, the system first collects ambient light data E (in lux) using a light sensor, and simultaneously obtains the current screen brightness value B (in nits) of the learning device. The system calculates the lighting adaptation coefficient β using a preset lighting adaptation evaluation formula: β = |E / 50 - B / 100| / 10 + 1. Here, dividing the ambient light data E by 50 standardizes the ambient light to the recommended illuminance level, and dividing the screen brightness B by 100 standardizes the screen brightness to the recommended brightness level. The difference between the two, divided by 10 and added by 1, is used to control the adaptation coefficient within a reasonable range. For example, when the ambient light is 300 lux and the screen brightness is 200 nits, the lighting adaptation coefficient β = |300 / 50 - 200 / 100| / 10 + 1 = 1.4. When the ratio of ambient light to screen brightness is close to the recommended standard, the light adaptation coefficient is close to 1, indicating a suitable lighting environment. When the difference between the two is large, the light adaptation coefficient will increase accordingly, indicating that an unsuitable lighting environment will exacerbate eye strain. Finally, the system multiplies the eye strain index F by the light adaptation coefficient β to obtain the corrected eye strain index F' = F × β. This correction method based on lighting environment can more comprehensively assess the user's eye strain risk and provide a more accurate triggering basis for subsequent protective measures. Compared with prediction results based solely on visual load and usage time, the correction method considering light adaptation factors can better reflect the impact of the actual usage environment on eye strain.
[0056] Step 105: When the corrected visual fatigue index exceeds the fatigue index threshold, perform the preset telephoto training operation on the user.
[0057] The preset far-focus training operation is a visual accommodation training method actively triggered by the system when it detects a risk of eye strain. This operation aims to guide the user's gaze from near to far objects, helping the ciliary muscles of the eyes to relax appropriately and alleviating eye strain caused by prolonged close-range use. The far-focus training operation includes displaying a training guidance interface on the learning device screen, guiding the user to shift their gaze to distant objects and maintain fixation for a certain period, thus achieving active accommodation of the visual system. This training operation, by changing the user's viewing distance and focus, can effectively alleviate eye strain symptoms and prevent further aggravation of eye strain.
[0058] Specifically, during implementation, the system needs to monitor the corrected visual fatigue index in real time and compare it with a preset fatigue index threshold. When the corrected visual fatigue index exceeds the threshold, it indicates that the user's visual system is already in a state of high fatigue, requiring timely protective measures. The system automatically triggers a preset telephoto training operation, guiding the user to shift their focus from the near screen to a distant object to alleviate accommodative fatigue. Specifically, a telephoto training prompt interface is displayed on the learning device's screen, guiding the user to take visual breaks. This telephoto training triggering mechanism based on the visual fatigue index can intervene before the user's visual fatigue reaches a dangerous level, preventing further aggravation of visual fatigue through proactive visual accommodation training. Compared to rest reminders at fixed intervals, this intelligent triggering method based on the visual fatigue index is more targeted, determining the training timing according to the user's actual visual load, thus improving the protective effect.
[0059] Based on the above embodiments, as an optional embodiment, step 105, performing a preset telephoto training operation on the user, may further include the following steps: Step 401: Determine the intensity level of the telephoto training based on the corrected visual fatigue index, and select the corresponding virtual distant scene from the preset scene library based on the intensity level.
[0060] Specifically, the system first determines the intensity level of far-focus training based on the corrected visual fatigue index F'. Specifically, the system divides the visual fatigue index range into multiple intervals, each corresponding to a different training intensity level: when F'∈[3, 5), it corresponds to light training; when F'∈[5, 7), it corresponds to moderate training; and when F'≥7, it corresponds to intense training. After determining the intensity level, the system selects a corresponding virtual distant scene from a preset scene library. The scene library stores different types of distant scenes, such as natural landscapes and city skylines, each with a corresponding training intensity attribute. For example, when the visual fatigue index F'=6.5, the system will select a scene with moderate training intensity, such as a city distant scene containing multiple visual focal points and rich depth of field. This scene selection mechanism based on visual fatigue levels can provide more targeted training content according to the user's actual visual state.
[0061] Step 402: Adjust the screen imaging distance of the learning device to the preset long-distance focus range, and generate dynamic guide markers in the virtual distant scene. The movement trajectory of the dynamic guide markers is determined by the eye movement direction identified in the eye movement data.
[0062] Specifically, the system first needs to adjust the screen imaging distance of the learning device. By adjusting display parameters or using optical imaging devices, the visual focus of the screen is adjusted to a preset far-distance range, typically set to a visual equivalent distance of 3-6 meters. This far-distance focus setting can help relax the user's ciliary muscles and alleviate accommodative tension caused by prolonged close-range eye use. In the virtual far-view scene, the system generates dynamic guidance markers based on real-time collected eye movement data. Specifically, the system first extracts the eye movement direction θ from the eye movement data, which is obtained by calculating the displacement vector of the pupil center position. Then, the system designs the movement trajectory of the guidance markers based on the eye movement direction. For example, when the system detects that the user's eyeball is moving to the upper right (θ=45°), it generates a guidance marker that moves along the upper right trajectory, guiding the user's gaze to follow the marker while scanning the far-view scene. This dynamic guidance method based on eye movement characteristics makes the training process more in line with the user's natural visual habits, improving training comfort and acceptance. By combining screen imaging distance adjustment with dynamic guidance markers, the system can more effectively guide users in telephoto vision training, helping to alleviate symptoms of eye strain.
[0063] Based on the above embodiments, as an optional embodiment, step 402: generating dynamic guide markers in the virtual distant scene, this step may further include the following steps: Step 412: Based on eye movement data, determine the target guidance direction for eye movement.
[0064] Specifically, the system needs to analyze and process real-time collected eye movement data to extract characteristic information of eye movements in order to determine the target guidance direction. Specifically, the system calculates the principal direction component and movement tendency of the eye movement trajectory to obtain the target guidance direction. This direction determination method based on eye movement characteristics allows subsequent visual training to better align with the user's natural visual habits. For example, when the system detects an upward tendency in the user's eyes, it determines the target guidance direction as upward, conforming to the user's visual activity patterns. Determining the target guidance direction in this way reduces user discomfort during visual training and increases training acceptance.
[0065] Based on the above embodiments, as an optional embodiment, step 412: determining the target guidance direction of eye movement based on eye movement data, this step may further include the following steps: Step 4121: Reconstruct the trajectory of the eye movement data and extract the amplitude parameters and smoothness parameters of the eye movement in the horizontal, vertical and diagonal directions.
[0066] Specifically, the system first acquires the original coordinate sequence {(xi, yi, ti)} of the user's eye movements using an eye-tracking device. To obtain a continuous motion trajectory, the system uses a cubic spline interpolation algorithm to reconstruct the trajectory of these discrete points, resulting in a continuous trajectory function R(t) = [x(t), y(t)]. Subsequently, the system decomposes the reconstructed trajectory into three directional components: horizontal Rx(t) = x(t), vertical Ry(t) = y(t), and diagonal Rd(t) = [x(t) ± y(t)]. For each directional component, the system calculates its motion amplitude parameter A = {Ax, Ay, Ad} and motion smoothness parameter S = {Sx, Sy, Sd}. The motion amplitude parameter is obtained by calculating the maximum displacement in each direction: Ai = max|Ri(t) - Ri(t0)|, i ∈ {x, y, d}; the motion smoothness parameter includes the rate of change of velocity λi = |d²Ri(t) / dt²| and the pause frequency fi. This multi-dimensional parameter extraction method can comprehensively reflect the user's eye movement characteristics.
[0067] Step 4122: Compare the motion amplitude parameters with the preset standard motion amplitude range, and identify the direction where the motion amplitude is lower than the standard range as the target guidance direction for eye movement.
[0068] Specifically, the extracted motion amplitude parameters are compared and analyzed with preset standard motion amplitude ranges. The standard motion amplitude ranges are reference values obtained based on statistical analysis of a large amount of normal visual activity data: horizontal direction Ax_std = ±30°, vertical direction Ay_std = ±25°, and diagonal direction Ad_std = ±20°. The system calculates the amplitude ratio for each direction: ηi = Ai / Ai_std, i ∈ {x, y, d}. When the amplitude ratio ηi < 0.8 for a certain direction, it indicates that the motion amplitude in that direction is significantly lower than the standard level, and the system marks it as the target guidance direction. For example, when ηy = 0.6, it indicates insufficient motion ability in the vertical direction, and the vertical direction needs to be set as the training target. This comparison method based on standard ranges can assess the user's motion ability level in various directions.
[0069] Step 4123: Perform statistical analysis on motion smoothness parameters, and identify the directions in the motion smoothness parameters where the rate of change of motion speed is higher than the rate of change threshold or the frequency of pauses is higher than the frequency threshold as the target guidance direction for eye movement.
[0070] Specifically, the system performs in-depth statistical analysis of motion smoothness parameters. For each direction, the system calculates the mean μλi and standard deviation σλi of the rate of change of velocity, and the mean μfi and standard deviation σfi of the pause frequency. The system presets a velocity change rate threshold λ0 = 2.5 m / s³ and a pause frequency threshold f0 = 3 times / second. When the velocity change rate in a certain direction exceeds the threshold (λi > λ0) or the pause frequency exceeds the threshold (fi > f0), the system identifies that direction as the target guidance direction. For example, when the velocity change rate in the horizontal direction λx = 3.0 m / s³ > λ0, it indicates that the user's eye movements in the horizontal direction are not smooth enough, and the horizontal direction needs to be used as a training target. This motion quality-based analysis method can discover subtle problems in the user's eye movement control, providing a basis for developing precise training programs.
[0071] Step 422: Based on the target guidance direction, construct a motion trajectory template, which includes at least one of horizontal motion, vertical motion, diagonal motion, and circular motion.
[0072] Specifically, the system selects appropriate basic trajectory types from a pre-set motion trajectory template library and combines them to construct the trajectory based on the determined target guidance direction. The motion trajectory templates include: horizontal motion trajectory R1(t) = [x0 ± vt, y0], vertical motion trajectory R2(t) = [x0, y0 ± vt], diagonal motion trajectory R3(t) = [x0 ± vt, y0 ± vt], and circular motion trajectory R4(t) = [x0 + r × cos(ωt), y0 + r × sin(ωt)]. The system selects the appropriate trajectory type based on the target guidance direction and can combine multiple basic trajectories to form more complex motion paths. For example, when the target guidance direction is upward to the right, the system will preferentially select the diagonal motion trajectory R3(t) as the basic template. This direction-based trajectory construction method can provide standardized motion path guidance for subsequent visual training.
[0073] Step 432: Set the initial movement speed and dwell time of the dynamic guide marker according to the intensity level; generate the dynamic guide marker in the preset depth level of the virtual distant scene, and control the dynamic guide marker to perform directional movement according to the movement trajectory template.
[0074] Specifically, the initial motion parameters of the dynamic guide markers are first set according to the training intensity level. For light training, the initial motion speed is set to 5° / s, with a dwell time of 2 seconds; for moderate training, the initial speed is 8° / s, with a dwell time of 1.5 seconds; and for intense training, the initial speed is 12° / s, with a dwell time of 1 second. The system sets multiple depth levels in the virtual distant scene, typically including a near-field layer (3 meters), a mid-field layer (4.5 meters), and a distant layer (6 meters). The system generates dynamic guide markers within these preset depth levels and controls their movement according to the constructed motion trajectory template. For example, when using a circular motion trajectory R4(t), the system controls the guide markers to cyclically move between different depth levels, with the radius r varying with depth and the angular velocity ω determined based on the initial motion speed. This multi-layered spatial guidance method can better train the user's visual adjustment ability.
[0075] Step 442: Calculate the tracking accuracy of the dynamic guide marker for user eye tracking in real time. When the tracking accuracy is less than the accuracy threshold, reduce the movement speed of the dynamic guide marker and extend the dwell time. When the tracking accuracy is greater than or equal to the accuracy threshold, increase the movement speed of the dynamic guide marker and shorten the dwell time.
[0076] Specifically, the system needs to evaluate the user's training tracking performance in real time. Specifically, the system calculates the tracking accuracy P by measuring the spatial deviation and time delay between the user's eye position and the guide marker position. The formula for calculating tracking accuracy is: P = 1 - [(Δx² + Δy²)½ / D + Δt / T], where Δx and Δy are the spatial position deviations, D is the screen diagonal length, Δt is the time delay, and T is the standard response time (typically 200ms). When the tracking accuracy P is detected to be less than the preset accuracy threshold of 0.8, the system reduces the movement speed of the dynamic guide marker by 20% and increases the dwell time by 30%. For example, reducing the current speed from 8° / s to 6.4° / s increases the dwell time from 1.5 seconds to 1.95 seconds. Conversely, when the tracking accuracy P is greater than or equal to the accuracy threshold, the system increases the movement speed by 15% and reduces the dwell time by 20%. This dynamic adjustment mechanism based on tracking accuracy ensures that the training difficulty remains within the user's adaptive range, improving the targeting and effectiveness of the training.
[0077] Step 403: Monitor the user's eye movement response data in real time during telephoto training, and dynamically adjust the motion parameters and guidance frequency of the dynamic guide markers based on the eye movement response data.
[0078] Specifically, the system needs to collect and analyze the user's eye movement response data in real time during telephoto training. Specifically, the system collects parameters such as eye movement velocity (v), fixation duration (t), and pupil diameter (d) using an eye-tracking device. Based on this response data, the system evaluates the user's training engagement and visual accommodation ability: when the detected eye movement velocity (v) is lower than 80% of the guide marker's movement speed, the system reduces the guide marker's movement speed; when the fixation duration (t) is less than 70% of the expected value, the system extends the time the guide marker stays at each position; when the change in pupil diameter (d) is less than the expected range, the system increases the guide marker's movement amplitude. For example, when the detected eye movement velocity is 5° / s, and the set speed of the guide marker is 8° / s, the system adjusts the guide marker speed to 6° / s. Simultaneously, the system dynamically adjusts the guidance frequency, i.e., the time interval between guide marker appearances, based on the user's response. When the user can stably follow the current guidance rhythm, the system appropriately increases the guidance frequency; when the user experiences difficulty following, the system decreases the guidance frequency. This dynamic adjustment mechanism based on real-time response enables the training process to better adapt to individual differences and real-time conditions of users.
[0079] Step 404: When the user's eye movement response data is detected to reach the preset recovery standard, the screen imaging distance is gradually adjusted back to the standard focal length according to the preset step size, and eye protection prompts are displayed on the screen of the learning device.
[0080] Specifically, the first step is to determine whether the user's eye movement response data has reached the preset recovery standards. These standards include the following indicators: eye-tracking accuracy exceeding 90%, pupillary accommodation speed recovering to more than 1.2 times the pre-training level, and fixation stability improving by more than 20%. When these indicators are simultaneously met, the system considers the user's visual accommodation ability to have been effectively restored. Subsequently, the system begins to gradually adjust the screen imaging distance back to the standard focal length (approximately 40 cm) according to preset step sizes. Specifically, the system uses a segmented adjustment method: each adjustment step is 50 cm, with an interval of 10 seconds between adjacent adjustments. This gradual adjustment avoids sudden stress on the visual system. For example, when the imaging distance for telephoto training is 5 meters, the system will gradually adjust the imaging distance to 40 cm over 90 seconds through 9 adjustment steps. During the adjustment process, the system displays eye-care prompts on the learning device screen, including the current level of visual fatigue recovery and recommended rest periods, guiding the user to establish scientific eye habits. This training termination mechanism based on recovery criteria, combined with progressive focus adjustment, can ensure the stability of training results and reduce discomfort during the adaptation process of the visual system.
[0081] Reference Figure 2This application provides a farsightedness system based on multimodal scene perception, comprising: a data acquisition module, a visual fatigue index determination module, an index correction module, and a far-focus training module, wherein: The data acquisition module is used to acquire ambient light data, eye movement data, and screen interaction data during the user's use of the learning device; The visual fatigue index determination module is used to identify the user's learning scenario based on screen interaction data, and calculate the corresponding visual load quantification value based on the identified learning scenario and eye movement data; based on the visual load quantification value and the user's continuous eye use duration, it predicts the user's visual fatigue index within a preset time window. The index correction module is used to compare and analyze ambient light data with the screen brightness of the learning device, calculate the light adaptation coefficient, and correct the visual fatigue index based on the light adaptation coefficient. The telephoto training module is used to perform preset telephoto training operations on the user when the corrected visual fatigue index exceeds the fatigue index threshold.
[0082] Based on the above embodiments, the visual fatigue index determination module is also used to extract application type, page dwell time, scrolling frequency and click density from screen interaction data, and construct corresponding interaction feature vectors; calculate the matching degree between the interaction feature vectors and the scene feature vectors corresponding to each learning scene in the preset learning scene template library; and select the learning scene with the highest matching degree as the user's learning scene.
[0083] Based on the above embodiments, the visual fatigue index determination module is further used to extract blink frequency, fixation duration, pupil diameter change rate, and eye movement speed from eye movement data as eye fatigue feature parameters; obtain the corresponding scene weight coefficient from a preset weight configuration table according to the identified learning scene; perform weighted calculations on each eye fatigue feature parameter with the scene weight coefficient to obtain the weighted fatigue value corresponding to each eye fatigue feature parameter; normalize and sum the weighted fatigue values corresponding to each eye fatigue feature parameter to obtain the user's visual load quantification value.
[0084] Based on the above embodiments, the basic fatigue increment per unit time is calculated according to the visual load quantification value; the fatigue accumulation coefficient is determined based on the continuous eye use duration, wherein the fatigue accumulation coefficient is positively correlated with the continuous eye use duration; the basic fatigue increment is multiplied by the fatigue accumulation coefficient to obtain the actual fatigue increment; the actual fatigue increment is integrated over time within a preset time window to obtain the user's visual fatigue index within the preset time window.
[0085] Based on the above embodiments, the telephoto training module is also used to determine the intensity level of telephoto training according to the corrected visual fatigue index, and select the corresponding virtual distant scene from the preset scene library based on the intensity level; adjust the screen imaging distance of the learning device to the preset distance focus range, and generate dynamic guide markers in the virtual distant scene. The movement trajectory of the dynamic guide markers is determined by the eye movement direction identified in the eye movement data; monitor the user's eye movement response data in real time during the telephoto training process, and dynamically adjust the movement parameters and guidance frequency of the dynamic guide markers according to the eye movement response data; when the user's eye movement response data is detected to reach the preset recovery standard, gradually adjust the screen imaging distance back to the standard focal length according to the preset step size, and display eye protection prompt information on the screen of the learning device.
[0086] Based on the above embodiments, the telephoto training module is also used to determine the target guidance direction of eye movement based on eye movement data; construct a motion trajectory template based on the target guidance direction, the motion trajectory template including at least one of horizontal movement, vertical movement, diagonal movement and circular movement; set the initial movement speed and dwell time of the dynamic guide mark according to the intensity level; generate the dynamic guide mark in the preset depth level of the virtual distant scene, and control the dynamic guide mark to perform directional movement according to the motion trajectory template; calculate the following accuracy of the user's eye tracking of the dynamic guide mark in real time, and when the following accuracy is less than the accuracy threshold, reduce the movement speed of the dynamic guide mark and extend the dwell time, and when the following accuracy is greater than or equal to the accuracy threshold, increase the movement speed of the dynamic guide mark and shorten the dwell time.
[0087] Based on the above embodiments, the telephoto training module is also used to reconstruct the trajectory of eye movement data, extract the amplitude parameters and smoothness parameters of eye movement in the horizontal, vertical and diagonal directions; compare the amplitude parameters with the preset standard amplitude range, and identify the direction of the amplitude of movement below the standard range as the target guidance direction of eye movement; perform statistical analysis on the smoothness parameters, and identify the direction of the movement speed change rate above the change rate threshold or the pause frequency above the frequency threshold as the target guidance direction of eye movement.
[0088] It should be noted that the above embodiments of the apparatus are only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0089] This application also discloses an electronic device. (See reference...) Figure 3 , Figure 3 This is a schematic diagram of the structure of an electronic device disclosed in an embodiment of this application. The electronic device 300 may include: at least one processor 301, at least one network interface 304, a user interface 303, a memory 305, and at least one communication bus 302.
[0090] The communication bus 302 is used to enable communication between these components.
[0091] The user interface 303 may include a display interface and a camera interface. Optionally, the user interface 303 may also include a standard wired interface and a wireless interface.
[0092] The network interface 304 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface).
[0093] The processor 301 may include one or more processing cores. The processor 301 connects to various parts of the server using various interfaces and lines, and performs various server functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 305, and by calling data stored in the memory 305. Optionally, the processor 301 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 301 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface graphics, and applications; the GPU is responsible for rendering and drawing the content required for display; and the modem handles wireless communication. It is understood that the modem may also not be integrated into the processor 301 and may be implemented as a separate chip.
[0094] The memory 305 may include random access memory (RAM) or read-only memory. Optionally, the memory 305 may include a non-transitory computer-readable storage medium. The memory 305 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 305 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-described method embodiments, etc.; the data storage area may store data involved in the above-described method embodiments, etc. Optionally, the memory 305 may also be at least one storage device located remotely from the aforementioned processor 301. (Refer to...) Figure 3 The memory 305, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and an application program based on a multimodal scene perception farsight method.
[0095] exist Figure 3 In the illustrated electronic device 300, the user interface 303 is mainly used to provide an input interface for the user and acquire user input data; while the processor 301 can be used to call an application program stored in the memory 305 for a far-seeing method based on multimodal scene perception. When executed by one or more processors 301, the electronic device 300 performs one or more methods as described in the above embodiments. It should be noted that, for the foregoing method embodiments, for the sake of simplicity, they are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, because according to this application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0096] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0097] In the various embodiments provided in this application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some service interface; the indirect coupling or communication connection between apparatuses or units may be electrical or other forms.
[0098] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0099] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0100] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as USB flash drives, portable hard drives, magnetic disks, or optical disks.
[0101] The above are merely exemplary embodiments of this disclosure and should not be construed as limiting the scope of this disclosure. Any equivalent changes and modifications made in accordance with the teachings of this disclosure shall still fall within the scope of this disclosure. Other embodiments of this disclosure will be readily apparent to those skilled in the art upon consideration of the specification and practical disclosure.
[0102] This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not described in this disclosure. The specification and embodiments are to be considered exemplary only.
Claims
1. A farsight method based on multimodal scene perception, characterized in that, include: Acquire ambient light data, eye movement data, and screen interaction data during the user's use of the learning device; The user's learning scenario is identified based on the screen interaction data, and the corresponding visual load quantification value is calculated based on the identified learning scenario and the eye movement data. Based on the visual load quantification value and the user's continuous eye use duration, predict the user's visual fatigue index within a preset time window; The ambient light data is compared and analyzed with the screen brightness of the learning device to calculate the light adaptation coefficient, and the visual fatigue index is corrected based on the light adaptation coefficient. When the corrected visual fatigue index exceeds the fatigue index threshold, a preset telephoto training operation is performed on the user.
2. The farsightedness method based on multimodal scene perception according to claim 1, characterized in that, The process of identifying the user's learning scenario based on the screen interaction data includes: Extract the application type, page dwell time, scrolling frequency, and click density from the screen interaction data, and construct the corresponding interaction feature vector; Calculate the matching degree between the interaction feature vector and the scene feature vector corresponding to each learning scene in the preset learning scene template library; The learning scenario with the highest matching degree is selected as the user's learning scenario.
3. The farsightedness method based on multimodal scene perception according to claim 1, characterized in that, The calculation of the corresponding visual load quantification value based on the identified learning scene and the eye movement data includes: Blink frequency, fixation duration, pupil diameter change rate, and eye movement speed are extracted from the eye movement data as characteristic parameters of eye fatigue. Based on the identified learning scenario, the corresponding scenario weight coefficient is obtained from the preset weight configuration table; Each of the aforementioned eye fatigue feature parameters is weighted and calculated with the scene weight coefficient to obtain the weighted fatigue value corresponding to each of the aforementioned eye fatigue feature parameters; The weighted fatigue values corresponding to each of the aforementioned eye fatigue characteristic parameters are normalized and summed to obtain the quantified value of the user's visual load.
4. The farsightedness method based on multimodal scene perception according to claim 1, characterized in that, The step of predicting the user's visual fatigue index within a preset time window based on the visual load quantification value and the user's continuous eye use duration includes: Calculate the basic fatigue increment per unit time based on the aforementioned visual load quantification value; A fatigue accumulation coefficient is determined based on the continuous eye use duration, wherein the fatigue accumulation coefficient is positively correlated with the continuous eye use duration. Multiply the basic fatigue increment by the fatigue accumulation coefficient to obtain the actual fatigue increment; The actual fatigue increment is integrated over time within the preset time window to obtain the user's visual fatigue index within the preset time window.
5. The farsightedness method based on multimodal scene perception according to claim 1, characterized in that, The step of performing a preset telephoto training operation on the user includes: The intensity level of the telephoto training is determined based on the corrected visual fatigue index, and a corresponding virtual distant scene is selected from the preset scene library based on the intensity level. The screen imaging distance of the learning device is adjusted to a preset long-distance focus range, and a dynamic guide mark is generated in the virtual distant scene. The movement trajectory of the dynamic guide mark is determined by the eye movement direction identified in the eye movement data. The system monitors the user's eye movement response data during telephoto training in real time and dynamically adjusts the motion parameters and guidance frequency of the dynamic guidance marker based on the eye movement response data. When the user's eye movement response data is detected to reach the preset recovery standard, the screen imaging distance is gradually adjusted back to the standard focal length according to the preset step size, and eye protection prompts are displayed on the screen of the learning device.
6. The farsightedness method based on multimodal scene perception according to claim 5, characterized in that, The generation of dynamic guide markers in the virtual distant scene includes: Based on the eye movement data, the target guidance direction for eye movement is determined; Based on the target guidance direction, a motion trajectory template is constructed, which includes at least one of horizontal motion, vertical motion, diagonal motion, and circular motion. The initial movement speed and dwell time of the dynamic guide marker are set according to the intensity level; Dynamic guide markers are generated in a preset depth level of the virtual distant scene, and the dynamic guide markers are controlled to perform directional movement according to the motion trajectory template; The tracking accuracy of the user's eye tracking of the dynamic guide marker is calculated in real time. When the tracking accuracy is less than the accuracy threshold, the movement speed of the dynamic guide marker is reduced and the dwell time is extended. When the tracking accuracy is greater than or equal to the accuracy threshold, the movement speed of the dynamic guide marker is increased and the dwell time is shortened.
7. The farsightedness method based on multimodal scene perception according to claim 6, characterized in that, The step of determining the target guidance direction of eye movement based on the eye movement data includes: The eye movement data is reconstructed to extract the amplitude parameters and smoothness parameters of eye movement in the horizontal, vertical and diagonal directions; The motion amplitude parameter is compared with a preset standard motion amplitude range, and the direction where the motion amplitude is lower than the standard range is identified as the target guidance direction for eye movement. Statistical analysis is performed on the motion smoothness parameters, and the directions in which the rate of change of motion speed is higher than the rate of change threshold or the frequency of pauses is higher than the frequency threshold are identified as the target guidance directions for eye movements.
8. A farsight system based on multimodal scene perception, characterized in that, The system includes: The data acquisition module is used to acquire ambient light data, eye movement data, and screen interaction data during the user's use of the learning device; The visual fatigue index determination module is used to identify the user's learning scenario based on the screen interaction data, and calculate the corresponding visual load quantification value based on the identified learning scenario and the eye movement data; and predict the user's visual fatigue index within a preset time window based on the visual load quantification value and the user's continuous eye use duration. The index correction module is used to compare and analyze the ambient light data with the screen brightness of the learning device, calculate the light adaptation coefficient, and correct the visual fatigue index based on the light adaptation coefficient. The telephoto training module is used to perform preset telephoto training operations on the user when the corrected visual fatigue index exceeds the fatigue index threshold.
9. An electronic device, characterized in that, The device includes a processor, a memory, a user interface, and a network interface. The memory is used to store instructions, the user interface and the network interface are used to communicate with other devices, and the processor is used to execute the instructions stored in the memory to enable the electronic device to perform the far-seeing method based on multimodal scene perception as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed, perform the far-seeing method based on multimodal scene perception as described in any one of claims 1-7.