Vision health behavior intervention method, system and equipment based on reinforcement learning and medium

Through a reinforcement learning-based method, we collect user eye behavior data and individual characteristics, construct structured feature vectors, design a multi-dimensional operational behavior space, and use reinforcement learning algorithms to generate personalized intervention strategies. This solves the problems of traditional myopia prevention and control methods that ignore individual differences and lack operability, achieves precise and user-friendly vision health management, and improves the effectiveness of myopia prevention and control.

CN120708795APending Publication Date: 2025-09-2612 MM HEALTH TECH (HAINAN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510708145.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Traditional myopia prevention and control methods ignore individual differences, lack operability, make effect evaluation difficult, and the intervention strategies cannot be dynamically adjusted and lack data linkage, resulting in poor myopia prevention and control effects.

Method used

A reinforcement learning-based method is used to collect user eye behavior data and individual characteristics, construct a structured feature vector, design a multi-dimensional operational behavior space, and use the reinforcement learning algorithm to learn the mapping relationship between the state and the intervention action, generate an optimized intervention strategy, and combine the axial length change prediction model to build a reward function to provide personalized and operational vision health intervention suggestions, which are converted into user-understandable behavioral guidance through natural language templates.

Benefits of technology

It achieves personalized vision health intervention, improves the targetedness and effectiveness of intervention, enhances the feasibility and long-term persistence of user implementation, dynamically adjusts intervention strategies to improve the accuracy of effect evaluation and user compliance, and improves the effect of myopia prevention and control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708795A_ABST
    Figure CN120708795A_ABST
Patent Text Reader

Abstract

The invention provides a vision health behavior intervention method, system and device based on reinforcement learning and a medium, and the method comprises the steps: collecting eye use behavior data and individual features of a user, and constructing a structured feature vector as eye use behavior features after preprocessing; the method comprises the following steps of: pre-designing a multi-dimensional operable behavior space comprising an eye using posture, environment illumination, eye using scene switching, a time mode and spectrum balance, and carrying out discretization grading according to intervention intensity; a reinforcement learning algorithm is utilized to learn a mapping relation between the state of the user and an intervention action, an optimized intervention strategy is generated by maximizing an expected reward, and an eye axis change prediction model is adopted to construct a reward function; and abstract actions in the intervention strategy are converted into specific behavior guidance which can be understood by a user by using a natural language template, so that the problems that individual differences are neglected in vision prevention and control, suggestions are lack of operability, effect evaluation is difficult, the intervention strategy cannot be dynamically adjusted and data linkage is insufficient are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of the intersection of medical health and artificial intelligence, and in particular to a method, system, device and medium for visual health behavior intervention based on reinforcement learning. Background Art

[0002] In an era of deep integration between healthcare and artificial intelligence, myopia prevention and control has become a major public health issue that demands urgent resolution. In recent years, myopia rates have continued to rise, with the trend of younger and more severe myopia becoming increasingly severe, posing a serious threat to physical and mental health and future development.

[0003] Traditional myopia prevention and control methods suffer from widespread limitations. First, traditional prevention and control measures often rely on general guidelines based on demographic data. These one-size-fits-all vision care recommendations, such as "more distant vision, less near vision" and "two hours of outdoor activity daily," fail to tailor solutions to individual eye habits, living environments, and physiological characteristics. This results in significant discrepancies in the effectiveness of the same intervention measures across individuals. Second, existing prevention and control recommendations lack practical application and guidance. For example, simply advising users to "maintain a good eye distance" lacks practical, actionable standards tailored to each user, making it difficult for users to effectively implement and maintain long-term adherence. Furthermore, traditional intervention methods lack rapid feedback mechanisms and objective data support. Intervention effectiveness typically requires three to six months of follow-up. This lack of rapid feedback mechanisms results in ineffective interventions persisting for extended periods. Finally, users' eye behavior and vision are constantly changing, and traditional static intervention strategies cannot adapt to these changes. Furthermore, there is a lack of effective linkage between the extensive eye behavior data collected by wearable devices and intervention measures.

[0004] Therefore, a reinforcement learning-based vision health behavior intervention method is urgently needed to solve the technical problems of ignoring individual differences in vision prevention and control, lack of operability of suggestions, difficulty in effect evaluation, inability to dynamically adjust intervention strategies, and insufficient data linkage. Summary of the Invention

[0005] In order to overcome the problems existing in related technologies, the present disclosure provides a method, system, device and medium for vision health behavior intervention based on reinforcement learning to solve the technical problems of ignoring individual differences in vision prevention and control, lack of operability of suggestions, difficulty in effect evaluation, and inability to dynamically adjust intervention strategies and insufficient data linkage.

[0006] One or more embodiments of this specification provide a method for visual health behavior intervention based on reinforcement learning, including the following steps:

[0007] Collect users' eye behavior data and individual characteristics, and construct structured feature vectors after preprocessing as eye behavior features;

[0008] Pre-design a multi-dimensional actionable behavior space that includes eye posture, ambient lighting, eye scene switching, temporal mode, and spectral balance, and discretize and grade it according to the intensity of intervention;

[0009] A reinforcement learning algorithm is used to learn the mapping relationship between the user's state and intervention action, and an optimized intervention strategy is generated by maximizing the expected reward. The eye behavior characteristics are used as the state, the operable behavior is used as the intervention action, and the eye axis change prediction model is used to construct the reward function.

[0010] Natural language templates are used to transform the abstract actions in the intervention strategy into specific behavioral guidance that users can understand.

[0011] Preferably, the method further comprises the following steps:

[0012] Using a hierarchical reinforcement learning approach, the actionable behavior space decision-making in five dimensions is gradually refined and decomposed, including:

[0013] Identify the main dimensions requiring intervention;

[0014] Determine the intensity of the intervention on key dimensions;

[0015] Determine the dimensions of assistance and their intensity of intervention;

[0016] Combine to form a complete intervention action.

[0017] Preferably, it also includes implementing a progressive intervention strategy, specifically including the following steps:

[0018] Determine the user's current eye behavior baseline, determine the ideal behavior target based on the eye axis prediction model, and divide the change from the baseline to the target into N gradual transition stages. When the user achieves stable execution in the current stage and the duration exceeds the threshold, advance to the next stage until the Nth stage.

[0019] Preferably, the eye axis change prediction model is used to construct a reward function, and the reward function includes a direct behavior improvement reward, an eye axis control reward, a compliance reward, and a complexity penalty;

[0020] Calculating the direct behavior improvement reward based on the degree of improvement of the behavior indicator after the user executes the intervention strategy;

[0021] The eye axis prediction model is used to evaluate the effect of the current behavior on the eye axis and calculate the eye axis control reward;

[0022] calculating the compliance reward based on the user's degree of completion of the intervention recommendation;

[0023] The complexity penalty is obtained by calculating the intervention complexity.

[0024] Preferably, a multi-layer feedback evaluation mechanism is also included, specifically including short-term behavioral feedback, mid-term physiological indicator feedback, and long-term effect evaluation;

[0025] The short-term behavioral feedback is used to evaluate changes in eye behavior within 24 hours after the intervention;

[0026] The mid-term physiological indicator feedback is used to evaluate the effect of the intervention on the axial length based on the axial length prediction model;

[0027] The long-term effect evaluation is used to assess the long-term intervention effect through regular axial length measurement and refractive examination.

[0028] Preferably, the individual characteristics include the user's age, gender, family history of myopia, occupation and current axial length status.

[0029] One or more embodiments of this specification provide a vision health behavior intervention system based on reinforcement learning, including a data acquisition module, a behavior space construction module, an intervention strategy generation module, and an intervention module;

[0030] The data collection module is used to collect the user's eye behavior data and individual characteristics, and construct a structured feature vector after preprocessing as the eye behavior feature;

[0031] The behavior space construction module pre-designs a multi-dimensional operational behavior space that includes eye posture, ambient lighting, eye scene switching, time mode, and spectral balance, and discretizes and grades it according to the intervention intensity;

[0032] The intervention strategy generation module is used to use a reinforcement learning algorithm to learn the mapping relationship between the user's state and the intervention action, and to generate an optimized intervention strategy by maximizing the expected reward, wherein the eye behavior characteristics are used as the state, the operable behavior is used as the intervention action, and the eye axis change prediction model is used to construct the reward function;

[0033] The intervention module is used to convert the abstract actions in the intervention strategy into specific behavioral guidance that can be understood by the user using a natural language template.

[0034] One or more embodiments of the present specification provide a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned reinforcement learning-based vision health behavior intervention method when executing the computer program.

[0035] One or more embodiments of this specification provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the above-mentioned reinforcement learning-based vision health behavior intervention method.

[0036] The present disclosure provides a method, system, device and medium for intervention in vision health behavior based on reinforcement learning. The advantage is that by collecting the user's eye behavior data and individual characteristics, constructing a structured feature vector after pre-processing, it can provide exclusive vision health intervention suggestions based on the actual situation of each user, fully considering the differences between individuals in eye habits, living environment and physiological characteristics, so that the intervention measures are more in line with the actual needs of users, significantly improving the pertinence and effectiveness of intervention, effectively controlling excessive growth of the eye axis, and realizing a fundamental change from "one size fits all" to "tailor-made"; pre-designing a multi-dimensional operational behavior space, and discretizing and grading it according to the intervention intensity, structuring and quantifying intervention dimensions such as eye posture, ambient lighting, and eye scene switching, so that the intervention suggestions are no longer vague and general, but have clear implementation standards and specific action guidelines, which greatly enhances the user's The feasibility of implementation and long-term persistence solve the problem that traditional prevention and control recommendations are difficult to implement; the reinforcement learning algorithm is used to learn the mapping relationship between the user's status and intervention actions, and the optimized intervention strategy is generated by maximizing the expected reward. Among them, the axial length change prediction model is used to construct the reward function, which can continuously explore the intrinsic relationship between user behavior and axial length change, and dynamically adjust the intervention strategy according to the actual effect. It can not only accurately evaluate the intervention effect, but also continuously improve the accuracy and effectiveness of the intervention strategy, and avoid the long-term use of ineffective intervention measures; natural language templates are used to convert the abstract actions in the intervention strategy into specific behavioral guidance that users can understand, which lowers the user's understanding threshold for complex technical intervention plans, makes the intervention suggestions more understandable, and is easy for users to accept and implement, effectively improves the user's intervention compliance, and thus improves the practicality and user experience of the entire myopia prevention and control intervention system. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate one or more embodiments of this specification or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0038] Figure 1 A flowchart of a method for visual health behavior intervention based on reinforcement learning provided in one or more embodiments of this specification;

[0039] Figure 2A schematic diagram of a closed-loop intervention system for eye axis control based on reinforcement learning provided in one or more embodiments of this specification;

[0040] Figure 3 A schematic diagram of the structure of a vision health behavior intervention system based on reinforcement learning provided in one or more embodiments of this specification;

[0041] Figure 4 A schematic diagram of the structure of a computer device provided in one or more embodiments of this specification. DETAILED DESCRIPTION

[0042] In order to enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the technical solutions in one or more embodiments of this specification will be clearly and completely described below in conjunction with the drawings in one or more embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this invention document.

[0043] The present invention will be described in detail below with reference to specific implementation methods and the accompanying drawings.

[0044] Method Example

[0045] According to an embodiment of the present invention, a method for visual health behavior intervention based on reinforcement learning is provided, such as Figure 1 FIG. 1 is a flow chart of a method for visual health behavior intervention based on reinforcement learning according to an embodiment of the present invention. The method for visual health behavior intervention based on reinforcement learning according to an embodiment of the present invention includes the following steps:

[0046] S110. Collect the user's eye behavior data and individual characteristics, and construct a structured feature vector after preprocessing as the eye behavior feature.

[0047] Wearable devices are used to continuously collect users' actual eye behavior data and individual characteristics. The eye behavior data includes multi-dimensional indicators such as eye distance, angle, and ambient lighting. After preprocessing, a structured feature vector is constructed.

[0048] First, we constructed a multi-granularity feature signal system to systematically organize the characteristics of eye habits at different time scales. The system consists of four time granularity levels:

[0049] At the second-level granularity, each feature cycle (1 second) collects the following six dimensions of raw data: left eye angle, right eye angle, left eye distance, right eye distance, indoor and outdoor environment judgment (Boolean value), screen viewing status judgment (Boolean value)

[0050] At the minute-level granularity, statistical calculations are performed on the second-level data to construct a 16-dimensional feature vector:

[0051] Percentage of time the left eye is at normal distance = seconds the left eye is at normal distance / 60 seconds

[0052] Percentage of time the right eye is at normal distance = seconds the right eye is at normal distance / 60 seconds

[0053] Percentage of time the left eye is in normal viewing angle = number of seconds the left eye is in normal viewing angle / 60 seconds

[0054] The proportion of time the right eye is in normal viewing angle = the number of seconds the right eye is in normal viewing angle / 60 seconds

[0055] Percentage of outdoor time = seconds spent outdoors / 60 seconds

[0056] Screen time percentage = number of seconds spent looking at the screen / 60 seconds

[0057] The mean and variance of light intensity in different frequency bands (r, g, b, ir):

[0058] Average light intensity in R band = ∑(light intensity in R band) / 60

[0059] R band light intensity variance = sqrt(∑(R band light intensity - R band light intensity mean) 2 / 60)

[0060] Mean light intensity of G band = ∑(light intensity of G band) / 60

[0061] G band light intensity variance = sqrt(∑(G band light intensity - G band light intensity mean) 2 / 60)

[0062] Average light intensity of band B = ∑(light intensity of band B) / 60

[0063] B-band light intensity variance = sqrt(∑(B-band light intensity - B-band light intensity mean) 2 / 60)

[0064] Average IR band light intensity = ∑(IR band light intensity) / 60

[0065] IR band light intensity variance = sqrt(∑(IR band light intensity - IR band light intensity mean) 2 / 60)

[0066] Strobe mean = ∑(strobe value) / 60

[0067] Strobe variance = sqrt(∑(strobe value - strobe mean) 2 / 60)

[0068] At the hourly granularity, statistical calculations are performed on 60 minute-level data to construct a 16-dimensional feature vector with the same structure. At the daily granularity, statistical calculations are performed on 24 hour-level data to construct a 16-dimensional feature vector with the same structure.

[0069] Then the feature space is regularized: the multi-granularity features are normalized uniformly:

[0070] X_norm=(X-μ_X) / σ_X;

[0071] Where μ_X is the mean of the feature dimension, σ_X is the standard deviation, and X represents a 16-dimensional feature vector. Multi-granularity feature fusion:

[0072] The different granularity features are hierarchically integrated to form the core component of the state vector:

[0073] ●Minute-level time series window: Take the minute-level features of the last 60 minutes to form a matrix F_minute∈

[0074] R^(60×16)

[0075] ●Hourly time series window: Take the hourly features of the last 24 hours to form a matrix F_hour∈

[0076] R^(24×16)

[0077] ●Daily time series window: Take the daily features of the last 7 days to form a matrix F_day∈R^(7×16) Feature extractor: Process multi-granularity time series features

[0078] ●Minute-level LSTM: processes F_minute and outputs h_minute

[0079] Hourly LSTM: processes F_hour and outputs h_hour

[0080] ●Daily LSTM: processes F_day and outputs h_day

[0081] Multi-head attention layer: Fusion of feature representations of different granularities:

[0082] h_multi_scale=Multi Head Attention([h_minute, h_hour, h_day]);

[0083] Static feature encoder: Processing user static features

[0084] The user's individual feature vector contains:

[0085] Age (normalized) age_norm

[0086] Gender (one hot encoding) gender_one hot

[0087] Initial axial length (normalized) axial_length_norm

[0088] Myopia family history (binary) family_history

[0089] All static features are connected to form a vector F_static∈R^k (where k is the total dimension of the static features):

[0090] h_static = MLP(F_static);

[0091] Historical intervention encoder: handles historical intervention features

[0092] Encoding user history intervention execution:

[0093] Last intervention type (one-hot encoding) last_intervention_type

[0094] Last intervention compliance (normalized) last_compliance_rate

[0095] Last intervention effect score (normalized) last_effect_score

[0096] Concatenate to form a vector F_history∈R^m (where m is the total dimension of historical intervention features):

[0097] h_history = MLP(F_history);

[0098] The complete status is:

[0099] Feature Fuser: Fuse all features

[0100] h_fusion=Concat([h_multi_scale, h_static, h_history]);

[0101] The multi-granularity features, static features and historical intervention features are fused through the feature extractor and attention mechanism to obtain the final state vector:

[0102] S=f_fusion([F_minute, F_hour, F_day], F_static, F_history);

[0103] Among them, f_fusion is the fusion function, which consists of the attention network and the fully connected layer.

[0104] Q-value network: predict the Q-value of each action

[0105] Q(s,a)=MLP(h_fusion)[a];

[0106] S120. Pre-design a multi-dimensional operational behavior space that includes eye posture, ambient lighting, eye scene switching, time mode, and spectral balance, and discretize and grade it according to the intensity of intervention.

[0107] The system constructs a structured five-dimensional behavior space, as follows:

[0108] 1. Eye posture dimension (P):

[0109] P1: Maintain current eye posture

[0110] P2: Slight adjustment (eye distance +5cm, angle ±5°)

[0111] P3: Moderate adjustment (eye distance +10cm, angle ±5°)

[0112] 2. Ambient lighting dimension (L):

[0113] L1: Maintain the current lighting environment

[0114] L2: Slightly enhanced (increased by 100-300 lux)

[0115] L3: Moderate adjustment (adjust light source position, increase 500 lux)

[0116] 3. Scene switching dimension (S):

[0117] S1: Maintain the current scene

[0118] S2: Short rest (20 seconds of looking into the distance, 3 times per hour)

[0119] S3: Moderate rest (5 minutes of activity, once every hour)

[0120] 4. Time pattern dimension (T):

[0121] T1: Maintain the current time structure

[0122] T2: Pomodoro Technique (25 minutes of work + 5 minutes of rest)

[0123] T3: Long cycle alternation (45 minutes of eye use + 15 minutes of outdoor activities)

[0124] 5. Spectral balance dimension (B):

[0125] B1: Maintain current spectral exposure

[0126] B2: Reduce blue light (turn on eye protection mode, filter blue light)

[0127] B3: Increase exposure to natural light (+30 minutes per day)

[0128] Each intervention action is represented as a five-dimensional vector a = [p, l, s, t, b], where each component takes the value {1, 2, 3}. The complete behavior space contains 3^5 = 243 possible intervention combinations.

[0129] S130. Utilize a reinforcement learning algorithm to learn the mapping relationship between the user's state and the intervention action, and generate an optimized intervention strategy by maximizing the expected reward, wherein the eye behavior characteristics are used as the state, the operable behavior is used as the intervention action, and an eye axis change prediction model is used to construct a reward function.

[0130] The system uses reinforcement learning methods such as Q-learning or Policy Gradient to learn the mapping relationship between the user's eye behavior characteristics (learning state) and operational behaviors (intervention actions), and generates an optimized intervention strategy by maximizing the expected reward (eye axis control effect). Among them, the eye axis change prediction model is used to construct the reward function. The eye axis change prediction model is a multi-scale Transformer prediction model constructed in the patent (application number: 202510493118.0), such as Figure 2 , which is a schematic diagram of the closed-loop intervention system for eye axis control based on reinforcement learning provided in this embodiment.

[0131] Reward function design

[0132] The system reward function design utilizes the eye axis change prediction model and consists of the following components:

[0133] 1. Direct behavior improvement reward: calculated based on the degree of improvement in the user's behavior indicators after the intervention.

[0134] r_behavior=w_1*(metric_new-metric_old) / metric_old;

[0135] 2. Eye axis control reward: Use the eye axis prediction model to evaluate the impact of the current behavior on the eye axis

[0136] r_axial=w_2*(1-predicted_axial_growth / baseline_growth);

[0137] 3. Compliance Rewards: Calculated based on the user's completion of the intervention recommendations

[0138] r_adherence=w_3*completion_rate;

[0139] 4. Complexity Penalty: Preventing the system from generating overly complex intervention plans

[0140] r_complexity=-w_4*intervention_complexity;

[0141] The comprehensive rewards are:

[0142] r=r_behavior+r_axial+r_adherence+r_complexity;

[0143] Among them, r_behavior represents the direct behavior improvement reward, w_1 represents the behavior improvement reward weight, metric_new represents the behavioral indicator value after the intervention (calculated by the difference from the standard axial change), metric_old represents the behavioral indicator value before the intervention (calculated by the difference from the standard axial change), r_axial represents the axial control reward, w_2 represents the axial control reward weight, predicted_axial_growth represents the predicted axial growth, baseline_growth represents the baseline axial growth, r_adherence represents the adherence reward, w_3 represents the adherence reward weight, completion_rate represents the intervention recommendation completion rate, r_complexity represents the complexity penalty, w_4 represents the complexity penalty weight, intervention_complexity represents the intervention complexity, and r represents the comprehensive reward.

[0144] Initialization and learning process:

[0145] During the cold start phase when data is insufficient, the system initializes the Q-value function based on the knowledge of ophthalmologists:

[0146] Define state-action expert rules:

[0147] For example: If the eye distance is insufficient and the lighting is insufficient, first improve the lighting (L3) and then adjust the eye distance (P2).

[0148] Converted to initial Q value:

[0149] Q_init(s,a)=expert_score(rule_matches(s,a));

[0150] Q-learning update rule:

[0151] The system uses experience replay and target network technology to optimize Q-value learning:

[0152] Experience collection: Interaction data<s,a,r,s'> Store in playback buffer D

[0153] Batch sampling: randomly sample mini-batch B from D

[0154] Calculate the target Q value:

[0155] y_j=r_j+γ*max_a'Q'(s'_j,a')

[0156] Where Q' is the target network

[0157] Calculate the loss:

[0158] L(θ)=1 / |B|*∑_j(y_j-Q(s_j,a_j;θ))

[0159] Gradient update:

[0160]

[0161] Target network update: Update once every N steps

[0162] θ'=τθ+(1-τ)θ'

[0163] Among them, s represents the current state, a represents the current action, r represents the reward, s' represents the next state, Q(s,a) represents the state-action value function, Q'(s',a') represents the state-action value function of the target network, D represents the replay buffer, B represents the mini-batch sample, y_j represents the target Q value, γ represents the discount factor, L(θ) represents the loss function, θ represents the current network parameters, θ' represents the target network parameters, α represents the learning rate, τ = 0.01 represents the soft update coefficient, and N represents the target network update period.

[0164] Explore-Exploit Strategy

[0165] The system uses a decaying ε-greedy strategy to balance exploration and exploitation:

[0166] ε=ε_min+(ε_max-ε_min)*exp(-decay_rate*episode);

[0167] Among them, ε_max=0.9, ε_min=0.05, and decay_rate=0.001.

[0168] Optimization and personalization mechanism:

[0169] The system optimizes itself through continuous exploration and utilization of user feedback, establishes an intervention effect verification mechanism, and continuously adjusts and optimizes the intervention strategy by comparing the changes in the axial length before and after the intervention, thereby achieving continuous improvement in system performance.

[0170] In addition, the system incorporates a personalized adaptation mechanism that dynamically adjusts the behavior space and reward function based on individual characteristics such as user age and current eye length, ensuring that intervention recommendations are tailored to the user's actual situation. The system also incorporates a gradual intervention mechanism to avoid overly drastic behavioral adjustments and improve user acceptance and compliance.

[0171] S140: Using a natural language template, convert the abstract actions in the intervention strategy into specific behavioral guidance that can be understood by the user.

[0172] The system uses natural language template technology to convert the abstract intervention strategies generated by the reinforcement learning algorithm into specific and executable behavioral guidance suggestions, allowing users to easily understand and implement intervention plans.

[0173] Template conversion mechanism:

[0174] The system decomposes the five-dimensional vector a = [p, l, s, t, b] in the multi-dimensional actionable behavior space into specific actions in each dimension, and then generates complete intervention recommendations through natural language templates:

[0175] For example: vector a = [2,3,1,2,1] conversion result

[0176] "Please increase your reading distance by 5 cm" (Eye Posture P2)

[0177] "Adjust the position of the desk lamp to increase the brightness to above 500 lux" (ambient lighting L3)

[0178] "Use the Pomodoro Technique, work for 25 minutes and then take a 5-minute break" (Time Mode T2)

[0179] Contextual expression and personalized language adaptation:

[0180] The system makes contextual adjustments to expressions based on user characteristics and environmental constraints:

[0181] 1. For users of different age groups:

[0182] Child user: "Hey kids, let's play a game! Put the book a little further away than it is now, about the length of your hand, and then ask your parents to help turn up the lamp a little brighter."

[0183] User: "To protect your eyesight, please place the tablet 5 cm further away and increase the ambient light. This can be achieved by adjusting the position of the desk lamp."

[0184] Adult users: "It is recommended to adjust your work posture, place the monitor about an arm's length away from your eyes, and increase the ambient light to above 500 lux. Consider adding an auxiliary desk lamp."

[0185] 2. For different scenarios:

[0186] Study scenario: "It is recommended to take a break every 25 minutes while studying. During the break, look into the distance or do eye exercises."

[0187] Entertainment scenario: "When using electronic devices for entertainment, it is recommended to turn on eye protection mode and limit continuous use to no more than 45 minutes."

[0188] 3. Adjust based on user historical compliance:

[0189] For users with high adherence: provide more complex and comprehensive intervention recommendations

[0190] For users with low compliance: Simplify the recommendations and focus on the most important 1-2 improvement points

[0191] The method provided in this embodiment collects the user's eye behavior data and individual characteristics, constructs a structured feature vector after pre-processing, and can provide exclusive vision health intervention suggestions based on the actual situation of each user. It fully considers the differences between individuals in eye habits, living environment and physiological characteristics, so that the intervention measures are more in line with the actual needs of users, significantly improves the pertinence and effectiveness of the intervention, effectively controls the excessive growth of the eye axis, and realizes the transformation from "one size fits all" to "tailor-made" vision health management; pre-designs a multi-dimensional operational behavior space including eye posture, ambient lighting, eye scene switching, time mode and spectral balance, and discretizes and grades it according to the intervention intensity, structures and quantifies the intervention dimensions such as eye posture, ambient lighting, eye scene switching, etc., so that the intervention suggestions are no longer general and vague, but have clear implementation standards and specific action guidelines, which greatly enhances the feasibility and long-term effectiveness of user implementation. Long-term persistence solves the problem that traditional prevention and control recommendations are difficult to implement; a reinforcement learning algorithm is used to learn the optimal mapping relationship between user eye behavior characteristics (states) and actionable behaviors (intervention actions), and an optimized intervention strategy is generated by maximizing expected rewards. Among them, an axial length change prediction model is used to construct a reward function, which can continuously explore the intrinsic relationship between user behavior and axial length change, and dynamically adjust the intervention strategy according to the actual effect. It can not only accurately evaluate the intervention effect, but also continuously improve the accuracy and effectiveness of the intervention strategy, and avoid the long-term use of ineffective intervention measures; a natural language template is used to convert the abstract actions in the intervention strategy into specific behavioral guidance that users can understand, which lowers the user's understanding threshold for complex technical intervention plans, makes the intervention suggestions more understandable, and is easy for users to accept and implement, effectively improves the user's intervention compliance, and thus improves the practicality and user experience of the entire myopia prevention and control intervention system.

[0192] Actionable behavior space

[0193] In one embodiment, the actionable behavior space includes five dimensions: eye posture, ambient lighting, eye scene switching, temporal mode, and spectral balance. Each dimension is designed with specific actionable behavior options and is discretely graded according to the intervention intensity, as follows:

[0194] Eye posture dimension (P):

[0195] P1: Maintain current eye posture

[0196] P2: Slight adjustment (eye distance +5cm, angle +-5°)

[0197] P3: Moderate adjustment (eye distance +10cm, angle +-5°)

[0198] Ambient lighting dimension (L):

[0199] L1: Maintain the current lighting environment

[0200] L2: Slightly enhanced (increased by 100-300 lux)

[0201] L3: Moderate adjustment (adjust light source position, increase 500 lux)

[0202] Scene switching dimension (S):

[0203] S1: Maintain the current scene

[0204] S2: Short rest (20 seconds of looking into the distance, 3 times per hour)

[0205] S3: Moderate rest (5 minutes of activity, once every hour)

[0206] Time pattern dimension (T):

[0207] T1: Maintain the current time structure

[0208] T2: Pomodoro Technique (25 minutes of work + 5 minutes of rest)

[0209] T3: Long cycle alternation (45 minutes of eye use + 15 minutes of outdoor activities)

[0210] Spectral balance dimension (B):

[0211] B1: Maintain current spectral exposure

[0212] B2: Reduce blue light (turn on eye protection mode, filter blue light)

[0213] B3: Increase exposure to natural light (+30 minutes per day)

[0214] Action code:

[0215] Each intervention action is represented as a five-dimensional vector a = [p, l, s, t, b], where each component takes the value {1, 2, 3}. The complete behavior space contains 3^5 = 243 possible intervention combinations.

[0216] The method provided in this embodiment divides the operational behavior space into five dimensions, such as eye posture, and sets specific behavioral options to solve the problem that traditional prevention and control recommendations are difficult to implement and improve operability. By discretely grading the intervention intensity, it adapts to the needs of different users and provides standardized data for reinforcement learning, helping to accurately explore the optimal strategy, promote intelligent and precise myopia prevention and control, and enhance the effectiveness of intervention.

[0217] Hierarchical reinforcement learning method

[0218] In one embodiment, the following steps are further included:

[0219] A hierarchical reinforcement learning method is used to decompose the actionable behavior space decision-making in five dimensions into a gradually refined process, including: determining the main dimension that needs intervention, determining the intervention intensity on the main dimension, determining the auxiliary dimension and its intervention intensity, and combining them to form a complete intervention action.

[0220] The specific process of decomposing the five-dimensional behavior space decision into gradually refined ones includes:

[0221] Level 1: Identify the main dimensions requiring intervention

[0222] dimension_q=DQN_dimension(state);

[0223] primary_dimension=argmax(dimension_q);

[0224] Level 2: Determine the intensity of intervention on the main dimensions

[0225] intensity_q=DQN_intensity(state, primary_dimension);

[0226] primary_intensity=argmax(intensity_q);

[0227] The third level: determine the auxiliary dimension and its intervention intensity

[0228] secondary_q=DQN_secondary(state, primary_dimension, primary_intensity);

[0229] secondary_actions=top_F(secondary_q, F=2);

[0230] Combined to form a complete intervention action:

[0231] a=combine(primary_dimension, primary_intensity, secondary_actions);

[0232] The method provided in this embodiment uses a hierarchical reinforcement learning method to decompose the five-dimensional operational behavior space. By determining the main intervention dimension and intensity, auxiliary dimension and intensity, and combining them to form a complete intervention action, it can not only focus on key issues to improve intervention efficiency, accurately control the intensity to balance the effect and burden, but also achieve multi-dimensional collaborative optimization, flexibly customize personalized plans, and comprehensively improve the effectiveness and pertinence of myopia prevention and control.

[0233] Progressive intervention strategy

[0234] In one embodiment, the method further includes implementing a progressive intervention strategy, specifically comprising the following steps:

[0235] A gradual intervention strategy is adopted to avoid overly drastic behavioral adjustments and improve user acceptance and compliance. First, a baseline behavioral assessment is conducted to determine the user's current eye behavior baseline. Then, based on the eye axis prediction model, the ideal behavioral target is determined. The change from baseline to target is divided into N gradual transition stages:

[0236] intervention_i=baseline+(target-baseline)*(i / N);

[0237] When the user reaches stable execution in the current stage and the duration exceeds the threshold, it advances to the next stage until the Nth stage.

[0238] Intervention recommendations are presented in the following ways:

[0239] The system presents intervention suggestions in various forms based on the user's device and scenario:

[0240] Real-time reminders: When unhealthy eye behavior is detected, a gentle reminder will be issued through the wearable device.

[0241] Scheduled push: Personalized eye care suggestions will be pushed at fixed time every day.

[0242] Progress feedback: Regularly demonstrate the improvement of user behavior and eye movement control effects.

[0243] Visual display: Intuitively display the improvement trend of users' eye behavior through charts.

[0244] Using natural language templates to translate abstract actions into user-friendly, concrete behavioral guidance lowers the barrier for users to understand complex technical interventions, making intervention recommendations more accessible and easier for users to accept and implement. This effectively improves user compliance, thereby enhancing the practicality and user experience of the entire myopia prevention and control intervention system. Combined with a progressive intervention strategy, this not only allows users to easily understand the intervention content but also allows them to gradually cultivate healthy eye habits, ultimately achieving long-term and effective myopia prevention and control results.

[0245] The method provided in this embodiment effectively reduces the difficulty for users to understand and implement the intervention strategy by implementing a progressive intervention strategy, thereby enhancing user acceptance. Ideal behavioral goals are determined based on the eye axis prediction model, giving the intervention a scientific basis and clear direction. The goal achievement process is divided into N gradual transition stages, and stable execution and duration thresholds are set. This step-by-step approach fully considers the laws of user behavior change, avoiding user fear due to overly high goals or excessive changes. It can help users gradually develop healthy eye habits over a longer period of time, enhance the long-term effectiveness and sustainability of the intervention strategy, and effectively improve users' eye health.

[0246] Reward function construction

[0247] In one embodiment, the eye axis change prediction model is used to construct a reward function, and the reward function includes direct behavior improvement reward, eye axis control reward, compliance reward and complexity penalty.

[0248] The direct behavior improvement reward is based on the degree of improvement in the user's behavior indicators after the user implements the intervention strategy:

[0249] r_behavior=w_1*(metric_new-metric_old) / metric_old;

[0250] Among them, r_behavior represents the direct behavior improvement reward, w_1 represents the behavior improvement reward weight, 0.3, metric_new represents the behavior indicator value after intervention, metric_old represents the behavior indicator value before intervention, and metric includes key indicators such as the proportion of normal eye distance and lighting qualification rate.

[0251] The eye axis control reward uses the eye axis prediction model to evaluate the impact of the current behavior on the eye axis:

[0252] r_axial=w_2*(1-predicted_axial_growth / baseline_growth);

[0253] Among them, r_axial represents the axial control reward, w_2 represents the axial control reward weight, which is 0.4, predicted_axial_growth represents the predicted axial growth, baseline_growth represents the baseline axial growth, and predicted_axial_growth is the expected axial growth under the current behavior given by the axial prediction model.

[0254] The compliance reward is based on the user's completion of the intervention recommendations:

[0255] r_adherence=w_3*completion_rate;

[0256] Completion_rate calculation method:

[0257] completion_rate=Σ(action_i*weight_i*completion_i) / Σ(action_i*weight_i);

[0258] Where r_adherence represents the adherence reward, w_3 represents the adherence reward weight, which is 0.15, completion_rate represents the intervention completion rate, action_i represents whether the i-th dimension is selected as the intervention dimension (1 for yes, 0 for no), weight_i represents the importance weight of the i-th dimension (determined based on the eye axis prediction model), and completion_i represents the execution completion of the i-th dimension intervention, which is calculated as follows:

[0259] Eye posture dimension (P): the improvement ratio of the normal eye posture duration

[0260] completion_P=(normal_posture_ratio_after-normal_posture_ratio_before) / (target_ratio-normal_posture_ratio_before);

[0261] Among them, completion_P represents the execution completion of the eye posture dimension.

[0262] normal_posture_ratio_after represents the normal posture ratio after intervention.

[0263] normal_posture_ratio_before represents the normal posture time ratio before intervention, and target_ratio represents the target posture time ratio.

[0264] Ambient lighting dimension (L): the improvement ratio of the duration of lighting meeting the standard

[0265] completion_L=(adequate_light_ratio_after-adequate_light_ratio_before) / (target_ratio-adequate_light_ratio_before);

[0266] Among them, completion_L represents the execution completion degree of the ambient lighting dimension, adequate_light_ratio_after represents the proportion of time that the lighting meets the standard after intervention, and adequate_light_ratio_before represents the proportion of time that the lighting meets the standard before intervention.

[0267] Scene switching dimension (S): rest frequency ratio

[0268] completion_S=actual_breaFs / recommended_breaFs;

[0269] Among them, completion_S represents the execution completion of the scene switching dimension, actual_breaFs represents the actual number of breaks, and recommended_breaFs represents the recommended number of breaks.

[0270] Time pattern dimension (T): Time pattern follows proportion

[0271] completion_T=time_following_pattern / total_activity_time;

[0272] Among them, completion_T represents the execution completion degree of the time pattern dimension, time_following_pattern represents the duration of following the time pattern, and total_activity_time represents the total activity duration.

[0273] Spectral balance dimension (B): Spectral adjustment execution ratio

[0274] completion_B=time_with_adjusted_spectrum / total_screen_time;

[0275] Among them, completion_B represents the execution completion of the spectrum balance dimension, time_with_adjusted_spectrum represents the duration of spectrum adjustment, and total_screen_time represents the total screen usage time.

[0276] The complexity penalty is used to prevent the system from generating overly complex intervention plans:

[0277] r_complexity=-w_4*intervention_complexity;

[0278] intervention_complexity=Σ(action_level_i-1) / max_possible_complexity;

[0279] Among them: r_complexity represents complexity penalty, w_4 represents complexity penalty weight of 0.15, intervention_complexity represents intervention complexity, action_level_i represents the level of action in the i-th dimension, max_possible_complexity represents the maximum possible complexity, the value is 10, action_level_i represents the level of action in the i-th dimension (from 1 to 3), max_possible_complexity = 5*3 = 15 (the maximum complexity when all five dimensions select the highest level 4).

[0280] In addition, the conditional complexity of the intervention is also considered:

[0281] conditional_complexity=number_of_conditions / max_conditions;

[0282] Among them, conditional_complexity represents conditional complexity, max_conditions represents the maximum number of conditions, and number_of_conditions represents the number of conditions (such as "if...then...") included in the intervention recommendation.

[0283] The final complexity is:

[0284] intervention_complexity=0.7*action_complexity+0.3*conditional_complexity;

[0285] The comprehensive rewards are:

[0286] r=r_behavior+r_axial+r_adherence+r_complexity;

[0287] Among them, the weights w_1~w_4 are set to 0.3, 0.4, 0.15, and 0.05 respectively through initial experiments.

[0288] The method provided in this embodiment, by constructing a reward function that includes a variety of reward and punishment mechanisms, incentivizes and constrains the generation and execution of intervention strategies from multiple dimensions. Direct behavior improvement rewards focus on the improvement of user behavior indicators, allowing users to instantly perceive the positive feedback of their own behavior changes; axial length control rewards use axial length prediction models to scientifically measure the impact of current behavior on axial length, guiding users to pay attention to long-term eye health; compliance rewards are based on completion, effectively stimulating users' enthusiasm for implementing intervention suggestions; complexity penalties avoid the system from generating complicated intervention plans, ensuring the operability and practicality of the plans. The synergistic effect of multiple mechanisms can not only promote users to actively improve their eye behavior and improve the implementation effect of intervention strategies, but also ensure the rationality and effectiveness of intervention plans, ultimately achieving scientific management and precise intervention of users' eye health.

[0289] In one embodiment, a multi-layer feedback evaluation mechanism is further included, specifically including short-term behavioral feedback, mid-term physiological indicator feedback, and long-term effect evaluation;

[0290] The short-term behavioral feedback is used to evaluate changes in eye behavior within 24 hours after the intervention:

[0291] Before and after comparison indicators:

[0292] Normal eye distance time ratio change rate Δd = (d_after - d_before) / d_before;

[0293] Normal lighting environment duration ratio change rate Δl = (l_after - l_before) / l_before;

[0294] The change rate of outdoor activity duration Δo = (o_after - o_before) / o_before;

[0295] Short-term score calculation:

[0296] score_short=w_d*Δd+w_l*Δl+w_o*Δo;

[0297] The mid-term physiological index feedback is used to evaluate the impact of intervention on axial length based on the axial length prediction model: Predicting axial length changes:

[0298] pred_axial_change=axial_prediction_model(behavior_features);

[0299] Among them, pred_axial_change represents the predicted axial change, and behavior_features represents behavioral features.

[0300] Eye axis control effect score:

[0301] score_axial=1-pred_axial_change / baseline_change.

[0302] Among them, score_axial represents the axial control effect score, and baseline_change represents the baseline axial change.

[0303] The long-term effect evaluation is used to evaluate the long-term intervention effect through regular axial length measurement and refractive examination:

[0304] Long-term intervention effects were assessed by regular axial length measurement and refractive examinations:

[0305] Axial eye growth rate: Comparison of axial eye growth rate before and after intervention

[0306] growth_reduction=(growth_rate_before-growth_rate_after) / growth_rate_before;

[0307] Among them, growth_reduction represents the reduction ratio of the axial growth rate, growth_rate_before represents the axial growth rate before intervention, and growth_rate_after represents the axial growth rate after intervention.

[0308] Myopia progression reduction rate:

[0309] myopia_control_rate=(progression_before-progression_after) / progression_before;

[0310] Among them, myopia_control_rate represents the myopia progression slowing rate, progression_before represents the myopia progression speed before intervention, and progression_after represents the myopia progression speed after intervention.

[0311] The method provided in this embodiment builds a comprehensive monitoring system covering the short, medium and long term. Short-term behavioral feedback can timely capture changes in eye behavior within 24 hours after intervention, allowing users to quickly understand the effects of their own behavioral adjustments and facilitate timely correction of deviations; mid-term physiological indicator feedback uses the axial length prediction model to scientifically predict the impact of intervention on axial length development and provide early warning of potential risks; long-term effect evaluation uses regular axial length measurement and refractive examinations to intuitively present the results of long-term intervention with actual data, providing a reliable basis for the optimization of subsequent intervention strategies. Multi-dimensional feedback evaluation cooperates with each other, which can not only track the user's status in real time and enhance the user's sense of control and confidence in the intervention process, but also ensure the dynamic adjustment and continuous optimization of the intervention strategy, thereby effectively improving the accuracy and effectiveness of eye health management.

[0312] System Example

[0313] According to an embodiment of the present invention, a vision health behavior intervention system based on reinforcement learning is provided. Figure 3 As shown, this is a structural diagram of the vision health behavior intervention system based on reinforcement learning provided in this embodiment. The vision health behavior intervention system based on reinforcement learning according to the embodiment of the present invention includes a data acquisition module 31, a behavior space construction module 32, an intervention strategy generation module 33 and an intervention module 34.

[0314] The data collection module 31 is used to collect the user's eye behavior data and individual characteristics, and construct a structured feature vector after preprocessing as the eye behavior feature.

[0315] The behavior space construction module 32 is used to pre-design a multi-dimensional operational behavior space including eye posture, ambient lighting, eye scene switching, time mode and spectral balance, and to discretize and grade it according to the intervention intensity.

[0316] The intervention strategy generation module 33 is used to use the reinforcement learning algorithm to learn the mapping relationship between the user's state and the intervention action, and to generate an optimized intervention strategy by maximizing the expected reward, wherein the eye behavior characteristics are used as the state, the operable behavior is used as the intervention action, and the eye axis change prediction model is used to construct the reward function.

[0317] The intervention module 34 is configured to use a natural language template to convert the abstract actions in the intervention strategy into specific behavioral guidance that can be understood by the user.

[0318] The system provided by this embodiment collects the user's eye behavior data and individual characteristics through the data collection module 31, and constructs a structured feature vector after pre-processing. It can provide exclusive vision health intervention suggestions based on the actual situation of each user, and fully considers the differences between individuals in eye habits, living environment and physiological characteristics, so that the intervention measures are more in line with the actual needs of users, significantly improve the pertinence and effectiveness of intervention, effectively control excessive growth of the eye axis, and realize the transformation from "one size fits all" to "tailor-made" vision health management; the behavior space construction module 32 pre-designs a multi-dimensional operational behavior space including eye posture, ambient lighting, eye scene switching, time mode and spectral balance, and discretizes and grades it according to the intervention intensity, and structures and quantifies the intervention dimensions such as eye posture, ambient lighting, eye scene switching, etc., so that the intervention suggestions are no longer vague and general, but have clear implementation standards and specific action guidelines, which greatly enhances the feasibility and effectiveness of user execution. Long-term persistence solves the problem that traditional prevention and control recommendations are difficult to implement; the intervention strategy generation module 33 uses a reinforcement learning algorithm to learn the mapping relationship between the user's eye behavior characteristics (state) and operational behaviors (intervention actions), and generates an optimized intervention strategy by maximizing the expected reward. Among them, the axial length change prediction model is used to construct a reward function, which can continuously explore the intrinsic relationship between user behavior and axial length change, and dynamically adjust the intervention strategy according to the actual effect. It can not only accurately evaluate the intervention effect, but also continuously improve the accuracy and effectiveness of the intervention strategy, and avoid the long-term use of ineffective intervention measures; the intervention module 34 uses a natural language template to convert the abstract actions in the intervention strategy into specific behavioral guidance that users can understand, which lowers the user's understanding threshold for complex technical intervention plans, makes the intervention suggestions more understandable, and is easy for users to accept and implement, effectively improves the user's intervention compliance, and thus improves the practicality and user experience of the entire myopia prevention and control intervention system.

[0319] In one embodiment, a decomposition module is further included, which is used to decompose the actionable behavior space decision of the five dimensions into a step-by-step refinement process by using a hierarchical reinforcement learning method, specifically including:

[0320] Identify the main dimensions that require intervention, determine the intervention intensity on the main dimensions, determine the auxiliary dimensions and their intervention intensity, and combine them to form a complete intervention action.

[0321] The embodiment of the present invention is a system embodiment corresponding to the above-mentioned method embodiment. The specific operations of the processing steps of each module can be understood by referring to the description of the method embodiment, and will not be repeated here.

[0322] like Figure 4As shown, the present invention also provides a computer-readable storage medium having a computer program stored thereon, which implements the vision health behavior intervention method based on reinforcement learning in the above-mentioned embodiment when the computer program is executed by a processor, or implements the vision health behavior intervention method based on reinforcement learning in the above-mentioned embodiment when the computer program is executed by a processor.

[0323] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (SynchlinF) DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct RAMbus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.

[0324] Each embodiment in this specification is described in a progressive manner. The same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device or system embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, refer to the partial description of the method embodiments. The device and system embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the scheme of this embodiment. A person of ordinary skill in the art can understand and implement it without making any creative efforts.

[0325] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present invention, and the contents not described in detail in the specification of the present invention belong to the common knowledge of those skilled in the art.

Claims

1. A vision health behavior intervention method based on reinforcement learning, characterized in that: The following steps are involved: Collect users' eye behavior data and individual characteristics, and construct structured feature vectors after preprocessing as eye behavior features; Pre-design a multi-dimensional actionable behavior space that includes eye posture, ambient lighting, eye scene switching, temporal mode, and spectral balance, and discretize and grade it according to the intensity of intervention; A reinforcement learning algorithm is used to learn the mapping relationship between the user's state and intervention action, and an optimized intervention strategy is generated by maximizing the expected reward. The eye behavior characteristics are used as the state, the operable behavior is used as the intervention action, and the eye axis change prediction model is used to construct the reward function. Natural language templates are used to transform the abstract actions in the intervention strategy into specific behavioral guidance that users can understand.

2. The method for visual health behavior intervention based on reinforcement learning according to claim 1, characterized in that: The following steps are also included: Using a hierarchical reinforcement learning approach, the actionable behavior space decision-making in five dimensions is gradually refined and decomposed, including: Identify the main dimensions requiring intervention; Determine the intensity of the intervention on key dimensions; Determine the dimensions of assistance and their intensity of intervention; Combine to form a complete intervention action.

3. The method for visual health behavior intervention based on reinforcement learning according to claim 1, characterized in that: A progressive intervention strategy was also designed, which included the following steps: Determine the user's current eye behavior baseline, determine the ideal behavior target based on the eye axis prediction model, and divide the change from the baseline to the target into N gradual transition stages. When the user achieves stable execution in the current stage and the duration exceeds the threshold, advance to the next stage until the Nth stage.

4. The method for visual health behavior intervention based on reinforcement learning according to claim 1, characterized in that: The reward function is constructed by using the eye axis change prediction model, and the reward function includes direct behavior improvement reward, eye axis control reward, compliance reward and complexity penalty; Calculating the direct behavior improvement reward based on the degree of improvement of the behavior indicator after the user executes the intervention strategy; The eye axis prediction model is used to evaluate the effect of the current behavior on the eye axis and calculate the eye axis control reward; calculating the compliance reward based on the user's degree of completion of the intervention recommendation; The complexity penalty is obtained by calculating the intervention complexity.

5. The method for visual health behavior intervention based on reinforcement learning according to claim 1, characterized in that: It also includes a multi-layered feedback evaluation mechanism, specifically including short-term behavioral feedback, mid-term physiological indicator feedback, and long-term effect evaluation; The short-term behavioral feedback is used to evaluate changes in eye behavior within 24 hours after the intervention; The mid-term physiological indicator feedback is used to evaluate the effect of the intervention on the axial length based on the axial length prediction model; The long-term effect evaluation is used to assess the long-term intervention effect through regular axial length measurement and refractive examination.

6. The method for visual health behavior intervention based on reinforcement learning according to claim 1, characterized in that: The individual characteristics include user age, gender, family history of myopia, occupation and current axial length status.

7. A vision health behavior intervention system based on reinforcement learning, characterized in that: It includes data collection module, behavior space construction module, intervention strategy generation module and intervention module; The data collection module is used to collect the user's eye behavior data and individual characteristics, and construct a structured feature vector after preprocessing as the eye behavior feature; The behavior space construction module is used to pre-design a multi-dimensional operational behavior space that includes eye posture, ambient lighting, eye scene switching, temporal mode, and spectral balance, and discretize and grade it according to the intervention intensity; The intervention strategy generation module is used to use a reinforcement learning algorithm to learn the mapping relationship between the user's state and the intervention action, and to generate an optimized intervention strategy by maximizing the expected reward, wherein the eye behavior characteristics are used as the state, the operable behavior is used as the intervention action, and the eye axis change prediction model is used to construct the reward function; The intervention module is used to convert the abstract actions in the intervention strategy into specific behavioral guidance that can be understood by the user using a natural language template.

8. The vision health behavior intervention system based on reinforcement learning according to claim 7, characterized in that: It also includes a decomposition module for using a hierarchical reinforcement learning method to decompose the decision-making of the five-dimensional actionable behavior space into a gradually refined process, including: Identify the main dimensions requiring intervention; Determine the intensity of the intervention along key dimensions; Determine the dimensions of assistance and their intensity of intervention; Combine to form a complete intervention action.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the vision health behavior intervention method based on reinforcement learning as described in any one of claims 1 to 6 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the vision health behavior intervention method based on reinforcement learning as described in any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Eye using parameter-based eye axis length change prediction method and device and medium

    CN120514319A