Method and system for identifying risk factors related to asthenopia and constructing risk prediction model
By acquiring and cleaning visual fatigue data, combining Logistic regression and random forest models, identifying and predicting risk factors of visual fatigue, the problem of difficult to effectively identify and predict visual fatigue risks in the prior art is solved, and the scientific research and prevention level of visual fatigue is improved.
Patent Information
- Application Number
- CN202411543533.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-31
- Publication Date
- 2025-05-02
AI Technical Summary
The prior art is difficult to effectively identify and predict risk factors of visual fatigue, which leads to frequent occurrence of visual fatigue problems and affects quality of life and health.
By obtaining the user's visual fatigue data, performing data cleaning and Logistic regression analysis, the risk factors of visual fatigue were screened out, and a random forest model was constructed for risk prediction.
Effectively identify the risk factors of visual fatigue, improve the scientific research level of visual fatigue, help deeply explore the risk factors of visual fatigue, and improve early identification and prevention of high-risk groups of visual fatigue.
Smart Images

Figure CN119920480A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to information processing technology, and in particular to a method and system for identifying risk factors related to visual fatigue and constructing a risk prediction model. Background Art
[0002] With the development of science and technology, video display terminals (VDT) have entered the public eye. Their media include mobile phones, laptops, e-readers, etc. They have the advantages of timely communication, wide information availability and paperless operation. However, a problem of visual fatigue caused by excessive and improper use of VDT has also emerged [1]. Visual fatigue can cause a series of symptoms including dry eyes, headache, burning sensation, tearing, blurred vision, dry eyes, etc., and can progress to myopia, esotropia and other eye diseases. Visual fatigue not only affects the quality of life, interpersonal relationships and sleep quality, but also increases the risk of unhealthy diet and reduces work efficiency and academic performance.
[0003] At present, the research on visual fatigue is mainly focused on prevention, control and treatment. Popular methods include modifying myopia glasses, inventing related eye protection lamps, and developing various eye care medicines. As people use visual display terminals (VDT) for a long time, visual fatigue problems occur frequently. Therefore, the current risk warning technology for visual fatigue is particularly important. Summary of the invention
[0004] In order to overcome the above technical defects, the present invention provides a method for identifying risk factors related to visual fatigue and constructing a risk prediction model.
[0005] In order to solve the above problems, the present invention is implemented according to the following technical solutions:
[0006] In a first aspect, the present invention provides a method for identifying risk factors related to visual fatigue and constructing a risk prediction model, comprising:
[0007] Acquiring user visual fatigue data, wherein the user visual fatigue data includes visual fatigue scales of multiple users;
[0008] Performing data cleaning on the visual fatigue data to obtain data in Excel format after data cleaning;
[0009] Based on Logistic regression analysis, the risk factors of visual fatigue were screened for the data in Excel format, and multiple risk factors of visual fatigue were obtained.
[0010] In combination with the first aspect, the present invention further provides a first specific implementation of the first aspect, specifically, the Logistic regression analysis includes:
[0011] Construct a generalized linear model, the model formula is:
[0012] Where p is the probability of an event occurring; β0 is the intercept term; β 1、 β2 to β n are all regression coefficients; x1, x2 to x n are all characteristic variables.
[0013] In combination with the first aspect, the present invention also provides a second specific implementation manner of the first aspect, specifically, the risk factors for visual fatigue include age, sleep quality, stress perception level and myopia.
[0014] In combination with the first aspect, the present invention further provides a third specific implementation of the first aspect, specifically, performing data cleaning on the visual fatigue data, specifically comprising:
[0015] Identifying missing values of the visual fatigue data;
[0016] Performing multiple interpolation on the missing values of the visual fatigue data to obtain interpolated visual fatigue data;
[0017] Convert visual fatigue data into Excel format data.
[0018] In combination with the first aspect, the present invention further provides a fourth specific implementation of the first aspect, specifically, using Excel format data and visual fatigue risk factors to construct a data training set and an internal validation set;
[0019] Based on the data training set and the internal validation set, a visual fatigue risk prediction model is constructed using machine learning.
[0020] In combination with the first aspect, the present invention also provides a fifth specific implementation scheme of the first aspect. Specifically, the visual fatigue risk prediction model adopts a random forest model, the number of trees in the random forest model is 300, and the prediction type of the random forest model is probability.
[0021] Preferably, the visual fatigue risk prediction model adopts a random forest model, the number of trees in the random forest model is 300, and the prediction type of the random forest model is probability.
[0022] In a second aspect, the present invention further provides a system for identifying risk factors related to visual fatigue and constructing a risk prediction model, comprising:
[0023] An acquisition module, which is used to acquire user visual fatigue data, wherein the user visual fatigue data includes visual fatigue scales of multiple users;
[0024] A data cleaning module, which is used to clean the visual fatigue data, and obtain data in Excel format after data cleaning;
[0025] Based on Logistic regression analysis, the risk factors of visual fatigue were screened for the data in Excel format, and multiple risk factors of visual fatigue were obtained.
[0026] Compared with the prior art, the present invention has the following beneficial effects:
[0027] The present invention provides a method for identifying risk factors related to visual fatigue and constructing a risk prediction model. The method comprises acquiring user visual fatigue data, wherein the user visual fatigue data comprises visual fatigue scales of multiple users; performing data cleaning on the visual fatigue data to obtain data in Excel format after data cleaning; and screening the data in Excel format for visual fatigue risk factors based on Logistic regression analysis to obtain multiple visual fatigue risk factors.
[0028] The present invention provides relevant technical means for effectively identifying risk factors of visual fatigue, filling the gap in analysis technology of risk factors of visual fatigue. The method provides new tools and means for scientific research on visual fatigue, which helps to deeply explore the risk factors of visual fatigue. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The specific embodiments of the present invention are further described in detail below in conjunction with the accompanying drawings, wherein:
[0030] Figure 1 It is a flow chart of a method for identifying risk factors related to visual fatigue and constructing a risk prediction model of the present invention. DETAILED DESCRIPTION
[0031] The following specifically illustrates the implementation mode of the present invention in conjunction with the accompanying drawings. The embodiments are provided for illustrative purposes only and are not to be construed as limitations of the present invention. The accompanying drawings are provided for reference and illustration only and do not constitute limitations on the scope of patent protection of the present invention, because many changes may be made to the present invention without departing from the spirit and scope of the present invention.
[0032] With the development of science and technology, video display terminals (VDT) have entered the public eye. Their media include mobile phones, laptops, e-readers, etc. They have the advantages of timely communication, wide information availability and paperless operation. However, a problem of visual fatigue caused by excessive and improper use of VDT has also emerged [1]. Visual fatigue can cause a series of symptoms including dry eyes, headaches, burning sensations, tearing, blurred vision, dry eyes, etc., and can progress to myopia, esotropia and other ophthalmic diseases. Visual fatigue not only affects the quality of life, interpersonal relationships and sleep quality, but also increases the risk of unhealthy diet, reduces work efficiency and academic performance. At present, research on visual fatigue mainly focuses on prevention and treatment. Popular methods include modifying myopia glasses, inventing related eye protection lamps and developing various eye health care drugs. As people use visual display terminals (VDT) for a long time, visual fatigue problems occur frequently. Therefore, the current risk warning technology for visual fatigue is particularly important.
[0033] For this purpose, refer to Figure 1 The embodiment of the present invention provides a flow chart of a method for identifying risk factors related to visual fatigue and constructing a risk prediction model. The present invention provides new tools and means for analyzing risk factors for visual fatigue and constructing a risk prediction model for visual fatigue.
[0034] Embodiment 1
[0035] The method of the present invention can be performed by an information processing system, which can be implemented in the form of hardware and / or software, and the system can be configured in a computer. Figure 1 As shown, the method includes:
[0036] S100: Acquire user visual fatigue data, where the user visual fatigue data includes visual fatigue scales of multiple users;
[0037] S200: Cleaning the visual fatigue data to obtain data in Excel format;
[0038] S300: Based on Logistic regression analysis, the risk factors of visual fatigue were screened for data in Excel format, and multiple risk factors of visual fatigue were obtained.
[0039] The present invention provides relevant technical means for effectively identifying risk factors of visual fatigue, fills the gap in the analysis technology of risk factors of visual fatigue, and provides new tools and means for scientific research on visual fatigue, which is helpful to deeply understand the risk factors affecting visual fatigue of college students. Specifically, the present invention provides a detailed description of each step.
[0040] S100: Acquire user visual fatigue data, where the user visual fatigue data includes visual fatigue scales of multiple users.
[0041] In a specific implementation, obtaining user visual fatigue data is the first step to identify risk factors and build a visual fatigue risk prediction model. These data usually include visual fatigue scales of multiple users. Specifically, the visual fatigue scale is designed to explore factors related to visual fatigue. Specifically, the visual fatigue scale collects data through self-assessment of respondents (i.e., users).
[0042] In a specific implementation, the visual fatigue scale developed by Wenzhou Medical University is used for the assessment of visual fatigue of the present invention. The visual fatigue scale consists of 20 items, each of which has a specific scoring system and is divided into four levels: 0, 1, 2 and 3. These levels represent none, occasionally, frequently and always, respectively. A frequency score of ≥16 points can be diagnosed. The severity of the symptoms is proportional to the score obtained-the higher the score, the more severe the symptoms. Specifically, the visual fatigue scale is shown in Table 1.
[0043] Table 1 Visual fatigue scale
[0044]
[0045]
[0046] In another preferred implementation, the user visual fatigue data includes a visual fatigue scale and a stress perception scale for multiple users. In a preferred implementation, the stress perception scale is used to measure the perceived stress level of the participant (ie, the user) by assigning a total score from 0 to 50. The scores are divided into three groups: low perceived stress (score ≤ 13), medium perceived stress (score between 14-26), and high perceived stress (score ≥ 27). Specifically, the stress perception scale is shown in Table 2.
[0047] Table 2 Stress Perception Scale
[0048]
[0049]
[0050] In another preferred implementation of the visual fatigue risk factor analysis, in addition to collecting the user's visual fatigue data, a stress perception scale is also included. This comprehensive approach can more comprehensively evaluate the factors that affect visual fatigue, because stress and emotional state are known to affect the occurrence and severity of visual fatigue. Combining the data of these two scales can more accurately identify the risk factors of visual fatigue, because stress and emotional state may be associated with the symptoms of visual fatigue.
[0051] S200: Cleaning the visual fatigue data to obtain data in Excel format.
[0052] In the specific implementation, in order to ensure the quality of the data, the collected visual fatigue data needs to be cleaned, which includes processing missing values, outliers and duplicate records. The cleaned data is usually saved in Excel format to prepare for subsequent data analysis and model building.
[0053] In a specific implementation, the visual fatigue data is cleaned, specifically including:
[0054] S210: Identify missing values of the visual fatigue data.
[0055] S220: performing multiple interpolation on the missing values of the visual fatigue data to obtain interpolated visual fatigue data.
[0056] S230: Convert the visual fatigue data into Excel format data.
[0057] In a specific implementation, steps S210, S220, and S230 may be implemented in the following manner:
[0058] S21: Call rstudio software, import the visual fatigue data into rstudio software, and analyze the structure, content and quality of the visual fatigue data.
[0059] In a preferred implementation, the visual fatigue data in CSV format is read using the read.csv() function, or the data in Excel format is read using the readxl::read_excel() function. The visual fatigue scale and the stress perception scale can both be in Excel format.
[0060] Analyze the structure of the data: Use the str() function to view the structure of the data, including the variable type and variable name. Use the summary() function to obtain summary statistics of the data.
[0061] Analyze the content of the data: Use the head() function to view the first few rows of the data to understand the data content. Use the tail() function to view the last few rows of the data.
[0062] Analyze the quality of the data: Regarding missing values: Use the is.na() function combined with the sum() function to check for missing values. Outliers: Use a boxplot to identify outliers.
[0063] S22: Calculate missing data values and perform multiple imputation or deletion.
[0064] In the present invention, multiple imputation is a statistical method used to deal with the problem of missing data. The multiple imputation method implemented by the mice package uses the following steps:
[0065] (1) Missing data mechanism: When dealing with missing data, we first need to understand the mechanism of missing data. MICE assumes that the data is Missing At Random (MAR) or Missing Completely At Random (MCAR). This means that the missing data is not completely random (MCAR), but is related to some characteristics of the observed data (MAR).
[0066] (2) Chained Equations: The mice package uses chained equations (also known as Fully Conditional Specification, FCS) for imputation. Each missing data variable is imputed based on other variables. These equations are iterated in sequence to generate multiple imputed data sets.
[0067] Mice uses chained equations for imputation, where the missing value of each variable is predicted using the conditional distribution of the other variables.
[0068] (3) Interpolation process: For each missing value of a variable, the mice package predicts and interpolates the missing value based on the values of other variables. This process usually involves the following steps:
[0069] 3.1 Initialization: Perform preliminary interpolation of missing values. In specific implementation, simple methods such as mean interpolation and regression interpolation can be used.
[0070] 3.2 Model building: Build a regression model for each variable containing missing values to predict the missing values of the variable. The choice of model usually depends on the data type and distribution of the variable. In general, the model uses a linear regression model, a logistic regression model, or a random forest model.
[0071] 3.3 Cyclic interpolation: Iteratively update the missing values of each variable, each time considering the latest interpolated values of other variables. This process is repeated many times until convergence.
[0072] (4) Generate multiple data sets: In order to reflect the uncertainty of interpolation, mice generates multiple interpolation data sets (usually 5 to 10). This allows the uncertainty of interpolation and the stability of the model to be evaluated.
[0073] (5) Analyze and merge results: After the imputation is completed, the data analyst will analyze each imputed data set separately and merge the results. The merging process uses Rubin's rule to summarize the results and evaluate the uncertainty of the imputation. The Rubin rule formula is as follows:
[0074] ①Estimated mean:
[0075]
[0076] ② Estimate variance:
[0077] Estimate within imputation variance:
[0078] Estimate between imputation variance:
[0079] Total variance:
[0080] S23: Check the data distribution, use the box plot method to identify and remove outliers. The outlier evaluation criteria are:
[0081] <Q1 - 1.5 * IQR or > Q3 + 1.5 * IQR.
[0082] Where Q1 is the first quartile, Q3 is the third quartile, and IQR is the interquartile range.
[0083] S24: Perform one - hot encoding on non - numerical binary categorical variables, use label encoding for ordered variables, and convert all categorical variables to factors.
[0084] S25: After data cleaning, export it as an excel format file.
[0085] S300: Based on Logistic regression analysis, screen the excel format data for visual fatigue risk factors to obtain multiple visual fatigue risk factors.
[0086] In specific implementation, based on Logistic regression analysis, screening the excel format data for visual fatigue risk factors includes the following steps:
[0087] S310: Construct a generalized linear model.
[0088] Specifically, it includes the following steps:
[0089] 1. Prepare data:
[0090] Ensure that the data frame contains visual fatigue events (as binary response variables) and related predictor variables.
[0091] 2. Define the model formula:
[0092] Construct the model formula, such as y ~ x1 + x2 + x3, where y is the binary response variable, and x1, x2, x3 are predictor variables.
[0093] 3. Use the glm() function:
[0094] Use the glm() function to build a Logistic regression model and set the family parameter to binomial to perform Logistic regression.
[0095] 4. Model fitting:
[0096] Call the summary() function to view the detailed results of the model, including regression coefficients, standard errors, z values, and P values.
[0097] In a specific implementation, the Logistic regression analysis includes:
[0098] Construct a generalized linear model, the model formula is:
[0099] Where p is the probability of an event occurring; β0 is the intercept term; β 1、 β2 to β n are all regression coefficients; x1, x2 to x n are all characteristic variables.
[0100] S320: Construct a stepwise regression model.
[0101] Specifically, the purpose of stepwise regression: Stepwise regression is used to identify the variables that contribute the least to the model and remove them from the model to simplify the model. The specific steps include the following:
[0102] 1. Use stepwise regression to select variables:
[0103] Use the step() function to perform stepwise regression and select the optimal model based on the AIC criterion.
[0104] 2. Calculation of AIC:
[0105] AIC is a criterion for measuring the quality of a model, and the calculation formula is: AIC = -2·log(L)+2k; where L is the likelihood function of the model and k is the number of parameters in the model. This helps identify and delete variables that have the least impact on the model, and simplifies the model by comparing model performance indicators (such as AIC).
[0106] 3. Model comparison:
[0107] Compare the AIC values of different models and select the model with the smallest AIC value as the final model.
[0108] S330: Automatic regression selection. Combine single-factor and multi-factor regression to optimize the model, select the optimal model, and filter variables according to the preset threshold.
[0109] In another implementation, the following steps are used:
[0110] 1. Use the generalized linear model of S310 for univariate analysis: From the generalized linear model, perform univariate logistic regression analysis on each variable to screen out statistically significant variables.
[0111] 2. Use the stepwise regression model of S320 as a reference: Specifically, the stepwise regression model can be used as a reference model in the automatic regression selection process. The model in the automatic selection process is similar to the stepwise regression model.
[0112] 3. Combining univariate and multivariate regression: The variables selected in the univariate analysis can be used to build a multivariate model. Use the variables in the stepwise regression model as a starting point, and then further optimize it through the automatic regression selection method.
[0113] 4. Use a preset threshold to filter variables: In a multifactor model, set a p-value threshold (e.g. 0.05) and select variables with p-values less than this threshold. These variables are considered to contribute significantly to the model.
[0114] 5. Construct the final model: Based on the results of the automatic regression selection, construct the final Logistic regression model. This model should include those variables that have a significant impact on the risk of visual fatigue.
[0115] 6. Model evaluation: Evaluate the final model, including checking the model coefficients, significance level, AIC value, confusion matrix and ROC curve.
[0116] In the final implementation, based on Logistic regression analysis, the Excel format data was screened for visual fatigue risk factors to obtain multiple visual fatigue risk factors. The visual fatigue risk factors include age, sleep quality, stress perception level and myopia. The function of the present invention is to identify reliable risk factors for visual fatigue, screen characteristic factors for visual fatigue and predict the risk of visual fatigue, so as to improve the early identification of people at high risk of visual fatigue.
[0117] The above is a detailed implementation process of risk factor identification in the method of the present invention. Specifically, the present invention is further supplemented with the following examples:
[0118] The present invention obtains the research data of visual fatigue of college students, and obtains 1782 copies in total. After removing duplicates and outliers, a total of 1741 copies of valid data are obtained. The missing values of age, height and weight are used to perform multiple interpolation on the data using the Mice package, and other missing data rows are deleted. Among the research subjects, there are 400 males and 1341 females, with an average age of 19.94±1.10 years old. The prevalence of visual fatigue in this study is 47.9% (n=907), among which females (49.1%) are slightly higher than males (43.8%).
[0119] The present invention firstly adopts the univariate logistic regression method to incorporate single variables one by one to screen the risk factors of visual fatigue, and obtains ten covariates closely related to visual fatigue, namely: age, exercise, non-medical students, staying up late, sleep quality, stress perception level, learning stage, myopia and time using electronic screens.
[0120] Then the above positive factors were extracted and included in the multivariate logistic regression model to obtain the correlation between covariates and visual fatigue. In the multivariate logistic regression model, the factors most correlated with visual fatigue were age, myopia, medical student status, staying up late, sleep quality and perceived stress level.
[0121] Table 3 Univariate and multivariate logistic regression analysis to identify risk factors for amblyopia.
[0122]
[0123]
[0124] Ten risk factors associated with visual fatigue were screened by univariate logistic regression, and then these 10 factors were included in the multivariate logistic regression model to explore their relationship with visual fatigue.
[0125] The present invention combines the currently popular theory of the hazards of electronic screens and the impact of excessive eye use on visual fatigue, and incorporates characteristics such as height, exercise, electronic screen time, paper reading time and learning stage on the basis of the original model.
[0126] In another preferred implementation, the method of the present invention further provides a step of constructing a risk prediction model, the method comprising:
[0127] S400: Use Excel format data and visual fatigue risk factors to construct data training sets and internal validation sets.
[0128] In the specific implementation, a random sampling method is used to divide the dataset into a data training set and an internal validation set with a ratio of 80:20.
[0129] Furthermore, feature selection is performed on the data training set, and features that are highly correlated with visual fatigue are selected as feature factors. Categorical variables are encoded, such as converting gender and age groups into numerical data.
[0130] Import the Excel format data and visual fatigue risk factors into a data analysis library, such as R or Python's Pandas library. In Python, you can use the train_test_split function from the sklearn.model_selection module to split the dataset.
[0131] S500: Based on the data training set and the internal validation set, a visual fatigue risk prediction model is constructed using machine learning.
[0132] Specifically, step S500 is constructed by the following implementation method:
[0133] S510: Create a task.
[0134] The S300 screening risk factors were extracted as feature factors, and the machine learning type was specified as a binary classification problem. The data training set and internal validation set were divided into a ratio of 8:2.
[0135] S510: Learner selection:
[0136] The learner selected is the random forest algorithm (RF). RF has a high prediction accuracy, a strong tolerance to outliers and noise, can process high-dimensional data (the number of variables is much larger than the number of observations), effectively analyzes nonlinear, collinear and interactive data, and can give variable importance measures (VIM) while analyzing the data. The calculation method of VIM is based on the Gini index and the out-of-bag (OOB) error rate.
[0137] Specifically, the model configuration of the random forest model includes:
[0138] If variable Xj appears M times in the i-th tree, then the importance of variable Xj in the i-th tree is:
[0139]
[0140] VARIABLE X j In RF, Gini importance is defined as (n is the number of classification trees in RF):
[0141]
[0142] The variable Xj in the VIM of the i-th tree j The permutation importance of is:
[0143]
[0144] Significance test of variables:
[0145]
[0146] S530: Model parameter tuning
[0147] This step includes defining the parameters and search range to be tuned, selecting a tuning strategy (such as grid search, random search), setting up cross-validation, and performing tuning. Several key tuning hyperparameters:
[0148] (1) ntree: The number of trees in the forest. More trees usually improve the performance of the model, but the computational cost also increases.
[0149] (2) mtry: The number of features that each tree considers when splitting a node. Smaller values make the tree more random and generally improve the generalization ability of the model.
[0150] (3) nodesize: The minimum number of samples for a terminal node. Smaller values make the tree more complex, while larger values simplify the tree.
[0151] (4) max_depth: Maximum depth of the tree (supported in some implementations). Used to limit the maximum depth of the tree to prevent overfitting.
[0152] In a specific implementation, the visual fatigue risk prediction model adopts a random forest model, the number of trees in the random forest model is 300, and the prediction type of the random forest model is probability.
[0153] 4.4 Model Performance Evaluation
[0154] The training task model of the present invention belongs to a classification model, and the performance evaluation of the model is related to the confusion matrix. The matrix shows the detailed information of the model prediction results, including the number of true positives (TP), false positives (FP), true negatives (TN) and false negatives (FN). The following are the commonly used evaluation indicators for binary classification models in machine learning:
[0155] (1) Accuracy: The proportion of samples with correct predictions to the total number of samples. The formula is: Accuracy = Total number of predictions / Number of correct predictions.
[0156] (2) Precision: The ratio of the number of samples correctly predicted as positive to the number of samples predicted as positive. The formula is: Precision = TP / (TP+FP).
[0157] (3) Recall: The ratio of the number of samples correctly predicted as positive to the number of samples actually positive. The formula is: Recall = TP / (TP+FN).
[0158] (4) Specificity: The ability of the model to correctly predict negative examples among negative examples. The formula is: Specificity = TN / (TN+FN).
[0159] (5) F1 Score: The harmonic mean of precision and recall. The formula is: F1 = 2*precision*recall / (precision+recall).
[0160] ROC curve (Receiver Operating Characteristic Curve): describes the relationship between the true positive rate (True Positive Rate) and the false positive rate (False Positive Rate). The ROC curve can be used to calculate the AUC (area under the curve). The AUC represents the overall performance of the model. The AUC value ranges from 0.5 (random guessing) to 1 (perfect classification).
[0161] The above is a detailed implementation process of the risk prediction model construction in the method of the present invention. Specifically, the present invention is further supplemented with the following examples:
[0162] The present invention divides the data training set and internal validation set into a ratio of 8:2. The Ranger random forest algorithm is used for classification, and the feature importance is evaluated based on the impurity method. The number of trees in the random forest is specified to be 300, and the prediction type of the model is probability.
[0163] At the same time, the present invention sets the random seed to 121 to ensure the repeatability of the results, so that each time the model is run, operations such as initialization and data segmentation can produce the same random results, so that the experimental results can be reproduced.
[0164] The present invention sets up a hyperparameter tuning process, in which a random search method is used to optimize two hyperparameters of the model (mtry and max.depth), the classification error rate of the model is evaluated under 5-fold cross validation, and the accuracy of the model is improved by generating multiple interpolation data sets. The search process is performed within a maximum of 40 evaluations, and each hyperparameter combination attempted is within the specified search space.
[0165] The present invention sets an optimal parameter value and trains the model based on it. We set the best hyperparameter value obtained by tuning to the learner and obtain the importance score of each feature in the model. The results are as follows: the out-of-bag error calculated during the return training process is 0.21.
[0166] Table 4 Importance scores of different features in the model
[0167]
[0168] Then the present invention uses the design_points feature selection strategy, selects features based on the training model matrix, uses 5-fold cross validation to evaluate the effect of feature selection, sets multiple evaluation criteria (AUC value, accuracy, sensitivity and specificity) to measure the effect of feature selection, and disables value checking to improve operating efficiency.
[0169] Table 5 Effectiveness evaluation of different models
[0170]
[0171]
[0172] According to the results shown in Table 5, the model that finally included age, electronic screen time, myopia, stress perception level, sleep time, staying up late, and learning stage features performed best (AUC value 0.731, accuracy 0.676).
[0173] On another implementation technical details, on implementation technical details:
[0174] Data cleaning: You can use the Pandas library to clean data, including removing duplicate records, correcting erroneous data, and processing missing values.
[0175] Feature Selection: Features can be selected using feature importance scores from random forests.
[0176] Model training: You can use the Random Forest Classifier or RandomForest Regressor in the scikit-learn library for model training.
[0177] Cross-validation: You can use the cross_val_score function in the scikit-learn library to perform cross-validation.
[0178] Model evaluation: You can use the accuracy_score, recall_score, f1_score and other functions in the scikit-learn library to evaluate the model.
[0179] Model optimization: You can use GridSearchCV or Randomized SearchCV to perform grid search and random search optimization of model parameters.
[0180] Result analysis: Analyze the importance of the features output by the model, which can be further displayed using a visualization library such as Matplotlib.
[0181] The present invention also provides a system for identifying risk factors related to visual fatigue and constructing a risk prediction model, which is used to implement the process of the above method, and the system includes:
[0182] An acquisition module, which is used to acquire user visual fatigue data, wherein the user visual fatigue data includes visual fatigue scales of multiple users;
[0183] A data cleaning module, which is used to clean the data of the visual fatigue scale, and obtain data in Excel format after data cleaning;
[0184] Based on Logistic regression analysis, the risk factors of visual fatigue were screened for the data in Excel format, and multiple risk factors of visual fatigue were obtained.
[0185] For other structures of the method and system for identifying risk factors related to visual fatigue and constructing a risk prediction model described in this embodiment, refer to the prior art.
[0186] The above description is only a preferred embodiment of the present invention and does not limit the present invention in any form. Therefore, any modification, equivalent change and modification made to the above embodiment according to the technical essence of the present invention without departing from the technical solution of the present invention shall still fall within the scope of the technical solution of the present invention.
Claims
1. A method for identifying risk factors related to visual fatigue and constructing a risk prediction model, characterized in that: include: Acquiring user visual fatigue data, wherein the user visual fatigue data includes visual fatigue scales of multiple users; Performing data cleaning on the visual fatigue data to obtain data in Excel format after data cleaning; Based on Logistic regression analysis, the risk factors of visual fatigue were screened for the data in Excel format, and multiple risk factors of visual fatigue were obtained.
2. The method for identifying risk factors and constructing a risk prediction model related to visual fatigue according to claim 1, characterized in that: The Logistic regression analysis includes: Construct a generalized linear model, the model formula is: Where p is the probability of an event occurring; β0 is the intercept term; β 1、 β2 to β n are all regression coefficients; x1, x2 to x n are all characteristic variables.
3. The method for identifying risk factors and constructing a risk prediction model related to visual fatigue according to claim 2, characterized in that: The risk factors for visual fatigue include age, sleep quality, perceived stress level and myopia.
4. The method for identifying risk factors and constructing a risk prediction model related to visual fatigue according to claim 1, characterized in that: The visual fatigue data is cleaned, specifically including: Identifying missing values of the visual fatigue data; Performing multiple interpolation on the missing values of the visual fatigue data to obtain interpolated visual fatigue data; Convert visual fatigue data into Excel format data.
5. The method for identifying risk factors and constructing a risk prediction model related to visual fatigue according to claim 1, characterized in that: Also includes: Use Excel format data and visual fatigue risk factors to build data training sets and internal validation sets; Based on the data training set and the internal validation set, a visual fatigue risk prediction model is constructed using machine learning.
6. The method for identifying risk factors and constructing a risk prediction model related to visual fatigue according to claim 5, characterized in that: include: The visual fatigue risk prediction model adopts a random forest model, the number of trees in the random forest model is 300, and the prediction type of the random forest model is probability.
7. A system for identifying risk factors related to visual fatigue and constructing a risk prediction model, characterized in that: include: An acquisition module, which is used to acquire user visual fatigue data, wherein the user visual fatigue data includes visual fatigue scales of multiple users; A data cleaning module, which is used to clean the visual fatigue data, and obtain data in Excel format after data cleaning; Based on Logistic regression analysis, the risk factors of visual fatigue were screened for the data in Excel format, and multiple risk factors of visual fatigue were obtained.
Citation Information
Cited By
Missing data completion method and system for pancreatic cystic lesion classification
CN122291083A