Method and system for predicting myopia of children

By establishing a problem database of vision-influencing factors and myopia prediction model, collecting and analyzing children's vision impact data, the problem of insufficient prevention methods for children in the existing technology is solved, and personalized and accurate myopia prediction is achieved in a timely interventional diagnosis and treatment, improving children's vision health.

CN119943402APending Publication Date: 2025-05-06BEIJING TONGREN HOSPITAL AFFILIATED TO CAPITAL MEDICAL UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510076727.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing children have insufficient measures to prevent myopia, and they cannot intervene in myopia diagnosis and treatment in a timely manner. In addition, traditional myopia examinations rely on subjective experience, making it difficult to make personalized and accurate predictions.

Method used

By setting up the vision impact database and myopia prediction model, establish a database of vision impact factors problem, extract questionnaire items using feature matching algorithm, combine and adjust the data collection questionnaire, collect the visual impact data of the subject being tested, and visual prediction model is used to predict vision.

Benefits of technology

It has achieved personalized and accurate prediction of myopia in children, and can intervene in myopia diagnosis and treatment in a timely manner, reduce the occurrence of high myopia, and improve children's vision health.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943402A_ABST
    Figure CN119943402A_ABST
Patent Text Reader

Abstract

The invention specifically discloses a method and a system for predicting myopia of children. The method comprises the following steps: setting a vision influence database and a myopia prediction model; analyzing and processing the data in the vision influence database to establish a vision influence factor problem library; extracting a plurality of questionnaire items from the eyesight influence factor question library by using a feature matching algorithm, and performing combination adjustment to obtain a data acquisition questionnaire; according to the data collection problem, collecting vision influence data corresponding to a subject, and inputting the vision influence data into the myopia prediction model for vision prediction. The system can guide intervention of child myopia, intervenes myopia diagnosis and treatment in time according to the vision change trend, reduces occurrence of high myopia, and improves the vision health of children.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of children's myopia prediction, and more specifically, to a method and system for children's myopia prediction. Background Art

[0002] With the increase in the time and frequency of using electronic products, excessive use of eyes, unhygienic use of eyes, lack of physical exercise and outdoor activities, etc., children's vision is greatly affected. Myopia in children has spread to varying degrees around the world, and myopia has become a global public health problem. Studies have shown that school age is the peak period for the onset of myopia, and the incidence of myopia reaches its peak in the primary school stage. Although the incidence of myopia in children and adolescents is so high and the consequences are so serious, there are still few means to prevent the occurrence and development of myopia at this stage. Traditional myopia examinations require professional institutions, and myopia predictions are basically subjective and empirical judgments. It is difficult to make personalized and accurate and reliable prediction trends, and their effects are controversial. The key reason is that the exact mechanism of myopia is still not completely clear. Therefore, how to comprehensively analyze the risk factors for myopia and explore the independent risk factors for myopia in school-age children can conveniently and effectively predict the trend of vision changes in the subjects and provide an effective basis for timely intervention in myopia diagnosis and treatment, which is of great significance. Summary of the invention

[0003] The present invention provides a method and system for predicting myopia in children, which solves the problem that existing means for preventing myopia in children are insufficient and myopia diagnosis and treatment cannot be timely intervened. The method and system can guide the intervention of myopia in children, timely intervene in myopia diagnosis and treatment according to the trend of vision changes, reduce the occurrence of high myopia, and improve children's vision health.

[0004] To achieve the above objectives, the present invention provides the following technical solutions:

[0005] A method for predicting myopia in children, comprising:

[0006] Setting up a vision impact database and myopia prediction model;

[0007] Analyzing and processing the data in the vision impact database to establish a vision impact factor question database;

[0008] Extracting multiple questionnaire items from the vision influencing factor question bank using a feature matching algorithm to combine and adjust to obtain a data collection questionnaire;

[0009] The vision impact data corresponding to the examinee is collected through the data collection questions and input into the myopia prediction model for vision prediction.

[0010] Preferably, it also includes:

[0011] When combining questionnaires, data association rules are established for questions of the same dimension, and association analysis is performed on questions of the same dimension to mine the corresponding vision impact data of the examinees, obtain the degree of relationship between different variables or individuals in the received vision impact data of the examinees, and the individual behavior patterns of the examinees.

[0012] Preferably, the myopia prediction model is provided with a data collection layer, a data correlation analysis layer and a myopia prediction layer;

[0013] The data collection layer is used to receive the visual impact data of the examinee to obtain the examination data and questionnaire data of the examinee;

[0014] The data correlation analysis layer is provided with a clustering algorithm to perform cluster analysis on the inspection data and questionnaire data to obtain data correlation;

[0015] The myopia prediction layer is provided with a prediction algorithm to perform vision prediction based on the correlation of the subject's vision impact data, and then generate a myopia prediction report.

[0016] Preferably, analyzing and processing the data in the vision impact database to establish a vision impact factor question library includes:

[0017] The data in the vision-influencing database are divided into different clusters through feature selection, distance measurement and clustering algorithm, and then classified into eight dimensions to explore the potential structure in the data, and the data are quantified according to the results of the classification analysis to obtain the vision-influencing factor question library;

[0018] Among them, the eight dimensions include: children's pregnancy and delivery conditions, children's health status, children's parents and grandparents' family conditions, children's dietary conditions, children's sleep conditions, children's daily eye use conditions, children's eye environment and children's vision health intervention behaviors.

[0019] Preferably, collecting the visual acuity impact data corresponding to the examinee through the data collection questions includes:

[0020] Define the data missing types in the visual acuity impact data obtained from the examinee, the data missing types including: incomplete random missing, random missing and non-random missing;

[0021] According to the type of missing data in the visual acuity impact data obtained from the examinees, the missing data are completed by selecting at least one of deleting rows, deleting columns, filling with the previous or next value, taking the average, taking the median value, using constant filling, linear interpolation, modeling prediction, and maximum likelihood estimation.

[0022] Preferably, the inputting the myopia prediction model to perform vision prediction includes:

[0023] The myopia prediction model is used to perform correlation analysis on the visual acuity impact data corresponding to the examinee, and to perform data classification and prediction to generate a myopia prediction report corresponding to the examinee.

[0024] Preferably, the data classification and prediction includes:

[0025] Classify vision impact data using hierarchical clustering, random forest, and / or K-means clustering methods;

[0026] Predictions are made on vision impact data using decision trees, Bayesian belief networks, neural networks, and / or genetic algorithms.

[0027] Preferably, the prediction of vision impact data using a decision tree, a Bayesian belief network, a neural network and / or a genetic algorithm includes:

[0028] The prior probability of obtaining the visual acuity impact data of the examinee is calculated by the Bayesian belief network method;

[0029] Class prediction is performed through the Naive Bayes classifier;

[0030] Logistic regression model was used to handle data with binomial dependent variables and multiple classifications.

[0031] Preferably, the method of predicting the vision-affecting data using a decision tree, a Bayesian belief network, a neural network and / or a genetic algorithm further comprises:

[0032] The collected historical vision impact data are divided into 80%, 15%, and 5%, among which 80% of the data is used as a training data set, 15% of the data is used as a test data set, and 5% of the data is used as an independent sample set to train the myopia prediction model and verify the prediction effect.

[0033] The present invention also provides a system for predicting children's myopia, using the above method for predicting children's myopia, including: a user end and a server;

[0034] The server is provided with a vision impact database and a myopia prediction model, and a vision impact factor question library is established based on the vision impact database;

[0035] The user terminal uses a feature matching algorithm to extract multiple questionnaire items from the vision influencing factor question library of the server and adjusts the items in combination to obtain a data collection questionnaire, and obtains the vision influencing data of the examinee through the data collection questionnaire;

[0036] The user terminal pre-processes the acquired vision impact data of the examinee and sends the pre-processed data to the server;

[0037] The server receives the vision impact data corresponding to the examinee, and inputs the myopia prediction model to perform vision prediction, so as to obtain the vision prediction result of the examinee, and then sends it to the corresponding user terminal;

[0038] The server also stores the received vision impact data of the examinee into the vision impact database, and updates the historical data of the vision impact database;

[0039] The myopia prediction model was trained and tested using updated historical data.

[0040] The present invention provides a method and system for predicting myopia in children, by setting a vision impact database and a myopia prediction model, and establishing a vision impact factor question library based on the vision impact database, extracting a data collection questionnaire from the vision impact factor question library to collect the vision impact data of the examinee, and performing vision prediction by the myopia prediction model. The method solves the problem that the existing prevention means for myopia in children are insufficient and the diagnosis and treatment of myopia cannot be intervened in time, can guide the intervention of myopia in children, intervene in myopia diagnosis and treatment in time according to the trend of vision changes, reduce the occurrence of high myopia, and improve the vision health of children. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the specific embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below.

[0042] Figure 1 It is a schematic diagram of a method for predicting myopia in children provided by the present invention.

[0043] Figure 2 It is a schematic diagram of a flow chart of myopia prediction provided by an embodiment of the present invention.

[0044] Figure 3 It is a schematic diagram of the architecture of a myopia prediction model provided by an embodiment of the present invention.

[0045] Figure 4 It is a schematic diagram of a machine self-learning process of a myopia prediction model provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0046] In order to enable persons skilled in the art to better understand the solutions of the embodiments of the present invention, the embodiments of the present invention are further described in detail below in conjunction with the accompanying drawings and implementation modes.

[0047] In view of the current problem that children's myopia cannot be diagnosed and treated in time, the present invention provides a method and system for predicting children's myopia, which solves the problem that the existing prevention measures for children's myopia are insufficient and the diagnosis and treatment of myopia cannot be intervened in time. It can guide the intervention of children's myopia, intervene in myopia diagnosis and treatment in time according to the trend of vision changes, reduce the occurrence of high myopia, and improve children's vision health.

[0048] like Figure 1 As shown, a method for predicting myopia in children comprises:

[0049] S1: Set up the vision impact database and myopia prediction model.

[0050] S2: Analyze and process the data in the vision-influencing database to establish a vision-influencing factor question library.

[0051] S3: extracting multiple questionnaire items from the vision influencing factor question bank using a feature matching algorithm to combine and adjust them to obtain a data collection questionnaire.

[0052] S4: Collect the vision impact data corresponding to the examinee through the data collection questions, and input it into the myopia prediction model to perform vision prediction.

[0053] In actual applications, students’ basic information needs to be entered at each screening, and each year is analyzed independently, so it is impossible to conduct continuous tracking and screening of students. In addition, because the mobility of school-age children and adolescents will fluctuate greatly with the promotion of school, the student groups screened in different years will change greatly, and the continuity analysis between the corresponding data will be distorted. At the same time, the existing electronic questionnaires are single in form, and questionnaires need to be added in the background. Questionnaires can only be added after students and parents enter basic information. Although simple jumps can be added between questions in the questionnaire, it is impossible to verify the questionnaire questions one by one according to the basic information. Moreover, the questionnaire templates for each year are based on the same set of frameworks, and it is impossible to configure exclusive personalized questionnaires based on student characteristics. By setting up a vision impact database, and constructing a vision impact factor question library based on the vision impact database, and using a feature matching algorithm to extract multiple questionnaire items from the vision impact factor question library, so as to combine and adjust to obtain a data collection questionnaire, the questionnaire form can be personalized and data analysis can be achieved.

[0054] In one embodiment, if Figure 2As shown in the figure, the vision impact data include: age, gender, axial length, refraction, eye distance, far vision reserve value, genetic factors, environmental factors and eye intensity. The vision impact data of the examinee are collected, and the vision is predicted by the myopia prediction model. This method tracks and analyzes the vision, refraction, axial length and hyperopia reserve of children and adolescents through the vision impact database, combines various environmental influencing factors, and further studies and clarifies the close correlation between various influencing factors and the occurrence and development of myopia through multi-factor regression analysis, and explores the incidence of myopia in different environments. To accurately predict the occurrence of myopia, it is necessary to simulate and deduce variables such as environment, eye habits, and eye use time for individuals, and combine AI analysis to further study and clarify the correlation between various influencing factors and factors such as axial length and hyperopia reserve, so as to provide a theoretical basis for the pre-prevention of myopia prevention and control. It solves the problem that the existing prevention methods for myopia in children are insufficient and the diagnosis and treatment of myopia cannot be intervened in time, can guide the intervention of myopia in children, intervene in myopia diagnosis and treatment in time according to the trend of vision changes, reduce the occurrence of high myopia, and improve children's vision health.

[0055] The method also includes: when combining questionnaires, establishing data association rules for questions of the same dimension, performing association analysis on questions of the same dimension, so as to perform mining operations on the vision impact data corresponding to the examinee, and obtain the degree of relationship between different variables or individuals in the received vision impact data of the examinee, as well as the individual behavior patterns of the examinee.

[0056] like Figure 3 As shown, the myopia prediction model is provided with a data collection layer, a data correlation analysis layer and a myopia prediction layer. The data collection layer is used to receive the subject's vision impact data to obtain the subject's examination data and questionnaire data. The data correlation analysis layer is provided with a clustering algorithm to perform cluster analysis on the examination data and questionnaire data to obtain data correlation. The myopia prediction layer is provided with a prediction algorithm to perform vision prediction based on the correlation of the subject's vision impact data, and then generate a myopia prediction report.

[0057] Specifically, the myopia prediction model uses K-means clustering, random forest and other methods to calculate visual acuity, refraction, axial length, and questionnaire information. In the process of machine learning, the samples are continuously trained to understand the hidden rules in the samples, and then a more accurate judgment is made on unknown samples based on these rules, and the data correlation is discovered. The weights of factors affecting myopia are further analyzed, and the predictive performance of the model is tested by continuously adding influencing factors, that is, the test set needs to be added for cyclic verification, so as to train a precise intervention model.

[0058] Further, analyzing and processing the data in the vision impact database to establish a vision impact factor question library includes:

[0059] The data in the vision-influencing database are divided into different clusters through feature selection, distance measurement and clustering algorithm, and then classified into eight dimensions to explore the potential structure in the data. The data are quantified according to the results of classification analysis to obtain the vision-influencing factor question library. The eight dimensions include: children's pregnancy and delivery status, children's health status, children's parents and grandparents' family status, children's diet status, children's sleep status, children's daily eye status, children's eye environment and children's vision health intervention behavior.

[0060] In one embodiment, the questions in the question bank are matched by student characteristics to automatically combine questionnaires. When establishing the question bank, firstly, the data in the influencing factor data set is divided into different clusters through feature selection, distance measurement and clustering algorithm, and classified into the current eight dimensions, the potential structure in the data set is explored, and the data is quantified according to the results of cluster analysis; when combining questionnaires, association analysis is required between questions of the same dimension, and association rules are established to eliminate noise data generated by logical contradictions or excessive jumps when students answer, and denoising is performed from the source of the data; after data collection, the basic information of students and vision collection data are used as dependent variables, and the quantified influencing factor data are used as independent variables, and regression analysis is performed according to feature selection and multiple collinearity tests. According to the data distribution, the hypothesis is further deduced, and the influencing factor question bank is iterated. An example of collected questionnaires is shown in Table 1:

[0061]

[0062]

[0063] Furthermore, collecting the visual acuity impact data corresponding to the examinee through the data collection questions includes:

[0064] The data missing types in the acquired vision impact data of the examinee are defined, and the data missing types include: incomplete random missing, random missing and non-random missing.

[0065] According to the type of missing data in the visual acuity impact data obtained from the examinees, the missing data are completed by selecting at least one of deleting rows, deleting columns, filling with the previous or next value, taking the average, taking the median value, using constant filling, linear interpolation, modeling prediction, and maximum likelihood estimation.

[0066] Specifically, incomplete data is missing values, which will cause sample loss and interfere with analysis results. Therefore, algorithms can be used to delete, replace or fill missing values. The specific process is as follows:

[0067] Step 1: Define the missing type:

[0068] Missing Completely at Random (MCAR): This means that the missing data is random and does not depend on any incomplete or complete variables. The appearance of null values ​​is completely unrelated to known or unknown features in the data set.

[0069] Missing at random (MAR): refers to the fact that the missing data is not completely random, that is, the missing data depends on other complete variables.

[0070] Missing Not at Random (MNAR): refers to the situation where the missing data depends on the incomplete variable itself.

[0071] Step 2: Select the appropriate algorithm for deletion or interpolation according to the type:

[0072] Delete rows: Only for time series with missing completely at random (MCAR). If the missing values ​​only account for a small part of the dataset, deleting rows is a perfect solution. However, it is not suitable for cases where the proportion of missing values ​​in the dataset is too high.

[0073] Delete columns: Generally speaking, when the proportion of null values ​​is higher than 60%, you can consider deleting the column. Of course, you can choose the proportion based on the actual situation. 30% is also acceptable when there is sufficient data.

[0074] Previous or Next Value: Only used for Missing Completely at Random (MCAR) When dealing with time series problems, you can use the previous or next value to fill in missing values.

[0075] Mean: Only used for Missing Completely at Random (MCAR) Because the mean is sensitive to outliers, it is not a good choice in most cases.

[0076] Median: used only for missing completely at random (MCAR) Similar to the mean, but more robust to outliers.

[0077] Mode value: Only used for Missing Completely at Random (MCAR) By choosing the most common value, you can be sure that most of the time you are filling in the null values ​​correctly. But be careful with multimodal distributions, because for these, using the mode is no longer a viable option.

[0078] Filling with constant value: (Only for Missing Not at Random (MNAR)) As we have seen before, the missing values ​​in the Missing Not at Random (MNAR) case actually contain a lot of information about the actual value. So, it is feasible to fill the missing values ​​with a constant value (unlike other types of values).

[0079] Linear interpolation: only for time series with missing completely at random (MCAR) In time series with trend and almost no seasonality problems, we can estimate the missing values ​​by linearly interpolating the values ​​before and after the missing values.

[0080] Modeling prediction: Take the missing attributes as the prediction target, divide the data set into two categories according to whether it contains missing values ​​of specific attributes, and use the existing machine learning algorithm to predict the missing values ​​of the prediction data set. The fundamental flaw of this method is that if other attributes are irrelevant to the missing attributes, the prediction result is meaningless; however, if the prediction result is quite accurate, it means that the missing attribute does not need to be included in the data set; the general situation is somewhere in between.

[0081] Maximum likelihood estimation: Under the condition that the missing type is missing at random, assuming that the model is correct for the complete sample, the marginal distribution of the observed data can be used to perform maximum likelihood estimation of the unknown parameters (Little and Rubin). This method is also called maximum likelihood estimation that ignores missing values. The calculation method commonly used in practice for maximum likelihood parameter estimation is expectation maximization (EM). This method is more attractive than deleting cases and single value interpolation. It has an important premise: it is suitable for large samples. The number of valid samples is sufficient to ensure that the maximum likelihood estimate is asymptotically unbiased and obeys a normal distribution. However, this method may fall into local extreme values, the convergence speed is not very fast, and the calculation is very complicated.

[0082] In one embodiment, as shown in Table 2, the information of the examinee is shown, in which the weight of the examinee 8 and the height of the examinee 18 are missing, and the filling steps are as follows:

[0083] (1) Calculate the average BMI "x1" of all subjects except subjects 8 and 18;

[0084] (2) First fill in the blank for subject 8, and use x1 and subject 8’s height to reversely infer subject 8’s weight a1;

[0085] (3) Calculate the average BMI "x2" of the subjects with the backfill value a1 of subject 8 added;

[0086] (4) If x1=x2, the data converges, and a1 is the final backfill value of the weight of the subject 8. If x1≠x2, the data does not converge, and proceed to step 5;

[0087] (5) Using x2 and the height of the subject 8, reversely calculate the weight a2 of the subject 8;

[0088] (6) Calculate the average BMI "x3" of the subjects with the backfill value a1 of subject 8 added;

[0089] (7) Repeat step 4 until convergence, and obtain the final backfill value a of the subject 8's weight, and fill it in;

[0090] (8) Calculate the average BMI "y1" of all subjects except subject 18;

[0091] (9) Fill in the blank for subject 18, and use y1 and the weight of subject 18 to reversely infer the height b1 of subject 18;

[0092] (10) Calculate the average BMI "y2" of the subjects by adding the backfill value b1 of subject 18; (11) If x1=x2, the data converges, and b1 is the final backfill value of the height of subject 8.

[0093]

[0094] Further, the inputting the myopia prediction model to perform vision prediction includes:

[0095] The myopia prediction model is used to perform correlation analysis on the visual acuity impact data corresponding to the examinee, and to perform data classification and prediction to generate a myopia prediction report corresponding to the examinee.

[0096] Further, the data classification and prediction includes:

[0097] Classify the vision impact data using hierarchical clustering, random forest and / or K-means clustering methods. Predict the vision impact data using decision trees, Bayesian belief networks, neural networks and / or genetic algorithms.

[0098] Furthermore, the method of using a decision tree, a Bayesian belief network, a neural network and / or a genetic algorithm to predict the vision impact data includes:

[0099] The prior probability of obtaining the subject's vision impact data is calculated using the Bayesian belief network method.

[0100] Class prediction is done using the Naive Bayes classifier.

[0101] Logistic regression model was used to handle data with binomial dependent variables and multiple classifications.

[0102] Specifically, the process of performing mining operations on the received vision impact data of the examinee to obtain prediction input data includes:

[0103] (1) Correlation analysis:

[0104] It is used to mine and discover important and interesting knowledge between item sets in a large amount of data. If there is a certain regularity between the values ​​of two or more variables, it is called association. In order to reflect the usefulness and certainty of the discovered rules when the association function is unknown or uncertain, the rules generated by association analysis must meet the minimum support threshold and the minimum confidence threshold.

[0105] Association rules are used to analyze and discover the degree of relationship between different variables or individuals in a database, and use these rules to find individual behavior patterns.

[0106] (2) Classification and prediction:

[0107] Classification is to describe a category and summarize the relevant features, and extract models that describe important data classes. There are many classification methods in data mining, mainly decision trees and decision rules, Bayesian belief networks, neural networks, and genetic algorithms. Prediction is to predict future data trends by establishing a continuous value function model. The main prediction methods include regression analysis, time series analysis, etc. Various classification models can also be used for prediction, but they mainly predict classification labels.

[0108] a. Decision Tree:

[0109] A decision tree is a tool that uses a binary tree diagram to represent processing logic. It is a method of classifying data. The goal of a decision tree is to predict or explain the response results for a categorical dependent variable. There are two main steps: first, a decision tree is built using a batch of known sample data; then, the built decision tree is used to predict the data. The process of building a decision tree can be seen as the process of generating data rules. Therefore, the decision tree realizes the visualization of data rules, and its output results are also easy to understand.

[0110] b. Bayesian Belief Network:

[0111] Simple Bayesian classification is mainly based on Bayesian Theorem to predict the classification results.

[0112] Bayes' Theorem: P(X), P(H), and P(X|H) can be calculated from given data and are prior probabilities. Bayes' Theorem provides a method to calculate the posterior probability P(H|X) from P(X), P(H), and P(X|H).

[0113] Bayes’ theorem is: P(A∩B) = P(A)*P(B|A)=P(B)*P(A|B).

[0114] Secondly, apply the naive Bayes classifier for category prediction:

[0115] hnb(x)=c∈Υarg max P(c)i=1∏dP(xi∣c).

[0116] Among them, hnb(x): represents the function of the naive Bayes classifier, which accepts a feature vector x as input and outputs a category c.

[0117] c: represents a possible category label, belonging to one of the category set Υ.

[0118] arg max: An operator that selects the value of c that maximizes the following expression.

[0119] P(c): The prior probability of category c, that is, the probability of category c appearing before observing any features.

[0120] i=1: Usually an index, indicating that iteration starts from the first feature.

[0121] ∏: The mathematical symbol for product, meaning to multiply the following terms together.

[0122] d: represents the total number of features.

[0123] P(xi|c): conditional probability, i.e. the probability of feature xi appearing given category c. xi is the i-th feature in feature vector x.

[0124] i=1 to d: Iterate over all features in feature vector x.

[0125] c. Logistic regression model:

[0126] The logistic regression model is a probability model that is suitable for research objects with two dependent variables and multiple classifications. It establishes a regression model with influencing factors as independent variables.

[0127] (3) Clustering:

[0128] Clustering is to divide the records in the database into multiple classes or clusters when the class to be divided is unknown, so that the objects in the same class have a high degree of similarity and the differences between different classes are large. It is a prerequisite for concept description and deviation analysis. Clustering methods in data mining include partitioning methods, hierarchical methods, density-based methods, grid-based methods, and model-based methods.

[0129] a. Hierarchical clustering method:

[0130] The method uses the distance matrix as the classification standard, treats each of the n samples as a class; calculates the distance between the n samples to form a distance matrix; merges the two classes with the closest distance into a new class; calculates the distance between the new class and the current classes; and merges and calculates again until there is only one class.

[0131] b. K-means clustering method:

[0132] Assign n samples to one of the K clusters, then calculate the mean of each cluster, and then assign each sample to the cluster with the closest mean, recalculate the mean of each cluster, and repeat the matching until the mean is closest to the sample data.

[0133] In one embodiment, k can be set to 3, sample a data is 20, sample b data is 19, sample c data is 1, sample d data is 10, sample e data is 15, sample f data is 3, sample g data is 21, sample h data is 2, and sample i data is 13. Randomly combine 3 samples, calculate the average value, compare with the group of samples, and repeat the matching until the average value is closest to the sample data, as shown in Table 3:

[0134]

[0135] (5) Sequential pattern mining:

[0136] Sequence pattern mining refers to mining and modeling the regularities or trends that appear frequently relative to time or other sequences. The sequence here generally refers to time series databases and sequence databases. The similarity between sequence analysis and association rules is that each sample in the sample data they use contains an item set or state set. The difference is that sequence analysis studies the transition between item sets (or states), while the association rule model studies the correlation between item sets.

[0137] Furthermore, the method of predicting the vision-affecting data using a decision tree, a Bayesian belief network, a neural network and / or a genetic algorithm also includes:

[0138] The collected historical vision impact data are divided into 80%, 15%, and 5%, among which 80% of the data is used as a training data set, 15% of the data is used as a test data set, and 5% of the data is used as an independent sample set to train the myopia prediction model and verify the prediction effect.

[0139] In practical applications, samples are divided into 80%, 15%, and 5%, of which 80% is used as training data, 15% as test data, and 5% as an independent sample set. The training set is used to train samples, and the test set is used to verify the prediction effect of the model. Models with better generalization ability can be selected, and both will enter the model construction stage. Independent sample sets are mainly used to independently test the prediction of several models on samples, which can better reflect the adaptability of the model.

[0140] In one embodiment, if Figure 4 As shown, the algorithm of the myopia prediction model is first defined and selected, and then the operation parameters of the algorithm are set, and the training samples are input for the machine learning process until the error range is reached or the minimum number of training times is reached.

[0141] It can be seen that the present invention provides a method for predicting myopia in children, by setting up a vision impact database and a myopia prediction model, and establishing a vision impact factor question library based on the vision impact database, extracting a data collection questionnaire from the vision impact factor question library to collect the vision impact data of the examinee, and performing vision prediction by the myopia prediction model. The method solves the problem that the existing prevention methods for myopia in children are insufficient and the diagnosis and treatment of myopia cannot be intervened in time, can guide the intervention of myopia in children, intervene in myopia diagnosis and treatment in time according to the trend of vision changes, reduce the occurrence of high myopia, and improve the vision health of children.

[0142] Correspondingly, the present invention also provides a system for predicting myopia in children, using the above-mentioned method for predicting myopia in children, including: a user end and a server.

[0143] The server is provided with a vision impact database and a myopia prediction model, and a vision impact factor question library is established based on the vision impact database.

[0144] The user terminal uses a feature matching algorithm to extract multiple questionnaire items from the vision-influencing factor question library of the server and adjusts the items to obtain a data collection questionnaire, and obtains the vision-influencing data of the examinee through the data collection questionnaire. The user terminal pre-processes the obtained vision-influencing data of the examinee and sends it to the server.

[0145] The server receives the vision impact data corresponding to the examinee and inputs it into the myopia prediction model to perform vision prediction, so as to obtain the vision prediction result of the examinee, and then sends it to the corresponding user terminal. The server also stores the received vision impact data of the examinee in the vision impact database, updates the historical data of the vision impact database, and uses the updated historical data to train and test the myopia prediction model.

[0146] In one embodiment, the process of obtaining the visual impact data of the examinee is first for the examinee to input his basic information through the user terminal held by him. Specifically, a unique identity identification code similar to an ID card can be generated through the basic information of the examinee, and the format is as follows: 2-digit household registration province code, 2-digit household registration city code, 2-digit household registration area and county code, 8-digit date of birth code, 3-digit sequence code (representing the order in which the examinees born on the same date in the same area are entered into the system, ranging from 000 to 999), 1-digit gender code (1 represents male; 2 represents female). When entering basic information, in addition to the household registration location, date of birth, and gender, the name must also be entered. When there are examinees born on the same date in the same area, the sequence code is determined by the name and the system entry time. Examples are shown in Table 4:

[0147]

[0148] It can be seen that the present invention provides a system for predicting myopia in children, wherein a vision impact database and a myopia prediction model are set up on a server, and a vision impact factor question library is established based on the vision impact database. The user terminal extracts a data collection questionnaire from the vision impact factor question library to collect the vision impact data of the examinee, and the myopia prediction model performs vision prediction. The system solves the problem that the existing prevention methods for myopia in children are insufficient and the diagnosis and treatment of myopia cannot be timely intervened, can guide the intervention of myopia in children, timely intervene in myopia diagnosis and treatment according to the trend of vision changes, reduce the occurrence of high myopia, and improve children's vision health.

[0149] The above describes in detail the structure, features and effects of the present invention based on the embodiments shown in the drawings. The above is only a preferred embodiment of the present invention, but the present invention is not limited to the scope of implementation shown in the drawings. Any changes made in accordance with the concept of the present invention, or modifications to equivalent embodiments with equivalent changes, which still do not exceed the spirit covered by the description and drawings, should be within the protection scope of the present invention.

Claims

1. A method for predicting myopia in children, characterized in that: include: Setting up a vision impact database and myopia prediction model; Analyzing and processing the data in the vision impact database to establish a vision impact factor question database; Extracting multiple questionnaire items from the vision influencing factor question bank using a feature matching algorithm to combine and adjust to obtain a data collection questionnaire; The vision impact data corresponding to the examinee is collected through the data collection questions and input into the myopia prediction model for vision prediction.

2. The method for predicting myopia in children according to claim 1, characterized in that: Also includes: When combining questionnaires, data association rules are established for questions of the same dimension, and association analysis is performed on questions of the same dimension to mine the corresponding vision impact data of the examinees, obtain the degree of relationship between different variables or individuals in the received vision impact data of the examinees, and the individual behavior patterns of the examinees.

3. The method for predicting myopia in children according to claim 2, characterized in that: The myopia prediction model is provided with a data collection layer, a data correlation analysis layer and a myopia prediction layer; The data collection layer is used to receive the visual impact data of the examinee to obtain the examination data and questionnaire data of the examinee; The data correlation analysis layer is provided with a clustering algorithm to perform cluster analysis on the inspection data and questionnaire data to obtain data correlation; The myopia prediction layer is provided with a prediction algorithm to perform vision prediction based on the correlation of the subject's vision impact data, and then generate a myopia prediction report.

4. The method for predicting myopia in children according to claim 3, characterized in that: The analyzing and processing of the data in the vision affecting database to establish a vision affecting factor question library includes: The data in the vision-influencing database are divided into different clusters through feature selection, distance measurement and clustering algorithm, and then classified into eight dimensions to explore the potential structure in the data, and the data are quantified according to the results of the classification analysis to obtain the vision-influencing factor question library; Among them, the eight dimensions include: children's pregnancy and delivery conditions, children's health status, children's parents and grandparents' family conditions, children's dietary conditions, children's sleep conditions, children's daily eye use conditions, children's eye environment and children's vision health intervention behaviors.

5. The method for predicting myopia in children according to claim 4, characterized in that: The collecting of the visual acuity impact data corresponding to the examinee through the data collection questions includes: Define the data missing types in the visual acuity impact data obtained from the examinee, the data missing types including: incomplete random missing, random missing and non-random missing; According to the type of missing data in the visual acuity impact data obtained from the examinees, the missing data are completed by selecting at least one of deleting rows, deleting columns, filling with the previous or next value, taking the average, taking the median value, using constant filling, linear interpolation, modeling prediction, and maximum likelihood estimation.

6. The method for predicting myopia in children according to claim 5, characterized in that: The inputting the myopia prediction model to perform vision prediction comprises: The myopia prediction model is used to perform correlation analysis on the visual acuity impact data corresponding to the examinee, and to perform data classification and prediction to generate a myopia prediction report corresponding to the examinee.

7. The method for predicting myopia in children according to claim 6, characterized in that: The data classification and prediction include: Classify vision impact data using hierarchical clustering, random forest, and / or K-means clustering methods; Predictions are made on vision impact data using decision trees, Bayesian belief networks, neural networks, and / or genetic algorithms.

8. The method for predicting myopia in children according to claim 7, characterized in that: The method of using a decision tree, a Bayesian belief network, a neural network and / or a genetic algorithm to predict the vision impact data includes: The prior probability of obtaining the visual acuity impact data of the examinee is calculated by the Bayesian belief network method; Class prediction is performed through the Naive Bayes classifier; Logistic regression model was used to handle data with binomial dependent variables and multiple classifications.

9. The method for predicting myopia in children according to claim 8, characterized in that: The method of predicting the vision impact data by using a decision tree, a Bayesian belief network, a neural network and / or a genetic algorithm also includes: The collected historical vision impact data are divided into 80%, 15%, and 5%, among which 80% of the data is used as a training data set, 15% of the data is used as a test data set, and 5% of the data is used as an independent sample set to train the myopia prediction model and verify the prediction effect.

10. A system for predicting myopia in children, using the method for predicting myopia in children according to any one of claims 1 to 9, characterized in that: include: Client and server; The server is provided with a vision impact database and a myopia prediction model, and a vision impact factor question library is established based on the vision impact database; The user terminal uses a feature matching algorithm to extract multiple questionnaire items from the vision influencing factor question library of the server and adjusts the items in combination to obtain a data collection questionnaire, and obtains the vision influencing data of the examinee through the data collection questionnaire; The user terminal pre-processes the acquired vision impact data of the examinee and sends the pre-processed data to the server; The server receives the vision impact data corresponding to the examinee, and inputs the myopia prediction model to perform vision prediction, so as to obtain the vision prediction result of the examinee, and then sends it to the corresponding user terminal; The server also stores the received vision impact data of the examinee into the vision impact database, and updates the historical data of the vision impact database; The myopia prediction model was trained and tested using updated historical data.