Biomarker for predicting myopia progress speed of teenagers and application of biomarker

By using a combination of genetic variations rs76371606, rs10811260, rs7900033, rs73309782, rs35548593, rs201190960 and rs3785874 as biomarkers, the problem that the existing technology cannot effectively predict the progression of myopia in adolescents is solved, and the effects of early screening and auxiliary diagnosis are achieved.

CN120648793APending Publication Date: 2025-09-16SUZHOU CENT FOR DISEASE CONTROL & PREVENTION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511116528.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-09-16

Smart Images

  • Figure CN120648793A_ABST
    Figure CN120648793A_ABST
Patent Text Reader

Abstract

The invention relates to a biomarker composition for predicting the progression speed of myopia of teenagers. The biomarker composition comprises more than two of the following genetic variations: rs76371606, rs10811260, rs7900033, rs73309782, rs35548593, rs201190960 and rs3785874. The invention also relates to a method for predicting the progression speed of myopia of teenagers. According to the invention, the correlation between the biomarker composition and the development speed of juvenile myopia is clarified for the first time, and the biomarker composition disclosed by the invention can be used in the fields of early screening and auxiliary diagnosis of juvenile myopia susceptibility, research and development of related new drugs and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of biological detection, and in particular to a biomarker for predicting the progression of myopia in adolescents and its application. Background Art

[0002] Myopia is a refractive error that mainly occurs in children and adolescents. Due to excessive elongation of the eye axis, parallel light rays enter the eyeball and focus in front of the retina, resulting in blurred vision.

[0003] The etiology of myopia is complex, with both genetic and environmental factors playing a significant role in its development. For children and adolescents with a genetic predisposition, timely identification of high-risk individuals is crucial to improve the effectiveness of myopia intervention. To date, the NHGRI-EBI GWAS Catalog (https: / / www.ebi.ac.uk / gwas / ) has reported approximately 560 susceptibility loci that have achieved genome-wide significance (P < 5.0 × 10-8) for myopia and related traits in genome-wide association studies (GWAS), but most variants remain unidentified. In 2020, a GWAS-Meta-analysis identified 336 new genetic markers associated with refractive error, the strongest of which was rs12193446 on the LAMA2 gene. In 2022, MTOR and PDGFRA genotyping studies found that rs1057079 and rs1064261 on the MTOR gene were associated with mild / moderate myopia. Recent GWAS identified LILRB2 (lead SNP: rs7247538) located at 19q13.42 as a new candidate locus for pathological myopia. Overall, some polygenic risk score (PRS) models constructed using genetic variants have shown high predictive accuracy in predicting myopia risk.

[0004] Currently, there are few clinical studies on Chinese adolescents who are susceptible to myopia, so there is an urgent need to study biomarkers that can predict the speed of myopia progression in this group of people. Summary of the Invention

[0005] To solve the above technical problems, the present invention includes the following aspects: The first aspect of the present invention provides a biomarker composition for predicting the progression of myopia in adolescents, wherein the biomarker composition comprises two or more of the following genetic variations: rs76371606, rs10811260, rs7900033, rs73309782, rs35548593, rs201190960 and rs3785874.

[0006] Preferably, the biomarker composition comprises three or more, four or more, five or more, or six or more of the following genetic variations: rs76371606, rs10811260, rs7900033, rs73309782, rs35548593, rs201190960, and rs3785874.

[0007] More preferably, the biomarker composition comprises a combination of the following genetic variations: rs76371606, rs10811260, rs7900033, rs73309782, rs35548593, rs201190960, and rs3785874.

[0008] Preferably, the rs76371606 is a C / A polymorphism, the reference allele is A, and the alternative allele is C; the rs10811260 is an A / G polymorphism, the reference allele is G, and the alternative allele is A; the rs7900033 is an A / G polymorphism, the reference allele is G, and the alternative allele is A; the rs73309782 is a C / T polymorphism, the reference allele is T, and the alternative allele is C; the rs35548593 is a TA / T polymorphism, the reference allele is T, and the alternative allele is TA; the rs201190960 is a C / CA polymorphism, the reference allele is CA, and the alternative allele is C; the rs3785874 is a T / G polymorphism, the reference allele is G, and the alternative allele is T.

[0009] Preferably, the standard for the speed of myopia progression in adolescents is as follows: adolescents whose myopia degree deepens by ≥1.0D within 2 years are defined as myopia progression, and adolescents whose myopia degree does not deepen or myopia degree deepens by <1.0D are defined as myopia non-progression.

[0010] Preferably, the adolescent is an Asian adolescent. More preferably, the adolescent is an East Asian adolescent. Further preferably, the adolescent is a Chinese adolescent.

[0011] A second aspect of the present invention provides a detection kit for predicting the progression of myopia in adolescents, wherein the kit comprises a reagent for detecting a biomarker composition, wherein the biomarker composition comprises two or more of the following genetic variations: rs76371606, rs10811260, rs7900033, rs73309782, rs35548593, rs201190960 and rs3785874.

[0012] Preferably, the biomarker composition comprises three or more, four or more, five or more, or six or more of the following genetic variations: rs76371606, rs10811260, rs7900033, rs73309782, rs35548593, rs201190960, and rs3785874.

[0013] More preferably, the biomarker composition comprises a combination of the following genetic variations: rs76371606, rs10811260, rs7900033, rs73309782, rs35548593, rs201190960, and rs3785874.

[0014] The third aspect of the present invention provides the use of the above-mentioned biomarker composition in predicting the progression rate of myopia in adolescents.

[0015] Preferably, the use is for non-diagnostic and non-therapeutic purposes.

[0016] A fourth aspect of the present invention provides the use of the above-mentioned detection kit in the preparation of a reagent for predicting the progression rate of myopia in adolescents.

[0017] The technical effects produced by the present invention are: This study, published in Nature Communications, for the first time, clarifies the correlation between the genetic variants rs76371606, rs10811260, rs7900033, rs73309782, rs35548593, rs201190960, and rs3785874 and the progression of myopia in Chinese adolescents. The biomarker combination of the present invention can be used for early screening, auxiliary diagnosis, and related drug development for adolescent myopia susceptibility. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 is a Manhattan plot of the discovery phase of the myopia progression GWAS of the present invention; Figure 2 This is the Locuszoom image of the myopia progression-associated variant rs76371606; Figure 3 This is the Locuszoom image of the myopia progression-associated variant rs10811260; Figure 4 This is the Locuszoom image of the myopia progression-associated variant rs7900033; Figure 5 This is the Locuszoom image of the myopia progression-associated variant rs73309782; Figure 6 This is the Locuszoom image of the myopia progression-associated variant rs35548593; Figure 7 This is the Locuszoom image of the myopia progression-associated variant rs201190960; Figure 8 This is the Locuszoom image of the myopia progression-associated variant rs3785874; Figure 9 This is the receiver operating characteristic curve for predicting the progression of myopia in adolescents using the seven related variations of the present invention. DETAILED DESCRIPTION

[0019] Experimental Example 1: Genome-wide association study of the progression of myopia in Chinese adolescents 1. Test method 1.1 Test subjects A total of 294 first-year high school students were recruited from five middle schools in Suzhou. Baseline data were collected in September 2020, and the last follow-up was in March 2023. All participants underwent annual refraction examinations by a professional ophthalmologist (fully automated computer ophthalmometer: Topcon RM 800), and all participants rested outdoors for at least half an hour before refraction. Adolescents whose myopia worsened by 1.0 D or more during the two-year follow-up were defined as myopia progression cases, while those whose myopia did not worsen or whose myopia worsened by less than 1.0 D were defined as myopia non-progression controls. Participants had no congenital eye disease, family history of glaucoma, history of ocular trauma, organic lesions of the cornea or fundus, or any prior ocular treatment or surgery. This study employed a two-stage design, consisting of a discovery phase (screening cohort) and a replication phase (validation cohort). The discovery cohort consisted of 89 cases and 87 controls from three middle schools in Suzhou. The replication cohort consisted of 59 cases and 59 controls from the remaining two middle schools in Suzhou. Case and control subjects were frequency matched by sex and baseline age (±1 year).

[0020] The research procedures of this trial were in accordance with the principles of the Declaration of Helsinki and were approved by the Ethics Committee of Suzhou Center for Disease Control and Prevention. All participants were informed of the content and purpose of the study and signed written informed consent.

[0021] 1.2 DNA extraction and genotyping Genomic DNA was extracted from human peripheral blood according to standard laboratory methods. Genotyping was performed on 294 individuals in the discovery and replication phases using the Infinium Asian Screening Array-24 v1.0 array (Illumina, San Diego, California, USA). This array, optimized for East Asian genomic characteristics, contains approximately 660,000 fixed marker loci. After hybridization of DNA samples with probes on the array, the array was scanned using the Illumina iScan system, generating high-resolution images. The images were read and further processed using Genome Studio software to generate genotype data for each sample.

[0022] 1.3 Quality Control in the Discovery Phase of Genome-Wide Association Studies Quality control was performed using PLINK 1.9 and R 4.4.1. (1) Sample quality control: Samples whose sex inferred from X chromosome genotyping was inconsistent with the reported sex were excluded; samples with a call rate lower than 97% were excluded; samples with an individual heterozygosity rate that deviated from the overall mean ± 3 standard deviations were excluded; recessive kinship between samples was estimated by consanguinity, and samples with low call rates in sample pairs with PI_HAT > 0.1875 were removed; principal component analysis (PCA) was performed using the smartpca program in EIGENSOFT v7.2.1 software (https: / / hsph.harvard.edu / research / price-lab / software / ). Genotype data from the 1000 Genomes Project Phase 3 were downloaded (http: / / www.1000genomes.org), and CHB / CHS (N = 208) population data were extracted.

[0023] Genotype data from myopia progression cases and controls were combined with the CHB / CHS data for ancestral PCA analysis. PCA analysis was performed on myopia progression cases and controls, excluding the MHC region (chr6: 25 MB-34 MB, hg19). The top 10 PCs were calculated for each iteration, and five iterations were performed. Samples with significant population stratification, defined as one or more PCs with a deviation of more than 6 standard deviations from the overall mean, were removed. The impact of different PCs on the GWAS statistical summary results was evaluated, and the top seven PCs were ultimately selected as covariates.

[0024] (2) The quality control of genetic variants was as follows: variants with a call rate lower than 97% were excluded; variants with a significant difference in missing data rates between cases and controls were excluded ( P <1.0×10 -5); remove variants that significantly deviate from the Hardy-Weinberg equilibrium in the control group ( P <1.0×10 -5 Variants with a minor allele frequency (MAF) < 0.01 were removed. After rigorous quality control, a total of 465,936 autosomal variants (89 cases and 87 controls) were included in the subsequent imputation analysis.

[0025] 1.3 Genotype filling The data were staged using SHAPEIT2. Genotype imputation was performed using Minimac3 using the 1000 Genomes Project Phase 3, Release 5 reference panel. Post-implication quality control was performed as follows: variants with an imputation quality score below 0.7 and variants with a mean average force (MAF) below 0.01 were removed. A total of 6,569,222 autosomal variants were used for association analysis.

[0026] 1.4 Correlation Analysis Genetic association analysis was performed using logistic regression additive models in PLINK 1.9 for the discovery phase data, adjusting for sex, baseline age, and the first seven PCs as covariates. QQ plots and Manhattan plots were generated using the CMplot package in R 4.4.1 (https: / / github.com / YinLiLin / CMplot).

[0027] 1.5 Replication Studies and Meta-Analysis This study was replicated in an independent sample and quality controlled using the same procedures and standards as the discovery phase. After rigorous quality control, 458,439 variants from 59 cases and 59 controls were retained and genotyped. P <1.0×10 -4 Variants were included in the replication association analysis. Meta-analysis of discovery and replication data was performed using the inverse variance-weighted fixed-effect model in METAL (version 2011-03-25) (https: / / genome.sph.umich.edu / wiki / METAL). Regional association maps of significant variants were generated using the LocusZoom website (http: / / locuszoom.org / ).

[0028] 1.6 Linkage disequilibrium fractional regression analysis In GWAS studies, linkage disequilibrium score regression (LDSC) analysis aims to distinguish polygenic effects from the influence of confounding factors. A calculated intercept closer to 1 indicates that genomic expansion is primarily due to polygenicity rather than confounding factors such as population stratification. Reference LD scores for East Asian populations were downloaded from the open digital repository (https: / / zenodo.org / records / 10515792) and summary statistics for the discovery-phase GWAS were analyzed.

[0029] 1.7 Functional Annotation The functional roles of the significant variants were investigated using HaploReg v4.2 (https: / / pubs.broadinstitute.org / mammals / haploreg / haploreg.php) and RegulomeDB 2.2 (https: / / www.regulomedb.org / regulome-search / ). On March 16, 2025, the associations between the discovered variants and gene expression regulation were investigated using expression quantitative trait locus (eQTL) and splicing quantitative trait loci (sQTL) data in the GTEx Portal v10 online database (https: / / www.gtexportal.org / home / ).

[0030] 1.8 MAGMA analysis To explore the enrichment of myopia progression GWAS associations in specific functions or pathways, gene set analysis was performed on the discovery-phase GWAS data using MAGMA v1.1 (https: / / cncr.nl / research / magma / ). GO / KEGG terms were downloaded from the MSigDB v2024.1 database (https: / / www.gsea-msigdb.org / gsea / index.jsp).

[0031] 2. Test results 2.1 General demographic characteristics A total of 148 cases and 146 controls passed quality control in the two phases. There were no statistically significant differences in gender, baseline age, and baseline refractive status between the two groups ( P>0.05). The mean baseline age of the case group was 15.64±0.34 years, with 47.97% male and 52.03% female. The mean baseline age of the control group was 15.66±0.34 years, with 48.63% male and 51.37% female. The participants' baseline refractive status and refractive progression during follow-up are shown in Table 1.

[0032] Table 1 Demographic and ocular characteristics of participants

[0033] Note: a OD: Right eye; OS: Left eye.

[0034] 2.2 Genome-wide association study and replication analysis In the discovery phase of the GWAS, 6,569,222 autosomal variants from 89 myopia progression cases and 87 healthy controls were subjected to rigorous quality control and imputed for further association analysis. GC ) was 1.063, showing only a slight inflation of the test statistic, and the intercept obtained from LDSC analysis was 1.003 (95% CI: 0.988–1.018), which was not significantly different from 1, indicating that the inflation was mainly due to polygenicity, and no obvious population stratification was found in the population. Association analysis found 443 suggestive associations ( P <1.0×10 -4 ) of the variation ( Figure 1 , where the X-axis shows the chromosome where the genetic variant is located and the Y-axis shows the -log10 of all variants ( P ) value. The blue horizontal line is indicative P Threshold ( P =1.0×10 -4 ), the red horizontal line indicates the genome-wide significant P Threshold ( P =5.0×10 -8 ).

[0035] To validate the findings of the discovery phase, 443 patients with myopia progression were enrolled in an independent sample of 59 cases and 59 controls. P <1.0×10 -4 A replication study was conducted on the variants of the discovery and replication cohorts. A meta-analysis of the association results of the discovery and replication cohorts showed that 7 new loci were suggestively associated with myopia progression, including FSTL5 in the 4q32.2 region (lead variant: rs76371606, P =1.32×10 -5, OR=3.06); 9p24.3 region SMARCA2 (lead variant: rs10811260, P =2.93×10 -5 , OR=2.32); 10p13 region CCDC3 (lead variant: rs7900033, P =4.32×10 -6 , OR=2.59); 12q13.13 region GALNT6 (lead variant: rs73309782, P =1.32×10-5, OR=3.45); 12q23.3 region CRY1 (lead variant: rs35548593, P =5.09×10 -5 , OR=2.39); ULK2 in the 17p11.2 region (lead variant: rs201190960, P =1.74×10 -6 , OR=0.38); and MYL4 in the 17q21.32 region (lead variant: rs3785874, P =7.34×10 -5 , OR=0.49) (see Figure 2-8 ,in Figure 2-5 and 8 in the lead single nucleotide polymorphism (SNP) and Figure 6-7 The lead insertion-deletion variant (Indel) in the chromatin is represented by purple, and the color coding of other variants indicates the LD relationship with the lead variant as follows: red indicates r 2 ≥0.8, orange indicates 0.6≤r 2 <0.8, green means 0.4≤r 2 <0.6, light blue indicates 0.2≤r 2 <0.4, dark blue indicates r 2 <0.2, grey indicates unknown).

[0036] Table 2 Association results of the 7 sites identified in Experimental Example 1

[0037] Table 2 (Continued) Association results of the 7 sites identified in Experimental Example 1

[0038] aAbbreviations: Chr, chromosome; Pos, position; OR, odds ratio; AF:allele frequency (case AF / control AF). b A1: alternative allele; A2: reference allele. c I 2 heterogeneity index of meta-analysis 2.3 Functional Evaluation rs10811260 is located in an intergenic region, while the remaining six variants are distributed in introns. These noncoding variants may have significant effects on gene expression. The HaploReg v4.2 database predicts that one variant is located in a promoter histone mark region, four in an enhancer histone mark region, and one in a deoxyribonuclease hypersensitive region. All seven variants alter transcription factor binding motifs (Table 3), suggesting potential gene regulatory functions. Furthermore, the RegulomeDB 2.2 database predicts that MYL4 rs3785874, FSTL5 rs76371606, and CRY1 rs35548593 have strong supportive evidence for regulatory functions, with ranking scores of 3a, 4, and 5, respectively, and corresponding model scores of 0.60, 0.61, and 0.62, suggesting that these variants may affect transcription factor binding. However, no significant evidence of transcription factor binding was found for the other variants (Table 3). The GTExPortal v10 database showed that two variants (GALNT6 rs73309782, MYL4 rs3785874) significantly affected gene expression. rs73309782 affected the expression of ACVR1B, a downstream protein of GALNT6, and rs3785874 affected the expression of MYL4 and its downstream protein EFCAB13-DT (Table 3).

[0039] Table 3 Functional annotations of the seven variants identified in Experiment 1

[0040] Table 3 (continued) Functional annotations of the seven variants identified in Experiment 1

[0041] aAbbreviations: ADRL, adrenal gland; BLD, blood; BRN, brain; ESDR, embryonic stem cells derived; GI, gastrointestinal tract; HRT, heart; MUS, muscle; SPLN, spleen; THYM, thymus. b BLD, ESDR, ADRL, HRT, GI, MUS, THYM, SPLN. c Pax-4, Pax-5, TCF12, Zec. d Foxp1, HDAC2, Hoxa4, Zfp105, p300. e Ranking score: 3a, TF binding + any motif + chromatin accessibilitypeak; 4, TF binding + chromatin accessibility peak; 5, TF binding or chromatin accessibility peak; 6, Motif hit; 7, Other. f Model score: range from 0 to 1, with 1 being most likely to be aregulatory variant. 2.4 GO / KEGG gene enrichment analysis GO / KEGG gene enrichment analysis was performed using MAGMA v1.1. The GWAS results in the discovery phase enriched GOBP_RETICULOPHAGY ( P = 3.41×10 -5 ), GOBP_NEGATIVE_REGULATION_OF_PROTEIN_MATURATION ( P = 9.49 × 10 -5 ), GOBP_REGULATION_OF_RETROGRADE_TRANSPORT_ENDOSOME_TO_GOLGI ( P=0.00021), GOCC_PEPTIDASE_INHIBITOR_COMPLEX ( P = 0.00092), GOCC_EUKARYOTIC_TRANSLATION_INITIATION_FACTOR_3_COMPLEX_EIF3M ( P = 0.0011), KEGG_REGULATION_OF_AUTOPHAGY ( P = 0.0011) and other GO / KEGG terms, suggesting that the risk variants may promote myopia progression by regulating these pathways.

[0042] 2.5 ROC Curve Analysis Seven variants, including rs76371606, rs10811260, rs7900033, rs73309782, rs35548593, rs201190960, and rs3785874, obtained by the present invention, were selected and their diagnostic performance for the progression of myopia in adolescents was analyzed using a receiver operating characteristic (ROC) curve. The experimental results showed that the seven variants had a strong performance in predicting the progression of myopia in adolescents, with an area under the ROC curve (AUC) value of 0.802 (see Figure 9 ), suggesting that the biomarkers of the present invention have good clinical distinguishing efficacy.

[0043] Although specific embodiments of the present invention have been described, it will be appreciated by those skilled in the art that various changes and modifications may be made to the present invention without departing from the scope or spirit of the present invention. Therefore, the present invention is intended to cover all such changes and modifications that fall within the scope of the appended claims and their equivalents.

Claims

1. A biomarker composition for predicting the progression of myopia in adolescents, characterized in that: The biomarker composition comprises two or more of the following genetic variations: rs76371606, rs10811260, rs7900033, rs73309782, rs35548593, rs201190960, and rs3785874.

2. The biomarker composition according to claim 1, characterized in that The biomarker composition comprises three or more, four or more, five or more, or six or more of the following genetic variations: rs76371606, rs10811260, rs7900033, rs73309782, rs35548593, rs201190960, and rs3785874.

3. The biomarker composition according to claim 2, characterized in that The biomarker composition comprises a combination of the following genetic variations: rs76371606, rs10811260, rs7900033, rs73309782, rs35548593, rs201190960, and rs3785874.

4. A detection kit for predicting the progression of myopia in adolescents, the kit comprising reagents for detecting a biomarker composition, wherein the biomarker composition comprises two or more of the following genetic variations: rs76371606, rs10811260, rs7900033, rs73309782, rs35548593, rs201190960 and rs3785874.

5. The detection kit according to claim 4, characterized in that The biomarker composition comprises three or more, four or more, five or more, or six or more of the following genetic variations: rs76371606, rs10811260, rs7900033, rs73309782, rs35548593, rs201190960, and rs3785874.

6. The detection kit according to claim 5, characterized in that The biomarker composition comprises a combination of the following genetic variations: rs76371606, rs10811260, rs7900033, rs73309782, rs35548593, rs201190960, and rs3785874.

7. Use of the biomarker composition according to any one of claims 1 to 3 in predicting the progression of myopia in adolescents.

8. Use of the detection kit according to any one of claims 4 to 6 in the preparation of a reagent for predicting the progression of myopia in adolescents.