A method for constructing a glaucoma genetic information database
By constructing a glaucoma genetic information database, the problem of scattered genetic information has been solved, and the systematic integration and access control of gene datasets have been realized, thereby enhancing the application value of glaucoma genetic research and clinical diagnosis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NINGXIA HUI AUTONOMOUS REGION PEOPLES HOSPITAL
- Filing Date
- 2026-02-28
- Publication Date
- 2026-06-02
AI Technical Summary
In existing technologies, glaucoma genetic information is scattered and lacks a systematically integrated database platform, making it difficult for researchers to obtain comprehensive genetic information. Furthermore, there is a lack of platforms that can combine genetic information query, functional analysis, and access control, making it difficult to meet the needs of precision medicine.
A glaucoma genetic information database was constructed. Genetic information was collected through literature retrieval and high-throughput gene sequencing, and bioinformatics analysis was performed to form a gene dataset. The database backend architecture and frontend website platform were built, integrating basic gene information query, protein interaction network analysis, and user role access control mechanism.
This system integrates glaucoma genetic information, providing strong support for clinical diagnosis and genetic counseling, improving data security and relevance, and promoting the integration of glaucoma genetic research and clinical practice.
Smart Images

Figure CN122135795A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of database technology, and more specifically, to a method for constructing a glaucoma genetic information database. Background Technology
[0002] In the research and clinical practice of glaucoma, genetic factors play a crucial role in its pathogenesis, diagnosis, and prognosis. However, current glaucoma genetic information is fragmented, lacking a systematically integrated database platform. On the one hand, relevant genetic information is scattered across various literature and databases, making it difficult for researchers and clinicians to efficiently obtain comprehensive genetic information related to glaucoma. On the other hand, glaucoma patients in different regional populations exhibit specific gene variations, and existing studies lack systematic collection and analysis of gene variation data for specific regional populations, resulting in limitations in our understanding of the genetic mechanisms of glaucoma.
[0003] Meanwhile, at the data application level, there is a lack of platforms that can combine gene information query, functional analysis and access control, which makes it difficult to meet the needs of glaucoma genetic research and clinical practice in the context of precision medicine. Summary of the Invention
[0004] In view of the shortcomings of the existing technology, the purpose of this invention is to provide a method for constructing a glaucoma genetic information database.
[0005] To achieve the above objectives, the present invention provides the following technical solution:
[0006] A method for constructing a glaucoma genetic information database, the method comprising the following steps:
[0007] Genetic information on glaucoma-related genes was collected through literature retrieval and database queries;
[0008] Sequencing results obtained by high-throughput gene sequencing of the genomes of glaucoma patients;
[0009] Bioinformatics analysis was performed on the sequencing results to screen out mutations and variations associated with glaucoma;
[0010] The mutation and variation data and genetic information are used to form a preliminary gene dataset;
[0011] The database backend architecture is built using a database management system based on the gene dataset, and a frontend website platform with data retrieval function is configured.
[0012] The front-end website platform integrates gene basic information query, gene ontology enrichment analysis, and protein interaction network analysis for gene datasets.
[0013] A user role-based access control mechanism is set up on the front-end website platform, and data retrieval analysis and access management functions are integrated for clinical diagnostic assistance and genetic counseling.
[0014] Preferably, the genetic information includes the chromosomal location of the gene, the gene name, the gene description, the gene structure, the known mutation types and their role in the pathogenesis of glaucoma.
[0015] Preferably, the sequencing results obtained by high-throughput gene sequencing of glaucoma patients specifically include the following steps:
[0016] When performing high-throughput genome sequencing on glaucoma patients, sample groups with significant characteristics are selected by stratification based on differences in patients' clinical manifestations;
[0017] During high-throughput genome sequencing, the methylation modification status of the genomic DNA of the sample group is detected, and the modification status data is correlated with the base sequence data obtained from sequencing.
[0018] After high-throughput genome sequencing is completed, low-quality sequence fragments are removed by dynamic threshold screening, and variant sites are identified in the screened sequences in combination with the clinical disease progression information of the samples.
[0019] Mark specific variant regions associated with disease stages;
[0020] The identified variant sites are compared with the epigenetic modification characteristics of the sample group to screen out regions with sequence variations and epigenetic modification abnormalities. These regions are then used as variant regions for glaucoma to generate sequencing results that include sequence variations, epigenetic modifications, and clinical relevance information.
[0021] Preferably, bioinformatics includes quality control of sequencing data, sequence alignment, variant identification, and gene mutations and variations associated with the glaucoma phenotype.
[0022] Preferably, the sequencing results are subjected to bioinformatics analysis to screen for mutations and variations associated with glaucoma, specifically including the following steps:
[0023] We constructed a multi-omics analysis model that integrates clinical phenotypes of glaucoma, and mapped the gene sequence data obtained from sequencing with clinical phenotype data on changes in intraocular pressure and progression of optic nerve damage in patients.
[0024] Differential weights are assigned to different types of variant sites;
[0025] A deep learning-based variant function prediction module is introduced. This module is trained on the functional characteristics of a large number of known glaucoma pathogenic variants and performs functional pathogenicity classification on the selected variant sites.
[0026] By combining association mapping, differential weighting, and functional pathogenicity grading, mutation and variant data related to the pathogenesis of glaucoma were screened out.
[0027] Preferably, the mutation and variation data and genetic information are used to form a preliminary gene dataset, which specifically includes the following steps:
[0028] The gene location, mutation type, and clinical manifestations and family inheritance patterns of the mutation data are mapped in a matrix.
[0029] By performing time-series matching of mutation and variation data at different disease stages with corresponding genetic information change trajectories, associated combinations that exhibit characteristic changes as the disease progresses are identified.
[0030] The association combinations are divided into core association groups and secondary association groups according to the strength of association. The core association group consists of data combinations with strong associations that are directly related to the incidence of glaucoma, while the secondary association group consists of data combinations with weak associations that are indirectly related to the incidence of glaucoma.
[0031] The core association groups and secondary association groups are integrated in a hierarchical structure to form a gene dataset that includes association strength indicators and temporal variation features.
[0032] Preferably, a user role-based access control mechanism is set up on the front-end website platform, specifically including the following steps:
[0033] When setting up a user role-based access control mechanism on the front-end website platform, construct a dynamic role-based access control system;
[0034] User roles are segmented to clarify the core business objectives and data access scope of different roles, and basic permission thresholds are set for each role.
[0035] A dynamic permission adjustment module is introduced, which monitors user operation behavior and data access frequency in real time;
[0036] Establish an access control tracking system to record users' data access, query, and analysis operations throughout the entire process;
[0037] The dynamic role-based permission mapping system, the dynamic permission adjustment module, and the permission operation traceability system are integrated to form a complete access control mechanism.
[0038] Preferably, the search terms provided by the gene basic information query include the chromosome location of the gene, the gene name, the gene description, and the gene structure.
[0039] Preferably, the access control function specifically grants the database administrator full permissions to add, modify, delete data, and manage users.
[0040] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method for constructing a glaucoma genetic information database.
[0041] Compared with the prior art, the present invention has the following beneficial effects:
[0042] This invention constructs a glaucoma genetic information database, which systematically integrates glaucoma-related genetic information and specific gene variation data, providing strong support for clinical diagnostic assistance and genetic counseling. The database backend and frontend website platform enable functions such as basic gene information query and protein-protein interaction network analysis, facilitating in-depth research into the genetic mechanisms of glaucoma. A user role-based access control mechanism ensures the security and relevance of data use, enhancing the application value of the data in clinical and research scenarios and promoting the integrated development of glaucoma genetic research and clinical practice.
[0043] This invention systematically integrates scattered glaucoma genetic information and specific gene variation data by combining literature collection with high-throughput sequencing, forming a structured gene dataset. Based on this dataset, a database backend architecture and a frontend website platform are constructed. This platform integrates functions such as basic gene information query, gene ontology enrichment analysis, and protein-protein interaction network analysis, providing a professional data support platform for research on the genetic mechanisms of glaucoma. It enables the platform to be effectively applied in clinical diagnostic assistance and genetic counseling scenarios, providing a practical genetic information tool for precision medicine practices in glaucoma. Attached Figure Description
[0044] Figure 1 This is a schematic diagram illustrating the steps of constructing a glaucoma genetic information database according to the present invention.
[0045] Figure 2 This is a schematic diagram of the structure of the electronic device provided in an embodiment of the present invention;
[0046] Figure 3 This is a heatmap of differential gene expression between the glaucoma patient group and the normal control group in an embodiment of the present invention;
[0047] Figure 4 This is a graph showing the results of GO / KEGG pathway enrichment analysis conducted on the differentially expressed glaucoma gene set screened in this embodiment of the invention;
[0048] Figure 5 This is a diagram showing the results of immune microenvironment analysis in the optic nerve damage area of glaucoma in this embodiment of the invention;
[0049] Figure 6This is a diagram showing the protein-protein interaction network and drug enrichment analysis of key candidate genes in embodiments of the present invention.
[0050] Figure 7 This is a graph showing the intergroup comparison and ROC curve analysis of the key gene combinations in this invention.
[0051] Figure 8 This is a graph showing the enrichment status of the GSEA gene set in an embodiment of the present invention.
[0052] 610. Processor; 620. Communication interface; 630. Memory; 640. Communication bus. Detailed Implementation
[0053] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0054] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0055] Secondly, the term "an embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places throughout this specification does not necessarily refer to the same embodiment, nor is it a single embodiment or an embodiment selectively excluded from other embodiments.
[0056] Reference Figures 1-8 As shown.
[0057] The embodiments further illustrate the method for constructing a glaucoma genetic information database proposed in this invention.
[0058] A method for constructing a glaucoma genetic information database, the method comprising the following steps:
[0059] Genetic information on glaucoma-related genes was collected through literature retrieval and database queries;
[0060] Sequencing results obtained by high-throughput gene sequencing of the genomes of glaucoma patients;
[0061] Bioinformatics analysis was performed on the sequencing results to screen out mutations and variations associated with glaucoma;
[0062] The mutation and variation data and genetic information are used to form a preliminary gene dataset;
[0063] The database backend architecture is built using a database management system based on the gene dataset, and a frontend website platform with data retrieval function is configured.
[0064] The front-end website platform integrates gene basic information query, gene ontology enrichment analysis, and protein interaction network analysis for gene datasets.
[0065] A user role-based access control mechanism is set up on the front-end website platform, and data retrieval analysis and access management functions are integrated for clinical diagnostic assistance and genetic counseling.
[0066] The database backend architecture is built using a database management system based on gene datasets to store various glaucoma-related gene data. Simultaneously, a front-end website platform with data retrieval capabilities is configured, allowing users to easily search for data. For example, if a user wants to find a specific gene related to glaucoma, they can quickly retrieve its relevant information through the front-end platform.
[0067] The front-end website platform integrates multiple functions, including basic gene information query, gene ontology enrichment analysis, and protein interaction network analysis. Taking basic gene information query as an example, users can find information such as the chromosome location of a gene, gene name, gene description, and gene structure. Gene ontology enrichment analysis helps users understand the biological processes and molecular functions in which the gene participates, such as analyzing which signaling pathways a glaucoma-related gene mainly participates in. Protein interaction network analysis can show the interaction relationships between the protein encoded by the gene and other proteins, thereby helping researchers understand the mechanism of gene action at the protein level.
[0068] A user role-based access control mechanism is implemented on the front-end website platform. Different users have different permissions. For example, database administrators have full permissions for adding, modifying, and deleting data, as well as managing users, while ordinary clinicians may only have permissions for data querying and analysis. By integrating data retrieval and analysis with access control functions, the platform can be used for clinical diagnostic assistance and genetic counseling. For instance, clinicians can use the platform to query patients' gene variation information and, combined with the analysis functions within their permissions, provide a basis for the diagnosis and genetic counseling of glaucoma patients, determining the patient's genetic risk and possible pathogenesis.
[0069] Genetic information includes the chromosomal location of a gene, gene name, gene description, gene structure, known mutation types, and their role in the pathogenesis of glaucoma.
[0070] The sequencing results obtained by high-throughput gene sequencing of glaucoma patients include the following steps:
[0071] When performing high-throughput genome sequencing on glaucoma patients, sample groups with significant characteristics are selected by stratification based on differences in patients' clinical manifestations;
[0072] During high-throughput genome sequencing, the methylation modification status of the genomic DNA of the sample group is detected, and the modification status data is correlated with the base sequence data obtained from sequencing.
[0073] After high-throughput genome sequencing is completed, low-quality sequence fragments are removed by dynamic threshold screening, and variant sites are identified in the screened sequences in combination with the clinical disease progression information of the samples.
[0074] The formula for calculating the dynamic threshold is: ,in, For dynamic thresholds, The mean of the sequence quality values. This represents the standard deviation of the sequence quality values.
[0075] After performing high-throughput genome sequencing on a group of glaucoma patient samples, the mean quality value of the sequencing sequences in this sample group was statistically analyzed. =32, standard deviation of sequence quality value =2.5.
[0076] Calculate the dynamic threshold using the formula: =32+2×2.5=37.
[0077] The calculation results indicate that fragments with a quality value below 37 in the sequencing sequence will be identified as low-quality sequences and removed, thereby ensuring the reliability of data for subsequent identification of variant sites.
[0078] Mean sequence quality value in another set of high-throughput sequencing data from patients with advanced glaucoma =30, standard deviation of sequence quality value =3.
[0079] Substitute into the formula to calculate the dynamic threshold: =30+2×3=36.
[0080] In other words, sequence fragments with a quality value less than 36 in this data set will be screened out. Because the gene variation characteristics of samples from patients with advanced glaucoma are more complex, it is necessary to strictly control data quality to accurately identify specific variation regions and ensure that the screened data can accurately reflect the gene variation related to the advanced stage of the disease.
[0081] Mark specific variant regions associated with disease stages;
[0082] The identified variant sites are compared with the epigenetic modification characteristics of the sample group to screen out regions with sequence variations and epigenetic modification abnormalities. These regions are then used as variant regions for glaucoma to generate sequencing results that include sequence variations, epigenetic modifications, and clinical relevance information.
[0083] When performing high-throughput genome sequencing on glaucoma patients, the patients are first stratified based on differences in their clinical manifestations, selecting sample groups with significant characteristics. For example, patients with different intraocular pressure levels and degrees of optic nerve damage can be grouped separately to ensure sample diversity and representativeness. During high-throughput genome sequencing, the methylation modification status of the genomic DNA in the sample groups is detected. Then, the modification status data is correlated with the base sequence data obtained from sequencing for analysis, allowing for simultaneous understanding of gene sequence information and epigenetic modifications. After sequencing, low-quality sequence fragments are removed using dynamic threshold screening. Then, the filtered sequences are analyzed in conjunction with the clinical disease progression information of the samples to identify variant sites. Specific variant regions associated with disease stages are also marked, such as different variant regions appearing in the early, middle, and late stages of glaucoma. The identified variant sites are compared with the epigenetic modification characteristics of the sample group to screen out regions that simultaneously possess sequence variations and epigenetic modification abnormalities. These regions are then used as variant regions for glaucoma, ultimately generating sequencing results that include sequence variations, epigenetic modifications, and clinical association information. These results can provide detailed data support for studying the pathogenesis and progression of glaucoma. For example, by analyzing the association between specific variant regions and methylation modifications, we can gain a deeper understanding of the role of genes in the occurrence and development of glaucoma.
[0084] Bioinformatics includes quality control of sequencing data, sequence alignment, variant identification, and gene mutations and variations associated with glaucoma phenotypes.
[0085] Bioinformatics analysis of sequencing results and screening for mutations and variants associated with glaucoma include the following steps:
[0086] We constructed a multi-omics analysis model that integrates clinical phenotypes of glaucoma, and mapped the gene sequence data obtained from sequencing with clinical phenotype data on changes in intraocular pressure and progression of optic nerve damage in patients.
[0087] Differential weights are assigned to different types of variant sites;
[0088] The formula for calculating the weight is: , The weight of the i-th type of variant site, The score represents the association between the i-th type of variant site and the incidence of glaucoma, where n is the total number of variant site types.
[0089] Sequencing of a sample from a patient with primary open-angle glaucoma (POAG) identified three core associated variant sites: missense variants in the MYOC gene, nonsense variants in the OPTN gene, and variants at the WDR36 gene splice site. Bioinformatics analysis yielded the following association scores for the three types of sites: =8.2、 =7.5、 =6.3.
[0090] Calculate the sum of correlation scores =8.2 + 7.5 + 6.3 = 22. Calculate the weights of each type of locus and the weight of the MYOC gene variant. ≈0.37; OPTN gene variant weight ≈0.34; WDR36 gene variant weight ≈0.29.
[0091] The MYOC gene variant, due to its highest weight, has been included in the glaucoma pathogenesis core association group and is a variant type of priority for clinical diagnosis.
[0092] A deep learning-based variant function prediction module is introduced. This module is trained on the functional characteristics of a large number of known glaucoma pathogenic variants and performs functional pathogenicity classification on the selected variant sites.
[0093] By combining association mapping, differential weighting, and functional pathogenicity grading, mutation and variant data related to the pathogenesis of glaucoma were screened out.
[0094] To explore the protein-protein interactions of newly discovered key glaucoma genes and provide clues for the discovery of potential drug targets, we used the integrated protein-protein interaction (PPI) network analysis module and the drug-gene interaction database (DGIdb) to conduct a comprehensive analysis of the key candidate genes listed in Table 1. The results are as follows: Figure 6 As shown.
[0095]
[0096] Table 1: Key candidate gene characteristics identified in this study
[0097] Further combining fold change, statistical significance, and association with glaucoma biological function, we identified a group of "key candidate genes" most likely driving disease development, the characteristics of which are summarized in Table 1. For example, PLEC (reticulin) exhibited extremely high statistical significance (B-value = 11.62) and the largest fold change in expression (logFC = -0.89), suggesting that its loss of function in maintaining cellular mechanical stability may be closely related to optic nerve damage in glaucoma. BCL2L1, as a key regulator of apoptosis, was significantly downregulated in the patient group (logFC = -0.79), directly pointing to the activation of the ganglion cell apoptosis pathway. Genes such as FLNA, ITGA3 / 4 / 6, and LAMA5 collectively depict a molecular pathological atlas of glaucoma characterized by cytoskeleton disorder, abnormal cell-matrix adhesion, and enhanced apoptosis signaling. These tabular data are the core content of the database and the starting point for subsequent functional validation, biomarker development, and drug target screening.
[0098] Figure 6 Part A in the diagram represents a panoramic view of the PPI network, where nodes represent proteins and edges represent interactions. Core nodes in the network (such as FLNA and BCL2L1) exhibit high connectivity, suggesting their crucial positions within the biological network and their potential to influence multiple biological processes through extensive interactions with other proteins.
[0099] Figure 6 Part B in the image highlights a protein network that directly interacts with BCL2L1. BCL2L1, a key anti-apoptotic protein, is downregulated in glaucoma patients (see [link to image]). Figure 3 This may relatively enhance the activity of its interacting partners (such as pro-apoptotic proteins like BAX and BAD), thereby promoting retinal ganglion cell apoptosis. This network relationship provides a new perspective for understanding the molecular mechanisms of ganglion cell loss in glaucoma.
[0100] Figure 6 Part C of the study, combined with a drug-gene interaction database, performed drug enrichment analysis on key nodes in the PPI network. The prediction results showed potential targets for BCL2L1 (small molecule inhibitors targeting the BCL2 family, such as ABT-263), FLNA (compounds that regulate cytoskeleton stability), and the integrin family (ITGA6, etc., corresponding to integrin inhibitors such as silengiptide). For example, using BCL-2 inhibitors such as ABT-263 may protect neurons by regulating the apoptosis threshold. This analysis demonstrates the application value of this database in drug target prediction and the discovery of clues for new drug development, providing theoretical guidance for subsequent experimental validation and the development of new strategies for glaucoma treatment.
[0101] First, a multi-omics analysis model integrating clinical phenotypes of glaucoma is constructed. This model maps sequenced gene data with clinical phenotype data such as changes in intraocular pressure (IOP) and progression of optic nerve damage. For example, it correlates sequence variations of a gene with the clinical process of persistently elevated IOP, thus establishing a link between gene data and clinical phenotypes. Then, different types of variant sites are assigned differential weights; for instance, variant sites located in gene coding regions may have higher weights because they have a more direct impact on protein function. Next, a deep learning-based variant function prediction module is introduced. This module is trained on the functional characteristics of a large number of known glaucoma pathogenic variants and classifies the selected variant sites according to their functional pathogenicity, such as high pathogenicity, intermediate pathogenicity, and low pathogenicity. Finally, combining the association mapping results, differential weights, and functional pathogenicity classifications, mutational data related to the pathogenesis of glaucoma are screened. For example, a variant site that is strongly associated with IOP changes, has a high weight, and is classified as high pathogenic is identified as key data closely related to the pathogenesis of glaucoma. This data helps researchers gain a deeper understanding of the genetic pathogenesis of glaucoma.
[0102] The process of forming a preliminary gene dataset from mutation and variation data and genetic information includes the following steps:
[0103] The gene location, mutation type, and clinical manifestations and family inheritance patterns of the mutation data are mapped in a matrix.
[0104] By performing time-series matching of mutation and variation data at different disease stages with corresponding genetic information change trajectories, associated combinations that exhibit characteristic changes as the disease progresses are identified.
[0105] The association combinations are divided into core association groups and secondary association groups according to the strength of association. The core association group consists of data combinations with strong associations that are directly related to the incidence of glaucoma, while the secondary association group consists of data combinations with weak associations that are indirectly related to the incidence of glaucoma.
[0106] The core association groups and secondary association groups are integrated in a hierarchical structure to form a gene dataset that includes association strength indicators and temporal variation features.
[0107] To investigate changes in the immune microenvironment during optic nerve injury in glaucoma, we correlated gene expression data with clinical immune indicators and used the CIBERSORT algorithm to infer the proportions of various infiltrating immune cells in the samples. The results are as follows: Figure 5As shown in the figure, the left side lists the main immune cell types involved in the analysis, including Type 1 T helper cells, Type 2 T helper cells, CD56bright natural killer cells, Natural killer T cells, Activated B cells, Activated CD4+ T cells, Activated CD8+ T cells, Central memory CD4+ T cells, Central memory CD8+ T cells, Effector memory CD4+ T cells, Effector memory CD8+ T cells, and Memory B cells, providing complete cell type labeling for subsequent immune cell proportion analysis.
[0108] Figure 5 Part A of the graph is a stacked bar chart, showing the relative proportions of the various immune cells mentioned above in the glaucoma group and the control group. By comparing the immune cell composition of the two groups, the immune microenvironment characteristics of the optic nerve damage area in glaucoma patients can be visually observed. The results showed that the proportions of microglia and T cell subsets were significantly increased in the glaucoma group, suggesting that these cells may be involved in the neuroinflammatory response in glaucoma.
[0109] Figure 5 Part B of the chart is a box plot, which quantitatively compares the proportions of key differentially expressed immune cells (such as microglia, M0 macrophages, and CD8+ T cells) between the glaucoma group and the control group, and marks statistical significance (p < 0.05, p < 0.01). The box plot visually displays the median, quartile range, and outlier distribution of the data, validating the trends observed in the stacked bar chart.
[0110] Figure 5Section C of the graph is a correlation heatmap, showing the correlation between key characteristic genes (such as microglia marker genes like C1QA, TYROBP, and AIF1, and T cell marker genes like CD3D and CD8A) and the infiltration levels of various immune cells. Red indicates a positive correlation, blue indicates a negative correlation, and the intensity of the color represents the strength of the correlation. The results show that microglia marker genes are significantly positively correlated with the proportion of microglia infiltration and negatively correlated with some T cell subsets; while T cell marker genes are positively correlated with their corresponding T cell subsets. This analysis reveals a complex immune microenvironment characterized by dysregulation of innate immunity (microglia activation) and adaptive immunity (T cell infiltration) in the optic nerve injury area of glaucoma, providing a new perspective for understanding the neuroimmune regulatory mechanism of glaucoma.
[0111] This demonstrates the database's ability to integrate multi-dimensional data and reveal disease immune mechanisms. It showcases changes in the immune microenvironment related to glaucoma at multiple levels, fully demonstrating the unique application value of this database in integrating transcriptomics and immunomics data to reveal disease immune mechanisms.
[0112] First, a matrix-based association mapping is performed between the gene location, mutation type, and clinical manifestations and familial inheritance patterns in the genetic information corresponding to mutation and variant data. For example, a specific variant type of a glaucoma-related gene is correlated with clinical manifestations such as elevated intraocular pressure and optic nerve atrophy in patients, as well as the inheritance pattern of multiple cases in the family, forming a clear association matrix. Then, time-series matching is performed on the mutation and variant data at different disease stages and the corresponding genetic information change trajectories. For example, gene variants discovered in the early stages of glaucoma are observed to change their corresponding genetic information as the disease progresses to the middle and late stages, marking those association combinations that show characteristic changes with disease progression. For instance, a variant that only slightly alters the condition in the early stages may cause significant gene functional abnormalities in the later stages. These association combinations are divided into core association groups and secondary association groups according to the strength of the association. The core association group consists of strongly correlated data combinations directly related to the onset of glaucoma, such as a specific gene variant directly causing optic nerve damage. The secondary association group consists of weakly correlated data combinations indirectly related to the onset of glaucoma, such as a gene variant affecting the expression of other pathogenic genes. By integrating core association groups and secondary association groups in a hierarchical structure, a gene dataset containing association strength indicators and temporal variation characteristics is formed. Such a dataset can clearly show the relationship between different gene data and the onset of glaucoma in terms of association strength and time course, providing comprehensive and hierarchical information support for studying the genetic mechanism and disease progression of glaucoma.
[0113] To verify the validity and reliability of the data integrated in this database, we used the database's built-in analysis tools to extract retinal tissue transcriptome sequencing data from 50 patients with primary open-angle glaucoma (POAG) and 50 matched healthy controls, and performed a systematic differentially expressed gene (DEG) analysis. The analysis results are presented in heatmap format. Figure 3 .
[0114] Figure 3 Part A of the graph is a volcano plot of differential gene expression from the dataset GSE151631. This dataset compares the retinal transcriptomes of glaucoma patients and normal controls. The horizontal axis represents the fold change after transformation, and the vertical axis represents the p-value for significance after transformation. Red dots represent genes significantly upregulated in the patient group (e.g., MYOC, OPTN), blue dots represent genes significantly downregulated (e.g., BCL2L1, FLNA), and gray dots represent genes with no significant difference. The graph shows a significant increase in both the number and magnitude of upregulated genes, suggesting potential activation of certain pathways during glaucoma pathogenesis.
[0115] Figure 3 Part B in the diagram is a differential gene expression volcano plot of another independent dataset, GSE112155. The differential gene distribution pattern in this dataset differs from that in Part A, with more pronounced downregulated genes. By comparing the differential expression profiles of the two datasets, genes exhibiting robust changes across datasets can be identified; these genes are core driver genes for glaucoma. This database effectively overcomes batch effects and population heterogeneity inherent in single datasets by integrating multiple datasets.
[0116] Figure 3 Part C of the diagram presents a pie chart showing the proportional distribution of major retinal cell types inferred from a single-cell deconvolution algorithm in the normal control group and the keratoconus group. Keratoconus is an eye disease with certain comorbid mechanisms with glaucoma. The figure shows that, compared with the normal group, the proportion of fibroblast-like cells is significantly increased and the proportion of corneal epithelial cells is decreased in the keratoconus group. This change in cellular composition provides clues to understanding the cellular microenvironment remodeling in ocular degenerative diseases and suggests that this database can support cross-disease comparative analysis.
[0117] Figure 3Section D in the figure is a heatmap focusing on the expression of genes related to the JAK-STAT signaling pathway. The figure shows the expression patterns of the STAT family and its upstream and downstream genes (such as BCL2, PTEN, PTPN11, PAH, etc.) in glaucoma patients and normal controls in the GSE151631 dataset. Red indicates high expression, and blue indicates low expression. The results show that STAT3, STAT5A / B, etc., are consistently highly expressed in the patient group, while the expression of the negative regulator PTPN11 is downregulated, suggesting that abnormal activation of the JAK-STAT signaling pathway may be involved in the pathogenesis of glaucoma. This analysis demonstrates the database's ability to perform fine-grained analysis of specific pathways.
[0118] Figure 3 Part E of the graph is a heatmap of JAK-STAT signaling pathway-related gene expression based on another independent dataset, GSE112155. The graph shows the expression patterns of gene sets that are the same as or similar to those in Part D within this dataset. By comparing the two heatmaps, D and E, the consistency and differences in the expression changes of JAK-STAT pathway-related genes across different research cohorts can be observed. The results show that although the overall expression background may differ between the two datasets, the high expression trend of key genes such as STAT3 in the patient group is reflected in both datasets, enhancing the reliability of this finding. The dual-dataset comparative analysis function of this database effectively supports the research needs of cross-dataset validation.
[0119] Figure 3 Part F in the figure presents a box plot showing the comparison of STAT3 gene expression levels in two groups of samples from the GSE151631 dataset. STAT3 is a core transcription factor in the JAK-STAT pathway. The figure visually demonstrates that the expression level of STAT3 in the glaucoma patient group was significantly higher than that in the normal control group (p < 0.01). Figure 3 The trends observed in the D heatmap are highly consistent, verifying the accuracy of the heatmap analysis.
[0120] Figure 3 The G portion of the graph, presented as a box plot, shows the comparison of PTPN11 gene expression levels in two groups of samples from the GSE151631 dataset. PTPN11 encodes the SHP-2 protein, a negative regulator of the JAK-STAT pathway. The results showed that the expression level of PTPN11 in the glaucoma patient group was significantly lower than that in the normal control group (p<0.05), which may be one of the molecular mechanisms leading to the hyperphosphorylation of signaling molecules such as STAT3 and abnormal activation of the pathway.
[0121] This analysis not only visually reflects the characteristic transcriptomic changes associated with glaucoma, but also confirms the high reproducibility of the data integrated in the database and the reliability of the analytical tools, providing a solid data foundation for further in-depth exploration of biological mechanisms.
[0122] Using the Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analysis tools integrated into this database, we performed functional annotation and pathway enrichment analysis on the differentially expressed glaucoma genes selected above. The results are as follows: Figure 4 As shown. Figure 4 Part A of the results shows the enrichment analysis of GO biological processes (BP), which reveals that these differentially expressed genes are significantly enriched in processes such as extracellular matrix organization, collagen fibril organization, and response to oxidative stress. Figure 4 Part B is the enrichment analysis of GO cellular components (CC). The results show that the proteins encoded by differentially expressed genes are mainly located in the extracellular matrix, collagen trimer, and basement membrane. Figure 4 Part C is the GO molecule function (MF) enrichment analysis, which mainly enriches in "extracellular matrix structural constituent" and "growth factor binding". Figure 4 Part D is a bubble chart of KEGG pathway enrichment analysis. The results show that the most significantly enriched pathways include "ECM-receptor interaction", "protein digestion and absorption", and "PI3K-Akt signaling pathway". Figure 4Part F in the diagram represents the network diagram of KEGG pathway enrichment analysis. The diagram shows differentially enriched KEGG pathways, including those related to "focal adhesion," "ECM-receptor interaction," "regulation of actin cytoskeleton," and "small cell lung cancer." The diagram reveals that genes such as FLNA, ITGA3, ITGA4, ITGA6, BCL2L1, and PDGFA are closely associated with these pathways, which collectively participate in key pathological processes of glaucoma, including extracellular matrix remodeling, cell adhesion, cytoskeleton rearrangement, and apoptosis regulation. Figure 4 Part E, presented in the form of a directed acyclic graph (DAG), more intuitively demonstrates the hierarchical relationships and statistical significance of the enrichment results of biological processes and molecular functions. These results systematically reveal that abnormal remodeling of the extracellular matrix, trabecular meshwork dysfunction, and related signaling pathway perturbations are key molecular events in the pathogenesis of glaucoma. This is highly consistent with current understanding of the pathogenesis of glaucoma and strongly demonstrates the powerful function of this database in revealing the biological mechanisms of diseases.
[0123] Setting up a user role-based access control mechanism on the front-end website platform includes the following steps:
[0124] When setting up a user role-based access control mechanism on the front-end website platform, construct a dynamic role-based access control system;
[0125] User roles are segmented to clarify the core business objectives and data access scope of different roles, and basic permission thresholds are set for each role.
[0126] A dynamic permission adjustment module is introduced, which monitors user operation behavior and data access frequency in real time;
[0127] Establish an access control tracking system to record users' data access, query, and analysis operations throughout the entire process;
[0128] The dynamic role-based permission mapping system, the dynamic permission adjustment module, and the permission operation traceability system are integrated to form a complete access control mechanism.
[0129] The basic gene information query provides search options including the chromosome location of the gene, gene name, gene description, and gene structure.
[0130] When setting up a user role-based access control mechanism on a front-end website platform, the first step is to construct a dynamic role-based access control system. Next, user roles are subdivided, clarifying the core business objectives and data access scope of different roles. Basic access thresholds are then set for each role. For example, the core business objective of a database administrator is to manage data and users, so their data access scope covers all data, and their basic access threshold is full permissions for adding, modifying, deleting, and managing users. Conversely, the core business objective of a clinician is to use data for diagnostic assistance, and their data access scope mainly includes gene data related to clinical diagnosis. Their basic access threshold is permissions for querying basic gene information and performing some analysis functions. Then, a dynamic access control module is introduced. This module monitors user actions and data access frequency in real time. If a clinician frequently accesses a certain type of in-depth analysis data in a short period due to research needs, the dynamic access control module may increase their access permissions to that type of data within a certain range. Simultaneously, an access control operation traceability system is built to record the entire process of user data access, querying, and analysis operations. For example, it records when a database administrator adds a gene data entry or when a doctor queries gene mutation information for a specific patient. The dynamic role-based access control system, the dynamic access control module, and the access control operation traceability system are integrated to form a complete access control mechanism. The basic gene information query provides search options including the chromosome location of the gene, gene name, gene description, and gene structure. For example, users can search for the gene name "MYOC" to find that it is located on chromosome 1, the gene description is that it is involved in the formation of the extracellular matrix of the trabecular meshwork, and the gene structure contains information on multiple exons and introns.
[0131] The access control function specifically grants database administrators full permissions to add, modify, and delete data, as well as manage users.
[0132] The determination formula is: In this context, F represents the comprehensive screening score, α, β, and γ are weighting coefficients, M represents the association mapping matching degree, W represents the differential weight value, and P represents the pathogenicity grade score. When F ≥ F0 (F0 is the preset qualified threshold), the data is determined to be the target mutation variant.
[0133] Taking the CYP1B1 gene mutation site of a suspected congenital glaucoma patient as an example, the association mapping matching degree M=9.2, the mutation site weight W=0.35, and the pathogenicity grade score P=8.8.
[0134] Substituting the values into the formula, the overall score F = 0.3 × 9.2 + 0.4 × 0.35 + 0.3 × 8.8 = 5.54. Since F = 5.54 < 6.0, this variant site is not included in the core target data for the time being. Further analysis combined with clinical information revealed that this patient is a sporadic case with no evidence of familial co-segregation. Further expansion of the sample is needed for validation to avoid missing potential pathogenic sites or misidentifying non-pathogenic sites.
[0135] To evaluate the clinical diagnostic potential of the selected key genes, we used this database to perform single-gene diagnostic efficacy analysis on each candidate gene and further constructed a multi-gene combination diagnostic model. The results are as follows: Figure 7 As shown.
[0136] Figure 7 The table section presents the diagnostic efficacy metrics for 10 key candidate genes (BCL2L1, BMP1, BSG, ITGA3, ITGA4, ITGA6, LAMA5, NT5E, and PLEC). The table lists the 1-specificity (False Positive Rate, FPR) and sensitivity (True Positive Rate, TPR) for each gene. As can be seen from the table, at the set diagnostic threshold, all 10 genes exhibited ideal diagnostic performance: 1-specificity was 0.00 for all genes (i.e., a false positive rate of 0, meaning that the normal control group was correctly identified without misdiagnosis), and sensitivity was 0.80 for all genes (i.e., a true positive rate of 80%, meaning that 80% of glaucoma patients could be correctly identified). This result indicates that these genes, when used alone as biomarkers, possess high diagnostic accuracy.
[0137] Figure 7 Part A of the figure is a box plot, comparing the distribution of the comprehensive risk scores of the aforementioned 10 key genes between the glaucoma patient group and the normal control group. Machine learning algorithms (such as logistic regression and random forest) were used to integrate the expression values of these 10 genes into a comprehensive risk score. The figure shows that the median and distribution range of the risk score in the glaucoma patient group were significantly higher than those in the normal control group (p < 0.001), indicating that this gene combination can effectively distinguish patients from healthy individuals.
[0138] Figure 7 Part B of the figure shows the receiver operating characteristic (ROC) curve. The area under the curve (AUC) of the diagnostic prediction model based on the comprehensive risk score reached 0.95 (95% confidence interval: 0.91-0.98), which is significantly higher than the diagnostic efficacy of a single gene, indicating that the multi-gene combination model has higher diagnostic accuracy and discriminative ability. The AUC values of each single gene are also marked in the figure for reference.
[0139] Figure 7Part C in the figure is the calibration curve, which shows that the predicted probability of the model has good consistency with the actual observation. The calibration curve is close to the ideal diagonal, and the Hosmer-Lemeshow test p > 0.05 further verifies the reliability and stability of the model.
[0140] Figure 7 Part D in the diagram represents the decision curve analysis (DCA), which demonstrates the net benefit of using this diagnostic model at different threshold probabilities. The results show that, across a wide range of threshold probabilities, the net benefit of using this model is higher than the extreme strategies of "all treatment" or "no treatment," and also superior to single-gene models, demonstrating the model's clinical practical value.
[0141] Figure 7 Part E in the diagram represents the clinical impact curve, which estimates the number of people identified as high-risk by the model at different risk thresholds and the number of true positives among them, visually demonstrating the potential impact of the model on clinical decision-making.
[0142] Figure 7 A- Figure 7 E systematically evaluated the diagnostic value and clinical applicability of key gene combinations from multiple dimensions, including single-gene diagnostic efficacy, comprehensive risk score comparison, ROC curve, calibration curve, decision curve, and clinical impact curve. This analysis demonstrates that this database can not only mine key gene features but also support a complete diagnostic tool development process, from single-gene assessment to multi-gene integrated modeling, providing a practical tool for early screening and accurate diagnosis of glaucoma.
[0143] To reveal biological trends at a more macroscopic pathway level and avoid information loss that may result from relying solely on differential gene thresholds, we used the Gene Set Enrichment Analysis (GSEA) tool, an integrated database tool, to analyze genome-wide expression profiles. The results are as follows: Figure 8 As shown.
[0144] In addition, through differential expression analysis and multi-dimensional screening processes in this database, we systematically identified a batch of genes that are abnormally expressed in glaucoma, and a partial list is shown in Table 2.
[0145]
[0146] Table 2: List of some abnormally expressed genes identified in this study
[0147] These genes cover multiple functional categories, including the cytoskeleton (such as FLNA and PLEC), apoptosis (such as BCL2L1), cell adhesion (such as ITGA3 / 4 / 6 and LAMA5), and metabolism (such as NT5E), providing a wealth of candidate targets for subsequent research. Researchers can directly use this list to perform one-click enrichment analysis, PPI network construction, or drug database lookup on this platform.
[0148] Figure 8 The table section shows a list of significantly enriched Reactome and KEGG pathways identified through GSEA analysis. The table lists key metrics such as pathway name, rank in the ordered dataset, and entropy score. As can be seen from the table, several pathways related to metabolism, signal transduction, and disease showed a significant enrichment trend in the glaucoma patient group, including:
[0149] The Reactome pathway includes the synthesis of bile acids and bile salts, the metabolism of bile acids and bile salts, glycolysis, and the metabolism of amino acids and their derivatives.
[0150] KEGG pathway: Dilated cardiomyopathy
[0151] The enrichment of these pathways suggests that the pathogenesis of glaucoma may involve cross-talk between energy metabolism disorders, lipid metabolism abnormalities, and cardiomyopathy-related signaling pathways, providing new clues for a deeper understanding of the systemic metabolic effects of glaucoma.
[0152] Figure 8 Part A of the diagram shows the enrichment plot of the Reactome pathway for bile acid and bile salt synthesis. The horizontal axis represents a list of genes ordered by the correlation between gene expression and phenotype, and the vertical axis represents the running enrichment score. The curve shows a peak at the pathway gene set location and shifts to the left, indicating that this pathway gene set is generally upregulated in the glaucoma patient group (Normalized Enrichment Score, NES > 0, FDR q-val < 0.05).
[0153] Figure 8Part B of the diagram shows an enrichment map of the Reactome "bile acid and bile salt metabolism" pathway. Similar to Part A, this pathway was also significantly enriched in the patient group, suggesting that abnormalities in the bile acid metabolism pathway may be related to the pathological process of glaucoma.
[0154] Figure 8 Part C of the diagram shows an enrichment map of the KEGG "dilated cardiomyopathy" pathway. This pathway was significantly enriched in the patient group, suggesting that cardiomyopathy-related signaling molecules may be aberrantly expressed in optic nerve tissue, reflecting a potential molecular association between glaucoma and other systemic diseases.
[0155] Figure 8 Part D of the diagram shows an enrichment map of the Reactome glycolysis pathway. This pathway was significantly enriched in the upregulated direction in the patient group, suggesting a possible shift in energy metabolism from oxidative phosphorylation to glycolysis (the Warburg effect) in glaucoma optic nerve tissue to adapt to the hypoxic or mitochondrial dysfunctional microenvironment.
[0156] Figure 8 Part E in the diagram shows an enrichment map of the Reactome "amino acid and its derivative metabolism" pathway. The enrichment of this pathway reflects the reprogramming of amino acid metabolism during the pathogenesis of glaucoma, which may be closely related to processes such as neurotransmitter synthesis and redox balance.
[0157] Figure 8 Part F in the figure presents the expression patterns of the top-ranking enriched genes in the aforementioned pathways in the glaucoma patient group and the normal control group in the form of a heatmap. As can be seen from the figure, these genes exhibit a clear differential expression trend between the two groups, further validating the reliability of the enrichment analysis.
[0158] Figure 4 Part G in the diagram is a network diagram of the enrichment results, showing the shared gene relationships between significantly enriched pathways. Nodes in the diagram represent pathways, node size represents enrichment significance, and lines indicate overlapping genes between pathways. This network diagram reveals the interrelationships between pathways such as bile acid metabolism, glycolysis, and amino acid metabolism, suggesting that synergistic changes in these metabolic pathways may jointly participate in the pathogenesis of glaucoma.
[0159] This systematically revealed glaucoma-related metabolic reprogramming and signal transduction abnormalities at the pathway level. These results not only validated... Figure 2 The enrichment analysis results also expanded our understanding of metabolic disorders in glaucoma, fully demonstrating the powerful function of this database in analyzing pathway activity at the whole genome level, and providing a theoretical basis for exploring metabolic therapeutic targets for glaucoma.
[0160] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements a method for constructing a glaucoma genetic information database.
[0161]
[0162] Table 3: Key Indicators for Glaucoma Gene Data Collection and Analysis
[0163] Table 3 shows the key indicators for glaucoma gene data collection and analysis. This glaucoma genetic information database construction trial included 150 glaucoma patient samples, divided into 5 groups based on clinical manifestation differences, providing a representative sample basis for subsequent data collection and analysis. In the high-throughput gene sequencing stage, sequencing coverage reached 98.2%, ensuring the comprehensiveness of gene sequence data. Meanwhile, low-quality sequences accounted for only 2.8%, and quality control methods such as dynamic threshold screening ensured the reliability of the sequencing data. In the variant site identification stage, a total of 320 variant sites were detected, of which 86 were specific variant sites related to the course of glaucoma, providing core targets for subsequent targeted analysis. Bioinformatics analysis showed that highly pathogenic variant sites accounted for 35%, and these sites were verified to be closely related to the pathogenesis of glaucoma through association mapping, weight allocation, and deep learning prediction. During the dataset construction process, 128 core association groups were formed. These data are directly and strongly correlated with the pathogenesis of glaucoma, providing core data support for the database. The front-end platform's functional verification results show that the gene query accuracy rate reaches 99.1%, and the running time of various analysis functions does not exceed 15 seconds, which not only ensures the accuracy of data query but also meets the needs of efficient use in clinical and research scenarios.
[0164] like As shown, the electronic device may include a processor 610, a communication interface 620, a memory 630, and a communication bus 640, wherein the processor 610, the communication interface 620, and the memory 630 communicate with each other through the communication bus 640. The processor 610 can call logical instructions in the memory 630 to execute a method for constructing a glaucoma genetic information database.
[0165] Furthermore, the logical instructions in the aforementioned memory 630 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory, random access memory, magnetic disks, or optical disks.
[0166] On the other hand, the present invention also provides a computer program product, the computer program product including a computer program that can be stored on a non-transitory computer-readable storage medium, and when the computer program is executed by a processor, the computer is able to execute a method for constructing a glaucoma genetic information database.
[0167] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform a method for constructing a glaucoma genetic information database.
[0168] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0169] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0170] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for constructing a glaucoma genetic information database, characterized in that, The method includes the following steps: Genetic information on glaucoma-related genes was collected through literature retrieval and database queries; Sequencing results obtained by high-throughput gene sequencing of the genomes of glaucoma patients; Bioinformatics analysis was performed on the sequencing results to screen out mutations and variations associated with glaucoma. The mutation and variation data and genetic information are used to form a preliminary gene dataset; The database backend architecture is built using a database management system based on the gene dataset, and a frontend website platform with data retrieval function is configured. The front-end website platform integrates gene basic information query, gene ontology enrichment analysis, and protein interaction network analysis for gene datasets. A user role-based access control mechanism is set up on the front-end website platform, and data retrieval analysis and access management functions are integrated for clinical diagnostic assistance and genetic counseling.
2. The method for constructing a glaucoma genetic information database according to claim 1, characterized in that, The genetic information includes the chromosomal location of the gene, gene name, gene description, gene structure, known mutation types, and their role in the pathogenesis of glaucoma.
3. The method for constructing a glaucoma genetic information database according to claim 1, characterized in that, The sequencing results obtained by high-throughput gene sequencing of glaucoma patients include the following steps: When performing high-throughput genome sequencing on glaucoma patients, sample groups with significant characteristics are selected by stratification based on differences in patients' clinical manifestations; During high-throughput genome sequencing, the methylation modification status of the genomic DNA of the sample group is detected, and the modification status data is correlated with the base sequence data obtained from sequencing. After high-throughput genome sequencing is completed, low-quality sequence fragments are removed by dynamic threshold screening, and variant sites are identified in the screened sequences in combination with the clinical disease progression information of the samples. Mark specific variant regions associated with disease stages; The identified variant sites are compared with the epigenetic modification characteristics of the sample group to screen out regions with sequence variations and epigenetic modification abnormalities. These regions are then used as variant regions for glaucoma to generate sequencing results that include sequence variations, epigenetic modifications, and clinical relevance information.
4. The method for constructing a glaucoma genetic information database according to claim 1, characterized in that, Bioinformatics includes quality control of sequencing data, sequence alignment, variant identification, and gene mutations and variations associated with glaucoma phenotypes.
5. The method for constructing a glaucoma genetic information database according to claim 1, characterized in that, Bioinformatics analysis of sequencing results and screening for mutations and variants associated with glaucoma include the following steps: We constructed a multi-omics analysis model that integrates clinical phenotypes of glaucoma, and mapped the gene sequence data obtained from sequencing with clinical phenotype data on changes in intraocular pressure and progression of optic nerve damage in patients. Differential weights are assigned to different types of variant sites; A deep learning-based variant function prediction module is introduced. This module is trained on the functional characteristics of a large number of known glaucoma pathogenic variants and performs functional pathogenicity classification on the selected variant sites. By combining association mapping, differential weighting, and functional pathogenicity grading, mutation and variant data related to the pathogenesis of glaucoma were screened out.
6. The method for constructing a glaucoma genetic information database according to claim 1, characterized in that, The process of forming a preliminary gene dataset from mutation and variation data and genetic information includes the following steps: The gene location, mutation type, and clinical manifestations and family inheritance patterns of the mutation data are mapped in a matrix. By performing time-series matching of mutation and variation data at different disease stages with corresponding genetic information change trajectories, associated combinations that exhibit characteristic changes as the disease progresses are identified. The association combinations are divided into core association groups and secondary association groups according to the strength of association. The core association group consists of data combinations with strong associations that are directly related to the incidence of glaucoma, while the secondary association group consists of data combinations with weak associations that are indirectly related to the incidence of glaucoma. The core association groups and secondary association groups are integrated in a hierarchical structure to form a gene dataset that includes association strength indicators and temporal variation features.
7. The method for constructing a glaucoma genetic information database according to claim 1, characterized in that, Setting up a user role-based access control mechanism on the front-end website platform includes the following steps: When setting up a user role-based access control mechanism on the front-end website platform, construct a dynamic role-based access control system; User roles are segmented to clarify the core business objectives and data access scope of different roles, and basic permission thresholds are set for each role. A dynamic permission adjustment module is introduced, which monitors user operation behavior and data access frequency in real time; Establish an access control tracking system to record users' data access, query, and analysis operations throughout the entire process; The dynamic role-based permission mapping system, the dynamic permission adjustment module, and the permission operation traceability system are integrated to form a complete access control mechanism.
8. The method for constructing a glaucoma genetic information database according to claim 1, characterized in that, The search terms provided by the gene basic information query include the chromosome location of the gene, gene name, gene description, and gene structure.
9. The method for constructing a glaucoma genetic information database according to claim 1, characterized in that, The access control function specifically grants database administrators full permissions to add, modify, and delete data, as well as manage users.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method for constructing a glaucoma genetic information database as described in any one of claims 1 to 9.