A prediction method, device and equipment for a genetic disease
By obtaining biological information, using generative artificial intelligence screening model and multiomic data analysis, combined with hereditary disease prediction model, the problem of low prediction accuracy of hereditary disease is solved, efficient screening and early intervention of hereditary disease is achieved, and the accuracy of disease diagnosis and prognosis is improved.
Patent Information
- Application Number
- CN202411187576.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-28
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2044-08-28
AI Technical Summary
The prior art cannot effectively improve the prediction accuracy of hereditary diseases and cannot accurately predict them based on biomarkers.
By obtaining the biological information of the target population, a generative artificial intelligence screening model is used to screen out biomarkers related to hereditary diseases, and combining multiomic data analysis and hereditary disease prediction model to predict hereditary disease risks.
It improves the screening and prediction efficiency of hereditary diseases, achieves early detection and early intervention, improves the accuracy of disease diagnosis and prognosis, has broad market demand and good application prospects.
Smart Images

Figure CN119170282B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical data analysis, and in particular, to a method, device, and equipment for predicting genetic diseases. Background Art
[0002] Genetic diseases are a class of diseases caused by changes in genetic material, including single-gene genetic diseases, polygenic inheritance, and chromosomal abnormality genetic diseases, etc. The occurrence and development of these diseases involve complex genetic and environmental factors. With the in-depth study of biomedical research, many biomarkers related to genetic diseases have been discovered, and these biomarkers can be used as the basis for disease diagnosis and prognosis.
[0003] However, these biomarkers are not completely inherited. Therefore, the prediction of later genetic diseases cannot be completely judged based on these biomarkers.
[0004] Therefore, how to improve the prediction accuracy of genetic diseases is a technical problem that needs to be solved urgently at present. Summary of the Invention
[0005] In view of the above problems, the present invention provides a method, device, and equipment for predicting genetic diseases that overcome the above problems or at least partially solve the above problems.
[0006] In a first aspect, the present invention provides a method for predicting genetic diseases, including:
[0007] Obtain the target biological information of the target population;
[0008] Based on the target biological information, use a screening model to screen out the target biomarkers related to genetic diseases;
[0009] Based on the target biological information and the target biomarkers, use multi-omics data analysis to obtain the target correlation result between the target biomarkers and genetic diseases;
[0010] Based on the target biomarkers and the target correlation result, use a genetic disease prediction model to predict the genetic disease risk of the target population.
[0011] Preferably, the construction method of the screening model is as follows:
[0012] Collect the first historical biological information of the first historical population, where the first historical population includes a part of the population with genetic diseases and a part of the population without genetic diseases, and the first historical biological information is marked with the first historical biomarkers related to genetic diseases;
[0013] Based on the first historical biological information, train a neural network model to obtain a screening model.
[0014] Preferably, the biological information includes:
[0015] genome, epigenome, transcriptome, and proteome;
[0016] The multi-omics data analysis includes: genome data analysis, epigenomic data analysis, transcriptomic data analysis, and proteomic data analysis.
[0017] Preferably, when performing genome data analysis, the multi-omics data analysis is performed based on the target biological information and the target biomarker to obtain the correlation result between the target biomarker and the genetic disease, including:
[0018] Based on bioinformatics software and databases, determine the gene location, quantity, structure, and function of the DNA sequence of the whole-genome sequencing data in the target biological information;
[0019] Based on the gene location, quantity, structure, and function, determine the expression of the target biomarker in a preset tissue or environment by analyzing the transcriptome under different conditions;
[0020] Based on the gene location, quantity, structure, and function, compare the corresponding genomes in the parental and offspring populations to determine the evolutionary relationship, gene family expansion and contraction, as well as the innovation and loss of function between the corresponding genomes of the target biomarker;
[0021] Based on the gene location, quantity, structure, and function, as well as the characteristics of biological elements such as known transcription start sites, splicing sites, promoters, enhancers, and miRNA binding sites, predict new functional elements of the target biomarker;
[0022] Based on the gene location, quantity, structure, and function, determine whether the target biomarker participates in a preset biological process or pathway by performing enrichment analysis on the first gene where the target biomarker is located;
[0023] Based on the gene location, quantity, structure, and function, determine the effect of the target biomarker on a preset phenotype and its correlation with a preset disease by performing genetic effect analysis on the second gene where the target biomarker is located;
[0024] Based on the gene location, quantity, structure, and function, predict whether the target biomarker has an impact on protein structure and function by simulating protein structure and function.
[0025] Preferably, when performing epigenetic data analysis, based on the target biological information and the target biomarker, multi-omics data analysis is used to obtain the correlation results between the target biomarker and genetic diseases, including:
[0026] Based on the target biological information and the target biomarker, any one or more of DNA methylation, histone modification, polycomb proteins, chromatin remodeling, and histone regulation are used to alter the epigenome to determine the effects of the target biomarker represented by the alteration on gene expression, cell differentiation, and the growth and development process.
[0027] Preferably, when performing transcriptome data analysis, based on the target biological information and the target biomarker, multi-omics data analysis is used to obtain the correlation results between the target biomarker and genetic diseases, including:
[0028] Based on the target biological information and the target biomarker, a preset process of transcriptome data analysis is used to obtain genes with differential expression;
[0029] Enrichment analysis is performed on the genes with differential expression to determine the related biological processes and pathways regulated by the target biomarker.
[0030] Preferably, when performing protein data analysis, based on the target biological information and the target biomarker, multi-omics data analysis is used to obtain the correlation results between the target biomarker and genetic diseases, including:
[0031] By comparing the expression of the proteome in different individuals of the same population, differentially expressed proteins are determined;
[0032] Network analysis is performed on the differentially expressed proteins to determine the interconnections of the differentially expressed proteins in different individuals;
[0033] Based on the differentially expressed proteins, key nodes are determined through functional annotation and enrichment analysis;
[0034] Based on the relevant associations and the key nodes, the key proteins, biological functions, and signaling pathways of the target biomarker are determined.
[0035] Preferably, the method for constructing the genetic disease prediction model includes:
[0036] Collect the second historical biomarkers of the second historical population and the historical correlation results between the second historical biomarkers and genetic diseases;
[0037] Input the second historical biomarker and the historical correlation result into a machine learning model for training to obtain a genetic disease prediction model. The machine learning model is specifically any one of the following:
[0038] Logistic regression model, random forest model, and support vector machine model.
[0039] In a second aspect, the present invention also provides a prediction device for genetic diseases, including:
[0040] An acquisition module, configured to acquire target biological information of a target population;
[0041] A screening module, configured to screen out biomarkers related to genetic diseases based on the target biological information by using a screening model;
[0042] An analysis module, configured to obtain a target correlation result between the target biomarker and the genetic disease by using multi-omics data analysis based on the target biological information and the target biomarker;
[0043] A prediction module, configured to predict the genetic disease risk of the target population by using a genetic disease prediction model based on the target biomarker and the target correlation result.
[0044] In a third aspect, the present invention also provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method described in the first aspect is implemented.
[0045] One or more technical solutions in the embodiments of the present invention have at least the following technical effects or advantages:
[0046] The present invention provides a prediction method for genetic diseases, including: acquiring target biological information of a target population; screening out target biomarkers related to genetic diseases based on the target biological information by using a screening model, obtaining a target correlation result between the target biomarker and the genetic disease by using multi-omics data analysis based on the target biological information and the target biomarker; predicting the genetic disease risk of the target population by using a genetic disease prediction model based on the target biomarker and the target correlation result. By analyzing the biomarker and its correlation result with the genetic disease and combining the correlation result with the biomarker, the risk of genetic diseases can be accurately predicted. Description of the Drawings
[0047] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the following detailed description of the preferred embodiments. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0048] Figure 1 The schematic flow chart of the steps of the method for predicting genetic diseases in the embodiments of the present invention is shown;
[0049] Figure 2 The schematic code diagram for predicting and annotating repetitive sequences of genes in the embodiments of the present invention is shown;
[0050] Figure 3 The schematic structural diagram of the device for predicting genetic diseases in the embodiments of the present invention is shown;
[0051] Figure 4 The schematic structural diagram of the computer device for implementing the method for predicting genetic diseases in the embodiments of the present invention is shown. Detailed Embodiments
[0052] The exemplary embodiments of the present disclosure will be described in more detail below with reference to the drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.
[0053] Embodiment 1:
[0054] The embodiments of the present invention provide a method for predicting genetic diseases, including:
[0055] S101, obtaining the target biological information of the target population;
[0056] S102, based on the target biological information, using a screening model to screen out the target biomarkers related to genetic diseases;
[0057] S103, based on the target biological information and the target biomarkers, using multi-omics data analysis to obtain the target correlation results between the target biomarkers and genetic diseases;
[0058] S104, based on the target biomarkers and the target correlation results, using a genetic disease prediction model to predict the genetic disease risk of the target population.
[0059] First, a biomarker is an indicator of key events related to the pathogenesis occurring within an organism, and it represents any measurable changes in the physiology, biochemistry, immunology, and genetics of organs, cells, and sub-cells of the body caused by exposure to various environmental factors. It can be used for disease diagnosis, determining the disease stage, or evaluating the safety and effectiveness of new drugs or new therapies in a pre-specified population.
[0060] The biological information here includes: genome, epigenome, transcriptome, and proteome.
[0061] There are many types of biomarkers, such as proteins, nucleic acids, lipids, metabolites, etc.
[0062] To accurately screen for biomarkers, a screening model needs to be constructed first. The biomarker screening of the present invention is based on generative artificial intelligence. Specifically, a generative artificial algorithm is used to perform deep learning and pattern recognition on the biological information of a large-scale population to screen out novel biomarkers closely related to genetic diseases. Using such a method can greatly improve the efficiency and accuracy of biomarker screening.
[0063] Among them, the method for constructing the screening model is as follows:
[0064] Collect the first historical biological information of the first historical population, where the first historical population includes a group of individuals with some genetic diseases and a group of individuals without genetic diseases, and the first historical biological information is marked with historical biomarkers related to genetic diseases; based on the first historical biological information, train a neural network model to obtain the screening model.
[0065] Examples of biomarkers related to genetic diseases are as follows: Tumor markers for lymphoma:
[0066] Serum lactate dehydrogenase (LD or LDH), an intracellular enzyme, is released into the blood when cells are damaged or destroyed. In lymphoma patients, the LDH level may increase.
[0067] Serum β2-microglobulin (β2-MG), a plasma protein, is usually filtered by the kidneys and excreted in the urine. In lymphoma patients, the level of β2-MG may increase.
[0068] Serum alpha-fetoprotein (AFP), a protein produced during fetal development, sometimes also increases in malignant tumors such as lymphoma, but its increase is not specific and is commonly used for diagnosing other types of tumors.
[0069] Antinuclear antibody (ANA), an autoantibody, may be positive in some lymphoma patients.
[0070] Serum β-2-globulin may also be elevated in some lymphoma patients.
[0071] Leukocyte alkaline phosphatase (WBC-AP), an enzyme, is normally present in white blood cells, but its level may be elevated in certain lymphoma patients.
[0072] During the model training process, the training data used includes: some groups with hereditary diseases who carry disease biomarkers; some groups with hereditary diseases who do not carry disease biomarkers; some groups without hereditary diseases who carry disease biomarkers; some groups without hereditary diseases who do not carry disease biomarkers.
[0073] These training data are input into the neural network model for training, and thus a screening model is obtained. This screening model can screen out disease biomarkers related to hereditary diseases from biological information.
[0074] Although there are these biomarkers to label lymphoma patients, however, some biomarkers do not necessarily lead to disease and are not necessarily inherited, so further analysis is needed.
[0075] Since biological information contains a lot of privacy information, before training the screening model, for the collected training data, that is, the sample data, it is necessary to ensure the security of the collection and storage of biological information. Therefore, a federated learning blockchain network is adopted, integrating a centralized and decentralized big data management and collaborative analysis platform, which uses an enterprise-level blockchain solution (Hyperledger Fabric) as the underlying blockchain platform. The encryption technology adopted by this platform can effectively ensure the security and privacy of data, and ensure that biological data is not tampered with or leaked.
[0076] After obtaining the screening model, next, perform S101 to obtain the target biological information of the target population.
[0077] Specifically, the target population can be diseased users or non-diseased users. The obtained target biological information can be obtained by analyzing blood to get the genome, epigenome, transcriptome, and proteome.
[0078] Among them, by using high-throughput sequencing technologies on blood, such as next-generation sequencing (NGS) and third-generation sequencing technologies, DNA sequencing is achieved to obtain the genome. By analyzing epigenetic markers such as DNA methylation and histone modification to analyze the distribution of markers in the genome, the epigenome is obtained. By extracting RNA from cells and performing sequencing analysis, the gene expression under specific conditions can be obtained to get the transcriptome. By detecting and analyzing plasma in blood, the proteome is obtained.
[0079] Next, perform S102. Based on the target biological information, use a screening model to screen out target biomarkers related to genetic diseases. Among them, the screening model is used to screen out biomarkers related to genetic diseases from biological information.
[0080] This step is the application of the screening model. First, input the biological information obtained in S101 into the screening model, and then the screening model outputs target biomarkers related to genetic diseases.
[0081] Next, in order to accurately predict the genetic diseases of the target population, perform S103. Based on the target biological information and target biomarkers, use multi-omics data analysis to obtain the target correlation results between the target biomarkers and genetic diseases.
[0082] After obtaining the target biomarkers of the target population, through multi-omics analysis, obtain the target correlation results between the biomarkers and genetic diseases.
[0083] Among them, multi-omics analysis includes: genomic data analysis, epigenetic data analysis, transcriptomic data analysis, and proteomic data analysis.
[0084] Specifically, when performing genomic data analysis, it is specifically achieved through the following methods:
[0085] Based on bioinformatics software and databases, determine the gene positions, quantities, structures, and functions of the DNA sequences of the whole-genome sequencing data in the target biological information;
[0086] Based on the gene positions, quantities, structures, and functions, by analyzing the transcriptome under different conditions, determine the expression conditions of the target biomarkers in a preset tissue or environment, specifically including high expression and low expression. This is an analysis from the gene expression profile, and rich expression information can be obtained in a short time. Among them, the frequency of the appearance of the target biomarkers represents the expression level of the gene, and this expression level also represents the degree of correlation between the target biomarkers and genetic diseases.
[0087] Based on the gene positions, quantities, structures, and functions, compare the corresponding genomes in the parental and offspring populations to determine the evolutionary relationships, amplification and contraction situations of gene families, and innovation and loss situations of functions between the corresponding genomes of the target biomarkers. This is an analysis from genome comparison. This genome comparison analysis can reveal gene functions and disease molecular mechanisms, clarify species evolutionary relationships and the internal structure of genomes. Through such analysis, determine the probability of the biomarker being inherited, and whether it is dominant inheritance, recessive inheritance, or in a cured state or restored recessive state, etc.
[0088] Based on the gene position, quantity, structure and function, as well as the characteristics of biological elements such as known transcription start sites, splicing sites, promoters, enhancers, and miRNA binding sites, new functional elements of the target biomarker are predicted. This is prediction from functional elements, thereby enabling prediction of the response of the elements to different environmental signals, such as living temperature, light, etc., which in turn affects gene expression. This has important implications for an individual's growth and development, metabolic processes, and the occurrence and development of diseases.
[0089] Based on the gene position, quantity, structure and function, by performing enrichment analysis on the first gene where the target biomarker is located, it is determined whether the target biomarker is involved in a preset biological process or pathway. This is functional enrichment analysis. Its essence is cluster analysis, by interpreting the biological knowledge represented by the first gene where the target biomarker is located, thereby revealing its role inside or outside the cell.
[0090] Based on the gene position, quantity, structure and function, by performing genetic effect analysis on the second gene where the target biomarker is located, the influence of the target biomarker on a preset phenotype and the correlation with a preset disease are determined; this is genetic effect analysis from mutations. By analyzing the influence of the target biomarker on the morphological structure, physiological function, and genetic information of a biological individual, the influence on the evolution of the biological population and the stability of the ecosystem is analyzed, and thus the correlation of the target biomarker with the preset phenotype and preset disease is obtained.
[0091] Based on the gene position, quantity, structure and function, by simulating protein structure and function, it is predicted whether the target biomarker has an impact on protein structure and function. This is protein structure prediction and analysis. Since the structure of a protein determines its biological function, for target biomarkers of the protein type, the biological function can be predicted by predicting the protein structure, that is, predicting whether there are related diseases such as protein folding diseases.
[0092] Among them, when determining the gene position, quantity, structure and function, RepeatMasker (a commonly used tool for detecting repetitive sequences) can be used to predict and annotate repetitive sequences of genes. Among them, information such as the proportion of repetitive sequences, repetitive sequence types, and quantity is counted, as shown in Figure 2 the code.
[0093] When using epigenetic data analysis, it is specifically implemented in the following way:
[0094] Based on the target biological information and target biomarker, any one or more of DNA methylation, histone modification, polycomb proteins and chromatin remodeling, and histone regulation are used to alter the epigenome to determine the effects of the target biomarker represented by the alteration on gene expression, cell differentiation, and the processes of growth and development.
[0095] Among them, DNA methylation can silence gene expression or activate transcription, while histone modification can regulate the structure and function of chromatin, thereby affecting gene accessibility. Therefore, epigenetic data analysis examines how these changes participate in regulating cell functions and characteristics, as well as their roles in disease occurrence and development.
[0096] For example, hepatocellular carcinoma (HCC) is a highly invasive malignant disease. Epigenome analysis regulates gene expression without changing the DNA sequence. Abnormal epigenetic changes can alter the expression of oncogenes or tumor suppressor genes, leading to tumorigenesis. Abnormal epigenetic changes may result in phenotypic changes such as tumor cell growth, immune escape, metastasis, heterogeneity, and chemoresistance. By comprehensively reviewing genomic data, the roles of epigenetics in the tumor microenvironment (TME), immune cell infiltration characteristics, and the inflammatory response of HCC are determined.
[0097] Transcriptome data analysis is carried out in the following specific ways:
[0098] Based on the target biological information and target biomarker, the preset process of transcriptome data analysis is adopted to obtain genes with differential expression;
[0099] Enrichment analysis is performed on the genes with differential expression to determine the related biological processes and pathways regulated by the target biomarker.
[0100] Among them, the preset process includes: quality control, data preprocessing, and differential analysis. Quality control mainly performs quality control on sequencing data, evaluates data quality, and excludes low-quality data. Commonly used quality control tools include FastQC and MultiQC. Data preprocessing involves preprocessing operations such as removing sequence adapters (trimming) and filtering the raw sequencing data to ensure the accuracy and reliability of subsequent differential analysis. Differential analysis is to identify genes with significantly different expression levels in different samples, and commonly used methods include DESeq2 and edgeR, etc. The specific code is as follows:
[0101] # Import the original count file
[0102] countData<-read.table(" / path / to expression data",header=TRUE,row.names=1)
[0103] # Create sample information
[0104] colData <- data.frame(condition = c(rep("Control", 3), rep("Treatment", 3)), row.names = colnames(countData))
[0105] # Create DESeq object
[0106] dds <- DESeqDataSetFromMatrix(countData = countData, colData = colData, design = ~condition)
[0107] # Perform quality control and filtering
[0108] dds <- dds[rowSums(counts(dds)) >= 10, ]
[0109] dds <- DESeq(dds)
[0110] # Perform differential analysis
[0111] res <- results(dds)
[0112] # Output differential expression matrix
[0113] write.table(res, file = "deseq2_results.txt", sep = "\t", quote = FALSE)
[0114] Finally, perform enrichment analysis on the differentially expressed genes to determine the relevant biological processes and pathways regulated by the target biomarkers. The code for the Goseq enrichment analysis tool is as follows:
[0115] # Initialize the GSEABase package
[0116] library(GSEABase)
[0117] setwd(" / path / to / your / gene / ids")
[0118] # Load gene ids
[0119] gene.ids.common <- read.csv("geneID2Symbol.txt", header = F, sep = "\t")[, 1]
[0120] # Run the goseq function
[0121] over.list <- goseq(de, "hsa", "canonical", gene.ids = gene.ids.common)
[0122] print(over.list)
[0123] When analyzing protein data, it is specifically achieved through the following methods:
[0124] By comparing the expression levels of different proteins in different individuals within the same population, differentially expressed proteins are determined;
[0125] Perform network analysis on the differentially expressed proteins to determine their interconnections in different individuals;
[0126] Based on the differentially expressed proteins, through functional annotation and enrichment analysis, key nodes are determined;
[0127] Based on the interconnections and key nodes, the key proteins, biological functions, and signaling pathways of the target biomarker are determined.
[0128] By performing PCA analysis (principal component analysis) on the preprocessed data, the degree of dispersion of the data is obtained, thereby obtaining the separation effect between different groups. After removing individual values, differentially expressed proteins are screened, and pathway analysis is performed based on the differentially expressed proteins, thereby determining the correlation results between the target biomarker and genetic diseases.
[0129] The analysis of each correlation result analyzes the correlation between the biomarker and genetic diseases from different aspects. Considering the diversity of various components in the population organism, from multiple perspectives, various omics data such as genomics, metabolomics, and proteomics are analyzed. Thus, the target correlation results between the target biomarker of the target population and genetic diseases can be obtained. The target correlation results indicate the degree of correlation between each target biomarker and the corresponding genetic disease.
[0130] The above - mentioned analysis methods of these correlation results can be used alone or in combination. There is no limitation here.
[0131] These target correlation results can be used as an effective basis for predicting genetic diseases in the later stage.
[0132] Therefore, execute S104. Based on the target biomarker and the target correlation results, use a genetic disease prediction model to predict the risk of genetic diseases.
[0133] Prior to this, it is necessary to first construct a genetic disease prediction model, including:
[0134] Collect the second historical biomarkers of the second historical population and the historical correlation results between the second historical biomarkers and genetic diseases;
[0135] Input the second historical biomarkers and the historical correlation results into a machine learning model for training to obtain a genetic disease prediction model. The machine learning model is specifically any one of the following:
[0136] Logistic regression model, random forest model, and support vector machine model.
[0137] The historical correlation results between the second historical biomarkers of the second historical population and genetic diseases can be obtained according to S103, and the second historical biomarkers of the second historical population can be obtained in the same way as the first historical biomarkers were labeled in the first historical population.
[0138] After obtaining the second historical biomarkers of the second historical population and the historical correlation results between the second historical biomarkers and genetic diseases, input the second historical biomarkers and the historical correlation results into a machine learning model for training to obtain a genetic disease prediction model.
[0139] Specifically, input the historical biomarkers, the historical correlation results, and the corresponding historical disease status into any one of the logistic regression model, random forest model, and support vector machine model for training to obtain a genetic disease prediction model.
[0140] By combining the second historical biomarkers and the historical correlation results between the second historical biomarkers and genetic diseases, the risk of genetic diseases can be effectively predicted, avoiding the problem of low prediction accuracy that occurs when predicting only based on biomarkers.
[0141] The result output by the genetic disease prediction model is the prediction confidence for genetic diseases, that is, the probability of suffering from the genetic disease. Among them, it includes the prediction confidences for various genetic diseases. The prediction confidences for various genetic diseases are output in the order of confidence.
[0142] Finally, after obtaining the genetic disease prediction model, input the target biomarkers and the target correlation results of the target population collected by S102 and S103 into the genetic disease prediction model, so as to output the risk degree of the genetic diseases of the target population, that is, the prediction confidences for various genetic diseases.
[0143] Adopting the technical solution of the present invention can greatly improve the screening and prediction efficiency of genetic diseases, contribute to the early detection and early intervention of genetic diseases, thereby improving the prevention and treatment level of genetic diseases. Secondly, the present invention can effectively process high-dimensional and complex biological data, contribute to the discovery of more biological markers related to genetic diseases, thereby improving the accuracy of disease diagnosis and prognosis. Finally, it can also provide sufficient generalization ability and interpretability, contribute to improving the practicability and reliability of the genetic disease screening and prediction model. The present invention has broad market demand and good application prospects.
[0144] For example, if a parent has a genetic disease and is currently in the advanced stage, the offspring does not necessarily have the genetic disease and will not necessarily be in the advanced stage either. On the one hand, the target biological markers of the offspring can be screened out, and on the other hand, the target correlation result with the genetic disease can be obtained by analyzing the target biological markers. Then, the target biological markers and the target correlation result are input into the genetic disease prediction model, and finally, the genetic disease risk of the offspring is output, including the inheritance rate of various genetic diseases.
[0145] One or more technical solutions in the embodiments of the present invention at least have the following technical effects or advantages:
[0146] The present invention provides a method for predicting genetic diseases, including: obtaining the target biological information of the target population; based on the target biological information, using a screening model to screen out the target biological markers related to genetic diseases; based on the target biological information and the target biological markers, using multi-omics data analysis to obtain the target correlation result between the target biological markers and genetic diseases; based on the target biological markers and the target correlation result, using a genetic disease prediction model to predict the genetic disease risk of the target population. By analyzing the biological markers and their correlation results with genetic diseases and combining the correlation results with the biological markers, the risk of genetic diseases can be accurately predicted.
[0147] Embodiment 2:
[0148] Based on the same inventive concept, the embodiment of the present invention also provides a device for predicting genetic diseases, as Figure 3 shown, including:
[0149] An acquisition module 301, configured to acquire the target biological information of the target population;
[0150] A screening module 302, configured to screen out the target biological markers related to genetic diseases based on the target biological information by using a screening model;
[0151] An analysis module 303, configured to obtain a target correlation result between the target biomarker and the genetic disease by performing multi-omics data analysis based on the target biological information and the target biomarker;
[0152] A prediction module 304, configured to predict the genetic disease risk of a target population by using a genetic disease prediction model based on the target biomarker and the target correlation result.
[0153] In an alternative embodiment, the method for constructing the screening model is as follows:
[0154] Collect first historical biological information of a first historical population, where the first historical population includes a part of the population with genetic diseases and a part of the population without genetic diseases, and the first historical biological information is labeled with first historical biomarkers related to genetic diseases;
[0155] Train a neural network model based on the first historical biological information to obtain a screening model.
[0156] In an alternative embodiment, the biological information includes:
[0157] Genome, epigenome, transcriptome, and proteome;
[0158] The multi-omics data analysis includes: genome data analysis, epigenomic data analysis, transcriptome data analysis, and proteome data analysis.
[0159] In an alternative embodiment, the analysis module 303 is configured to:
[0160] Based on bioinformatics software and databases, determine the gene positions, quantities, structures, and functions of the DNA sequences in the whole-genome sequencing data of the target biological information;
[0161] Based on the gene positions, quantities, structures, and functions, determine the expression of the target biomarker in a preset tissue or environment by analyzing the transcriptome under different conditions;
[0162] Based on the gene positions, quantities, structures, and functions, compare the corresponding genomes in the parental and offspring populations to determine the evolutionary relationship, gene family expansion and contraction, and innovation and loss of functions between the corresponding genomes of the target biomarker;
[0163] Based on the gene positions, quantities, structures, and functions, and the characteristics of biological elements such as known transcription start sites, splicing sites, promoters, enhancers, and miRNA binding sites, predict new functional elements of the target biomarker;
[0164] Based on the gene position, quantity, structure and function, by performing enrichment analysis on the first gene where the target biomarker is located, determine whether the target biomarker participates in a preset biological process or pathway;
[0165] Based on the gene position, quantity, structure and function, by performing genetic effect analysis on the second gene where the target biomarker is located, determine the influence of the target biomarker on a preset phenotype and its correlation with a preset disease;
[0166] Based on the gene position, quantity, structure and function, by simulating protein structure and function, predict whether the target biomarker has an impact on protein structure and function.
[0167] In an alternative embodiment, the analysis module 303 is configured to:
[0168] Based on the target biological information and the target biomarker, use any one or more of DNA methylation, histone modification, polycomb proteins and chromatin remodeling, and histone regulation to alter the epigenome, so as to determine the influence of the target biomarker represented by the alteration on gene expression, cell differentiation, and the process of growth and development.
[0169] In an alternative embodiment, the analysis module 303 is configured to:
[0170] Based on the target biological information and the target biomarker, use a preset process for transcriptome data analysis to obtain genes with differential expression;
[0171] Perform enrichment analysis on the genes with differential expression to determine the related biological processes and pathways regulated by the target biomarker.
[0172] In an alternative embodiment, the analysis module 303 is configured to:
[0173] By comparing the expression of different proteins in different individuals of different populations, determine differentially expressed proteins;
[0174] Perform network analysis on the differentially expressed proteins to determine the interconnections of the differentially expressed proteins in different individuals;
[0175] Based on the differentially expressed proteins, determine key nodes through functional annotation and enrichment analysis;
[0176] Based on the relevant interconnections and the key nodes, determine the key proteins, biological functions, and signal pathways of the target biomarker.
[0177] In an alternative embodiment, the method for constructing the genetic disease prediction model includes:
[0178] Collect the second historical biomarkers of the second historical population and the historical correlation results between the second historical biomarkers and genetic diseases;
[0179] Input the second historical biomarkers and the historical correlation results into a machine learning model for training to obtain a genetic disease prediction model, and the machine learning model is specifically any one of the following:
[0180] Logistic regression model, random forest model, and support vector machine model.
[0181] In an alternative embodiment, the prediction module 304 is used for:
[0182] Input the target biomarkers and the target correlation results into the genetic disease prediction model, and output the risk level of genetic diseases of the target population.
[0183] Embodiment III:
[0184] Based on the same inventive concept, an embodiment of the present invention provides a computer device, as Figure 4 shown, including a memory 404, a processor 402, and a computer program stored on the memory 404 and executable on the processor 402. When the processor 402 executes the program, the steps of the above-mentioned method for predicting genetic diseases are implemented.
[0185] Among them, in Figure 4 the bus architecture (represented by bus 400), bus 400 may include any number of interconnected buses and bridges. Bus 400 links various circuits including one or more processors represented by processor 402 and memory represented by memory 404 together. Bus 400 can also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art. Therefore, they will not be further described herein. The bus interface 406 provides an interface between the bus 400 and the receiver 401 and the transmitter 403. The receiver 401 and the transmitter 403 can be the same element, that is, a transceiver, which provides a unit for communicating with various other devices on the transmission medium. The processor 402 is responsible for managing the bus 400 and general processing, while the memory 404 can be used to store data used by the processor 402 when performing operations.
[0186] Embodiment IV:
[0187] Based on the same inventive concept, an embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the above-mentioned method for predicting genetic diseases are implemented.
[0188] The algorithms and displays provided herein are not inherently related to any particular computer, virtual system, or other device. Various general-purpose systems can also be used in conjunction with the teachings based herein. The structure required to construct such systems will be apparent from the above description. In addition, the present invention is not directed to any particular programming language. It should be understood that the content of the present invention described herein can be implemented using various programming languages, and the description of the specific language above is for the purpose of disclosing the best mode of the present invention.
[0189] In the specification provided herein, a large number of specific details are set forth. However, it can be understood that the embodiments of the present invention can be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.
[0190] Similarly, it should be understood that in order to streamline this disclosure and assist in understanding one or more of the various inventive aspects, in the foregoing description of the exemplary embodiments of the present invention, the various features of the present invention are sometimes grouped together in a single embodiment, figure, or description thereof. However, the disclosed method should not be construed as reflecting an intention that the claimed invention requires more features than are expressly recited in each embodiment. Rather, as reflected in each embodiment, the inventive aspects lie in less than all the features of the preceding disclosed single embodiment. Thus, the claims following the detailed description are hereby expressly incorporated into this detailed description, with each claim standing on its own as a separate embodiment of the present invention.
[0191] Those skilled in the art can understand that the modules in the devices in the embodiments can be adaptively changed and disposed in one or more devices different from the embodiments. The modules or units or components in the embodiments can be combined into a module or unit or component, and in addition, they can be divided into multiple sub-modules or sub-units or sub-components. Except that at least some of such features and / or processes or units are mutually exclusive, any combination can be adopted to combine all the features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all the processes or units of any method or device so disclosed. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) can be replaced by an alternative feature that provides the same, equivalent, or similar purpose.
[0192] In addition, those skilled in the art can understand that although some embodiments herein include certain features included in other embodiments rather than other features, the combination of features of different embodiments means that it is within the scope of the present invention and forms different embodiments. For example, in the detailed description, any one of the claimed embodiments can be used in any combination.
[0193] Each component embodiment of the present invention can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art should understand that a microprocessor or a digital signal processor (DSP) can be used in practice to implement some or all of the functions of some or all of the components of the prediction device for genetic diseases and computer equipment according to the embodiments of the present invention. The present invention can also be implemented as a device or device program (for example, a computer program and a computer program product) for executing part or all of the methods described herein. Such a program for implementing the present invention can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.
[0194] It should be noted that the above embodiments illustrate the present invention rather than limit the present invention, and those skilled in the art can design alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in the claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present invention can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In the unit claims listing several devices, several of these devices can be embodied by the same item of hardware. The use of the words first, second, and third, etc. does not denote any order. These words can be interpreted as names.
Claims
1. A method for predicting a genetic disease, characterized in that, Including: Obtaining the target biological information of the target population, where the biological information includes: genome, epigenome, transcriptome, and proteome; Based on the target biological information, using a screening model to screen out target biomarkers related to genetic diseases; Based on the target biological information and the target biomarkers, using multi-omics data analysis to obtain the target correlation results between the target biomarkers and genetic diseases, where the multi-omics data analysis includes: genome data analysis, epigenome data analysis, transcriptome data analysis, and proteome data analysis; Based on the target biomarkers and the target correlation results, using a genetic disease prediction model to predict the genetic disease risk of the target population.
2. The method according to claim 1, wherein The construction method of the screening model is as follows: Collecting the first historical biological information of the first historical population, where the first historical population includes a part of the population with genetic diseases and a part of the population without genetic diseases, and the first historical biological information is marked with the first historical biomarkers related to genetic diseases; Based on the first historical biological information, training a neural network model to obtain a screening model.
3. The method according to claim 1, characterized in that When using genome data analysis, the step of based on the target biological information and the target biomarkers, using multi-omics data analysis to obtain the target correlation results between the target biomarkers and genetic diseases includes: Based on bioinformatics software and databases, determining the gene locations, quantities, structures, and functions of the DNA sequences in the whole-genome sequencing data of the target biological information; Based on the gene locations, quantities, structures, and functions, by analyzing the transcriptome under different conditions, determining the expression of the target biomarkers in a preset tissue or environment; Based on the gene locations, quantities, structures, and functions, comparing the corresponding genomes in the parental and offspring populations to determine the evolutionary relationships, gene family expansion and contraction, as well as innovation and loss of functions between the corresponding genomes of the target biomarkers; Based on the gene locations, quantities, structures, and functions, and the characteristics of biological elements such as known transcription start sites, splicing sites, promoters, enhancers, and miRNA binding sites, predicting new functional elements of the target biomarkers; Based on the gene locations, quantities, structures, and functions, by performing enrichment analysis on the first gene where the target biomarker is located, determining whether the target biomarker participates in a preset biological process or pathway; Based on the gene locations, quantities, structures, and functions, by performing genetic effect analysis on the second gene where the target biomarker is located, determining the influence of the target biomarker on a preset phenotype and its correlation with a preset disease; Based on the gene locations, quantities, structures, and functions, by simulating protein structures and functions, predicting whether the target biomarker has an impact on protein structures and functions.
4. The method according to claim 1, characterized in that When performing epigenetic data analysis, based on the target biological information and the target biomarker, multi-omics data analysis is used to obtain the target correlation result between the target biomarker and the genetic disease, including: Based on the target biological information and the target biomarker, any one or more of DNA methylation, histone modification, polycomb proteins, chromatin remodeling, and histone regulation are used to alter the epigenome to determine the effects of the target biomarker represented by the alteration on gene expression, cell differentiation, and the growth and development process.
5. The method according to claim 1, characterized in that, When performing transcriptome data analysis, based on the target biological information and the target biomarker, multi-omics data analysis is used to obtain the target correlation result between the target biomarker and the genetic disease, including: Based on the target biological information and the target biomarker, a preset process of transcriptome data analysis is used to obtain genes with differential expression; Enrichment analysis is performed on the genes with differential expression to determine the related biological processes and pathways regulated by the target biomarker.
6. The method according to claim 1, wherein When performing protein data analysis, based on the target biological information and the target biomarker, multi-omics data analysis is used to obtain the target correlation result between the target biomarker and the genetic disease, including: By comparing the expression of the proteome in different individuals of the same population, differentially expressed proteins are determined; Network analysis is performed on the differentially expressed proteins to determine the interconnections of the differentially expressed proteins in different individuals; Based on the differentially expressed proteins, key nodes are determined through functional annotation and enrichment analysis; Based on the interconnections and the key nodes, the key proteins, biological functions, and signal pathways of the target biomarker are determined.
7. The method according to claim 1, characterized in that, The method for constructing the genetic disease prediction model includes: Collecting the second historical biomarker of the second historical population and the historical correlation result between the second historical biomarker and the genetic disease; Inputting the second historical biomarker and the historical correlation result into a machine learning model for training to obtain a genetic disease prediction model, and the machine learning model is specifically any one of the following: Logistic regression model, random forest model, and support vector machine model.
8. A prediction device for a genetic disease, characterized in that, Including: An acquisition module for acquiring the target biological information of the target population, and the biological information includes: genome, epigenome, transcriptome, and proteome; A screening module for screening out biomarkers related to genetic diseases based on the target biological information by using a screening model; An analysis module for obtaining the target correlation result between the target biomarker and the genetic disease based on the target biological information and the target biomarker by using multi-omics data analysis, and the multi-omics data analysis includes: genome data analysis, epigenetic data analysis, transcriptome data analysis, proteome data analysis; A prediction module for predicting the genetic disease risk of the target population by using a genetic disease prediction model based on the target biomarker and the correlation result.
9. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method described in any one of claims 1 to 7.