Method for constructing function-specific core markers of complex diseases and related devices
By acquiring biological knowledge data to measure disease correlations, core markers of specific functions in complex diseases are screened out, solving the problems of excessively high number of markers and low accuracy, and achieving efficient and accurate construction of core markers.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHENZHEN HUADA GENE INST
- Filing Date
- 2024-03-07
- Publication Date
- 2026-05-15
AI Technical Summary
In existing technologies, the number of markers identified during the construction of core markers for complex diseases is too high, and their relevance to the disease cannot be accurately determined, resulting in complex construction and low accuracy.
By acquiring biological knowledge data related to the target disease, we extract candidate genes, target function set data, disease-associated marker set and disease transcriptomics data, and use these data to measure disease association and screen out selected genes as core markers of specific functions.
It reduces the cost of constructing core disease markers, improves the accuracy and ease of construction, and saves time and computing resources.
Smart Images

Figure CN119920318B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of biodetection technology, and in particular to a method for constructing core markers specific to complex diseases and related equipment. Background Technology
[0002] The integration of computer technology and biology in research on complex diseases necessitates the analysis of omics data to inform treatment decisions. However, related technologies require the screening of numerous markers to identify core markers specific to complex diseases. The sheer number of markers to be identified, coupled with the uncertainty of their relevance to the complex disease, makes the construction of core markers for complex diseases complex and inaccurate. Therefore, improving the accuracy and ease of constructing core markers specific to complex diseases has become a pressing technical challenge. Summary of the Invention
[0003] The main objective of this application is to propose a method and related equipment for constructing functionally specific core markers for complex diseases, aiming to make the construction of functionally specific core markers for complex diseases more accurate and simple.
[0004] To achieve the above objectives, a first aspect of this application proposes a method for constructing core markers specific to complex diseases, the method comprising:
[0005] Acquire biological knowledge data related to the target disease;
[0006] From the biological knowledge data, candidate genes associated with the target disease, target function set data, disease-associated marker set, and disease transcriptomics data are extracted; wherein, the target function set data is a function set data composed of the functional name associated with the target disease, RNA targeting relationship, and gene genetic variation results;
[0007] Based on the target function set data, the disease association marker set, and the disease transcriptomics data, the disease association of the candidate genes is measured to obtain disease association measurement data for each candidate gene.
[0008] The candidate genes are screened based on the disease association measurement data to obtain selected genes;
[0009] The target disease is labeled based on the selected gene to obtain the specific functional core marker of the target disease.
[0010] In some embodiments, the target functional set data includes: a functional gene set, a target gene set, and a loss-of-function intolerance set. The functional gene set represents a set of genes with functions related to the target disease. The target gene set represents a set of genes that have a targeting relationship with the RNA of the target disease. The loss-of-function intolerance set includes the degree of loss-of-function intolerance for each candidate gene. The step of evaluating the disease association of the candidate genes based on the target functional set data, the disease association marker set, and the disease transcriptomics data to obtain disease association metric data for each candidate gene includes:
[0011] The pathogenicity of each candidate gene is measured based on the set of disease-associated markers to obtain pathogenicity measurement data;
[0012] The functional contribution of each candidate gene is measured according to the functional gene set to obtain functional contribution measurement data.
[0013] Differential expression measurement data was obtained by measuring the differential expression of each candidate gene based on the disease transcriptomics data.
[0014] Gene targeting measurement is performed on each candidate gene according to the target gene set to obtain gene targeting measurement data.
[0015] Based on the set of loss-of-function intolerance levels, each candidate gene is measured for loss-of-function intolerance to obtain loss-of-function intolerance measurement data.
[0016] The pathogenicity measurement data, the functional contribution measurement data, the differential expression measurement data, the gene targeting measurement data, and the loss-of-function intolerance measurement data are used to construct disease association measurement data for each candidate gene.
[0017] In some embodiments, the differential expression measurement of each candidate gene based on the disease transcriptomics data to obtain differential expression measurement data includes:
[0018] The data categories for acquiring the disease transcriptomics data;
[0019] The disease transcriptomics data are preprocessed according to the data categories to obtain candidate transcriptomics data;
[0020] The candidate transcriptomics data are matrix transformed to obtain a transcriptomics matrix; wherein each row of the transcriptomics matrix represents a gene and each column represents a sample.
[0021] Differential expression was measured on the transcriptomics matrix to obtain the differential expression measurement data.
[0022] In some embodiments, the differential expression measurement of the transcriptomics matrix to obtain the differential expression measurement data includes:
[0023] The expression values of the disease group and the normal group were extracted from the transcriptomics matrix;
[0024] Obtain the differential expression values between the expression values of the disease group and the expression values of the normal group;
[0025] The differential expression values are merged to obtain the target expression value;
[0026] Based on the target expression value and the preset expression threshold, the grade data is determined;
[0027] The rank data is mean-processed to obtain the differential expression measurement data.
[0028] In some embodiments, the target gene set includes: a miRNA target gene set, a lncRNA target gene set, and an RNA relationship set; the step of performing gene targeting evaluation on each candidate gene based on the target gene set to obtain gene targeting metric data includes:
[0029] The candidate genes are targeted based on the miRNA target gene set to obtain first target measurement data.
[0030] Based on the set of LncRNA-targeting genes, the candidate genes are targeted to obtain second target measurement data.
[0031] Based on the miRNA target gene set, the LncRNA target gene set, and the RNA relationship set, the target relationship of the candidate genes is measured to obtain target relationship measurement data.
[0032] The first target measurement data, the second target measurement data, and the target relationship measurement data are concatenated to obtain the gene target measurement data.
[0033] In some embodiments, the step of measuring loss-of-function intolerance for each candidate gene based on the set of loss-of-function intolerance levels to obtain loss-of-function intolerance measurement data includes:
[0034] Obtain the gene name of the candidate gene;
[0035] Based on the gene name, the probability of loss of function intolerance is searched in the set of degree of loss of function intolerance to obtain the measurement data of loss of function intolerance.
[0036] In some embodiments, the step of screening the candidate genes based on the disease association measurement data to obtain selected genes includes:
[0037] The candidate genes are sorted in descending order based on the disease association measurement data to obtain the sorting number of each candidate gene;
[0038] The selected gene is extracted from the candidate genes according to the preset number of gene selections and the sorting number.
[0039] To achieve the above objectives, a second aspect of this application provides an apparatus for constructing core markers specific to complex diseases, the apparatus comprising:
[0040] The data acquisition module is used to acquire biological knowledge data related to the target disease;
[0041] The extraction module is used to extract candidate genes associated with the target disease, target function set data, disease-associated marker set, and disease transcriptomics data from the biological knowledge data; wherein, the target function set data is a function set data composed of the functional name associated with the target disease, RNA targeting relationship, and gene genetic variation results;
[0042] The association measurement module is used to measure the disease association of the candidate genes based on the target function set data, the disease association marker set, and the disease transcriptomics data, and to obtain the disease association measurement data for each candidate gene.
[0043] The screening module is used to screen the candidate genes based on the disease association measurement data to obtain selected genes;
[0044] The labeling module is used to label the target disease based on the selected gene to obtain the specific functional core marker of the target disease.
[0045] To achieve the above objectives, a third aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described in the first aspect.
[0046] To achieve the above objectives, a fourth aspect of the present application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in the first aspect.
[0047] This application proposes a method and related equipment for constructing function-specific core markers for complex diseases. It extracts biological knowledge data related to the target disease from existing biological databases / literature. This extracted biological knowledge data includes candidate genes associated with the target disease, target function set data, disease-associated marker sets, and disease transcriptomics data. By analyzing the association between each candidate gene and the target disease using the target function set data, disease-associated marker sets, and disease transcriptomics data, disease association measurement data is obtained. This disease association measurement data is then used to screen candidate genes that play a crucial role in the occurrence and development of the target disease, and these selected genes are then used as core markers for that function of the target disease. Therefore, by constructing function-specific core markers for the target disease using existing publicly available background knowledge, and by employing the target function set data, disease-associated marker sets, and disease transcriptomics data from this publicly available background knowledge to analyze the association between genes and the disease, the construction cost of function-specific core markers for the target disease is reduced while maintaining accuracy. Attached Figure Description
[0048] Figure 1 This is a flowchart of a method for constructing core markers specific to complex diseases provided in embodiments of this application;
[0049] Figure 2 yes Figure 1 The flowchart of step S103 in the process;
[0050] Figure 3 yes Figure 2 The flowchart of step S203 in the process;
[0051] Figure 4 yes Figure 3 The flowchart of step S304 in the process;
[0052] Figure 5 yes Figure 2 The flowchart of step S204 in the process;
[0053] Figure 6 yes Figure 2 The flowchart of step S205 in the document;
[0054] Figure 7 yes Figure 1 The flowchart of step S104 in the process;
[0055] Figure 8 This is a flowchart of a method for constructing core markers specific to complex diseases according to another embodiment of this application;
[0056] Figure 9 This is a schematic diagram of the structure of the device for constructing core markers specific to complex diseases provided in the embodiments of this application;
[0057] Figure 10 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0058] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0059] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0060] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0061] First, let's analyze some of the terms used in this application:
[0062] Gene sets are collections of genes associated with specific biological processes, functions, or pathways. These gene sets can be defined and collected through various methods and databases, such as gene expression data, gene function annotation, and literature mining.
[0063] Biomarkers are biochemical indicators / molecules that can mark changes or potential changes in the structure or function of systems, organs, tissues, cells, and subcellular structures. They have a very wide range of applications. Biomarkers can be used for disease diagnosis, disease staging, or to evaluate the safety and efficacy of new drugs or therapies in target populations.
[0064] Transcriptomics studies gene expression at the RNA level. The transcriptome is the sum of all RNA that a living cell can transcribe, and it is an important tool for studying cell phenotype and function. The transcription process, which uses DNA as a template to synthesize RNA, is the first step in gene expression and a key step in its regulation. Gene expression refers to the entire process by which the genetic information carried by a gene is transformed into a distinguishable phenotype.
[0065] Non-coding RNA (ncRNA) refers to RNA molecules produced during transcription that do not participate in protein coding but instead have other important biological functions. Unlike protein-coding mRNA, non-coding RNA is not translated into proteins but exists in cells as RNA molecules.
[0066] Messenger RNA (mRNA) is an RNA molecule synthesized during transcription. It carries the genetic information encoded in DNA and delivers this information to the cytoplasm to participate in protein synthesis.
[0067] Non-coding RNA sets (ncRNA sets) refer to a combination of various non-coding RNAs present in a cell. Non-coding RNAs can be classified according to length and function, and different types of non-coding RNAs have different regulatory roles and biological functions.
[0068] Variant sets refer to a collection of genomic variations found in a specific population. These variations can take the form of single nucleotide polymorphisms (SNPs), insertions / deletions (Indels), structural variations (SVs), or gene rearrangements.
[0069] Ischemic stroke (IS) is a brain injury caused by insufficient blood supply to the brain. It refers to a condition where blockage of blood vessels in the brain leads to ischemia and hypoxia in a specific region, resulting in brain cell damage or death. According to the TOAST etiological classification, ischemic stroke can be divided into five types: large-artery atherosclerosis, cardioembolism, small-vessel occlusion, acute stroke of other determined etiology, and stroke of undetermined etiology.
[0070] As mentioned above, complex diseases are typically caused by the interaction of multiple genes and environmental factors, and their pathogenesis is quite complex. Therefore, screening for specific core biomarkers to diagnose complex diseases is a challenging task. Traditionally, the screening of core biomarkers for complex diseases involves large-scale population genomic analysis to identify gene variants associated with disease risk, and then using these gene variants as core biomarkers. However, large-scale population genomic analysis requires the identification of a large number of genes, resulting in an excessively high number of biomarkers to be identified, and it is difficult to accurately determine whether they are associated with complex diseases. This makes the construction of core biomarkers for complex diseases complex and inaccurate.
[0071] Based on this, embodiments of this application provide a method and related equipment for constructing function-specific core markers for complex diseases. The aim is to acquire biological knowledge data related to the target disease, and extract candidate genes, target function set data, disease-associated marker set, and disease transcriptomics data required for core marker construction from the biological knowledge data. Then, based on the target function set data, disease-associated marker set, and disease transcriptomics data, the association degree between each candidate gene and the target disease is analyzed to obtain disease association measurement data. Selected genes are screened from the candidate genes using the disease association measurement data, and these selected genes are used as function-specific core markers for the target disease. Therefore, by fully utilizing publicly available data and literature, and then constructing function-specific core markers for the target disease based on this background knowledge, not only is the cost of constructing disease core markers reduced, but it also eliminates the need for pre-constructing large-scale markers, saving time and computational resources for constructing function-specific core markers for diseases.
[0072] The method and related equipment for constructing functionally specific core markers for complex diseases provided in this application are specifically described through the following embodiments. First, the method for constructing functionally specific core markers for complex diseases in this application is described.
[0073] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0074] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0075] The method for constructing core biomarkers for function-specific complex diseases provided in this application relates to the field of biodetection technology. This method can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, etc.; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application implementing the method for constructing core biomarkers for function-specific complex diseases, but is not limited to the above forms.
[0076] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0077] It should be noted that in all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user is obtained through pop-ups or redirection to confirmation pages. Only after obtaining the user's separate permission or consent is the necessary user-related data required for the proper functioning of these embodiments acquired.
[0078] Figure 1 This is an optional flowchart of a method for constructing core markers specific to complex diseases provided in the embodiments of this application. Figure 1 The method may include, but is not limited to, steps S101 to S105.
[0079] Step S101: Obtain biological knowledge data related to the target disease;
[0080] Step S102: Extract candidate genes associated with the target disease, target function set data, disease-associated marker set, and disease transcriptomics data from biological knowledge data; wherein, the target function set data is a set of function data composed of the functional name associated with the target disease, RNA targeting relationship, and gene genetic variation results;
[0081] Step S103: Based on the target function set data, disease association marker set and disease transcriptomics data, the disease association of candidate genes is measured to obtain the disease association measurement data for each candidate gene.
[0082] Step S104: Screen candidate genes based on disease association measurement data to obtain selected genes;
[0083] Step S105: Mark the target disease according to the selected gene to obtain the specific function core marker of the target disease.
[0084] Steps S101 to S105 of this embodiment involve acquiring biological knowledge data of the target disease, extracting candidate genes from the biological knowledge data, and evaluating the candidate genes' association with the target disease's target function set data, disease-associated marker set, and disease transcriptomics data. Based on the target function set data, disease-associated marker set, and disease transcriptomics data, the association between candidate genes and the target disease is evaluated to obtain disease association measurement data. Selected genes are then screened from the candidate genes using the disease association measurement data. The selected genes are then used to label the target disease, resulting in specific function core markers for the target disease. Therefore, by fully utilizing existing publicly available biological knowledge data related to the target disease, and analyzing the correlation between candidate genes and the target disease based on assessable candidate gene data within the biological knowledge data, genes closely related to the target disease can be screened as core markers. This eliminates the need for gene analysis on a large population, saving the cost of constructing core disease markers. Furthermore, the entire core marker construction process is simple and rapid, saving time and computational resources.
[0085] In step S101 of some embodiments, biological knowledge data can be obtained from publicly available literature. Biological knowledge can also be obtained through other means, not limited to this. It should be noted that the biological knowledge data refers to known biological background knowledge related to the target disease, and different biological background knowledge comes from different platforms. Communication with the publicly available platform is required to retrieve the biological background knowledge related to the target disease. By utilizing readily available and publicly available biological background knowledge to construct core markers for the target disease, it is unnecessary to perform gene identification on a large-scale population in advance, or to identify a large number of markers, thus saving costs and improving efficiency in constructing core markers for complex diseases.
[0086] It should be noted that the platforms for acquiring biological background knowledge in this embodiment include NCBI PubMed, Google Scholar, biological knowledge databases, and literature, and there are no restrictions on the platforms used. Specifically, biological knowledge databases include: biological function databases, disease marker databases, RNA knowledge databases, and large-scale population gene databases, and there are no specific restrictions on the biological knowledge databases used. Specifically, NCBI PubMed is a publicly available biomedical literature database covering all aspects of the field of biology, including basic science, clinical medicine, and epidemiology. Google Scholar is a free learning search engine that indexes academic papers, dissertations, research reports, and conference papers from various fields. Therefore, by establishing connections with publicly available platforms that possess knowledge in all aspects of the biological field to acquire biological knowledge data, the acquired biological knowledge data can be accurately constructed into core markers specific to the functions of complex diseases.
[0087] In step S102 of some embodiments, the biological knowledge data encompasses the biological background knowledge associated with the target disease and originates from multiple platforms, each capable of extracting different biological knowledge data. Therefore, candidate genes, target function sets, disease-associated marker sets, and disease transcriptomics data are extracted from the biological knowledge data, and these data originate from different platforms. It should be noted that candidate genes are genes recorded in the biological background knowledge as being associated with the target disease. The target function set data records genes related to a specific function, RNA relationship, or genetic variation of the target disease. The target function set data can be used to determine which disease-associated functions, RNA targeting relationships, and the degree of genetic variation of each candidate gene are involved in. The disease-associated marker set includes representative genes associated with the target disease from the biological background knowledge. The disease transcriptomics data is the sum of all RNA that a living cell can transcribe, used to study the expression of each candidate gene. Therefore, by extracting target function set data, disease-associated marker set, and disease transcriptomics data from biological background knowledge, we can more accurately screen out the core functional markers that characterize the target disease.
[0088] Specifically, the target functional set data includes a functional gene set, a target gene set, and a set of loss-of-function intolerance levels. The functional gene set contains genes with functions related to the target disease, i.e., genes carrying functions related to the target disease. The target gene set contains RNAs associated with the target disease and genes with a targeting relationship to those RNAs. The loss-of-function intolerance level set records the probability of loss-of-function intolerance for each candidate gene. This probability represents the likelihood of gene function loss during treatment of the target disease, and it is related to factors such as abnormal mutations, environmental exposure, and drug dosage. Therefore, by screening candidate genes from multiple perspectives as core markers of specific functions for the target disease, the selected core markers become more accurate research subjects for the treatment of the target disease, enabling targeted disease treatment and improving treatment efficacy.
[0089] In some embodiments, the functional gene set is obtained from a biological function database, which can be GO or KEGG, and the type of biological function database is not limited. It should be noted that GO and KEGG are two commonly used biological function databases for annotating and classifying gene functions. Therefore, extracting the functional gene set from GO or KEGG databases can characterize the genes corresponding to each function, and the same gene can have multiple functions. Further explanation is needed: the functional gene set is like a GO node or a KEGG pathway, and it is a collection of gene lists, where each gene list records genes with the same function. There is no structural relationship between gene lists within a GO node, while there is a structural relationship between gene lists within a KEGG node. Therefore, the functional set data extracted directly from the GO or KEGG biological function database is F = {f1, ..., f...} n}, where G is the union of all candidate genes, obtained by applying any functional set f i After removing the included structures and interactions, the set of functional genes containing only gene names is obtained. In a functional gene set, some genes may be involved in multiple functions, or even cover all functions. Therefore, the functional gene set is divided into a core functional gene set and a specific functional gene set. The core functional gene set contains genes that appear in all functional sets, while the specific functional gene set contains genes that have functions other than those in the core functional gene set. Therefore, the core functional gene set is denoted as G. core , that is The set of genes with specific functions is denoted as G. specific , that is And G = G core ∪G specific .
[0090] For example, if the target disease is ischemic stroke, abbreviated as IS, a literature search based on databases such as NCBI PubMed and Google Scholar reveals the functional nodes related to IS as GO:0034599 (cellular response to oxidative stress) in the Gene Ontology database and the pathway hsa04080 (Neuroactive ligand-receptor interaction-Homo sapiens (human)) in the KEGG database, i.e., F = {GO:0034599, hsa04080}. The union of all included genes is denoted as G, containing a total of 459 genes, with the core functional gene set G... core It contains 2 genes, and the set of genes with specific functions is denoted as G. specificIt contains 457 genes.
[0091] In some embodiments, the target gene set includes a set of RNA-targeting genes and a set of RNA relationships. It should be noted that a list of non-coding RNAs (miRNAs, lncRNAs) related to the target disease is extracted from literature and RNA knowledge databases, denoted as MI = {mi1, ..., mi...}. n}, and LN = {ln1, ..., ln n Furthermore, the RNA knowledge database can include databases such as HMDD, miR2Disease, and LncRNADisease v3.0, and there are no restrictions on the type of RNA knowledge database. Specifically, HMDD is a database that specifically collects and organizes disease information related to human microRNAs, miR2Disease is a database that collects information on the association between miRNAs and human diseases, and LncRNADisease is a database that collects information on the association between long non-coding RNAs (LncRNAs) and human diseases. Therefore, a list of non-coding RNAs related to the target disease is extracted from databases such as HMDD, miR2Disease, and LncRNADisease v3.0. Then, based on the list of non-coding RNAs (miRNAs and LncRNAs), the relationship set between miRNAs and target genes is obtained from databases such as miRTarBase and LncRNA2Target to obtain the RNA target gene set. Since there are two types of non-coding RNAs, the RNA target gene set is denoted as [missing information - likely a typo]. as well as It's important to note that miRTarBase is a database that collects information on the relationships between miRNAs and target genes, while LncRNA2Target is a database that collects information on the relationships between lncRNAs and target genes. Therefore, genes with target relationships between miRNAs and lncRNAs related to the target disease can be directly obtained from miRTarBase and LncRNA2Target to construct an RNA-targeting gene set. Simultaneously, LncBase is a database that collects information on the relationships between messenger RNAs and target genes. Therefore, obtaining the relationships between miRNAs and lncRNAs based on LncBase yields an RNA relationship set, and for any miRNA... i The set of LncRNAs regulated by it is denoted as
[0092] In some embodiments, the loss-of-function intolerance set is obtained by acquiring genetic variation data from large-scale population genome databases such as gnomAD, dbSNP, and 1KGP. It should be noted that gnomAD is a genome-wide variation frequency database, dbSNP is a public database that maintains information on single nucleotide and structural variations in humans and other organisms, and 1KGP is a genome sequencing project that constructs a detailed map of human genome variation encompassing diverse populations. Therefore, by extracting genetic variation data from these three gene variation-related databases, and then calculating the loss-of-function intolerance probability for each candidate gene based on this data, the loss-of-function intolerance set can be obtained. Thus, the loss-of-function intolerance set allows for more accurate selection of specific functional markers for target diseases.
[0093] In some embodiments, a disease-associated biomarker set is a list of disease markers. Biomarkers related to the target disease are extracted from literature or disease marker databases and denoted as Golden-sets. It should be noted that in this embodiment, the disease marker database is MarkerDB and HPO (Human Phenotype Ontology), and there are no specific limitations on the specific database. MarkerDB is a gene marker database that provides information and resources related to gene markers, while HPO is a standardized terminology and classification system used to describe human phenotypes. By extracting the disease-associated biomarker set of the target disease from MarkerDB and HPO, it is possible to accurately characterize the biomarkers related to the target disease as recorded in existing knowledge, thereby assisting in the more accurate construction of core biomarkers for the target disease.
[0094] For example, based on the HPO database, biological markers related to IS were obtained, resulting in a set of disease-associated markers consisting of 37 genes.
[0095] In some embodiments, disease transcriptomics data is a dataset composed of transcriptomics data of the target disease, obtained from databases such as NCBI GEO and EBI ArrayExpress, with no restrictions on the databases used for acquiring the disease transcriptomics data. It should be noted that both NCBI GEO and EBI ArrayExpress maintain and manage a public gene expression data repository. Therefore, expression data for each candidate gene can be obtained through NCBI GEO and EBI ArrayExpress, facilitating more accurate analysis of the correlation between candidate genes and the target disease, and enabling the selection of core functional markers that can characterize the target disease.
[0096] In summary, by selecting different databases or literature for data extraction based on different data, we can make full use of existing publicly available biological background knowledge to assist in the construction of specific functional core markers for target diseases, without the need for gene analysis and various experiments on the target disease, thus saving the cost of constructing specific functional core markers for target diseases.
[0097] Please see Figure 2 In some embodiments, step S103 may include, but is not limited to, steps S201 to S206:
[0098] Step S201: Measure the pathogenicity of each candidate gene based on the disease-associated marker set to obtain pathogenicity measurement data;
[0099] Step S202: Measure the functional contribution of each candidate gene according to the functional gene set to obtain functional contribution measurement data;
[0100] Step S203: Measure the differential expression of each candidate gene based on the disease transcriptomics data to obtain differential expression measurement data;
[0101] Step S204: Perform gene targeting measurement on each candidate gene according to the target gene set to obtain gene targeting measurement data;
[0102] Step S205: Measure the loss-of-function intolerance for each candidate gene according to the set of loss-of-function intolerance levels to obtain loss-of-function intolerance measurement data.
[0103] Step S206: Pathogenicity measurement data, functional contribution measurement data, differential expression measurement data, gene targeting measurement data, and loss-of-function intolerance measurement data are used to construct disease association measurement data for each candidate gene.
[0104] In step S201 of some embodiments, the disease-associated marker set records genes that cause the target disease. Therefore, the pathogenicity of the candidate gene is determined by checking whether it is included in the disease-associated marker set, thereby obtaining pathogenicity measurement data. It should be noted that the pathogenicity measurement data is used to measure the pathogenicity of each candidate gene, that is, to analyze whether the candidate disease is a pathogenic gene of the target disease.
[0105] Specifically, the pathogenicity of candidate genes is measured as shown in equation (1):
[0106]
[0107] In the formula, g j It is any candidate gene in G, and Golden-sets is a set of disease-associated markers.
[0108] For example, if the disease-associated marker set for IS contains 37 genes, and 13 of them contain G, then the pathogenicity measure of these 13 candidate genes is 1, and the pathogenicity measure of the remaining genes is 0.
[0109] In step S202 of some embodiments, the functional gene set consists of multiple lists of functional genes, and each list contains genes with the same function. Contribution metric data characterizes the contribution of candidate genes to the target disease-related functions, and performing functional contribution metric on candidate genes also determines the number of functions each candidate gene possesses. The more functions a candidate gene possesses, the higher its functional contribution metric data, indicating that the candidate gene influences multiple specific functions in the target disease.
[0110] Specifically, the functional contribution of candidate genes is measured as shown in equation (2):
[0111]
[0112]
[0113] Where 1≤i≤n, The value range is (0, 1). As can be seen from the above, by calculating the number of functions of each candidate gene and then using the ratio between the number of functions and the total number of functions as the functional contribution metric, the functional contribution of each candidate gene to the target disease can be accurately characterized.
[0114] In step S203 of some embodiments, disease transcriptomics data records the gene expression of each candidate gene, and differential expression measurement data characterizes the gene expression of candidate genes. If the gene expression is significant in the target disease, then the gene has a greater impact on the target disease. Therefore, by measuring the differential expression of candidate genes using disease transcriptomics data, the expression status of each candidate gene in the target disease can be obtained. Thus, candidate genes selected based on differential expression measurement data can serve as core markers of specific functions for the target disease.
[0115] In step S204 of some embodiments, the target gene set records the target genes of the target disease-related RNA, and the gene targeting measurement data characterizes the targeting of candidate genes to the target disease-related RNA. Based on the gene targeting measurement data, candidate genes can be screened to select genes that are more associated with the target disease as core markers.
[0116] In step S205 of some embodiments, loss-of-function intolerance metric data characterizes loss-of-function mutations in genes, which prevent the gene or gene product from performing its biological function normally. If a gene loses its original function, it may be a gene that causes the target disease. Therefore, genes screened based on loss-of-function intolerance metric data are more accurate as core functional markers for the target disease.
[0117] In steps S201 to S206 of this embodiment, the candidate genes are evaluated from five aspects to determine whether they have an impact on the target disease, in order to screen out specific functional core markers that can accurately characterize the target disease by measuring pathogenicity, functional contribution, differential expression, gene targeting, and loss-of-function intolerance.
[0118] Please see Figure 3 In some embodiments, step S203 may include, but is not limited to, steps S301 to S304:
[0119] Step S301: Obtain the data categories of disease transcriptomics data;
[0120] Step S302: Preprocess the disease transcriptomics data according to the data category to obtain candidate transcriptomics data;
[0121] Step S303: Perform matrix transformation on the candidate transcriptomics data to obtain a transcriptomics matrix; wherein each row in the transcriptomics matrix represents a gene, and each column represents a sample;
[0122] Step S304: Perform differential expression measurement on the transcriptomics matrix to obtain differential expression measurement data.
[0123] In steps S301 to S302 of some embodiments, the data categories of disease transcriptomics data include gene chips, deep sequencing data, and single-cell sequencing data, etc. Different preprocessing methods are applied to the disease transcriptomics data according to different data categories, achieving targeted preprocessing of the disease transcriptomics data to convert disease transcriptomics data of different categories into data in a unified format. It should be noted that, for the convenience of uniformly representing transcriptomics data as behavioral genes, they are listed as samples, and their expression values are in the range of 0 to 1.
[0124] Specifically, if the disease transcriptomics data is gene chip data, probe matching, multi-probe processing, missing value imputation, and sample screening are required. The preprocessing operations for gene chip data are not limited to these. If the disease transcriptomics data is deep sequencing data (bulk RNA-seq data), preset tools are used to filter low-quality reads, remove adapters, remove short sequences, and perform qualitative and quantitative analysis of RNA. It should be noted that the preset tools in this embodiment are fastp, Hisat2, and featurecounts; other embodiments are not limited to these. Specifically, fastp, Hisat2, and featurecounts are used to process common sequencing data and can perform corresponding preprocessing operations on the sequencing data. If the disease transcriptomics data is single-cell sequencing data (scRNA-seq / snRNA-seq data), tools such as seurat are used for quality control, filtering, and batch effect correction of the single-cell sequencing data. Seurat is software for single-cell transcriptomics data (e.g., an R language package), which can efficiently preprocess single-cell sequencing data. This embodiment does not specifically limit the tools used.
[0125] In step S303 of some embodiments, the preprocessed candidate transcriptomics data is converted into an expression profile in matrix form to obtain a transcriptomics matrix. Each row in the transcriptomics matrix represents the expression value of a candidate gene, and each row represents a sample. The value range of the expression value is [0, 1].
[0126] In step S304 of some embodiments, since the matrix records the expression value of each candidate gene in each sample, the expression difference of each candidate gene in the disease group and the normal group can be calculated based on the transcriptomics matrix to obtain differential expression measurement data, making the calculation of differential expression measurement data simpler and more accurate.
[0127] In steps S301 to S304 of this embodiment, preprocessing is performed on disease transcriptomics data of different data categories, and then the preprocessed transcriptomics data is converted into a transcriptomics matrix in matrix form, so that it is easier and more accurate to measure the differential expression of each candidate gene through the transcriptomics matrix.
[0128] Please see Figure 4 In some embodiments, step S304 may include, but is not limited to, steps S401 to S405:
[0129] Step S401: Extract the expression values of the disease group and the normal group from the transcriptomics matrix;
[0130] Step S402: Obtain the differential expression values between the disease group and the normal group;
[0131] Step S403: Merge the differential expression values to obtain the target expression value;
[0132] Step S404: Determine the grade data based on the target expression value and the preset expression threshold;
[0133] Step S405: Mean the rank data to obtain differential expression measurement data.
[0134] In step S401 of some embodiments, the transcriptomics matrix records the expression values of the disease group and the normal group. Since multiple transcriptomics matrices exist for each candidate gene, it is necessary to calculate the expression value of the disease group in each transcriptomics matrix as the disease group expression value, and then calculate the expression value of the normal group in each transcriptomics matrix as the normal group expression value. It should be noted that the disease group expression value is denoted as... The normal expression group is denoted as
[0135] In step S402 of some embodiments, the differential expression value characterizes the expression difference of the candidate gene between the disease group and the normal group. In this embodiment, it is necessary to take the logarithm of the ratio between the expression value of the disease group and the expression value of the normal group as the differential expression value. The differential expression value is characterized as shown in equation (3):
[0136]
[0137] In the formula, x represents any set of data from the k transcriptomics matrices.
[0138] In step S403 of some embodiments, for the same candidate gene, the differential expression values obtained from different data are denoted as... By analyzing differential expression values The differential expression values are then merged to obtain the target expression value. It should be noted that in this embodiment, the differential expression values are merged using methods such as Fisher's combined probability test.
[0139] Specifically, the target expression value is calculated as shown in equation (4):
[0140]
[0141] In step S404 of some embodiments, the target expression value is compared with a preset expression threshold to obtain rank data, and the rank data is defined as... Specifically, the grade data characterizes the gene expression level of candidate genes. If the target expression value is less than or equal to the expression threshold, the grade data is determined to be 1; if the target expression value is greater than the expression threshold, the grade data is determined to be 0.
[0142] The determination of the grade data is shown in equation (5):
[0143]
[0144] In step S405 of some embodiments, the hierarchical data for each transcriptome matrix is: Differential expression metrics were obtained by averaging the rank data corresponding to the K transcriptomics matrices, and the differential expression metrics are shown in Equation (6):
[0145]
[0146] In the formula, For differential expression measurement data, k is the transcriptome matrix.
[0147] For example, two datasets from the NCBI GEO database containing AIS (acute ischemic stroke) data and healthy controls were selected: GSE16561 (gene chip data) and GSE202518 (RNA-seq data) for differential expression analysis. If set... The threshold value is 0.05. After calculation, the merged value... The minimum value is 0.05843998, so all differential expression measurement data are assigned a value of 0.
[0148] In steps S401 to S405 of this embodiment, the differential expression values between the disease group and the normal group are obtained, and then the differential values are merged into the target expression value. Finally, the grade data corresponding to each sample column is calculated based on the target expression value and the preset expression threshold. The mean of the grade data is used as the differential expression measurement data, which makes the calculation of the differential expression measurement data of each candidate gene simple and can accurately characterize the gene expression of each candidate gene.
[0149] Please see Figure 5 In some embodiments, step S204 may include, but is not limited to, steps S501 to S504:
[0150] Step S501: Targeting candidate genes is measured based on the miRNA target gene set to obtain the first target measurement data;
[0151] Step S502: Targeting candidate genes is measured based on the LncRNA target gene set to obtain second target measurement data;
[0152] Step S503: Measure the target relationship of candidate genes based on the miRNA target gene set, the LncRNA target gene set, and the RNA relationship set to obtain target relationship measurement data;
[0153] Step S504: The first target measurement data, the second target measurement data, and the target relationship measurement data are spliced together to obtain gene target measurement data.
[0154] In step S501 of some embodiments, the first targeting metric data characterizes the targeting relationship between the candidate gene and the miRNA, that is, it determines whether there is a targeting relationship between the candidate gene and the miRNA related to the target disease. Specifically, the number of candidate genes in the miRNA target gene set is determined as the number of miRNAs targeted and regulated, and then the ratio between the number of miRNAs targeted and regulated and the total number of candidate genes is obtained as the first targeting metric data. Specifically, the first targeting metric data is calculated as shown in equation (7):
[0155]
[0156]
[0157] Where 1≤i≤n, First target measurement data The range of its value is [0, 1].
[0158] In step S502 of some embodiments, the second targeting metric data characterizes the targeting relationship between the candidate gene and the LncRNA, that is, it determines whether there is a targeting relationship between the candidate gene and the LncRNA related to the target disease. Specifically, the number of candidate genes in the LncRNA target gene set is determined as the number of LncRNA targeted regulation, and then the ratio between the number of Lnc targeted regulation and the total number of candidate genes is obtained as the second targeting metric data. Specifically, the second targeting metric data is calculated as shown in equation (8):
[0159]
[0160]
[0161] Where 1≤i≤n, Second target measurement data The range of its value is (0, 1).
[0162] In step S503 of some embodiments, after measuring the targeting relationship between candidate genes and miRNAs and lncRNAs, it is necessary to determine whether there is a competitive relationship between the miRNA-targeting gene set and the lncRNA-targeting gene set. It is necessary to determine whether the candidate gene appears in both the miRNA-targeting gene set and the lncRNA-targeting gene set, and also to determine whether there is a competitive relationship between them to obtain targeting relationship measurement data, as shown in equation (9):
[0163]
[0164] In the formula, This is a collection of miRNA-targeting genes. It is a set of LncRNA-targeting genes.
[0165] In step S504 of some embodiments, the first target measurement data, the second target measurement data, and the target relationship measurement data are concatenated into gene target measurement data, and the concatenation method includes, but is not limited to: summation, averaging, and weighted summation.
[0166] For example, 105 ncRNAs associated with IS were downloaded from the HMDD database. MI contains 100 elements, and LN contains 5 elements. For any miRNA among them... i and LncRNA ln i Based on databases such as miRTarBase and LncRNA2Target, the relationship sets between miRNAs and their target genes were obtained. The union of experimentally confirmed target genes for 46 miRNAs was 1992, meaning the miRNA target gene set contained 1992 genes. However, no experimentally confirmed target gene set was found for the 5 LncRNAs. Therefore, the first target measurement data for all genes... Second target measurement data All values are 0.
[0167] In steps S501 to S504 of this embodiment, the targeting relationship between each candidate gene and miRNAs and lncRNAs related to the target disease is analyzed, and it is then determined whether there is a competitive relationship between the two sets of target genes to which they belong, so as to construct gene targeting measurement data for each candidate gene, so as to accurately characterize the association between candidate genes and target diseases through gene targeting measurement data.
[0168] Please see Figure 6 In some embodiments, step S205 includes, but is not limited to, steps S601 to S602:
[0169] Step S601: Obtain the gene name of the candidate gene;
[0170] Step S602: Based on the gene name, perform a lookup of the probability of loss of function intolerance in the set of loss of function intolerance levels to obtain the measurement data of loss of function intolerance.
[0171] In steps S601 to S602 of some embodiments, by obtaining the gene name of each candidate gene, the corresponding loss-of-function intolerance probability is directly found in the loss-of-function intolerance set based on the gene name as loss-of-function intolerance measurement data, which is used to determine the degree of tolerance of the candidate gene to loss of function.
[0172] Specifically, for the set of functional intolerance levels, the probability of functional intolerance is calculated as shown in equation (10):
[0173]
[0174] in, Indicates candidate gene g j The frequency of mutations in the middle Indicates based on candidate gene g j The probability of loss of function intolerance is calculated by taking into account the size of the mutation, the expected distribution of the mutation type, and the functional impact of the mutation. The value range is [0, 1], where 0 indicates that the candidate gene is very tolerant to loss of function, and 1 indicates that the candidate gene is very intolerant to loss of function.
[0175] For example, based on the calculation of loss-of-function intolerance probability: the above 459 genes were obtained from the set of loss-of-function intolerance levels. Furthermore, in this embodiment, the set of intolerance levels for loss of function is from the literature [PMID: 30531870].
[0176] In steps S601 to S602 of this embodiment, the corresponding probability of loss of function intolerance is found directly from the set of loss of function intolerance based on the name of each candidate gene, so as to clarify the tolerance of each candidate gene to loss of function.
[0177] In step S206 as illustrated in some embodiments, the pathogenicity measurement data GG, functional contribution measurement data KG, differential expression measurement data CF, first target measurement data MIG, second target measurement data LNG, target relationship measurement data MLG, and loss-of-function intolerance measurement data VPG for each candidate gene are combined into disease association measurement data by means of averages or summation. It should be noted that there are no specific limitations on the method of combining multiple data points into disease association measurement data.
[0178] Please see Figure 7 In some embodiments, step S104 may include, but is not limited to, steps S701 to S702:
[0179] Step S701: Sort the candidate genes in descending order based on the disease association measurement data to obtain the sorting number of each candidate gene.
[0180] Step S702: Selected genes are extracted from candidate genes according to the preset number of genes to be selected and the sorting number.
[0181] In step S701 of some embodiments, candidate genes are sorted in descending order according to disease association measurement data to obtain a ranking number for each candidate gene, that is, the candidate genes are sorted in descending order of disease association measurement data. The ranking number characterizes the role of the candidate gene in the occurrence and development of the target disease.
[0182] In step S702 of some embodiments, candidate genes with a preset number of gene selections before the sorting number are obtained as selected genes in order to select genes that can represent the target disease.
[0183] For example, if there are 459 genes associated with IS disease, the pathogenicity measure data (GG), functional contribution measure data (KG), differential expression measure data (CF), first target measure data (MIG), second target measure data (LNG), target relationship measure data (MLG), and loss-of-function intolerance measure data (VPG) are concatenated to form the disease association measure data (T_score). The T_scores are then sorted in descending order to obtain the final score of the gene-disease relationship. Let x = 5, then the top 5% of genes in the T_score ranking are selected as the core functional markers specific to IS. The final result is shown in Table 1 below.
[0184]
[0185]
[0186] Table 1
[0187] In steps S701 to S702 of this embodiment, by selecting several candidate genes with the highest correlation as selected genes, genes that play an important role in the occurrence and development of the target disease are selected, making the construction of the specific functional core markers of the target disease more accurate.
[0188] Please refer to Figure 8 In this embodiment, a set of functional genes is extracted from a biological functional database named GO and KEGG. Disease-associated marker sets for the target disease were extracted from MarkerDB and HPO, denoted as "Golden-sets". Transcriptomics data for the disease were obtained from databases such as NCBI GEO and EBI ArrayExpress, denoted as omics-sets. A list of non-coding RNAs associated with the target disease was extracted from databases such as HMDD, miR2Disease, and LncRNADisease v3.0. Then, based on the list of non-coding RNAs (miRNAs, lncRNAs), a set of relationships between miRNAs and target genes was obtained from databases such as miRTarBase and LncRNA2Target to obtain the RNA-target gene set, denoted as... as well as Based on LncBase, the relationship between miRNAs and lncRNAs is obtained to form an RNA relationship set, and for any miRNA mi i The set of LncRNAs regulated by it is denoted as Genetic variation data were obtained from large-scale population genome databases such as gnomAD, dbSNP, and 1KGP to obtain a set of loss-of-function intolerance levels V = {v1, ..., v2}. nThe pathogenicity metric (GG) is determined by whether candidate genes are located in the disease-associated marker set. The functional contribution metric (KG) is calculated by determining the number of functions present for each candidate gene and then using the ratio between the number of functions and the total number of functions. Disease transcriptomics data from different data categories are preprocessed and then converted into a matrix form. Differential expression values between the disease group and the normal group are obtained from the transcriptomics matrix, and these differential values are then merged into target expression values. Finally, the rank data for each sample column is calculated based on the target expression value and a preset expression threshold, and the mean of the rank data is used as the differential expression metric (CF). The number of candidate genes in the miRNA target gene set is determined as the miRNA target regulation number. The ratio between the miRNA target regulation number and the total number of candidate genes is then used as the first target metric, MIG. Similarly, the number of candidate genes in the lncRNA target gene set is determined as the lncRNA target regulation number. The ratio between the lncRNA target regulation number and the total number of candidate genes is then used as the second target metric, LNG. A competitive relationship between the miRNA and lncRNA target gene sets is determined to obtain the target relationship metric, MLG. For each candidate gene, the corresponding loss-of-function intolerance probability is found from the loss-of-function intolerance degree set to obtain the loss-of-function intolerance metric, VPG. Finally, the pathogenicity metric (GG), functional contribution metric (KG), differential expression metric (CF), first target metric (MIG), second target metric (LNG), target relationship metric (MLG), and loss-of-function intolerance metric (VPG) are combined using a mean or summation method to form the disease association metric. Candidate genes are sorted in descending order of disease association measurement data to obtain the sorting number of each candidate gene. The number of candidate genes selected before the sorting number is determined as the selected genes. The specific function core marker is obtained by labeling the function of the target disease using the selected genes.
[0189] Please see Figure 9 This application also provides an apparatus for constructing core markers specific to complex diseases, which can realize the above-mentioned method for constructing core markers specific to complex diseases. The apparatus includes:
[0190] Data acquisition module 901 is used to acquire biological knowledge data related to the target disease;
[0191] Extraction module 902 is used to extract candidate genes associated with the target disease, target function set data, disease-associated marker set and disease transcriptomics data from biological knowledge data; wherein, the target function set data is a function set data composed of the functional name associated with the target disease, RNA targeting relationship and gene genetic variation results;
[0192] The association measurement module 903 is used to measure the disease association of candidate genes based on the target function set data, the disease association marker set and the disease transcriptomics data, and obtain the disease association measurement data for each candidate gene.
[0193] The screening module 904 is used to screen candidate genes based on disease association measurement data to obtain selected genes;
[0194] The labeling module 905 is used to label the target disease based on the selected gene to obtain the specific functional core marker of the target disease.
[0195] The specific implementation of the device for constructing core markers specific to complex diseases is basically the same as the specific implementation of the method for constructing core markers specific to complex diseases described above, and will not be repeated here.
[0196] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the aforementioned method for constructing core markers specific to complex diseases. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.
[0197] Please see Figure 10 , Figure 10 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes:
[0198] The processor 1001 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.
[0199] The memory 1002 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 1002 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1002, and the processor 1001 calls and executes the method for constructing core markers specific to complex diseases according to the embodiments of this application.
[0200] Input / output interface 1003 is used to implement information input and output;
[0201] The communication interface 1004 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0202] Bus 1005 transmits information between various components of the device (e.g., processor 1001, memory 1002, input / output interface 1003, and communication interface 1004);
[0203] The processor 1001, memory 1002, input / output interface 1003 and communication interface 1004 are connected to each other within the device via bus 1005.
[0204] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for constructing core markers specific to complex diseases.
[0205] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0206] The method and related equipment for constructing function-specific core markers for complex diseases provided in this application extract biological knowledge data related to the target disease from existing biological databases. The biological knowledge data extracted from different biological databases includes candidate genes associated with the target disease, target function set data, disease-associated marker set, and disease transcriptomics data. By analyzing the association between each candidate gene and the target disease using the target function set data, disease-associated marker set, and disease transcriptomics data, disease association measurement data is obtained. This disease association measurement data is then used to screen out selected genes from the candidate genes that play an important role in the occurrence and development of the target disease, and these selected genes are used as core markers for that function of the target disease. Therefore, constructing function-specific core markers for the target disease using existing publicly available background knowledge, and using the target function set data, disease-associated marker set, and disease transcriptomics data from the publicly available background knowledge to analyze the association between genes and the disease, makes the construction of function-specific core markers for the target disease cost-effective and accurate.
[0207] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0208] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0209] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0210] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0211] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0212] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0213] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0214] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0215] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0216] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0217] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. A method for constructing core markers specific to complex diseases, characterized in that, The method includes: Acquire biological knowledge data related to the target disease; wherein, the biological knowledge data is known biological background knowledge related to the target disease, and the biological knowledge data comes from multiple public platforms; From the biological knowledge data, candidate genes associated with the target disease, target function set data, disease-associated marker set, and disease transcriptomics data are extracted; wherein, the target function set data is a function set data composed of the functional name associated with the target disease, RNA targeting relationship, and gene genetic variation results; Based on the target function set data, the disease association marker set, and the disease transcriptomics data, the disease association of the candidate genes is measured to obtain disease association measurement data for each candidate gene. The candidate genes are screened based on the disease association measurement data to obtain selected genes; The target disease is labeled according to the selected gene to obtain the specific functional core marker of the target disease; The target functional set data includes: a functional gene set, a target gene set, and a set of loss-of-function intolerance levels. The functional gene set represents a set of genes with functions related to the target disease. The target gene set represents a set of genes that have a targeting relationship with the RNA of the target disease. The set of loss-of-function intolerance levels includes the degree of loss-of-function intolerance for each candidate gene.
2. The method according to claim 1, characterized in that, The step of evaluating the disease association of the candidate genes based on the target function set data, the disease association marker set, and the disease transcriptomics data to obtain disease association metric data for each candidate gene includes: The pathogenicity of each candidate gene is measured based on the set of disease-associated markers to obtain pathogenicity measurement data; The functional contribution of each candidate gene is measured according to the functional gene set to obtain functional contribution measurement data. Differential expression measurement data was obtained by measuring the differential expression of each candidate gene based on the disease transcriptomics data. Gene targeting measurement is performed on each candidate gene according to the target gene set to obtain gene targeting measurement data. Based on the set of loss-of-function intolerance levels, each candidate gene is measured for loss-of-function intolerance to obtain loss-of-function intolerance measurement data. The pathogenicity measurement data, the functional contribution measurement data, the differential expression measurement data, the gene targeting measurement data, and the loss-of-function intolerance measurement data are used to construct the disease association measurement data for each candidate gene.
3. The method according to claim 2, characterized in that, The differential expression measurement of each candidate gene based on the disease transcriptomics data, to obtain differential expression measurement data, includes: The data categories for acquiring the disease transcriptomics data; The disease transcriptomics data are preprocessed according to the data categories to obtain candidate transcriptomics data; The candidate transcriptomics data are matrix transformed to obtain a transcriptomics matrix; wherein each row of the transcriptomics matrix represents a gene and each column represents a sample. Differential expression was measured on the transcriptomics matrix to obtain the differential expression measurement data.
4. The method according to claim 3, characterized in that, The differential expression measurement of the transcriptomics matrix, to obtain the differential expression measurement data, includes: The expression values of the disease group and the normal group were extracted from the transcriptomics matrix; Obtain the differential expression values between the expression values of the disease group and the expression values of the normal group; The differential expression values are merged to obtain the target expression value; Based on the target expression value and the preset expression threshold, the grade data is determined; The rank data is mean-processed to obtain the differential expression measurement data.
5. The method according to claim 2, characterized in that, The target gene set includes: a miRNA target gene set, a lncRNA target gene set, and an RNA relationship set; the gene targeting evaluation of each candidate gene based on the target gene set to obtain gene targeting metric data includes: The candidate genes are targeted based on the miRNA target gene set to obtain first target measurement data. Based on the set of LncRNA-targeting genes, the candidate genes are targeted to obtain second target measurement data. Based on the miRNA target gene set, the LncRNA target gene set, and the RNA relationship set, the target relationship of the candidate genes is measured to obtain target relationship measurement data. The first target measurement data, the second target measurement data, and the target relationship measurement data are concatenated to obtain the gene target measurement data.
6. The method according to claim 2, characterized in that, The step of measuring loss-of-function intolerance for each candidate gene based on the set of loss-of-function intolerance levels to obtain loss-of-function intolerance measurement data includes: Obtain the gene name of the candidate gene; Based on the gene name, the probability of loss of function intolerance is searched in the set of degree of loss of function intolerance to obtain the measurement data of loss of function intolerance.
7. The method according to any one of claims 1 to 6, characterized in that, The step of screening candidate genes based on the disease association measurement data to obtain selected genes includes: The candidate genes are sorted in descending order based on the disease association measurement data to obtain the sorting number of each candidate gene; The selected gene is extracted from the candidate genes according to the preset number of gene selections and the sorting number.
8. A device for constructing core markers specific to complex diseases, characterized in that, The apparatus for constructing functionally specific core markers for complex diseases according to any one of claims 1 to 7 includes: The data acquisition module is used to acquire biological knowledge data related to the target disease; The extraction module is used to extract candidate genes associated with the target disease, target function set data, disease-associated marker set, and disease transcriptomics data from the biological knowledge data; wherein, the target function set data is a function set data composed of the functional name associated with the target disease, RNA targeting relationship, and gene genetic variation results; The association measurement module is used to measure the disease association of the candidate genes based on the target function set data, the disease association marker set, and the disease transcriptomics data, and to obtain the disease association measurement data for each candidate gene. The screening module is used to screen the candidate genes based on the disease association measurement data to obtain selected genes; The labeling module is used to label the target disease based on the selected gene to obtain the specific functional core marker of the target disease.
9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method for constructing core markers specific to complex diseases as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method for constructing core markers specific to complex diseases as described in any one of claims 1 to 7.