Method, device, and computer-readable recording medium for estimating polygenic risk scores (PRS) using deep learning and metalearning models
The combination of deep learning and metalearning models enhances PRS estimation accuracy by addressing the limitations of existing methods, effectively capturing SNP interactions for improved disease prediction.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-27
- Publication Date
- 2026-03-12
AI Technical Summary
Existing methods for estimating polygenic risk scores (PRS) using genome-wide association studies (GWAS) often fail to accurately reflect the interactions between single nucleotide polymorphisms (SNPs), leading to suboptimal disease prediction.
A method utilizing a deep learning model to assign weights to SNPs for PRS estimation, followed by a metalearning model to generate a final estimate, combining multiple primary estimates from the deep learning models to reduce overfitting and bias.
The approach achieves high-accuracy PRS estimation by accounting for SNP interactions, reducing overfitting and bias, thereby improving disease risk prediction.
Smart Images

Figure 2026508758000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a method, apparatus, and computer-readable medium for estimating polygenic risk scores using deep learning and metalearning models. [Background technology]
[0002] The polygenic risk score (PRS) method is one of the methods for measuring the risk of specific diseases due to congenital factors, and its influence may increase if multiple genetic factors are reflected in predictive models, etc.
[0003] Specifically, the polygenic risk score may refer to a value obtained through a process of modulating the influence value of genetic variants to reflect the characteristics of a specific disease, such as through a numerical process of weighting single nucleotide polymorphisms (SNPs) or specific SNPs.
[0004] Recent studies have shown that while individual SNPs generally show low disease association, certain combinations of SNPs can show high disease association, leading to various studies being conducted to identify optimal SNP combinations that can predict disease onset.
[0005] The above-mentioned background art is technical information that the inventor possessed for the purpose of deriving the present invention or that he acquired in the process of deriving the present invention, and is not necessarily publicly known art that was made public to the general public prior to the filing of the present invention. Summary of the Invention [Problem to be solved by the invention]
[0006] Some embodiments of the present disclosure aim to provide a method, device, and computer-readable recording medium for estimating a polygenic risk score using a deep learning model and a metalearning model. The problems to be solved by the present invention are not limited to the problems mentioned above, and other problems and advantages of the present invention not mentioned will be understood from the following description and will become more clearly understood from the examples of the present invention. Furthermore, it will be understood that the problems to be solved and advantages of the present invention can be achieved by the means set forth in the claims and combinations thereof. [Means for solving the problem]
[0007] As a technical means for achieving the above-mentioned technical problem, a first aspect of the present disclosure can provide a method for estimating a polygenic risk score (PRS), comprising the steps of: acquiring data related to a Genome-Wide Association Study (GNAS) of an individual for whom the polygenic risk score is to be estimated; acquiring a plurality of primary estimates for the individual's polygenic risk score from the data via a plurality of deep learning models; generating a second test set for a metalearning model using the plurality of primary estimates; and acquiring a final estimate for the polygenic risk score from the second test set via the metalearning model.
[0008] A second aspect of the present disclosure may provide an apparatus for estimating a polygenic risk score, the apparatus including at least one memory and at least one processor, wherein the processor acquires data related to a genome-wide association study (GWAS) of an individual for whom the polygenic risk score is to be estimated, acquires a plurality of primary estimates for the individual's polygenic risk score from the data via a plurality of deep learning models, generates a second test set for a metalearning model using the plurality of primary estimates, and acquires a final estimate for the polygenic risk score from the second test set via the metalearning model.
[0009] A third aspect of the present disclosure can provide a computer-readable recording medium having recorded thereon a program for causing a computer to execute the method according to the first aspect.
[0010] In addition, other methods, other systems for implementing the present invention, and computer-readable recording media storing computer programs for carrying out the methods may also be provided.
[0011] Other aspects, features, and advantages beyond those described above will become apparent from the following drawings, claims, and detailed description of the invention. [Effects of the Invention]
[0012] According to one embodiment of the present disclosure, polygenic risk scores can be estimated with high accuracy. [Brief explanation of the drawings]
[0013] [Figure 1] FIG. 1 is a diagram illustrating an example of a method for estimating a polygenic risk score using a deep learning model and a metalearning model according to one embodiment. [Figure 2]FIG. 1 is a diagram illustrating an example of the internal configuration of an apparatus for estimating a polygenic risk score using a deep learning model and a metalearning model according to one embodiment. [Figure 3] 1 is a flowchart illustrating an example of a method for estimating a polygenic risk score using a deep learning model and a metalearning model according to one embodiment. [Figure 4] FIG. 1 illustrates an example of a method for obtaining a first-order estimate of a polygenic risk score via multiple deep learning models according to one embodiment. [Figure 5] FIG. 1 is a diagram illustrating multiple deep learning models according to one embodiment. [Figure 6] FIG. 1 illustrates an example of how a final estimate of a polygenic risk score is obtained via a metalearning model according to one embodiment. [Figure 7] FIG. 2 is a diagram illustrating a metalearning model according to an embodiment. BEST MODE FOR CARRYING OUT THE INVENTION
[0014] According to one embodiment of the present disclosure, a method for estimating a polygenic risk score (PLS) includes obtaining data related to a genome-wide association study (GWAS) of an individual for whom the polygenic risk score is to be estimated, obtaining a plurality of primary estimates for the individual's polygenic risk score from the data via a plurality of deep learning models, generating a second test set for a metalearning model using the plurality of primary estimates, and obtaining a final estimate for the polygenic risk score from the second test set via the metalearning model. DETAILED DESCRIPTION OF THE INVENTION
[0015] The advantages and features of the present invention, as well as methods for achieving them, will become more apparent from the detailed description of the embodiments accompanied by the accompanying drawings. However, the present invention is not limited to the embodiments presented below, and can be realized in various forms, and it should be understood that the present invention includes all modifications, equivalents, and alternatives within the spirit and technical scope of the present invention. The embodiments presented below are provided to fully disclose the present invention and fully convey the scope of the invention to those skilled in the art to which the present invention pertains. In describing the present invention, if a detailed description of related publicly known technology is considered to obscure the gist of the present invention, such detailed description will be omitted.
[0016] The terms used in this application are merely used to describe specific embodiments and are not intended to limit the present invention. The singular expressions include the plural expressions unless the context clearly dictates otherwise. In this application, terms such as "comprise" or "have" are intended to specify the presence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and should be understood not to preclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0017] Some embodiments of the present disclosure may be represented by functional blocks and various processing steps. Some or all of these functional blocks may be implemented in a variety of hardware and / or software configurations that perform specific functions. For example, the functional blocks of the present disclosure may be implemented by one or more microprocessors or by circuit configurations for a given function. Furthermore, for example, the functional blocks of the present disclosure may be implemented in various programming or scripting languages. The functional blocks may be implemented by algorithms executed by one or more processors. Furthermore, the present disclosure may employ conventional techniques for electronic configuration, signal processing, and / or data processing. Terms such as "mechanism," "element," "means," and "configuration" may be used broadly and are not limited to mechanical and physical configurations.
[0018] Furthermore, the connecting lines or connecting members between components shown in the drawings are merely exemplary of functional and / or physical or circuit connections, and in an actual device the connections between components may be represented by various interchangeable or additional functional, physical, or circuit connections.
[0019] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The present invention will be described in detail below with reference to the accompanying drawings. However, the present invention may be embodied in many different forms and should not be construed as being limited to the examples set forth herein.
[0020] As used herein, the term "gene" refers to a segment of a nucleic acid sequence (also referred to herein as a "coding sequence" or "coding region") that encodes a protein or RNA, optionally together with regulatory regions, such as promoters, operators, terminators, etc., which may be located upstream or downstream of the coding sequence.
[0021] The term "genetic information" as used herein encompasses information obtained by genetic analysis of a subject, including, for example, information regarding genetic traits or genetic mutations associated with the onset of a particular disease. The genetic mutations may be in the form of, but are not limited to, missense mutations, frameshift mutations, nonsense mutations, splice mutations, nucleotide substitutions, insertions, or deletions. In a specific example, the genetic information may include single nucleotide polymorphisms (SNPs). The risk of developing a disease calculated based on such genetic information includes the congenital risk of developing the disease.
[0022] As used herein, the term "polymorphism" refers to the presence of two or more alleles at a single genetic locus, and a polymorphic site in which only a single nucleotide differs between individuals is called a single nucleotide polymorphism (SNP). Preferred polymorphic markers have two or more alleles that occur at a frequency of 1% or more, more preferably 5% or 10% or more, in a selected population.
[0023] FIG. 1 is a diagram illustrating an example of a method for estimating a polygenic risk score using a deep learning model and a metalearning model according to one embodiment.
[0024] Genome-Wide Association Studies (GWAS) are an exploratory method for finding traits (e.g., height, hair color, eye color, and risk of various diseases) associated with genetic variants.
[0025] Genome-wide association studies generally use a method in which the genetic information of cases (a group with the trait of interest, e.g., a patient group) and controls (a group without the trait, e.g., a normal group) is compared across the entire genome, and genetic variants with a higher frequency in cases are selected as genetic variants associated with the trait.
[0026] As an example, disease risk can be predicted by reflecting the presence or absence of mutations in specific genes known to play important roles in disease development, including many of the genetic mutations identified by genome-wide association studies.
[0027] As an example of a method for predicting the risk of developing a disease, disease-associated genetic mutations can be searched for by calculating a polygenic risk score (PRS).
[0028] Here, calculating a polygenic risk score may refer to a method for determining the relationship between a genetic variant and the onset of a specific disease, for example, a method for determining whether a genetic variant selected through the results of a genome-wide association study is associated with the specific disease, even if it does not directly affect the onset of the specific disease.
[0029] That is, the polygenic risk score calculation method is one of the methods for measuring the risk of a specific disease due to congenital factors, and the influence may be increased if multiple genetic factors are reflected in a prediction model, etc. Specifically, the polygenic risk score may refer to a value obtained through a process of modulating the influence value of genetic mutations to reflect the characteristics of a specific disease, such as through a numerical process in which weighting is given to single nucleotide polymorphisms (SNPs) or specific SNPs.
[0030] Here, single nucleotide polymorphism (SNP) is a type of genetic variation in which genetic base sequences show differences between individuals, and can refer to a position where a single base is different in the base sequence and two allelic base sequence (bi-alelic) variations occur at a frequency of 1% or more in a population.
[0031] In recent years, advances in genome analysis technologies such as genome-wide association studies and next-generation sequencing have led to the development of technologies that can analyze human genome variants, particularly SNP information.
[0032] On the other hand, predicting the risk of developing a disease may involve searching for disease-associated genetic variants. For example, an individual's polygenic risk score can be estimated from genetic variant data associated with the individual's single nucleotide polymorphisms via a predetermined machine learning modeling (algorithm).
[0033] Referring to FIG. 1, a PRS estimation device according to one embodiment of the present disclosure can estimate a polygenic risk score 170 from genome-wide association study results 110 via a PRS estimation model 120 including multiple deep learning models 130 and a metalearning model 150.
[0034] For example, the genome-wide association study results 110 can include genetic information about single nucleotide polymorphisms.
[0035] A general PRS estimation model can estimate PRS using highly associated single nucleotide polymorphisms obtained as a result of GWAS. In this case, the PRS estimation model assumes that each single nucleotide polymorphism acts independently, and estimates PRS, which may not reflect the interaction of single nucleotide polymorphisms.
[0036] On the other hand, in the case of the PRS estimation model 120 according to one embodiment of the present disclosure, PRS estimates can be obtained from GWAS results via the deep learning model 130 and the metalearning model 150.
[0037] Here, single nucleotide polymorphisms included in the GWAS results can be input as input data to the deep learning model 130. The deep learning model 130 assigns weights to each single nucleotide polymorphism to estimate PRSs, so that the deep learning model 130 can estimate PRSs that reflect the interactions of single nucleotide polymorphisms.
[0038] On the other hand, the deep learning model 130 can be composed of multiple deep learning models 130, as in the example shown in Figure 1. Each of the multiple deep learning models 130 can receive the GWAS results 110 as input and output a primary PRS estimate.
[0039] According to one embodiment of the present disclosure, the meta-learning model 150 included in the PRS estimation model 120 can output a final PRS estimate using the primary PRS estimate output from the deep learning model 130.
[0040] That is, the metalearning model 150 can perform a learning process on the weights that the deep-learning model 130 learned to output the primary PRS estimates, and output the final PRS estimates. For example, the primary PRS estimates obtained from the deep-learning model 130 can be used to generate a dataset for the metalearning model 150.
[0041] On the other hand, examples of algorithms that can realize the meta-learning model 150 include, but are not limited to, algorithms and / or methods (techniques) such as a logistic regression model, a support vector machine, a decision tree, a nearest-neighbor classifier, a neural network, a random forest, and a boosted tree.
[0042] A method for estimating a polygenic risk score using a deep learning model and a metalearning model according to one embodiment will now be described in detail with reference to FIGS.
[0043] FIG. 2 is a diagram illustrating an example of the internal configuration of a device that estimates a polygenic risk score using a deep learning model and a metalearning model according to one embodiment.
[0044] 2, a PRS estimation device 200 may include a communication unit 210, a processor 230, and a memory 250. Only components relevant to the embodiment are shown in the PRS estimation device 200 of FIG. 2. Therefore, a person skilled in the art will understand that the PRS estimation device 200 may further include other general-purpose components in addition to the components shown in FIG.
[0045] The communication unit 210 may include one or more components that enable wired / wireless communication with an external server or device. For example, the communication unit 210 may include a short-range communication unit (not shown) or a mobile communication unit (not shown) for communication between the PRS estimation apparatus 200 and the external device.
[0046] The memory 250 is hardware that stores various data processed within the PRS estimation device 200, and can store programs for processing and control by the processor 230.
[0047] Memory 250 can include random access memory (RAM), such as dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), CD-ROM, Blu-ray, or other optical disk storage, hard disk drive (HDD), solid state drive (SSD), or flash memory.
[0048] Processor 230 controls the overall operation of PRS estimation device 200. For example, processor 230 executes a program stored in memory 250 to overall control an input unit (not shown), a display (not shown), communication unit 210, memory 250, and the like.
[0049] The processor 230 may be implemented using at least one of application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, micro-controllers, microprocessors, and other electrical units for performing functions.
[0050] Processor 230 can control the operation of PRS estimation device 200 by executing a program stored in memory 250. As an example, processor 230 can perform at least a portion of a method for estimating a PRS using a deep learning model and a metalearning model, which will be described with reference to FIGS.
[0051] The processor 230 can obtain data related to a genome-wide association study of the individual for whom the polygenic risk score is to be estimated.
[0052] Here, data related to genome-wide association studies may include genetic information related to one or more single nucleotide polymorphisms (SNPs) of an individual.
[0053] The processor 230 can obtain multiple first-order estimates for the individual's polygenic risk score from the data via multiple deep learning models.
[0054] For example, processor 230 can select multiple p-values that indicate the significance levels of the data for the GWAS, generate a SNP list including one or more SNPs for the individual based on the selected p-values, and generate a first test set for each of the multiple deep learning models using the SNP list. Then, processor 230 can input the first test set into each of the multiple deep learning models to obtain multiple first estimates for the polygenic risk score.
[0055] Here, each of the multiple deep learning models may include a model trained using a first training set generated based on genetic information and PRS values for one or more SNPs of an individual.
[0056] The processor 230 may generate a second test set for the metalearning model using the plurality of first-order estimates.
[0057] The processor 230 can obtain a final estimate for the polygenic risk score from the second test set via the meta-learning model.
[0058] Here, the metalearning model can include multiple primary estimates for the polygenic risk score and a model trained using a second training set generated based on the PRS values.
[0059] The final estimate may also include a value calculated based on weights for each of the multiple primary estimates obtained from the multiple deep learning models.
[0060] FIG. 3 is a flowchart illustrating an example method for estimating polygenic risk scores using a deep learning model and a metalearning model according to one embodiment.
[0061] Referring to Figure 3, the method for estimating a polygenic risk score using a deep learning model and a metalearning model may include steps 310 to 370. However, the method is not limited thereto, and in addition to the steps shown in Figure 3, other general steps may also be included in the method for estimating a polygenic risk score using a deep learning model and a metalearning model.
[0062] First, in step 310, processor 230 may obtain data related to a genome-wide association study (GWAS) of an individual for whom a polygenic risk score is to be estimated.
[0063] Here, the data related to the genome-wide association study may include genetic information related to one or more single nucleotide polymorphisms of an individual. For example, the data related to the genome-wide association study may include genetic information related to a single nucleotide polymorphism that has been selected as being associated with the onset of a specific disease as a result of the genome-wide association study.
[0064] For example, data on genome-wide association studies for breast cancer may include markers for single nucleotide polymorphisms, such as Nature 551, 92-94 (2017), selected based on considerations such as race, sample size, and methodology.
[0065] Thereafter, in step 330, processor 230 may obtain multiple first-order estimates for the individual's polygenic risk score from data related to the genome-wide association study described above via multiple deep learning models.
[0066] Below, with reference to Figures 4 and 5, an example in which the PRS estimation device obtains multiple first-order estimates via multiple deep learning models and an example regarding multiple deep learning models will be described.
[0067] FIG. 4 illustrates an example method for obtaining a first-order estimate of a polygenic risk score via multiple deep learning models according to one embodiment.
[0068] Referring to FIG. 4, as described above, the processor 230 can obtain multiple first-order estimates for an individual's polygenic risk score from data 402 related to genome-wide association studies via multiple deep learning models 401.
[0069] For example, the processor 230 may obtain a first primary estimate 451 for an individual's polygenic risk score from data 402 relating to a genome-wide association study via a first deep learning model 431 included in a plurality of deep learning models 401.
[0070] Similarly, the processor 230 can obtain an Nth primary estimate 453 for an individual's polygenic risk score from data 402 related to genome-wide association studies via an Nth deep learning model 433 included in the plurality of deep learning models 401.
[0071] According to the above-described embodiment, the processor 230 can input the same input data, that is, data 402 related to genome-wide association studies, to each of the deep learning models included in the multiple deep learning models 401, and obtain multiple primary estimates for the polygenic risk score.
[0072] According to another embodiment, processor 230 can use data 402 related to the genome-wide association study to generate a test set that is input data for each deep learning model included in the plurality of deep learning models 401. Processor 230 can then obtain multiple first-order estimates for the individual's polygenic risk score from the test set via each deep learning model.
[0073] For example, processor 230 can select multiple p-values that indicate the significance level of the data for the GWAS. Processor 230 can generate an SNP list including one or more SNPs for the individual based on the selected p-values, and can use the generated SNP list to generate a first test set for each of the multiple deep learning models.
[0074] For example, the processor 230 may generate an SNP list including one or more SNPs that satisfy a condition for the P value in the data 402 related to the genome-wide association study based on the selected multiple P values.
[0075] As an example, the processor 230 can generate a first SNP list 411 using SNPs that satisfy a p-value < 0.0003, and can use the first SNP list 411 to generate a test set for the first deep learning model 431.
[0076] As another example, the processor 230 can generate an Nth SNP list 413 using SNPs that satisfy a p-value < 0.0002, and can use the Nth SNP list 413 to generate a test set for the Nth deep learning model 433.
[0077] The processor 230 can then input the first test set into each of a plurality of deep learning models to obtain a plurality of first estimates for the polygenic risk score.
[0078] 4, the processor 230 may obtain a first primary PRS estimate 451 from the first SNP list 411 via a first deep learning model 431. Further, the processor 230 may obtain an Nth primary PRS estimate 453 from the Nth SNP list 413 via an Nth deep learning model 433.
[0079] Meanwhile, in an embodiment, the processor 230 may generate N SNP lists and obtain N primary PRS estimates through N deep learning models. Here, the larger the value of N determined by the processor 230, the more likely it is that overfitting and bias in the PRS estimates will be prevented. However, the processor 230 may determine the value of N in consideration of the capacity of the memory 250 of the processor 230, the performance of the hardware, etc.
[0080] Meanwhile, in this disclosure, a test set may refer to a data set input to a deep learning model and a metalearning model to obtain a PRS estimate via the deep learning model and the metalearning model. That is, the test set may refer to input data for evaluating the performance of a PRS estimate obtained via the deep learning model and the metalearning model.
[0081] In the following, a training set in this disclosure may refer to a data set for training a deep learning model and a metalearning model.
[0082] FIG. 5 is a diagram illustrating multiple deep learning models according to one embodiment.
[0083] According to one embodiment, the method for estimating a PRS may include a training process and an estimation process, where the training process may refer to a process in which the processor 230 trains a deep learning model, and the estimation process may refer to a process in which the processor 230 obtains a PRS estimate through the trained deep learning model.
[0084] For example, the processor 230 can generate a first training set 530 based on the individual's genetic information for one or more SNPs and the PRS values corresponding to the genetic information, and the processor 230 can train the deep learning model 510 using the first training set 530.
[0085] That is, each of the deep learning models 510 included in the plurality of deep learning models may include a model trained using a first training set 530 generated based on genetic information 531 regarding one or more SNPs of an individual and PRS values 533 corresponding to the genetic information.
[0086] Here, the genetic information regarding SNPs can include a list of SNPs generated by selecting multiple P values as described above.
[0087] For example, the processor 230 may generate a first training set including the first SNP list described above in FIG. 4 as input data 531 and PRS values corresponding to the first SNP list as output data 533 to train a first deep learning model.
[0088] Similarly, the processor 230 may train an Nth deep learning model using a first training set 530, where the first training set 530 may include the Nth SNP list as input data 531 and the PRS values corresponding to the Nth SNP list as output data 533.
[0089] According to one embodiment, the deep learning model 510 can, during the training process, learn weights to be assigned to the SNPs included in the SNP list in order to output the PRS value 533. As an example, the first deep learning model 510 can, during the training process, learn weights to be assigned to the SNPs included in the first SNP list in order to output the PRS value 533.
[0090] Meanwhile, in the above-described estimation process, processor 230 may generate a first test set 550 based on data related to a genome-wide association study of an individual for whom a polygenic risk score is to be estimated. For example, first test set 550 may include, as input data 551, a list of SNPs generated by selecting genetic information or multiple P-values for one or more SNPs of the individual. For example, first test set 550 may include an Nth SNP list from the above-described first SNP list.
[0091] Subsequently, in the estimation process, the processor 230 can obtain a plurality of primary PRS estimates from the first test set 550 via the plurality of deep learning models as output data 553. At this time, each of the deep learning models 510 included in the plurality of deep learning models can output the primary PRS estimates 553 from the SNP list 551 using weights assigned to the SNPs included in the SNP list learned in the above-mentioned learning process.
[0092] Continuing with reference to FIG. 3, in the next step 350, the processor 230 may generate a second test set for the Meta-Learning model using the first-order estimates for the individual's PRS.
[0093] In a next step 370, the processor 230 may obtain a final estimate of the PRS from the second test set via the meta-learning model.
[0094] Here, the second test set may refer to a data set to be input to the metalearning model to obtain a final estimate for the PRS, as described above, i.e., the second test set may refer to input data to be input to the metalearning model to evaluate performance against the final estimate obtained through the metalearning model.
[0095] According to one embodiment, the meta-learning model can output a final estimate for the PRS based on a second test set generated using multiple primary estimates that are output data of the deep learning models, the second test set being different from the first test set of the multiple deep learning models.
[0096] As a result, the PRS estimation method that obtains PRS estimates through a meta-learning model outputs PRS estimates based on a dataset that is different from the test set of the deep learning model, which can reduce overfitting and bias compared to PRS estimation methods that obtain PRS estimates only through a deep learning model.
[0097] An example will now be described with reference to FIG. 6 in which processor 230 obtains a final estimate for the PRS from a second data set via a meta-learning model.
[0098] Additionally, the processor 230 may train the meta-learning model based on a dataset different from the first training set of the deep learning model. In this regard, an example of a meta-learning model is described below with reference to FIG.
[0099] FIG. 6 illustrates an example of how a final estimate of a polygenic risk score is obtained via a metalearning model according to one embodiment.
[0100] Referring to FIG. 6, the processor 230 may obtain a final PRS estimate 650 from the second test set 610 via the meta-learning model 630.
[0101] As mentioned above, the second test set 610 may refer to a data set generated using multiple first-order PRS estimates obtained from a deep learning model.
[0102] For example, the plurality of primary PRS estimates may include integers from 1 to 5. For example, processor 230 may obtain a first primary PRS estimate of "3" from a first SNP list via a first deep learning model. Further, processor 230 may obtain a second primary PRS estimate of "2" from a second SNP list via a second deep learning model.
[0103] As an example, assuming that processor 230 obtains five primary PRS estimates through first to fifth deep learning models, the first to fifth primary PRS estimates may be "3, 2, 3, 2, 3" in order.
[0104] In the above example, processor 230 may generate second test set 610 for meta-learning model 630 using the first through fifth primary PRS estimates: "3, 2, 3, 2, 3."
[0105] Also, as an example, processor 230 may obtain a final estimate 650 of PRS of "3" from a second test set 610 generated using "3, 2, 3, 2, 3" via meta-learning model 630.
[0106] According to one embodiment, the meta-learning model 630 may be input with a second test set generated using "3, 2, 3, 2, 3" and output a final PRS estimate of "3."
[0107] Here, the multiple primary PRS estimates "3, 2, 3, 2, 3" may be PRS estimates output from SNP lists generated based on different P values. For example, the first primary PRS estimate "3" and the second primary PRS estimate "2" may represent values output from the first SNP list and the second SNP list, respectively.
[0108] That is, even if the algorithms that construct the first and second deep learning models are the same, the input data input to each deep learning model are different, and the weights given to the SNPs included in the SNP list during the learning process of each deep learning model are different, so the first primary PRS estimate and the second primary PRS estimate may be PRSs estimated using different methods.
[0109] Therefore, the meta-learning model 630 receives as input data the second test set generated using various primary PRS estimates estimated in different ways as described above, which further prevents overfitting and bias in the final PRS estimates.
[0110] FIG. 7 is a diagram illustrating a metalearning model according to one embodiment.
[0111] 5, according to one embodiment, the method for estimating a PRS may include a training process and an estimation process, where the training process may include the processor 230 training the meta-learning model 710.
[0112] For example, processor 230 can generate second training set 730 based on a plurality of first-order estimates for the PRS and PRS values corresponding to the plurality of first-order estimates, and processor 230 can train meta-learning model 710 using second training set 730.
[0113] That is, the meta-learning model 710 may include a plurality of primary estimates 731 for the PRS and a model trained using a second training set 730 generated based on PRS values 733 corresponding to the plurality of primary estimates.
[0114] As an example, during the learning process, the metalearning model 710 can learn the weights to be assigned to the first through Nth primary PRS estimates 731 in order to output the PRS value 733. That is, during the learning process, the metalearning model 710 can learn the weights to be assigned to the output data of the deep-learning model 510.
[0115] Alternatively, in another example, the metalearning model 710 can learn the weights assigned to SNPs during the training process to enable a first deep learning model to output a first primary PRS estimate, or the weights assigned to SNPs during the training process to enable an Nth deep learning model to output an Nth primary PRS estimate. The metalearning model 710 can assign weights to each of the primary PRS estimates 731 using the weights assigned to SNPs during the training process.
[0116] Then, in the estimation process described above, the processor 230 can obtain a final PRS estimate 753 from the second test set 750 via the meta-learning model 710.
[0117] That is, the meta-learning model 710 can use a data set generated using the first to Nth primary PRS estimates as input data 751 and output a final PRS estimate as output data 753.
[0118] Here, according to one embodiment, the meta-learning model 710 can output a final PRS estimate based on the weights assigned to each of the primary PRS estimates included in the second test set 750 during the learning process described above.
[0119] For example, the final PRS estimate may be an estimate that is output based on a weight assigned to each of the output data of the plurality of deep learning models. For example, the final PRS estimate may include an estimate that is output based on a weight assigned to each of the plurality of primary PRS estimates output from the plurality of deep learning models.
[0120] Embodiments of the present invention may be implemented in the form of a computer program executable by various components on a computer, and such a computer program may be recorded on a computer-readable medium, which may include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specially configured to store and execute program instructions, such as ROM, RAM, and flash memory.
[0121] On the other hand, the computer program may be one specially designed and constructed for the present invention, or one that is well known and available to those skilled in the art of computer software. Examples of computer programs include not only machine code such as that produced by a compiler, but also high-level language code that can be executed by a computer using an interpreter, etc.
[0122] According to one embodiment, methods according to various embodiments of the present disclosure may be provided in a computer program product. The computer program product may be traded as a commodity between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., a compact disc read-only memory (CD-ROM)) or may be distributed online (e.g., downloaded or uploaded) via an application store (e.g., Play Store™) or directly between two user devices. In the case of online distribution, at least a portion of the computer program product may be at least temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store server, or an intermediary server.
[0123] Unless explicitly stated or stated to the contrary, the steps constituting the method of the present invention may be performed in any suitable order. The order in which the steps are described above is not necessarily intended to limit the scope of the present invention. The use of any examples or exemplary terms (e.g., etc.) in the present invention is merely for the purpose of explaining the present invention in detail, and the scope of the present invention is not limited by the examples or exemplary terms unless otherwise limited by the claims. Furthermore, those skilled in the art will recognize that various modifications, combinations, and variations can be made depending on design conditions and factors within the scope of the appended claims or their equivalents.
[0124] Therefore, the concept of the present invention should not be limited to the above-described embodiments, and it can be said that not only the scope of the claims described below, but also all scopes equivalent to or modified equivalently from the scope of the claims belong to the scope of the spirit of the present invention.
Claims
1. 1. A method for estimating a polygenic risk score (PRS), comprising: obtaining data on a genome-wide association study (GWAS) of the individual for whom the polygenic risk score is to be estimated; obtaining a plurality of first-order estimates for the individual's polygenic risk score from the data via a plurality of deep learning models; generating a second test set for the metalearning model using the plurality of primary estimates; and obtaining a final estimate for the polygenic risk score from the second test set via the metalearning model.
2. The data on the genome-wide association study 10. The method of claim 1, comprising genetic information regarding one or more Single Nucleotide Polymorphisms (SNPs) of the individual.
3. Each of the plurality of deep learning models 2. The method of claim 1, comprising a model trained using a first training set generated based on genetic information for one or more SNPs of the individual and PRS values corresponding to the genetic information.
4. obtaining the plurality of primary estimates for the polygenic risk score comprises: selecting a plurality of p-values representing the significance levels of the data related to the GWAS; generating a SNP list including one or more SNPs for the individual based on the selected P-values, and generating a first test set for each of the plurality of deep learning models using the SNP list; and 2. The method of claim 1, comprising inputting the first test set into each of the plurality of deep learning models to obtain the plurality of primary estimates for the polygenic risk score.
5. The meta-learning model is 2. The method of claim 1, further comprising: a model trained using the plurality of primary estimates for the polygenic risk score and a second training set generated based on PRS values corresponding to the plurality of primary estimates.
6. The final estimate is The method of claim 1 , further comprising outputting an estimate based on weights assigned to each of the plurality of primary estimates obtained from the plurality of deep learning models.
7. A computer-readable recording medium having recorded thereon a program for causing a computer to execute the method of claim 1.
8. 1. An apparatus for estimating a polygenic risk score, comprising: The device comprises: at least one memory; and at least one processor; The processor: obtaining genome-wide association study (GWAS) data for the individual for whom the polygenic risk score is to be estimated; obtaining a plurality of first-order estimates for the individual's polygenic risk score from the data via a plurality of deep learning models; generating a second test set for the metalearning model using the plurality of primary estimates; obtaining a final estimate for the polygenic risk score from the second test set via the meta-learning model.
9. The data on the genome-wide association study The device of claim 8 , comprising genetic information regarding one or more single nucleotide polymorphisms (SNPs) of the individual.
10. Each of the plurality of deep learning models The apparatus of claim 8 , comprising a model trained using a first training set generated based on genetic information for one or more SNPs of the individual and PRS values corresponding to the genetic information.
11. Obtaining the plurality of primary estimates for the polygenic risk score comprises: selecting a plurality of p-values representing the significance levels of the data related to the GWAS; generating a SNP list including one or more SNPs for the individual based on the selected plurality of P-values, and generating a first test set for each of the plurality of deep learning models using the SNP list; 9. The apparatus of claim 8, further comprising inputting the first test set to each of the plurality of deep learning models to obtain the plurality of primary estimates for the polygenic risk score.
12. The meta-learning model is 9. The apparatus of claim 8, further comprising: a model trained using the plurality of primary estimates for the polygenic risk score and a second training set generated based on PRS values corresponding to the plurality of primary estimates.
13. The final estimate is The apparatus of claim 8 , further comprising an output estimate based on weights assigned to each of the plurality of primary estimates obtained from the plurality of deep learning models.