A construction method, scoring method, device and electronic device for a sequence deep learning model of polygenic risk scores in the general population
The multigene risk scoring model constructed through LASSO regression screening and sequence neural network solves the problems of insufficient genetic information capture and neglecting linkage imbalance relationships in the prior art, and achieves more efficient and accurate risk scoring.
Patent Information
- Application Number
- CN202411335392.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-24
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2044-09-24
AI Technical Summary
The existing multigene risk scoring methods have shortcomings in capturing genetic information and predictive capabilities, especially ignoring the large number of potential weak-effect variants and linkage imbalance relationships.
LASSO regression was used to screen genetic variants, and a multigene risk scoring model was constructed in combination with sequence neural networks, integrating the genotype information, linkage imbalance correlation coefficient and distance of the upstream and downstream SNPs of each genetic variant, and enhancing the capture ability of the model.
It improves the sensitivity and accuracy of multigene risk scores, enables more comprehensive capture of genetic information and applies it in cohorts of different natural populations.
Smart Images

Figure CN119314549B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of gene detection and analysis, and in particular to a method for constructing a sequence deep learning model for polygenic risk scoring of a natural population, a scoring method, a device and an electronic device. Background Art
[0002] Polygenic Risk Scores (PRS) play an important role in the early screening of complex diseases and precision medicine. By comprehensively analyzing the numerous tiny variations in an individual's genetic code, a comprehensive and sophisticated risk assessment system is constructed, opening up a new path for the prediction and prevention of complex diseases such as chronic respiratory diseases and lung cancer.
[0003] In the process of constructing polygenic risk scores, the conventional practice is to identify genetic variants associated with chronic respiratory diseases based on the summary data of Genome-Wide Association Study (GWAS). However, the statistical significance threshold usually used in GWAS is very strict, which means that only some variants with strong signals will be included, while ignoring a large number of potential weak effect variants, which are not significant individually, but may have a significant contribution to the disease risk when combined.
[0004] In view of this, how to more comprehensively capture genetic information and improve the predictive ability of risk scores has become a research focus. Accurate polygenic risk scores are helpful for early screening of complex diseases in natural populations and precision medicine. It is crucial to develop an algorithm that can efficiently and accurately analyze polygenic risk scores in natural populations. This method requires powerful bioinformatics analysis capabilities to accurately capture the connection with clinical phenotypes from massive genomic variation data.
[0005] In view of this, the present invention is proposed. Summary of the invention
[0006] The objective of the present invention is to provide a construction method, a scoring method, a device and an electronic device for a sequence deep learning model of polygenic risk scores in the general population to solve the above technical problems. In order to more comprehensively capture genetic information and improve the predictive ability of risk scores, the present invention adopts the LASSO regression (Least Absolute Shrinkage and Selection Operator) method. LASSO regression is a variable screening technique that can not only handle high-dimensional data, but also automatically shrink (or completely eliminate) the coefficients that contribute less to the model through a regularization penalty term, thereby retaining important variations. This method allows more variations that are statistically less significant but may still have certain biological significance to be incorporated while controlling the model complexity, improving the sensitivity and accuracy of polygenic risk scores.
[0007] Compared with the strict threshold of genome-wide association studies, using LASSO regression to screen genetic variations related to specific diseases will screen out more genetic variations, and most of them are in linkage disequilibrium (LD). When it comes to polygenic risk scores, the effects of genetic variations in linkage disequilibrium (LD) may overlap. Moreover, different genetic structures and different LD relationships exist among different general populations, which are also the main reasons why the PRS model cannot be universal in each population. Traditional deep neural networks (DNNs) and convolutional neural networks (CNNs) usually ignore the role of LD when dealing with genetic variations.
[0008] The inventors found that by using a sequence neural network, the correlation in sequence data, including LD among multiple variations, can be effectively captured. By using a sequence neural network, genomic dataset of each general population cohort can be better utilized, and a more accurate polygenic risk score model can also be constructed.
[0009] In addition, the present invention innovatively provides a new genetic variation data frame integrating multiple SNPs of each genetic variation. Each genetic variation data frame includes the genotype information of multiple SNPs upstream and downstream of each target genetic variation, the LD correlation coefficient of each general population variation, and the distance from the target genetic variation. Incorporating the genotype, distance and correlation coefficient of the upstream and downstream sequences of the target variation into the deep learning model comprehensively can effectively enrich the variation information, effectively consider the LD relationship, improve the accuracy of the software, and help the model to be applied in multiple general population cohorts.
[0010] The present invention uses LASSO regression to screen genetic variations, which not only reduces the influence of strict thresholds, but also incorporates as many variations with possible effects as possible, ensuring the reliability of the subsequent sequence model.
[0011] The present invention is implemented as follows:
[0012] In a first aspect, the present invention provides a method for constructing a sequence deep learning model for polygenic risk scores in the general population, which includes the following steps:
[0013] S1: Use LASSO regression to screen genetic variant data from the preprocessed genotype data;
[0014] S2: Establish a genetic variant window: Calculate the linkage disequilibrium (LD) correlation coefficient of each variant in the general population for the genetic variant data after LASSO regression in step S1; Establish a genetic variant data frame of X SNPs for each genetic variant; where each genetic variant data frame includes the genotype information of X SNPs of each target genetic variant, the distance from the target genetic variant, and the LD correlation coefficient; The obtained dataset of n*X*Y is the genetic variant window, where n is the number of variants; The number of SNPs X in each genetic variant data frame ≤ 51; Y ≥ 3;
[0015] The X SNPs refer to the target variant, 0.5*(X - 1) genetic variants upstream of the target variant, and 0.5*(X - 1) genetic variants downstream of the target variant; and the X SNPs are located within 250 kb upstream and downstream of the target variant, and the correlation coefficient r2 with the target variant > 0.1;
[0016] S3: Construction of a sequence deep learning model for polygenic risk scores:
[0017] Perform dimensionless processing on the genetic variant window obtained in step S2, and then input the dimensionless data into a sequence neural network for model training until the model reaches the preset convergence condition, and the model is used for polygenic risk scores.
[0018] In a second aspect, the present invention also provides a method for obtaining polygenic risk scores in the general population based on a deep learning model, including: Scoring the SNP data to be scored using the sequence deep learning model constructed based on the above method for constructing a sequence deep learning model for polygenic risk scores.
[0019] In a third aspect, the present invention also provides a device for polygenic risk scores in the general population, including: an input module, a control module, and an output module;
[0020] The input module is configured to: input SNP sample data and / or sample genotype data;
[0021] The control module includes: a genotype data preprocessing module, a genetic variant data screening module, a genetic variant window construction module, a model construction module, and a data scoring module;
[0022] The genotype data preprocessing module is configured to: perform file conversion on SNP sample data and / or sample genotype data;
[0023] The genetic variation data screening module is configured to: screen genetic variation data using LASSO regression;
[0024] The genetic variation window construction module is configured to: calculate the linkage disequilibrium (LD) correlation coefficient of each natural population variation for the genetic variation data after LASSO regression; establish a genetic variation data frame of X SNPs for each genetic variation; where each genetic variation data frame includes the genotype information of X SNPs of each target genetic variation, the distance from the target genetic variation, and the LD correlation coefficient; obtaining a dataset of n*X*Y is the genetic variation window, where n is the number of variations; the number of SNPs X in each genetic variation data frame ≤ 51; Y ≥ 3; the X SNPs refer to the target variation, 0.5*(X - 1) genetic variations upstream of the target variation, and 0.5*(X - 1) genetic variations downstream of the target variation; and the X SNPs are located within 250 kb upstream and downstream of the target variation, and the correlation coefficient r2 with the target variation > 0.1;
[0025] The model construction module is configured to: dimensionlessize the obtained genetic variation window, and then input the dimensionlessized data into a sequence neural network for model training until the model reaches a preset convergence condition;
[0026] The data scoring module is configured to: score the SNP data to be scored based on the sequence deep learning model constructed by the model construction module;
[0027] The output module is configured to: output the polygenic risk score result.
[0028] In a fourth aspect, the present invention also provides an electronic device, including a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the steps of the above method are completed.
[0029] In a fifth aspect, the present invention also provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by the processor, the steps of the above method are completed.
[0030] The present invention has the following beneficial effects:
[0031] (1) The present invention innovatively provides a new genetic variant data frame integrating multiple SNPs for each genetic variant. Each genetic variant data frame includes the genotype information of multiple SNPs upstream and downstream of each target genetic variant, the LD correlation coefficient of each natural population variant, and the distance from the target genetic variant. Incorporating the genotypes, distances, and correlation coefficients of the upstream and downstream sequences of the target variant into the deep learning model comprehensively can effectively enrich the variant information, effectively consider the LD relationship, improve the accuracy of the software, and help the model to be applied in multiple natural population cohorts.
[0032] (2) By using a sequence neural network, the correlation in sequence data, including LD between multiple variants, can be effectively captured. By using the sequence neural network, the genomic datasets of each natural population cohort can be better utilized, and a more accurate polygenic risk score model can also be constructed.
[0033] (3) The present invention uses LASSO regression to screen genetic variants, which not only reduces the influence of strict thresholds but also incorporates as many variants with possible effects as possible, ensuring the reliability of the subsequent sequence model. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0035] Figure 1 It is the overall flowchart of the construction method of the sequence deep learning model for polygenic risk scoring of natural populations;
[0036] Figure 2 It is a schematic diagram of genetic variant window data and model development. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0037] Reference will now be provided in detail to the embodiments of the present invention, one or more examples of which are described below. Each example is provided by way of explanation and not limitation of the invention. In fact, it will be apparent to those skilled in the art that various modifications and variations can be made to the present invention without departing from the scope or spirit of the invention. For example, features described or illustrated as part of one embodiment can be used in another embodiment to yield a still further embodiment.
[0038] Complex diseases, also known as polygenic diseases, refer to diseases where the phenotypes of related traits do not follow an obvious Mendelian inheritance pattern. Their genetic mechanisms are relatively complex and are not controlled by a single pair of genes but are co-regulated by multiple genes and environmental factors. They are common chronic diseases that are difficult to cure medically and seriously affect the physical health of the people. Common chronic diseases mainly include cardiovascular and cerebrovascular diseases, diabetes, malignant tumors, chronic respiratory diseases, arthritis, obesity, etc.
[0039] In a first aspect, the present invention provides a method for constructing a sequence deep learning model for polygenic risk scores in natural populations, which comprises the following steps:
[0040] S1: Using LASSO regression to screen genetic variant data from the preprocessed genotype data;
[0041] S2: Establishing a genetic variant window: calculating the linkage disequilibrium (LD) correlation coefficient of each natural population variant for the genetic variant data after LASSO regression in step S1; establishing a genetic variant data frame of X SNPs for each genetic variant; wherein each genetic variant data frame includes the genotype information of X SNPs of each target genetic variant, the distance from the target genetic variant, and the LD correlation coefficient; obtaining a dataset of n*X*Y as the genetic variant window, where n is the number of variants; the number of SNPs X in each genetic variant data frame ≤ 51; Y ≥ 3;
[0042] The X SNPs refer to the target variant, 0.5*(X - 1) genetic variants upstream of the target variant, and 0.5*(X - 1) genetic variants downstream of the target variant; and the X SNPs are located within 250 kb upstream and downstream of the target variant, and the correlation coefficient r2 with the target variant > 0.1;
[0043] S3: Constructing a sequence deep learning model for polygenic risk scores:
[0044] Performing dimensionless processing on the genetic variant window obtained in step S2, and then inputting the dimensionless data into a sequence neural network for model training until the model reaches a preset convergence condition, and the model is used for polygenic risk scoring.
[0045] In the above model construction method, a new genetic variant data frame integrating multiple SNPs of each genetic variant is provided. Each genetic variant data frame includes the genotype information of multiple SNPs upstream and downstream of each target genetic variant, the LD correlation coefficient of each natural population variant, and the distance from the target genetic variant. Thus, comprehensively incorporating the genotype, distance, and correlation coefficient of the upstream and downstream sequences of the target variant into the deep learning model can effectively enrich variant information, effectively consider the LD relationship, improve the accuracy of the software, and help the model to be applied in multiple natural population cohorts.
[0046] The present invention uses a sequential neural network, which can effectively capture the correlations in sequential data, including the LD between multiple mutations. By using the sequential neural network, it is possible to better utilize the genomic datasets of various natural population cohorts and also construct a more accurate polygenic risk score model.
[0047] In addition, the present invention uses LASSO regression to screen genetic mutations, which not only reduces the influence of strict thresholds but also includes as many mutations with possible effects as possible, ensuring the reliability of subsequent sequence models.
[0048] The size of the genetic mutation window in the present invention takes into account up to 25 genetic mutation sites and genetic mutations upstream and downstream of the target mutation, thereby effectively enriching mutation information and facilitating the accurate analysis of polygenic risk.
[0049] When X SNPs are located within 250 kb upstream and downstream of the target mutation, the correlation coefficient r2 with the target mutation > 0.1, and the number of genetic mutations upstream or downstream of the target mutation is less than 0.5 * (X - 1); then the actual genetic mutations within 250 kb upstream and downstream are added to the genetic mutation window, and the remaining genetic mutations upstream or downstream of the target mutation are filled with identifiers different from the mutation genotypes (such as NA).
[0050] When X SNPs are located within 250 kb upstream and downstream of the target mutation, the correlation coefficient r2 with the target mutation > 0.1, and the number of genetic mutations upstream or downstream of the target mutation is greater than 0.5 * (X - 1); then 0.5 * (X - 1) genetic mutations within 250 kb upstream and downstream are added to the genetic mutation window.
[0051] In a preferred embodiment of the application of the present invention, these flanking mutations (upstream and downstream mutations) are located within 250 kb upstream and downstream and the correlation coefficient r 2 > 0.1.
[0052] In a preferred embodiment of the application of the present invention, the number of SNPs in each genetic mutation data frame: 11 ≤ X ≤ 51;
[0053] In a preferred embodiment of the application of the present invention, the number of SNPs in each genetic mutation data frame: 21 ≤ X ≤ 51. When X = 31, 15 SNPs upstream or downstream of the target mutation are added to the genetic mutation window.
[0054] In a preferred embodiment of the application of the present invention, if the number of genetic variations upstream of the target variation is less than 0.5*(X - 1), it is filled with an identifier different from the variant genotype; if the number of genetic variations downstream of the target variation is less than 0.5*(X - 1), it is filled with an identifier different from the variant genotype, such as NA. Ensure that all inputs have uniform traits.
[0055] In a preferred embodiment of the application of the present invention, the genotypes of each genetic variation data frame are numerically and / or alphabetically identified;
[0056] In a preferred embodiment of the application of the present invention, the genotype of each genetic variation data frame includes the major allele genotype, heterozygote, minor allele genotype, and an identifier different from the variant genotype;
[0057] In a preferred embodiment of the application of the present invention, when the SNP is upstream of the target genetic variation, the distance from the target genetic variation is set to a negative value; when the SNP is downstream of the target genetic variation, the distance from the target genetic variation is set to a positive value; when the SNP is the target genetic variation itself, the distance from the target genetic variation is set to 0.
[0058] In a preferred embodiment of the application of the present invention, when Y≥4, each genetic variation data frame further includes other variation characteristics of X SNPs of each target genetic variation, and the other variation characteristics are selected from at least one of the variant type and the variant length characteristic.
[0059] In a preferred embodiment of the application of the present invention, Y is 3 or 4. Y is the number of columns of the data frame, such as 3 columns or 4 columns. For example, the first column is genotype information, the second column is the distance from the target variation, and the third column is the linkage disequilibrium correlation coefficient. In this case, Y is 3. If each genetic variation data frame further includes the variant type of X SNPs of each target genetic variation, the fourth column can be set as the variant type of the SNP, and Y is 4. If each genetic variation data frame further includes the variant type and the variant length characteristic of X SNPs of each target genetic variation, the fourth column can be set as the variant type of the SNP, and the fifth column can be set as the variant length characteristic, and Y is 5.
[0060] In a preferred embodiment of the application of the present invention, using LASSO regression to screen genetic variation data includes: using LASSO regression to minimize the following loss function:
[0061]
[0062] where N is the number of samples; p is the number of genetic variations; y i is the target value of the i-th observation; x ij is the value of the j-th genetic variation of the i-th observation; β iis the coefficient of the j-th characteristic variable; β0 is the intercept term; λ is the regularization parameter;
[0063] In a preferred embodiment of the application of the present invention, the PLINK 1.90beta software is used to calculate the linkage disequilibrium (LD) correlation coefficient of each natural population variation;
[0064] In a preferred embodiment of the application of the present invention, the sequence neural network is selected from a recurrent neural network, a Transform, a recursive neural network, a gated recurrent unit, and a long short-term memory unit.
[0065] The recurrent neural network (RNN) includes, but is not limited to, GRU and LSTM.
[0066] Taking the training of the GRU model as an example, after inputting the dimensionless data, it passes through 3 layers of GRU plus an activation layer, and then through two fully connected layers and an activation layer to obtain the prediction result. The activation function includes, but is not limited to, the Sigmoid function, the Tanh activation function, and the ReLU function, etc. Then, through the optimizer, it gradually iterates to make the loss function value tend to the global minimum. The optimizer includes, but is not limited to, the stochastic gradient descent method, the SGD with momentum method, and the adaptive learning rate algorithm.
[0067] The polygenic risk score of the natural population refers to: scoring the risk of at least one polygenic disease among cardiovascular and cerebrovascular diseases, diabetes, malignant tumors, chronic respiratory diseases, arthritis, and obesity in the natural population;
[0068] In a preferred embodiment of the application of the present invention, the polygenic risk score of the natural population refers to: scoring the risk of chronic respiratory diseases in the natural population.
[0069] In a second aspect, the present invention also provides a method for obtaining the polygenic risk score of the natural population based on a deep learning model, including: scoring the SNP data to be scored by using the sequence deep learning model constructed based on the above-mentioned method for constructing the sequence deep learning model of the polygenic risk score.
[0070] In a third aspect, the present invention also provides a device for polygenic risk scoring, including: an input module, a control module, and an output module;
[0071] The input module is configured to: input SNP sample data and / or sample genotype data;
[0072] The control module includes: a genotype data preprocessing module, a genetic variation data screening module, a genetic variation window construction module, a model construction module, and a data scoring module;
[0073] The genotype data preprocessing module is configured to: perform file conversion on SNP sample data and / or sample genotype data;
[0074] The genetic variation data screening module is configured to: screen genetic variation data using LASSO regression;
[0075] The genetic variation window construction module is configured to: calculate the linkage disequilibrium (LD) correlation coefficient of each natural population variation for the genetic variation data after LASSO regression; establish a genetic variation data frame of X SNPs for each genetic variation; where each genetic variation data frame includes the genotype information of X SNPs of each target genetic variation, the distance from the target genetic variation, and the LD correlation coefficient; obtain a dataset of n*X*Y as the genetic variation window, where n is the number of variations; the number of SNPs X in each genetic variation data frame ≤ 51; Y ≥ 3; the X SNPs refer to the target variation, 0.5*(X - 1) genetic variations upstream of the target variation, and 0.5*(X - 1) genetic variations downstream of the target variation; and the X SNPs are located within 250 kb upstream and downstream of the target variation, and the correlation coefficient r2 with the target variation > 0.1;
[0076] The model construction module is configured to: dimensionless the obtained genetic variation window, and then input the dimensionless data into a sequence neural network for model training until the model reaches a preset convergence condition;
[0077] The data scoring module is configured to: score the SNP data to be scored based on the sequence deep learning model constructed by the model construction module;
[0078] The output module is configured to: output the polygenic risk score result.
[0079] In a fourth aspect, the present invention further provides an electronic device, including a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the steps of the above method are completed.
[0080] Specifically, the electronic device may further include a bus and a communication interface, and the memory, the processor, and the communication interface are directly or indirectly electrically connected to each other to achieve data transmission or interaction. For example, these components may be electrically connected to each other through one or more buses or signal lines. The processor can process information and / or data related to target recognition to perform one or more functions described in this application.
[0081] The memory can be, but is not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), Electric Erasable Programmable Read-Only Memory (EEPROM), etc.
[0082] The processor can be an integrated circuit chip with signal processing capabilities. The processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0083] In a fifth aspect, the present invention also provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the steps of the above method are completed.
[0084] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. For those not specified in the embodiments, the conventional conditions or the conditions recommended by the manufacturer are followed. For the reagents or instruments not specified by the manufacturer, they are all conventional products that can be obtained through commercial purchases.
[0085] The features and performance of the present invention will be further described in detail below in conjunction with the embodiments.
[0086] Accurate polygenic risk scores contribute to the early screening of complex diseases and precision medicine. Developing an algorithm that can efficiently and accurately analyze polygenic risk scores has become crucial. Such a method requires powerful bioinformatics analysis capabilities to precisely capture the connection with clinical phenotypes from a vast amount of genomic variant data.
[0087] Example 1
[0088] This embodiment provides a method for constructing a sequence deep learning model for polygenic risk scores in the general population, and the overall route is referred to Figure 1 as shown. The specific steps are as follows:
[0089] 1. Data preprocessing
[0090] In the training and use of the model of the present invention, vcf files, ped / map files, and raw files are allowed. Vcf is a file format for storing variant information in genomic sequences, including annotation information as the beginning part of the file and variant information as the main body. The map / ped files are related and usually used together. The map file records the variant position information, and the ped file records the pedigree and genotype information of the research subjects. The raw file is converted into a genotype file based on the additive effect. The first line of the file records the variant position information, and the remaining are genotype data of 0, 1, and 2. If it is a vcf file and ped / map files, they are converted into raw files using PLINK 1.90beta software [1] for subsequent model training and use.
[0091] 2. Screening genetic variant data using LASSO regression
[0092] This embodiment uses LASSO regression to screen important genetic variants, reduce the complexity of the subsequent model, and prevent overfitting. The optimization objective of Lasso regression is to minimize the following loss function:
[0093]
[0094] where N is the number of samples; p is the number of genetic variants; y i is the target value of the i-th observation; x ij is the value of the j-th genetic variant of the i-th observation; β j is the coefficient of the j-th feature variable. When it is 0, x ij is deleted, achieving the effect of screening genetic variants; β0 is the intercept term; λ is the regularization parameter, used to control the intensity of regularization;
[0095] 3. Establishing a genetic variant window:
[0096] Based on the genetic variant data after LASSO regression, the present invention uses PLINK 1.90beta software to calculate the LD correlation coefficient of each variant in the general population.
[0097] This embodiment proposes a window of 51 SNPs, that is, considering 25 genetic variant sites and genetic variants upstream of the target variant. These flanking variants are located within 250 kb upstream and downstream, and the correlation coefficient r with the target variant 2> 0.1. If the number of flank mutations is less than 25, NA is filled to ensure that all inputs have uniform traits.
[0098] Numerically identify the genotypes of each genetic variant data frame. In this embodiment, four values are set for each variant, which are 0, 1, 2, and -9 in sequence. 0 represents homozygous and the major allele genotype, 1 represents heterozygous, 2 represents homozygous and the minor allele genotype, and -9 represents NA. In other embodiments, the genotypes of each genetic variant data frame can also be alphabetically identified.
[0099] In addition, this embodiment also incorporates LD correlation coefficients as the basis for subsequent models to adapt to various natural populations.
[0100] In this embodiment, numerically identify the distance between the SNP and the target genetic variant and the LD correlation coefficient. When the SNP is upstream of the target genetic variant, the distance from the target genetic variant is set to a negative value; when the SNP is downstream of the target genetic variant, the distance from the target genetic variant is set to a positive value; when the SNP is the target genetic variant itself, the distance from the target genetic variant is set to 0. Set the upstream NA to -999999 bp and the downstream NA to 999999 bp. For the linkage disequilibrium correlation coefficient, NA is 0.
[0101] Finally, a 51 * 3 data frame is obtained for each variant (refer to Figure 2 ). The first column is genotype information, the second column is the distance from the target variant, and the third column is the linkage disequilibrium correlation coefficient. Therefore, an n * 51 * 3 data set will be obtained, where n is the number of variants and 3 is the number of columns.
[0102] 4. Construct a sequence deep learning model for polygenic risk scores of natural populations.
[0103] The deep learning network architecture is used to process genetic variant window data. First, dimensionless the genetic variant window data to ensure the reliability of subsequent models. If the number of variants n in the training data is greater than 50, the model input data will adopt a batch strategy to reduce the model training pressure. Sequence models such as RNN (such as GRU and LSTM) and Transform can be used for model training.
[0104] Taking the training of the GRU model as an example in this embodiment, after inputting the dimensionless data, the present invention passes through 3 layers of GRU plus an activation layer, and then through two fully connected layers and an activation layer to obtain the prediction result.
[0105] Activation functions include but are not limited to Sigmoid function, Tanh activation function, ReLU function, etc. In this embodiment, it is the Sigmoid function.
[0106] Then, the optimizer is gradually iterated to make the loss function value tend to the global minimum. The optimizers include, but are not limited to, the stochastic gradient descent method, the SGD with momentum method, and the adaptive learning rate algorithm. In this embodiment, the stochastic gradient descent method is used.
[0107] Embodiment 2
[0108] This embodiment provides a device for multi-gene risk scoring of the natural population, including: an input module, a control module, and an output module;
[0109] The input module is configured to: input SNP sample data and / or sample genotype data;
[0110] The control module includes: a genotype data preprocessing module, a genetic variation data screening module, a genetic variation window construction module, a model construction module, and a data scoring module;
[0111] The genotype data preprocessing module is configured to: perform file conversion on the SNP sample data and / or sample genotype data;
[0112] The genetic variation data screening module is configured to: screen the genetic variation data using LASSO regression;
[0113] The genetic variation window construction module is configured to: calculate the linkage disequilibrium (LD) correlation coefficient of each natural population variation for the genetic variation data after LASSO regression; establish a genetic variation data frame of 51 SNPs for each genetic variation; where each genetic variation data frame includes the genotype information of 51 SNPs of each target genetic variation, the distance from the target genetic variation, and the LD correlation coefficient; obtain an n*51*3 data set as the genetic variation window, where n is the number of variations; and the 51 SNPs are within 250 kb upstream and downstream of the target variation, and the correlation coefficient r2 with the target variation > 0.1;
[0114] The model construction module is configured to: dimensionlessize the obtained genetic variation window, and then input the dimensionlessized data into a sequence neural network for model training until the model reaches the preset convergence condition;
[0115] The data scoring module is configured to: score the SNP data to be scored based on the sequence deep learning model constructed by the model construction module;
[0116] The output module is configured to: output the multi-gene risk scoring result.
[0117] Embodiment 3
[0118] This embodiment provides a method for obtaining polygenic risk scores for the general population based on a deep learning model, which includes: scoring SNP data to be scored using a sequence deep learning model constructed based on the above-mentioned method for constructing a sequence deep learning model for polygenic risk scores.
[0119] In summary, the present invention innovatively provides a new genetic variant data frame integrating multiple SNPs of each genetic variant. Each genetic variant data frame includes genotype information of multiple SNPs upstream and downstream of each target genetic variant, LD correlation coefficients of each natural population variant, and the distance from the target genetic variant. Incorporating the genotypes, distances, and correlation coefficients of the upstream and downstream sequences of the target variant comprehensively into the deep learning model can effectively enrich variant information, effectively consider LD relationships, improve the accuracy of the software, and help the model be applied in multiple natural population cohorts.
[0120] The present invention uses a sequence neural network, which can effectively capture the correlations in sequence data, including LD between multiple variants. By using a sequence neural network, genomic datasets of various natural population cohorts can be better utilized, and a more accurate polygenic risk score model can also be constructed.
[0121] The present invention uses LASSO regression to screen genetic variants, which not only reduces the influence of strict thresholds but also incorporates as many variants with possible effects as possible, ensuring the reliability of the subsequent sequence model.
[0122] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
[0123] References:
[0124] [1] Chang CC, Chow CC, Tellier LC, Vattikuti S, Purcell SM, Lee JJ. Second-generation PLINK: rising to the challenge of larger and richer datasets. Gigascience. 2015 Feb 25;4:7.
Claims
1. A method for constructing a sequence deep learning model for polygenic risk scoring of natural populations, characterized in that: It includes the following steps: S1: LASSO regression is used to screen genetic variation data for preprocessed genotype data; S2: Establishing a genetic variation window: Calculate the linkage disequilibrium correlation coefficient of each natural population variation based on the genetic variation data after LASSO regression in step S1; Establish a genetic variation data frame of X SNPs for each genetic variation; Each of the genetic variation data frames includes the genotype information of the X SNPs of each target genetic variation, the distance to the target genetic variation and the LD correlation coefficient; The obtained n*X*Y data set is the genetic variation window, where n is the number of variations; The number of SNPs in each genetic variation data frame is X≤51; Y≥3; X SNPs refer to the target variation, 0.5*(X-1) genetic variations upstream of the target variation, and 0.5*(X-1) genetic variations downstream of the target variation; and the X SNPs are located within 250kb upstream and downstream of the target variation, and have a correlation coefficient r2 > 0.1 with the target variation; S3: Construction of sequence deep learning model for polygenic risk score: The genetic variation window obtained in step S2 is dimensionless, and then the dimensionless data is input into a sequence neural network for model training until the model reaches a preset convergence condition. The model is used for polygenic risk scoring.
2. The method for constructing a sequence deep learning model for polygenic risk scoring of natural populations according to claim 1, characterized in that: The number of SNPs in each of the genetic variation data frames: 11≤X≤51.
3. The method for constructing a sequence deep learning model for polygenic risk scoring of natural populations according to claim 2, characterized in that: The number of SNPs in each of the genetic variation data frames: 21≤X≤51.
4. The method for constructing a sequence deep learning model for polygenic risk scoring of natural populations according to claim 1, characterized in that: If the number of genetic variants upstream of the target variant is less than 0.5*(X-1), it is filled with an identifier that is different from the variant genotype; if the number of genetic variants downstream of the target variant is less than 0.5*(X-1), it is filled with an identifier that is different from the variant genotype.
5. The method for constructing a sequence deep learning model for polygenic risk scoring of natural populations according to claim 1, characterized in that: The genotype of each genetic variation data frame is numerically and / or alphabetically labeled.
6. The method for constructing a sequence deep learning model for polygenic risk scoring of natural populations according to claim 5, characterized in that: The genotype of each genetic variation data frame includes a major allele type, a heterozygote, a minor allele type, and an identifier that is distinguished from the variant genotype.
7. The method for constructing a sequence deep learning model for polygenic risk scoring of natural populations according to claim 6, characterized in that: When the SNP is located upstream of the target genetic variation, the distance to the target genetic variation is set to a negative value; when the SNP is located downstream of the target genetic variation, the distance to the target genetic variation is set to a positive value; when the SNP is the target genetic variation, the distance to the target genetic variation is set to 0.
8. The method for constructing a sequence deep learning model for polygenic risk scoring of natural populations according to claim 1, characterized in that: When Y≥4, each of the genetic variation data frames further includes other variation features of the X SNPs of each target genetic variation, and the other variation features are selected from at least one of variation type and variation length features.
9. The method for constructing a sequence deep learning model for polygenic risk scoring of natural populations according to claim 1, characterized in that: Using the LASSO regression to screen genetic variation data includes: using LASSO regression to minimize the following loss function: Where N is the number of samples; p is the number of genetic variants; is the target value of the ith observation; is the value of the jth genetic variant of the ith observation; is the coefficient of the jth characteristic variable; is the intercept term; is the regularization parameter.
10. The method for constructing a sequence deep learning model for polygenic risk scoring of natural populations according to claim 9, characterized in that: The linkage disequilibrium correlation coefficient of each natural population variant was calculated using PLINK 1.90 beta software.
11. The method for constructing a sequence deep learning model for polygenic risk scoring of natural populations according to claim 10, characterized in that: The sequence neural network is selected from a recurrent neural network, a Transform, a recursive neural network, a gated recursive unit, and a long short-term memory unit.
12. The method for constructing a sequence deep learning model for polygenic risk scoring of natural populations according to claim 10, characterized in that: The polygenic risk score for a natural population refers to scoring the risk of at least one polygenic disease among cardiovascular and cerebrovascular diseases, diabetes, malignant tumors, chronic respiratory diseases, arthritis and obesity in a natural population.
13. The method for constructing a sequence deep learning model for polygenic risk scoring of natural populations according to claim 12, characterized in that: The natural population polygenic risk score refers to: scoring the risk of chronic respiratory diseases in the natural population.
14. A method for obtaining polygenic risk scores of natural populations based on a deep learning model, characterized in that: include: The sequence deep learning model constructed based on the method for constructing a sequence deep learning model for polygenic risk scoring according to any one of claims 1 to 13 scores the SNP data to be scored.
15. A device for polygenic risk scoring of natural populations, characterized in that: include: Input module, control module and output module; The input module is configured to: input SNP sample data and / or sample genotype data; The control module includes: a genotype data preprocessing module, a genetic variation data screening module, a genetic variation window construction module, a model construction module and a data scoring module; The genotype data preprocessing module is configured to: perform file conversion on the SNP sample data and / or the sample genotype data; The genetic variation data screening module is configured to: screen genetic variation data using LASSO regression; The genetic variation window construction module is configured as follows: calculating the linkage disequilibrium correlation coefficient of each population variation for the genetic variation data after LASSO regression; establishing a genetic variation data frame of X SNPs for each genetic variation; wherein each of the genetic variation data frames includes the genotype information of the X SNPs for each target genetic variation, the distance to the target genetic variation and the LD correlation coefficient; obtaining a data set of n*X*Y, which is the genetic variation window, wherein n is the number of variations; the number of SNPs in each of the genetic variation data frames is X≤51; Y≥3; the X SNPs refer to the target variation, 0.5*(X-1) genetic variations upstream of the target variation and 0.5*(X-1) genetic variations downstream of the target variation; and the X SNPs are located within 250kb upstream and downstream of the target variation, and the correlation coefficient with the target variation is r2>0.1; The model building module is configured to: non-dimensionalize the obtained genetic variation window, and then input the non-dimensionalized data into a sequence neural network for model training until the model reaches a preset convergence condition; The data scoring module is configured to: score the SNP data to be scored based on the sequence deep learning model constructed by the model construction module; The output module is configured to output a polygenic risk score result.
16. An electronic device, characterized in that: The invention comprises a memory and a processor, and computer instructions stored in the memory and executed on the processor. When the computer instructions are executed by the processor, the steps of the method for constructing a sequence deep learning model for polygenic risk scores according to any one of claims 1 to 13 are completed, or the steps of the method for obtaining polygenic risk scores for natural populations based on a deep learning model according to claim 14 are completed.
17. A computer-readable storage medium, characterized in that: Used to store computer instructions, which, when executed by a processor, complete the steps of the method for constructing a sequence deep learning model for a polygenic risk score as described in any one of claims 1 to 13, or complete the steps of the method for obtaining a polygenic risk score for a natural population based on a deep learning model as described in claim 14.
Citation Information
Patent Citations
Multi-gene genetic risk score calculation method and system based on tissue specific regulatory network atlas
CN117409860A
Self-designed single-nucleotide polymorphism chip and method of computing polygenicrisk score for given populations using self-designed single-nucleotide polymorphism chip
US20230335218A1