Computer-based method and apparatus for analyzing genetic data

JP7885216B2Active Publication Date: 2026-07-06ゲノミクス リミテッド
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023533271
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-12-01
Filing Date
2021-11-26
Publication Date
2026-07-06
Estimated Expiration
2041-11-26

Smart Images

  • Figure 0007885216000021
    Figure 0007885216000021
  • Figure 0007885216000022
    Figure 0007885216000022
  • Figure 0007885216000023
    Figure 0007885216000023
Patent Text Reader

Abstract

A method of analyzing genetic data for an organism is disclosed that includes receiving a plurality of input units. Each input unit includes information about an association between a genetic variant in a region of the genome and a phenotype or phenotype combination. The method includes performing an iteration that includes determining, for each variant, whether the variant is causal to the phenotype or phenotype combination based on the input units. If the variant is causal to the phenotype or phenotype combination, a sampled effect size of the variant on the phenotype or phenotype combination is determined based on the input units and the information about the correlation between the variants in the region. For each variant, a predicted effect size of the variant on the phenotype or phenotype combination is determined based on an average of posterior effect sizes calculated across the iterations of the sampled effect sizes or using the sampled effect sizes.
Need to check novelty before this filing date? Find Prior Art

Claims

1. A computer-based method for analyzing genetic data about an organism, Receiving multiple input units, each input unit containing information about the relationship between multiple genetic variants in a target region of the organism's genome and one of multiple phenotypes or phenotypic combinations of the organism; For each of the aforementioned multiple genetic variants, Based on the aforementioned multiple input units, determine which of the aforementioned multiple phenotypes or phenotypic combinations the genetic variant is responsible for, and If it is determined that the genetic variant is the cause of two or more phenotypes or phenotypic combinations among the plurality of phenotypes or phenotypic combinations, then determining the sampled effect size of the genetic variant for each of the two or more phenotypes or phenotypic combinations based on the correlation of linkage disequilibrium (LD) between the plurality of input units and the plurality of genetic variants in the region of interest, the method comprising: (i) calculating the probability distribution of the effect size of the genetic variant for the two or more phenotypes or phenotypic combinations, wherein the probability distribution depends on the correlation between the effect sizes of the genetic variant for the two or more phenotypes or phenotypic combinations; and (ii) sampling the value of the effect size from the probability distribution. Performing one or more iterations including, For each genetic variant, the predicted effect size of the genetic variant for one or more of the two or more phenotypes or phenotype combinations is determined based on the average of the posterior effect sizes of the genetic variant for the input units calculated using the sampled effect sizes, or over a subset of the replicates of the sampled effect sizes of the genetic variant for the two or more phenotypes or phenotype combinations. A method that includes this.

2. Determining which of the aforementioned genetic variants is the cause of the aforementioned multiple phenotypes or phenotypic combinations is: The probability of the information from the plurality of input units, assuming that the genetic variant is not the cause of either the phenotype or the phenotypic combination, The probability of the information from the plurality of input units, assuming that the genetic variant is the cause of all of the phenotype or phenotypic combination, The probability of the information from the plurality of input units, assuming that the genetic variant is the cause of one or more subsets of the phenotype or phenotypic combination. Calculating multiple probabilities, The probabilistic determination of which of the above multiple phenotypes or phenotypic combinations the genetic variant is the cause of, based on the above multiple probabilities, The method according to claim 1, including the method described in claim 1.

3. (a) The probability of the information from the plurality of input units, assuming that the genetic variant is the cause of one or more of the phenotypes or phenotypic combinations, The proportion of the multiple genetic variants that are expected to be the cause, The plurality of input units, and Correlation between the effect size of the genetic variant and the phenotype or phenotypic combination Depends on, and (b) The probability of the information from the plurality of input units, assuming that the genetic variant is not the cause of either the phenotype or the phenotypic combination, is The proportion of the multiple genetic variants that are expected to be the cause, and The aforementioned multiple input units The method according to claim 2, which depends on either or both of the following:

4. For each of the one or more subsets of the phenotype or phenotypic combination, the probability of the information from the plurality of input units, assuming that the genetic variant is the cause of the subset of the phenotype or phenotypic combination, is: The proportion of the multiple genetic variants that are expected to be the cause, A subset of input units including the input unit which includes information about the relationship between the plurality of genetic variants and one of the subsets of phenotypes or phenotypic combinations, and Correlation between the effect size of the genetic variant and the phenotype or phenotypic combination The method according to claim 2 or 3, which depends on the method.

5. (a) The proportions of the multiple genetic variants that are expected to be the cause are predetermined, or (b) The proportion of the plurality of genetic variants that are expected to be causative is updated in each iteration, the method according to claim 3 or 4.

6. (a) The correlation between the effect size of the genetic variant and the phenotype or phenotype combination is predetermined, or The method according to any one of claims 3 to 5, wherein (b) the correlation between the effect size of the genetic variant for the phenotype is updated in each iteration.

7. The method according to any one of claims 2 to 6, wherein the input units are determined from each population, and each of the plurality of probabilities depends on one or more parameters that quantify the overlap in the populations between each pair of input units.

8. (a) The probability distribution is a multivariate normal distribution. (b) The sampling of the effect size values ​​is performed using a Monte Carlo Gibbs sampler, and (c) The method according to any one of claims 1 to 7, wherein the sampling of the effect size value in each iteration depends on the sampled effect size from one or more previous iterations.

9. (a) The correlation between the effect size of the genetic variant and the phenotype or phenotype combination is predetermined, or The method according to any one of claims 1 to 8, wherein (b) the correlation between the effect size of the genetic variant and the phenotype or phenotypic combination is updated in each iteration.

10. The method according to any one of claims 1 to 9, wherein determining the sampled effect size includes using a model of causal relationships between the plurality of phenotypes or combinations of phenotypes.

11. Each of the one or more iterations further includes subtracting a weighted effect size for each genetic variant determined to be causative from information about the association between each other genetic variant in each input unit and the phenotype or phenotypic combination, The weighted effect size is the sampled effect size of the genetic variant to the phenotype or phenotype combination of the input unit, weighted by the respective correlation coefficients between the genetic variant and each other genetic variant. The correlation coefficient is determined based on the correlation of linkage disequilibrium between the plurality of genetic variants in the region of interest. The method according to any one of claims 1 to 10.

12. The method according to any one of claims 1 to 11, wherein performing one or more iterations includes performing a predetermined number of iterations.

13. The method according to any one of claims 1 to 12, wherein each of the one or more iterations further comprises the step of evaluating a convergence parameter, and performing one or more iterations includes performing the iterations until a predetermined condition for the convergence parameter is met.

14. The method according to any one of claims 1 to 13, wherein the information relating to the association between each of the plurality of genetic variants and the phenotype or phenotypic combination includes, for each of the plurality of genetic variants, an estimate of the strength of the association between the genetic variant and the phenotype or phenotypic combination, and an error in the estimate of the strength of the association.

15. A method for determining a polygenic risk score for a target phenotype or combination of target phenotypes in a target individual, Receiving genetic information about the target region of the genome of the aforementioned target individual, To receive a predicted effect size for the target phenotype or target phenotype combination of multiple genetic variants in the target region, determined using a method for analyzing genetic data according to any one of claims 1 to 14, The polygene risk score is determined based on the genetic information and predictive effect size of the target individual. A method that includes this.

16. A device for analyzing genetic data about organisms, A receiving unit configured to receive multiple input units, each input unit containing information about the relationship between multiple genetic variants in a target region of the organism's genome and one of multiple phenotypes or phenotypic combinations of the organism, For each of the aforementioned multiple genetic variants, Based on the aforementioned multiple input units, determine which of the aforementioned multiple phenotypes or phenotypic combinations the genetic variant is responsible for, and If it is determined that the genetic variant is the cause of one or more of the phenotypes or phenotypic combinations, then determining the sampled effect size of the genetic variant for one or more phenotypes or phenotypic combinations based on the correlation of linkage disequilibrium (LD) between the plurality of input units and the plurality of genetic variants in the region of interest, comprising: (i) calculating the probability distribution of the effect size of the genetic variant for one or more phenotypes or phenotypic combinations, wherein the probability distribution depends on the correlation between the effect sizes of the genetic variant for the phenotypes or phenotypic combinations; and (ii) sampling the value of the effect size from the probability distribution. Perform one or more iterations including, For each genetic variant, the predicted effect size of the genetic variant for one or more of the phenotypes or phenotype combinations is determined based on the average of the posterior effect sizes of the genetic variant for the input units calculated using the sampled effect sizes, or over a subset of the iterations of the sampled effect sizes of the genetic variant for one or more of the phenotypes or phenotype combinations. A data processing unit configured as follows: A device equipped with the following features.

17. A computer program or computer-readable medium that includes instructions causing a computer to perform the method according to any one of claims 1 to 15 when the program is executed by the computer.