Calculation method of disease risk score and matching score, and matching system

The method calculates disease risk scores and genetic compatibility using multiple polymorphic loci data to prevent genetic diseases in offspring by optimizing mating selection.

JP7698354B1Active Publication Date: 2025-06-25SEEDNA INC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2024146892
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-08-28
Publication Date
2025-06-25
Estimated Expiration
2044-08-28

AI Technical Summary

Technical Problem

Existing genetic testing methods for animals focus on single genes and do not account for multiple alleles, disease risk quantification, or optimal mating to prevent genetic diseases in offspring.

Method used

An arithmetic method calculating disease risk scores based on genotype data from multiple polymorphic loci, considering chromosomal abnormalities, and a matching system evaluating genetic compatibility between individuals using polymorphic loci data to optimize mating.

Benefits of technology

Provides a simple score for disease risk and genetic compatibility, enabling effective prevention of genetic diseases in offspring by optimal mating selection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007698354000001_ABST
    Figure 0007698354000001_ABST
Patent Text Reader

Abstract

Provided are an arithmetic method and a matching system for calculating a disease risk based on data of genotypes of each of n polymorphic loci. 【Solution means】In a matching system 1, a disease risk score calculation means 113 of an information processing unit 11 calculates, based on data of genotypes of each of n polymorphic loci in a biological individual, a sum of values obtained by multiplying the number of disease-related alleles (effect alleles) at each polymorphic locus by the phenotype severity of the disease-related alleles at each polymorphic locus, as a disease risk score.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technique for calculating disease risk and genetic compatibility based on genetic information.

Background Art

[0002] Livestock such as cows and pigs and companion animals such as dogs and cats have been selectively bred for specific preferred traits to improve and establish breeds. As a result of such trait selection, the frequency of occurrence of specific gene mutations is higher in populations of animals with specific traits than in populations of wild animals. Therefore, there is a high possibility that individuals homozygous for the mutant allele will be born by mating between males and females having the same mutant allele. Such individuals are likely to develop or exacerbate diseases caused by the mutant allele.

[0003] Due to such problems, in recent years, the demand for genetic testing of animals has been increasing. However, generally, genetic testing for animals is performed by targeting a single gene (see, for example, Non-Patent Document 1), and attempts have not been made to simultaneously test multiple alleles and quantify the disease risk by synthesizing the results. In addition, attempts have not been made to perform genetic testing of males and females and consider the optimal matching that makes it difficult for offspring to develop genetic diseases.

Prior Art Documents

Non-Patent Documents

[0004]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] The first object of the present invention is to provide a technique for calculating a disease risk based on data of genotypes of each of n polymorphic loci. The second object is to provide a technique for scoring a combination of male and female in which a child is less likely to develop a genetic disease from a genetic background.

Means for Solving the Problems

[0006] The present invention for solving the above first or second problem is as follows.

[0007] [1] An arithmetic method including calculating a disease risk score by performing an operation essentially including the calculation according to the following formula (1) based on data of genotypes of each of n polymorphic loci in a biological individual.

Number

[0008] [2] The arithmetic method according to [1], wherein the operation essentially including the calculation according to the formula (1) is an operation represented by the following formula (2).

Number

[0009] [3] The arithmetic method according to [2], wherein c is a numerically determined value (DR value) determined in advance according to whether the disease-related allele at each polymorphic locus is dominant or recessive.

[0010] [4] wherein a is a numerical value (CAV value) predetermined according to the presence or absence of one or more chromosomal abnormalities selected from chromosomal number abnormalities, copy number variations in specific DNA regions, and chromosomal structural abnormalities, or a numerical value obtained by adding or multiplying a numerical value (CAV value) predetermined according to the presence or absence of one or more chromosomal abnormalities selected from chromosomal number abnormalities, copy number variations in specific DNA regions, and chromosomal structural abnormalities, and a numerical value indicating the severity of the phenotype of the chromosomal abnormality, The calculation method according to [2] or [3].

[0011] [5] The calculation method according to any one of [1] to [4], including ranking the disease risk scores of specific individuals of the organisms in the population of the organisms.

[0012] [6] Based on the genotype data of each of n polymorphic loci in each of two different organisms, performing an operation essentially including calculation by a mathematical formula selected from the following formulas (7), (18), and (19), and calculating the genetic compatibility between the organisms as a matching score. [Number] Formula (7) [Number] Formula (18) Formula (7) + Formula (18) ··· Formula (19) In Formula (7), ME k is the number of disease-related alleles at each polymorphic locus in the male or a numerical value proportional to the number when the two organisms are male and female. FE k is the number of disease-related alleles at each polymorphic locus in the female or a numerical value proportional to the number when the two organisms are male and female. S kis a numerical value indicating the severity of the phenotype of the disease-related allele at each polymorphic locus. However, the n polymorphic loci targeted by formula (7) are polymorphic loci in which at least one of the alleles that appear is a disease-related allele. In formula (18), R l is a numerical value determined according to the genotype similarity at each polymorphic locus of each of the two biological individuals. However, the m polymorphic loci targeted by formula (18) are polymorphic loci in which all of the alleles that appear are constitution / talent-related alleles and no disease-related alleles appear.

[0013] [7] The operation essentially including the calculation by the formula (7) is the operation expressed by the following formula (8), the operation method according to [6].

Number

[0014] [8] The operation method according to [7], wherein c is a numerical value (DR value) predetermined according to whether the disease-related allele at each polymorphic locus is dominant or recessive.

[0015] [9] The a is a numerical value (CAV value) predetermined by the presence or absence of one or more chromosomal abnormalities selected from chromosomal number abnormalities, copy number variations of specific DNA regions, and chromosomal structure abnormalities in either or both of the male and the female, or a numerical value obtained by adding or multiplying a numerical value (CAV value) predetermined by the presence or absence of one or more chromosomal abnormalities selected from chromosomal number abnormalities, copy number variations of specific DNA regions, and chromosomal structure abnormalities in either or both of the male and the female, and a numerical value indicating the severity of the phenotype of the chromosomal abnormality. The calculation method described in [8].

[0016]

[10] The calculation method according to any one of [6] to [9], including performing ranking of matching scores between specific biological individuals in a set of matching scores between any biological individuals in the population of the organisms.

[0017]

[11] A matching system for evaluating genetic compatibility, comprising a storage unit and an information processing unit, wherein the storage unit stores a database that is a set of genetic data regarding a plurality of biological individuals, and the information processing unit, calculates the genetic compatibility between two different biological individuals as a matching score based on the genetic data of the two biological individuals, or calculates the genetic compatibility between a specific biological individual and each of the plurality of biological individuals as a matching score based on the genetic data of the specific biological individual and the genetic data of the plurality of biological individuals stored in the storage unit, comprises a matching score calculation means, wherein the genetic data is data of genotypes of each of n polymorphic loci in a biological individual, and / or data regarding a disease risk score calculated from the data of genotypes of each of n polymorphic loci in a biological individual The matching system is as described above.

[0018]

[12] The information processing unit of the matching system according to

[11] comprises an extraction means for extracting a biological individual that shows a matching score satisfying a certain criterion in combination with the specific biological individual from among the plurality of biological individuals registered in the database.

[0019]

[13] The disease risk score is calculated by the calculation method according to any one of [1] to [4], The matching score is calculated by the calculation method described in any one of [6] to [9], and the matching system described in

[11] or

[12] .

Advantages of the Invention

[0020] According to the present invention, the disease risk of a biological individual or the genetic compatibility of male and female can be presented as a simple score.

Brief Description of the Drawings

[0021]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Modes for Carrying Out the Invention

[0022] Hereinafter, the present invention will be described in detail by way of examples. In the embodiments of the present invention, A (numerical value) to B (numerical value) means A or more and B or less. Also, the preferred embodiments, more preferred embodiments, etc. exemplified below can be used in appropriate combinations with each other regardless of expressions such as "for example", "one", "preferred", and "more preferred". In addition, the description of the numerical range is illustrative, and ranges appropriately combined with the upper and lower limits of each range and the numerical values of the examples can also be preferably used (for example, when described as A to B and C to D, combinations of A to D or C to B can be used). Furthermore, terms such as "containing" or "including" may be read as "consisting essentially of" or "consisting only of".

[0023] [Calculation of disease risk score] The present invention relates to an arithmetic method for calculating a disease risk score. Here, the "disease risk score" as used in this specification does not evaluate the disease risk based on the presence or absence of detection of a single disease-causing gene mutation. In the present invention, the "disease risk score" refers to a numerical value calculated by comprehensively considering the presence or absence of one or more chromosomal abnormalities selected from mutations in the responsible genes of single-gene diseases, chromosomal number abnormalities, copy number variations in specific DNA regions, and chromosomal structural abnormalities, and the genotype data of multiple polymorphic loci.

[0024] The present invention includes performing an operation that essentially includes calculation according to the following formula (1) based on the genotype data of each of n polymorphic loci in a biological individual. The genotype data may be obtained by any analysis method. Specifically, examples include nucleotide sequence analysis, mass spectrometry, digital PCR, SNV microarray, real-time PCR, etc. As a specific means of nucleotide sequence analysis, next-generation sequencer (NGS) can be mentioned. Next-generation sequencer is a sequencing method that enables parallel sequencing of a large number of cloned molecules and single nucleic acid molecules. In the present invention, any NGS system may be adopted. For example pyrosequencing (such as GS Junior (Roche)), sequencing by synthesis using reversible terminator dyes (such as MiSeq (Illumina)), sequencing by ligation (such as SeqStudio Genetic Analyzer (Thermo Fisher SCIENTIFIC)), ion semiconductor sequencing (such as Ion Proton System (Thermo Fisher SCIENTIFIC)), Sequencing by a CMOS (Complementary Metal Oxide Semiconductor) chip (such as the iSeq 100 System (Illumina)), sequencing by probe-anchor synthesis (cPAS) and improved DNA nanoball technology (DNBSEQ (BGI)), nanopore sequencing (PromethION (Oxford Nanopore Technologies)), and the like can be mentioned.

[0025] The organism to be evaluated is not particularly limited and may be a plant or an animal. Examples of plants include rice, wheat, barley, soybean, corn, etc., and examples of animals include humans, dogs, cats, cows, horses, chickens, mice, rats, sheep, etc. without limitation.

[0026] The disease risk score refers to a numerical value obtained by performing an operation essentially including the calculation according to the following formula (1) based on the genotype data of each of the n polymorphic loci in an individual organism.

[0027]

Number

[0028] In formula (1), n represents an arbitrary integer of 1 or 2 or more. In formula (1), E k is the number of disease-related alleles at each polymorphic locus or a numerical value proportional to the number. However, the n polymorphic loci targeted by formula (1) are polymorphic loci in which at least one of the alleles that appear is a disease-related allele.

[0029] Here, disease-related alleles widely include alleles that have a negative or positive impact on the onset of the disease. Examples of alleles that have a negative impact include alleles having a direct causal mutation of the disease, mutant alleles that lead to an increased risk of disease onset, mutant alleles that lead to disease aggravation, and mutant alleles that cause disease onset, increased risk of disease onset, and aggravation in combination with other factors. Examples of alleles that have a positive impact include alleles that confer resistance to diseases, such as alleles that result in a constitution that is less likely to contract or progress a disease.

[0030] Regarding the polymorphic loci discovered so far, they are linked to related diseases, allele frequencies, mapping information, etc. and are recorded in various databases. By searching these databases, polymorphic loci that are known to be related to diseases can be easily grasped. For example, the following can be exemplified as databases regarding single nucleotide polymorphic loci. dbSNP: https: / / www.ncbi.nlm.nih.gov / snp / SNPedia: https: / / www.snpedia.com / GWAS catalog: https: / / www.ebi.ac.uk / gwas /

[0031] In formula (1), S k is a numerical value indicating the severity of the phenotype. The severity of the phenotype is a numerical value assigned to each polymorphic locus, and indicates the magnitude of the impact brought about when a disease-related allele appears at that polymorphic locus. In particular, when the disease-related allele is an allele that has a negative impact, the severity of the phenotype indicates the severity of the disease caused by that disease-related allele.

[0032] The severity of the phenotype may be an evaluation value based on a two-stage or more stage evaluation. For example, as an embodiment of assigning a seven-stage evaluation according to the following criteria according to the magnitude of the impact brought about by the appearance of a disease-related allele. Note that the seven-stage evaluation described below is merely an example. The number of evaluation stages and evaluation criteria can be set as appropriate. -5 Causes embryonic lethality or dies before reaching adulthood. -4 Causes severe diseases or disorders and causes difficulties in daily life. -3 Causes a disease or disorder that requires continuous treatment. -2 Causes a mild disease or disorder that can be prevented from developing by preventive medicine. -1 Affects the constitution to a certain extent. +1 Resistance to disease

[0033] As described above, the severity of the phenotype can be set as appropriate. However, it is preferable to set it so that the absolute value of the severity of the phenotype increases as the impact when disease-related alleles appear becomes greater.

[0034] The severity of the phenotype may be assigned negative numerical values for negative impacts and positive numerical values for positive impacts, as in the above-described reference example. In this case, the larger the numerical value of the disease risk score calculated by formula (1), the better the score can be evaluated.

[0035] Alternatively, the severity of the phenotype may be assigned positive numerical values for negative impacts and negative numerical values for positive impacts, contrary to the above-described reference example. In this case, the smaller the numerical value of the disease risk score calculated by formula (1), the better the score can be evaluated.

[0036] It is also possible to adopt an embodiment in which the severity of the phenotype at each of the specific n polymorphic loci is defined in advance. The severity of the phenotype to be assigned to each polymorphic locus can be set as appropriate based on existing research reports, clinical records, etc. regarding specific disease-related alleles that appear at each polymorphic locus.

[0037] It is preferable to adopt an embodiment in which a table recording the severity of the phenotype linked to each of the specific n polymorphic loci is prepared in advance. The table may be stored in a database as electronic data and referred to during the calculation by formula (1).

[0038] In addition, the estimated value of the impact of gene mutations by a machine learning model may be used as the severity of the phenotype. Such a machine learning model is trained using a machine learning algorithm. By using data such as existing data and research results on disease-related alleles and their impacts, the frequency of occurrence of disease-related alleles, and the high evolutionary conservation of bases related to mutations, the model learns patterns and correlations and acquires the ability to predict the clinical impact of disease-related alleles. Since machine learning models for inferring the impact of gene mutations on gene functions are widely used in the fields of genetics and medicine, these existing models may be utilized.

[0039] Existing tools for inferring the impact of gene mutations on the functions of gene products include PolyPhen-2 (Polymorphism Phenotyping v2), SIFT (Sorting Intolerant From Tolerant), CADD (Combined Annotation Dependent Depletion), GERP (Genomic Evolutionary Rate Profiling), DeepVariant, and others. The severity of the phenotype of each polymorphic locus can be set using these tools.

[0040] Note that the summation operation in formula (1) is E1×S1+E2×S2+···+E k ×S k +···+E n-1 ×S n-1 +E n ×S n can be expressed as. Each term constituting this summation operation is an exponent corresponding to each polymorphic locus. In this specification, each term constituting this summation operation is referred to as a disease risk index. That is, the disease risk score can be expressed as the sum of the disease risk indices at each polymorphic locus.

[0041] The "operation essentially including the calculation according to formula (1)" includes, in addition to the summation operation itself represented by formula (1), operations such as adding an arbitrary constant or variable to the summation of formula (1), multiplying the summation of formula (1) by an arbitrary constant or variable, multiplying each term of formula (1) by an arbitrary constant or variable, etc., and all operations including formula (1) in part of the formula notation are included.

[0042] The "operation essentially including the calculation according to formula (1)" may be an operation expressed by the following formula (2).

Number

[0043] In formula (2), a may be an arbitrary constant (including 0) or an arbitrary variable. b may be an arbitrary constant (excluding 0) or an arbitrary variable. c may be an arbitrary constant (including 0) or an arbitrary variable.

[0044] When c is a variable corresponding to a polymorphic locus, formula (2) can be expressed by the following formula (3).

Number

[0045] As the variable c in formula (3), a numerical value (DR value) determined in advance according to whether the disease-related allele at each polymorphic locus is dominant or recessive can be adopted. In this case, formula (3) can be expressed by the following formula (4).

Number

[0046] The method of assigning the DR value is not particularly limited, but it is preferable to set the DR value assigned when the disease-related allele is dominant to a value larger than the DR value assigned when the disease-related allele is recessive. More specifically, when the disease-related allele is dominant, the DR value assigned is preferably set to twice the DR value assigned when the disease-related allele is recessive. For example, the DR value assigned when the disease-related allele is dominant can be set to 4, and the DR value assigned when the disease-related allele is recessive can be set to 2.

[0047] In addition, an operation may be performed to reflect the presence or absence of one or more chromosomal abnormalities selected from the number of chromosomes, copy number variations in specific DNA regions, and chromosomal structural abnormalities in the disease risk score (hereinafter, chromosomal number abnormalities, copy number variations in specific DNA regions, and chromosomal structural abnormalities may be collectively referred to as "chromosomal abnormalities"). Specifically, as "a" in any of formulas (2) to (4), a variable reflecting the presence or absence of chromosomal abnormalities may be adopted. As "a", any of the following numerical values can be selected. (i) A numerical value predetermined according to the presence or absence of chromosomal abnormalities (ii) A numerical value obtained by adding or multiplying a numerical value predetermined according to the presence or absence of chromosomal abnormalities and a numerical value indicating the severity of the phenotype of the chromosomal abnormality

[0048] In this specification, the "numerical value predetermined according to the presence or absence of chromosomal abnormalities" may be referred to as "Chromosome Abnormality Value (CAV)".

[0049] Generally, serious disorders and diseases often occur in individuals with chromosomal abnormalities. Therefore, regardless of the content of the second term (the sum of E k ×S k ×c), it is preferable to set the CAV so that the disease risk can be immediately evaluated as high when there is a chromosomal abnormality. That is, the CAV when there is a chromosomal abnormality is preferably set to a numerical value with a significantly larger absolute value than the numerical value assumed to be indicated by the second term (the sum of Ek × Sk × c).

[0050] On the other hand, when there is no chromosomal abnormality, the second term (E k ×S kIt is preferable to set the CAV so that the disease risk can be evaluated only by the sum of (×c). That is, when there is no chromosomal abnormality, the value of CAV is set to a value significantly smaller than the numerical value assumed to be indicated by the sum of (E k ×S k ×c) (for example, it may be 0).

[0051] The CAV may be set according to the type of chromosome. When the organism to be evaluated has m types of chromosomes, the CAV is set for each of the chromosomes from chromosome 1 to chromosome m, and "a" can be a variable represented by the following formula (5).

Number

[0052] Also, when adopting the above-mentioned (ii), "a" can be a numerical value represented by the following formula (6).

Number

[0053] In formula (6), CAV k is a numerical value preset according to the presence or absence of chromosomal abnormality of the k-th chromosome, and CS k is a numerical value indicating the severity of the phenotype of chromosomal abnormality in the k-th chromosome.

[0054] When the influence of chromosomal abnormality of the k-th chromosome is small, the value of CS k is set small, and when the influence of chromosomal abnormality of the k-th chromosome is serious (for example, embryonic lethal), the value of CS k can be set large.

[0055] As "a" in any one of formulas (2) to (4), an index reflecting the presence or absence of chromosomal abnormalities may be reflected. Examples of such an index include known indices such as Z-score or NCV (Normalized Chromosome Value). For the calculation methods of these indices, reference can be made to, for example, Japanese Patent No. 7331325.

[0056] In addition, as an embodiment, the genotyping information and the information regarding the presence or absence of chromosomal abnormalities obtained in the process of implementing an analysis technique for performing two or more types of genetic analyses among phenotype analysis, chromosomal abnormality analysis, and parentage testing disclosed in Japanese Patent No. 7331325 may be diverted to the calculation of a disease risk score.

[0057] As an embodiment, after calculating the disease risk score of a specific biological individual, ranking thereof may be performed. Specifically, ranking of the disease risk score of a specific biological individual in the population of the organism is executed.

[0058] The specific embodiment of ranking is not particularly limited. For example, as an embodiment, one or two or more reference values may be set in advance for the disease risk score, and ranking may be performed by comparing the disease risk score of the biological individual to be evaluated with the reference value. For example, it can be an embodiment where if it is equal to or higher than a certain reference value, it is ranked as rank A, and if it is less than a certain reference value and equal to or higher than a certain reference value, it is evaluated as rank B.

[0059] In addition, the disease risk score is calculated in advance for a population composed of a plurality of biological individuals (preferably having a statistically significant sample size), and one or two or more score thresholds for belonging to an arbitrary upper ratio are set. The ranking of the disease risk score of a specific organism to be evaluated in the population of biological individuals can be performed according to whether the disease risk score calculated for the specific organism to be evaluated is equal to or higher than any threshold value and / or less than a specific threshold value.

[0060] [Calculation of matching score] The present invention also relates to a calculation method including calculating the genetic compatibility between two biological individuals as a matching score.

[0061] As used herein, the "matching score" includes a score reflecting the magnitude of the disease risk of offspring produced by mating between a male biological individual and a female biological individual, a score reflecting the genetic similarity between individuals of the opposite sex or the same sex, and a score integrating these. These will be described in order below.

[0062] <Calculation of a score reflecting the magnitude of the disease risk of offspring produced by mating between a male biological individual and a female biological individual> First, an embodiment for calculating a score reflecting the magnitude of the disease risk of offspring produced by mating between a male biological individual and a female biological individual will be described. Specifically, this embodiment essentially includes performing an operation based on the genotype data of each of n polymorphic loci in each of the male biological individual and the female biological individual, and performing a calculation according to the following formula (7), and calculating the genetic compatibility between the male and the female as a matching score.

[0063]

Equation

[0064] In addition, regarding the organism to be evaluated, genotype data, polymorphic loci, disease-related alleles, and the severity of the phenotype in the method for calculating the matching score, the content described in the items for calculating the above-mentioned disease risk score is directly applicable.

[0065] In formula (7), ME k is the number of disease-related alleles at each polymorphic locus in the male or a numerical value proportional to the number. FE k is the number of disease-related alleles at each polymorphic locus in the female or a numerical value proportional to the number. S kis a numerical value indicating the severity of the phenotype of the disease-related allele at each polymorphic locus, and its specific embodiments are as described above. The n polymorphic loci targeted by formula (7) are polymorphic loci in which at least one of the alleles that appear is a disease-related allele.

[0066] "Operations essentially including the calculation by formula (7)" include, in addition to the sum operation itself represented by formula (7), operations of adding any constant or variable to the sum of formula (7), operations of multiplying the sum of formula (7) by any constant or variable, operations of multiplying each term of formula (7) by any constant or variable, and all operations including formula (7) in a part of the formula notation.

[0067] "Operations essentially including the calculation by formula (7)" may be operations expressed by the following formula (8).

Number

[0068] The a to c in formula (8) are as appropriate as the content described in the items for calculating the above-mentioned disease risk score.

[0069] When c is a variable corresponding to the polymorphic locus, formula (8) can be expressed by the following formula (9).

Number

[0070] The variable c of formula (9) k As, for the disease-related allele at each polymorphic locus, a numerically determined value (DR value) can be adopted according to whether it is dominant or recessive. In this case, formula (9) can be expressed by the following formula (10).

Number

[0071] DR in formula (10) k is as it is appropriate for the content described in the items for calculating the disease risk score described above.

[0072] Also, when the disease risk scores have been calculated for each of male and female, the following formula (11) is also included in the "operation essentially including the calculation by formula (7)".

Number

[0073] MDI in formula (11) k is a disease risk index corresponding to each polymorphic locus in male, and ME k ×S k is a numerical value calculated by the mathematical formula of ×c. FDI in formula (11) k is a disease risk index corresponding to each polymorphic locus in female, and FE k ×S k is a numerical value calculated by the mathematical formula of ×c. MDI k and FDI k in the multiplication of, since (S k ×c) 2 is included, it is divided by S k ×c. When formula (11) is expanded, it becomes the same as formula (8).

[0074] Also, an operation for reflecting the presence or absence of chromosomal abnormality in the matching score may be performed. Specifically, as "a" in any of formulas (8) to (11), a variable reflecting the presence or absence of chromosomal abnormality may be adopted. As "a", any of the following numerical values can be selected. (iii) A numerical value (CAV value) predetermined by the presence or absence of chromosomal abnormality in either or both of the male and the female (iv) A numerical value obtained by adding or multiplying a numerical value (CAV value) predetermined by the presence or absence of chromosomal abnormality in either or both of the male and the female and a numerical value indicating the severity of the phenotype of the chromosomal abnormality

[0075] Generally, when there is a chromosomal abnormality in either or both of the male and female individuals, the offspring born from the mating of such individuals often develop severe disorders or diseases. Therefore, regardless of the content of the second term (the sum of ME k ×FE k ×S k ×c) in formulas (8) to (11), it is preferable to set the CAV so that when there is a chromosomal abnormality in either or both of the male and female, the matching score can be immediately evaluated as being small. That is, the CAV when there is a chromosomal abnormality in either or both of the male and female is preferably set to a value that is significantly larger in absolute value than the numerical value assumed to be indicated by the second term (the sum of ME k ×FE k ×S k ×c).

[0076] On the other hand, when there is no chromosomal abnormality in either the male or the female, it is preferable to set the CAV so that the matching score can be evaluated only by the second term (the sum of ME k ×FE k ×S k ×c). That is, the value of the CAV when there is no chromosomal abnormality in either the male or the female is preferably set to a value that is significantly smaller (for example, it may be 0) than the numerical value assumed to be indicated by the second term (the sum of ME k ×FE k ×S k ×c).

[0077] In the case of (iii) above, "a" in formulas (8) to (11) may be in a form where the CAVs of the male and female are evaluated separately. That is, the value of "a" may be a variable represented by the formula MCAV + FCAV or MCAV × FCAV Here, "MCAV" is the CAV preset according to the presence or absence of chromosomal abnormality in the male, and "FCAV" is the CAV preset according to the presence or absence of chromosomal abnormality in the female.

[0078] ​CAV may be set according to the type of chromosome. When the organism to be evaluated has m types of chromosomes, CAV can be set for each of the chromosomes from chromosome 1 to chromosome m, and "a" can be a variable represented by the following formula (12), (13), or (14).

Number

[0079] In formula (12), CAV k is a numerical value preset according to the presence or absence of chromosomal abnormalities on the k-th chromosome in either or both males and females. When there are no chromosomal abnormalities in either male or female, when there are chromosomal abnormalities in either male or female, and when there are chromosomal abnormalities in both males and females, the value of CAV k may be assigned in advance step by step.

[0080]

Number

[0081]

Number

[0082] In formula (13) or (14), MCAV k is a numerical value preset according to the presence or absence of chromosomal abnormalities on the k-th chromosome in males, and FCAV k is a numerical value preset according to the presence or absence of chromosomal abnormalities on the k-th chromosome in females.

[0083] Note that in formula (14), when there are no chromosomal abnormalities on the k-th chromosome, it is preferable not to adopt "0" as the numerical value pre-assigned to MCAV k and FCAV k This is because when there are no chromosomal abnormalities in either male or female, but there are chromosomal abnormalities in the other, the value of formula (14) will become "0".

[0084] Also, when adopting the above (iv), "a" can be a numerical value represented by the following formula (15), formula (16), or formula (17).

Number

Number

Number

[0085] In formula (15), CAV k is a numerical value preset according to the presence or absence of chromosomal abnormalities on the k-th chromosome. In formula (16) and formula (17), MCAVk is a numerical value preset according to the presence or absence of chromosomal abnormalities on the k-th chromosome in males, and FCAVk is a numerical value preset according to the presence or absence of chromosomal abnormalities on the k-th chromosome in females. And in formulas (15) to (17), CS k is a numerical value indicating the severity of the phenotype of chromosomal abnormalities on the k-th chromosome.

[0086] When the influence of chromosomal abnormalities on the k-th chromosome is small, the value of CS k is set small, and when the influence of chromosomal abnormalities on the k-th chromosome is significant (for example, embryonic lethal), the value of CS k can be set large.

[0087] In addition, in formula (17), when there are no chromosomal abnormalities on the k-th chromosome, it is preferable not to adopt "0" as the numerical value pre-assigned to MCAV k and FCAV k This is because when there are no chromosomal abnormalities in either males or females, but there are chromosomal abnormalities in the other, the value of formula (17) will become "0".

[0088] As "a" in any of formulas (8) to (11), it may reflect an index reflecting the presence or absence of chromosomal abnormalities. Examples of such indices include known indices such as Z-score or NCV (Normalized Chromosome Value). For the calculation methods of these indices, reference can be made to, for example, Japanese Patent No. 7331325.

[0089] In addition, as an embodiment, it is also possible to use, for calculating the matching score, the genotyping information and the information regarding the presence or absence of chromosomal abnormalities obtained during the implementation process of an analysis technique for performing two or more types of genetic analyses among phenotype analysis, chromosomal abnormality analysis, and paternity testing, which are disclosed in Japanese Patent No. 7331325.

[0090] <Score reflecting the genetic similarity between opposite sexes or between the same sexes> This embodiment essentially includes an operation of performing a calculation according to the following formula (18) based on the genotype data of each of n polymorphic loci in each of two different biological individuals, and calculating the genetic compatibility between the biological individuals as a matching score.

[0091]

Equation

[0092] In formula (18), R l is a numerical value determined according to the similarity of genotypes at each polymorphic locus of the two biological individuals. However, the m polymorphic loci targeted by formula (18) are polymorphic loci where all the alleles that appear are constitution / talent-related alleles and no disease-related alleles appear.

[0093] Here, the "constitution / talent-related alleles" broadly include alleles related to the constitution and talents of an individual organism that are not related to diseases. For example, alleles that affect appearance (such as skin color, hair color, eye color, body length, body shape, etc.), athletic ability (such as muscle strength, endurance, explosive power, etc.), learning ability, memory, resistance to mental stress, and talents related to music and art are included in the "constitution / talent-related alleles".

[0094] As can be seen from the fact that ability tests, talent tests, etc. are provided in DTC (Direct to Consumer) genetic testing services, a large number of constitution / talent-related alleles have been reported in academic papers and the like. Also, in databases such as the GWAS catalog (https: / / www.ebi.ac.uk / gwas / ), a large number of constitution / talent-related alleles as defined in the present invention are also registered. It is possible to preset m specific polymorphic loci known to have these known constitution / talent-related alleles as the object of formula (18).

[0095] The "operation essentially including the calculation by formula (18)" includes, in addition to the summation operation itself represented by formula (18), operations such as adding an arbitrary constant or variable to the sum of formula (18), multiplying the sum of formula (18) by an arbitrary constant or variable, and multiplying each term of formula (18) by an arbitrary constant or variable, and all operations including formula (18) in part of the formula notation are included.

[0096] R in formula (18) l can be assigned according to a predetermined standard according to the similarity of the genotypes of the same polymorphic locus in two individual organisms. For example, suppose there are two types of constitution / talent-related alleles A and B that appear at a certain polymorphic locus. In this case, scores can be assigned as follows according to the combination of the genotype in one individual organism and the genotype in the other individual organism.

Table 1

[0097] The scores in Table 1 are merely examples. As shown in Table 1, it may be in a form where a higher score is set as the similarity of genotypes is higher. In this case, it can be evaluated that the higher the matching score calculated by Expression (18), the better the genetic compatibility of the two individuals. Alternatively, it may be in a form where a lower score is set as the similarity of genotypes is higher. In this case, it can be evaluated that the lower the matching score calculated by Expression (18), the better the genetic compatibility of the two individuals.

[0098] <Integrated Score> This embodiment essentially includes an operation that performs a calculation using a mathematical formula selected from the following Expression (19) based on the genotype data of each of n polymorphic loci in each of two different biological individuals, and calculates the genetic compatibility between the biological individuals as a matching score. Expression (7) + Expression (18) ··· Expression (19)

[0099] The above Expression (19) is represented by the sum of Expression (7) and Expression (18). The specific contents of Expression (7) and Expression (18) are as described above.

[0100] In one embodiment, essentially includes an operation that performs a calculation using a mathematical formula selected from the following Expression (20) based on the genotype data of each of n polymorphic loci in each of two different biological individuals, and calculates the genetic compatibility between the biological individuals as a matching score. Expression (8) + Expression (18) ··· Expression (20)

[0101] The above Expression (20) is represented by the sum of Expression (8) and Expression (18). The specific contents of Expression (8) and Expression (18) are as described above.

[0102] <Ranking of Matching Score> After calculating the matching scores between two biological individuals, ranking may be performed. Specifically, ranking of the matching scores in a combination of specific biological individuals in the set of matching scores between any two biological individuals is executed.

[0103] The specific embodiments of ranking are not particularly limited. For example, one or more reference values may be preset for the matching scores, and ranking may be performed by comparing the matching scores of the two biological individuals to be evaluated with the reference values. For example, an embodiment may be adopted in which if it is equal to or higher than a certain reference value, it is ranked as rank A, and if it is less than a certain reference value and equal to or higher than a certain reference value, it is ranked as rank B.

[0104] Regarding ranking in the case of calculating a score reflecting the magnitude of the disease risk of offspring produced by mating between a male biological individual and a female biological individual, the following two different embodiments will be described.

[0105] In the first embodiment regarding ranking of the matching scores, the matching scores are calculated in a one-to-one manner between a plurality of male populations and a plurality of female populations in advance. The number of biological individuals constituting the population is preferably a statistically significant sample number. One or more score thresholds for belonging to any upper ratio in the set of matching scores thus calculated are set. The ranking of the matching scores of the male-female combination to be evaluated in the said set can be performed according to whether the matching score calculated for the specific male and female to be evaluated is equal to or higher than any threshold value and / or less than a specific threshold value.

[0106] In a second embodiment regarding the ranking of matching scores, among specific male-female combinations to be evaluated, a biological individual of either sex is selected. Then, the matching score is calculated between the selected specific biological individual of that sex and each of a plurality of biological individuals of the opposite sex. A score threshold for belonging to an arbitrary upper percentage in the set of matching scores thus calculated is set to 1 or 2 or more. Based on whether the matching score calculated for a specific male and female to be evaluated is equal to or greater than any threshold and / or less than a specific threshold, the matching score of the male-female combination to be evaluated in the set can be ranked.

[0107] [Matching System] The present invention also relates to a matching system for evaluating genetic compatibility. The matching system of the present invention includes a storage unit and an information processing unit. FIG. 1 is a diagram showing an example of the hardware configuration of the matching system 1. As shown in FIG. 1, the matching system 1 includes an information processing unit 11, a storage unit 12, an input unit 13, an output unit 14, a peripheral interface 15, and a network interface 16.

[0108] The information processing unit 11 includes one or more processors such as a CPU or an MPU, and controls the overall operation processing of the matching system 1 by executing an OS and other applications.

[0109] The storage unit 12 is a storage device such as an HDD or an optical disk, or a semiconductor memory element such as an SSD, a ROM, or a RAM, and stores a program for causing a computer to execute the arithmetic method according to the present invention, and various data such as data and software used when the information processing unit 11 executes processing based on the program.

[0110] The input unit 13 is an input device such as a keyboard, a mouse, a touch panel, an OCR (Optical Character Reader), or a microphone, and inputs an operation request by the user to the information processing unit 11.

[0111] The output unit 14 is an output device such as a display, a projector, or a speaker, and displays the result of the processing by the information processing unit 11 and the like.

[0112] The peripheral interface 15 has hardware, firmware, and / or software that is required to communicate with peripheral devices including media devices such as magnetic disks or optical disk devices, other processing devices, or other input sources, which are used in relation to the present invention.

[0113] The network interface 16 includes known hardware, firmware, and / or software that enables the information processing unit 11 to communicate with other devices via a wired network or a wireless network, a local network or a wide area network, a private network or a public network. For example, such networks include known networks such as the World Wide Web, the Internet, or an enterprise internal network.

[0114] As shown in FIG. 2, the matching system 1 includes, as functional components, a genetic analysis means 112, a disease risk score calculation means 113, a matching score calculation means 114, an extraction means 115, and a ranking means 116. These are specifically realized by hardware such as the information processing unit 11 and the storage unit 12 through information processing by software stored in the storage unit 12.

[0115] The genetic analysis means 112 is a means for analyzing the raw data output by various analysis devices and analyzing the genotype of the evaluation target. Its specific mode is not particularly limited and can be realized by known analysis software. The matching system 1 is configured to be able to receive the raw data output by various analysis devices via the peripheral interface 15 or the network interface 16.

[0116] In this embodiment, the genotype data analyzed by the genetic analysis means 112 is stored in the database stored in the storage unit 12.

[0117] The disease risk score calculation means 113 is a means for calculating a disease risk score based on the genotype data of the subject to be evaluated. The specific mode of calculating the disease risk score is as described above.

[0118] The genotype data to be the target of the calculation by the disease risk score calculation means 113 may be the one generated by the genetic analysis means 112, or may be the one taken in from the outside via the peripheral interface 15 or the network interface 16.

[0119] In this embodiment, the disease risk score calculated by the disease risk score calculation means 113 is stored in the database stored in the storage unit 12.

[0120] The storage unit 12 stores a database that is a collection of genetic data on a plurality of biological individuals. The genetic data referred to here may be either or both of the following (a) and (b). (a) Data on the genotype of each of the n polymorphic loci in a biological individual (b) Data related to the disease risk score calculated from the data on the genotype of each of the n polymorphic loci in a biological individual

[0121] Regarding the specific mode of the genotype data referred to here, the description regarding the invention of the above calculation method may be applied as it is. Also, regarding the specific mode of the disease risk score, the description regarding the invention of the above calculation method can be applied as it is, but any evaluation value that qualitatively or quantitatively represents the disease risk calculated or inferred based on the genetic data may be adopted as the disease risk score.

[0122] The matching score calculation means 114 is a means for calculating the genetic compatibility as a matching score. The matching score calculation means 114 can execute an arithmetic process for calculating the genetic compatibility between a specific biological individual and each of a plurality of biological individuals as a matching score based on the genetic data of the specific biological individual and the genetic data of the plurality of biological individuals stored in the storage unit 12.

[0123] Regarding the specific mode of calculating the matching score by the matching score calculation means 114, the content described in the above item [Calculation of Matching Score] can be applied. As described in the same item, the matching score calculated by the matching score calculation means 114 includes a score reflecting the magnitude of the disease risk of the offspring generated by the mating of a male biological individual and a female biological individual, a score reflecting the genetic similarity between opposite sexes or same sexes, and a score integrating these.

[0124] Hereinafter, an embodiment in which the matching score calculation means 114 executes an arithmetic process for calculating the genetic compatibility between a male and a female biological individual as a matching score based on the genetic data of a specific male biological individual and the genetic data of a specific female biological individual will be specifically described.

[0125] An example is shown in FIG. 2 and will be described in more detail. In the example of FIG. 2, the specific biological individual is male, and the genetic data of x female biological individuals are stored in the database. In this embodiment, a matching score is calculated for each combination of the specific biological individual (male) and the x biological individuals (female) registered in the database. The method of calculating the matching score is not particularly limited, but preferably it is calculated by executing an operation essentially including the above formula (7).

[0126] Also, another example is shown in FIG. 3 and will be described in further detail. In the example of FIG. 3, a specific biological individual is male, and the database stores data (specifically, disease risk indices) regarding the disease risk scores of x female biological individuals. As described above, the disease risk index is a numerical value corresponding to each term of the summation operation essentially including formula (1). In this embodiment, a matching score is calculated for each combination of the specific biological individual (male) and each of the x biological individuals (female) registered in the database. The method for calculating the matching score is not particularly limited, but preferably it is calculated by executing the operation of formula (11) described above.

[0127] The extraction means 115 is means for extracting, from among the plurality of biological individuals registered in the database, a biological individual (provided that it is an individual of the opposite sex to the specific biological individual) that shows a matching score satisfying a certain criterion in combination with the specific biological individual.

[0128] Explaining by analogy with the examples of FIGS. 2 and 3, the extraction means 115 extracts a combination showing a matching score satisfying a certain criterion from among the matching scores calculated between a specific biological individual (male) and each of the x biological individuals (female) registered in the database. That is, a biological individual (female) whose genetic compatibility with the specific biological individual (male) satisfies a certain criterion is extracted.

[0129] The specific aspect of the “certain criterion” mentioned here is not particularly limited, and any criterion can be determined in advance. For example, it may be set to extract those showing a matching score equal to or higher than an arbitrary threshold value, or it may be set to extract those that fall within an arbitrary rank when sorted in descending order of the matching score.

[0130] In the present embodiment, the information processing unit 11 includes a ranking unit 116. The ranking unit 116 is a unit that executes ranking of the disease risk score calculated by the disease risk score calculation unit 113 or the matching score calculated by the matching score calculation unit 114. Regarding the specific mode of ranking, the matters described in the item of [Calculation of disease risk score] or [Calculation of matching score] can be applied.

[0131] The system configuration illustrated in FIG. 1 will be described with reference to FIG. 4 for another embodiment. In this embodiment, data of genotypes of n polymorphic loci in a biological individual is transmitted from the information terminal 21 held by the user 1 via the network NW, and the matching system 1 receives this via the network interface 16. Then, the disease risk score calculation unit 113 included in the information processing unit 11 calculates the disease risk score, and transmits the information from the network interface 16 to the user 1 information terminal 21 via the network NW.

[0132] In parallel with such processing, the information processing unit 11 adds the genotype data and / or information regarding the disease risk score received from the user 1 information terminal 21 to the database stored in the storage unit 12.

[0133] That is, in the present embodiment, in response to the request of the user 1 who wants to know the disease risk score of a specific biological individual, the genotype data and / or the disease risk score received from the user 1 can be added to and updated in the database.

[0134] Also, in the present embodiment, data of the genotypes of n polymorphic loci in a biological individual is transmitted from the information terminal 22 held by user 2 via the network NW, and the matching system 1 receives this via the network interface 16. The information processing unit 11, by the matching score calculation means 114, based on the genotype data of a specific biological individual received from user 2 and the genetic data of a plurality of biological individuals stored in the storage unit 12, calculates the genetic compatibility between the specific biological individual and each of the plurality of biological individuals (however, the individuals are of the opposite sex to the specific biological individual) as a matching score. Then, the extraction means 115 extracts, from among the plurality of biological individuals registered in the database, a biological individual (however, the individual is of the opposite sex to the specific biological individual) that shows a matching score satisfying a certain criterion in combination with the specific biological individual of user 2. Thereafter, the information processing unit 11 transmits information regarding the extracted biological individual from the network interface 16 to the user 2 information terminal 22 via the network NW.

[0135] Further, an embodiment may be adopted in which the genetic data of a plurality of biological individuals and specific information for identifying the same are associated and stored in the storage unit 12. Examples of the specific information include an individual number for identifying a biological individual and an inspection ID of an inspection performed to obtain the genetic data. In this embodiment, a signal designating the above-described specific information is transmitted from the information terminal 22 held by user 2 via the network NW, and the matching system 1 receives this via the network interface 16. The information processing unit 11 identifies the genetic data of a specific biological individual specified by user 2 based on the specific information from the data stored in the storage unit 12. Then, the matching score calculation means 114 calculates the genetic compatibility between the specific biological individual and each of the plurality of biological individuals (however, the individuals are of the opposite sex to the specific biological individual) as a matching score based on the genotype data of the specific biological individual specified by user 2 based on the specific information and the genetic data of the other plurality of biological individuals stored in the storage unit 12. The subsequent information processing flow is as described above.

[0136] In this embodiment, it is possible to meet the needs of User 2 who wants to efficiently search for opposite-sex biological individuals with excellent genetic compatibility with the biological individuals owned, managed, or specified by oneself.

[0137] Next, an embodiment will be described in which the matching score calculation means 114 calculates a score reflecting the genetic similarity between opposite sexes or the same sex.

[0138] An example is shown in FIG. 5 and will be described in more detail. In the example of FIG. 5, the specific biological individual is A1, regardless of its gender. The database stores the genetic data of biological individuals A2 to A x For each combination of the specific biological individual and each of the x - 1 biological individuals registered in the database, a matching score is calculated. The method for calculating the matching score is not particularly limited, but preferably it is calculated by executing an operation essentially including the above-described formula (18).

[0139] The genders of biological individuals A2 to A x may be the same as that of biological individual A1, may be of the opposite sex, or may include both the same sex and the opposite sex. The embodiment in which the genders of biological individuals A2 to A x and the gender of biological individual A1 are the same sex can be applied to, for example, a service for scoring the genetic compatibility (genetic similarity) of same-sex couples in humans.

[0140] The extraction means 115 is a means for extracting, from among the plurality of biological individuals registered in the database, a biological individual that exhibits a matching score satisfying a certain criterion in combination with the specific biological individual.

[0141] In the example of FIG. 5, the extraction means 115 extracts a combination that exhibits a matching score satisfying a certain criterion from among the matching scores calculated between the specific biological individual and each of the x - 1 biological individuals registered in the database. That is, a biological individual whose genetic compatibility (genetic similarity) with the specific biological individual satisfies a certain criterion is extracted.

[0142] The specific embodiments of the "certain criteria" mentioned here are applicable to the descriptions in the embodiments of FIGS. 2 and 3. Also, regarding the specific embodiments of the information processing unit 11 in the embodiment of FIG. 5, the descriptions in the embodiments of FIGS. 2 and 3 are applicable.

[0143] Even when adopting the embodiment shown in FIG. 5, the configuration shown in FIG. 4 can be applied. In this case, data of the genotypes of n polymorphic loci in a biological individual are transmitted from the information terminal 22 held by the user 2 via the network NW, and the matching system 1 receives this via the network interface 16. The information processing unit 11 calculates, by the matching score calculation means 114, the genetic compatibility between a specific biological individual and each of a plurality of biological individuals as a matching score based on the genotype data of the specific biological individual received from the user 2 and the genetic data of the plurality of biological individuals stored in the storage unit 12. Then, the extraction means 115 extracts a biological individual showing a matching score that satisfies certain criteria in combination with the specific biological individual of the user 2 from among the plurality of biological individuals registered in the database. Thereafter, the information processing unit 11 transmits information regarding the extracted biological individual from the network interface 16 to the user 2 information terminal 22 via the network NW.

[0144] Also, even when adopting the embodiment of FIG. 5, an embodiment can be applied in which the genetic data of a plurality of biological individuals and the specific information for identifying the same are associated and stored in the storage unit 12. The explanation of the specific information is as described above. In the present embodiment, it is possible to meet the needs of the user 2 who wants to efficiently search for a biological individual having excellent genetic compatibility (genetic similarity) with the biological individual held, managed, or designated by himself / herself.

Industrial Applicability

[0145] The present invention can be applied to the evaluation of disease risks and the evaluation of genetic compatibility. The present invention can be applied to the evaluation of disease risks in livestock and companion animals and the consideration of mating combinations. In addition, in conventional matching sites, items such as appearance and annual income have been the targets of matching evaluation. However, by applying the present invention, matching based on genetic compatibility can be realized. More specifically, the present invention can be applied to services that score genetic compatibility from the perspective of the disease risks that may occur in offspring, and services that score the genetic similarity of same-sex couples.

Claims

1. A calculation method comprising: a computer performing a calculation essentially including a calculation according to the following formula (1) based on genotype data for each of n polymorphic loci in an individual organism, thereby calculating a disease risk score. [0010] Formula (1) E k is the number of disease-associated alleles at each polymorphic locus or a numerical value proportional to said number. S k is a numerical value indicating the severity of the phenotype of a disease-associated allele at each polymorphic locus. However, the n polymorphic loci targeted by formula (1) are polymorphic loci in which at least one of the alleles that appears is a disease-associated allele.

2. The method according to claim 1 , wherein the operation essentially including a calculation according to the formula (1) is an operation expressed by the following formula (2). [0025] Formula (2) a is an arbitrary constant or an arbitrary variable. b is any constant (except 0) or any variable. c is an arbitrary constant or an arbitrary variable.

3. The calculation method according to claim 2 , wherein the c is a numerical value (DR value) that is predetermined depending on whether the disease-associated allele at each polymorphic locus is dominant or recessive.

4. The a is A numerical value (CAV value) that is determined in advance based on the presence or absence of one or more chromosomal abnormalities selected from chromosomal numerical abnormalities, copy number mutations in specific DNA regions, and chromosomal structural abnormalities, or a numerical value obtained by adding or multiplying a numerical value (CAV value) determined in advance based on the presence or absence of one or more chromosomal abnormalities selected from a chromosomal numerical abnormality, a copy number variation in a specific DNA region, and a chromosomal structural abnormality, and a numerical value indicating the severity of the phenotype of the chromosomal abnormality; The calculation method according to claim 2.

5. The method of claim 1 , further comprising: performing a ranking of disease risk scores for particular individual organisms in the population of organisms.

6. A calculation method comprising: a computer performing a calculation that essentially includes a calculation using a formula selected from the following formulas (7), (18), and (19) based on genotype data for each of n polymorphic loci in each of two different biological individuals, and calculating the genetic compatibility between the biological individuals as a matching score. [0030] Formula (7) [0045] Formula (18) Formula (7) + Formula (18)... Formula (19) In formula (7), M.E. k is the number of disease-associated alleles at each polymorphic locus in the male when the two biological individuals are male and female, or a numerical value proportional to said number. F.E. k is the number of disease-associated alleles at each polymorphic locus in the female when the two biological individuals are male and female, or a numerical value proportional to said number. S k is a numerical value indicating the severity of the phenotype of a disease-associated allele at each polymorphic locus. However, the n polymorphic loci targeted by formula (7) are polymorphic loci in which at least one of the alleles that appears is a disease-associated allele (effect allele). In formula (18), R l is a numerical value determined according to the degree of similarity between the genotypes at each polymorphic locus of the two individual organisms. However, the m polymorphic loci targeted by formula (18) are polymorphic loci in which all the alleles that appear are constitution / talent-related alleles and no disease-related alleles appear.

7. The method according to claim 6, wherein the operation essentially including a calculation according to the formula (7) is an operation expressed by the following formula (8). [0050] Formula (8) a is an arbitrary constant or an arbitrary variable. b is any constant (except 0) or any variable. c is an arbitrary constant or an arbitrary variable.

8. The calculation method according to claim 7, wherein the c is a numerical value (DR value) that is predetermined depending on whether the disease-associated allele at each polymorphic locus is dominant or recessive.

9. The a is a predetermined value based on the presence or absence of one or more chromosomal abnormalities selected from chromosomal numerical abnormalities, copy number mutations in specific DNA regions, and chromosomal structural abnormalities in either or both of the male and the female, or a numerical value obtained by adding or multiplying a predetermined numerical value based on the presence or absence of one or more types of chromosomal abnormalities selected from a chromosomal numerical abnormality, a copy number mutation in a specific DNA region, and a chromosomal structural abnormality in either or both of the male and the female, and a numerical value indicating the severity of the phenotype of the chromosomal abnormality. The calculation method according to claim 7.

10. The computation method according to claim 6 , further comprising: ranking matching scores between specific individual organisms in a set of matching scores between any individual organisms in the population of organisms.

11. A matching system for evaluating genetic compatibility comprising a memory unit and an information processing unit, the storage unit stores a database which is a collection of genetic data relating to a plurality of individual organisms; The information processing unit includes: Calculate the genetic compatibility of two different organisms as a matching score based on the genetic data of the two organisms; or calculating a genetic compatibility between the specific biological individual and each of the plurality of biological individuals as a matching score based on the genetic data of the specific biological individual and the genetic data of the plurality of biological individuals stored in the storage unit; A matching score calculation means is provided, The genetic data includes data on the genotype of each of n polymorphic loci in an individual organism, and / or data on a disease risk score calculated from data on the genotype of each of n polymorphic loci in an individual organism. and The matching score is calculated by the calculation method according to claim 6. Matching system.

12. The matching system according to claim 11 , wherein the information processing unit includes an extraction means for extracting, from the plurality of biological individuals registered in the database, biological individuals that exhibit a matching score that satisfies a certain criterion in combination with the specific biological individual.

13. The matching system according to claim 11, wherein the disease risk score is calculated by a calculation method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Diagnostic support system and method

    JP2007004211A

  • Method for detecting genetic risk of chronic nephropathy

    JP2010142187A

  • Method for analyzing digital data

    JP2019514148A

  • Systems and methods for interpreting data and providing recommendations to users based on data relating to the user's genetic data and gut microbiota composition

    JP2021508488A

  • Risk determination method and risk determination system for arrhythmia or arrhythmia-related disorder

    JP2023168188A