Genetic characteristic estimation device, control method, and program

The genetic characteristic estimation device and method improve the accuracy of genetic trait prediction by identifying key mutations associated with specific cell or organ types, thereby providing a more precise genetic characteristic index value.

JP7730475B2Active Publication Date: 2025-08-28NEC CORP +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2023529152
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-06-14
Publication Date
2025-08-28
Estimated Expiration
2041-06-14

AI Technical Summary

Technical Problem

Existing technologies for estimating genetic characteristics of organisms are limited in their accuracy and specificity, as they do not adequately distinguish between genetic mutations that significantly influence the organism's traits and those that do not.

Method used

A genetic characteristic estimation device and method that identifies a specific mutation of interest associated with the type of cell or organ, calculating a genetic characteristic index value based on the characteristics of this mutation, while minimizing the influence of other mutations.

Benefits of technology

This approach allows for a more accurate representation of the genetic characteristics of an organism by focusing on mutations with significant impact, enhancing the precision of trait prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007730475000008
    Figure 0007730475000008
  • Figure 0007730475000009
    Figure 0007730475000009
  • Figure 0007730475000010
    Figure 0007730475000010
Patent Text Reader

Abstract

This genetic feature estimation device (2000) acquires gene mutation information (30) and positional information (40). The gene mutation information (30) refers to information relating to a gene mutation occurring in a deoxyribonucleic acid (DNA) sequence in a target cell (20) obtained from a target organism (10). The positional information (40) assigns a position on the DNA sequence to a type of a cell or a type of an organ. The genetic feature estimation device (2000) identifies a gene mutation occurring at a position that is assigned to the type of the target cell (20) or the type of an organ containing the target cell (20) in the positional information (40) from among gene mutations shown in the gene mutation information (30), and then calculates a genetic feature index value that represents a genetic feature of the target organism (10) on the basis of a characteristic of the identified gene mutation.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a technique for estimating the genetic characteristics of an organism. [Background technology]

[0002] Technologies for estimating the genetic characteristics of organisms have been developed. For example, Patent Document 1 discloses a technology for predicting a trait of an evaluation target from the genetic mutations of the evaluation target, using a database that stores information on genetic mutations common to a group of samples that exhibit a common trait. The system in Patent Document 1 uses information from the database to calculate a score representing the degree of association between one or more genetic mutations possessed by the evaluation target and a specific trait, and predicts the trait based on this score. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] International Publication No. 2019 / 181022 Summary of the Invention [Problem to be solved by the invention]

[0004] An object of the present disclosure is to provide a new technique for estimating the genetic characteristics of an organism. [Means for solving the problem]

[0005] The genetic characteristic estimation device of the present disclosure includes an acquisition unit that acquires genetic mutation information regarding genetic mutations in a DNA (deoxyribonucleic acid) sequence possessed by a target cell obtained from a target organism, and position information in which positions on the DNA sequence correspond to cell types or organ types; and a calculation unit that identifies a target mutation, which is a genetic mutation at a position in the position information that corresponds to the type of the target cell or the type of organ containing the target cell, from among the genetic mutations indicated by the genetic mutation information, and calculates a genetic characteristic index value that represents the genetic characteristic of the target organism based on the characteristics of the target mutation.

[0006] The control method of the present disclosure is executed by a computer and includes an acquisition step of acquiring genetic mutation information on genetic mutations in a DNA (deoxyribonucleic acid) sequence of a target cell obtained from a target organism and position information in which positions on the DNA sequence correspond to cell types or organ types, and a calculation step of identifying a mutation of interest from the genetic mutations indicated by the genetic mutation information, which is a genetic mutation at the position in the position information that corresponds to the type of the target cell or the type of organ containing the target cell, and calculating a genetic characteristic index value representing a genetic characteristic of the target organism based on characteristics of the mutation of interest.

[0007] The non-transitory computer-readable medium of the present disclosure stores a program that causes a computer to execute the control method of the present disclosure. [Effects of the Invention]

[0008] According to the present disclosure, a new technique for estimating the genetic characteristics of an organism is provided. [Brief explanation of the drawings]

[0009] [Figure 1] FIG. 1 is a diagram illustrating an example of an outline of the operation of the genetic characteristic estimating device of the first embodiment. [Figure 2] 1 is a block diagram illustrating an example of the functional configuration of a genetic characteristic estimating apparatus according to a first embodiment. [Figure 3] FIG. 1 is a block diagram illustrating an example of a hardware configuration of a computer that realizes a genetic characteristic estimation device. [Figure 4] 1 is a flowchart illustrating the flow of processing executed by the genetic characteristic estimating device of the first embodiment. [Figure 5] FIG. 10 is a diagram illustrating an example of gene mutation information in a table format. [Figure 6] FIG. 10 is a diagram illustrating an example of location information in a table format. [Figure 7] FIG. 10 is a diagram illustrating an example of contribution degree information in a table format. DETAILED DESCRIPTION OF THE INVENTION

[0010] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. In each drawing, the same or corresponding elements are designated by the same reference numerals, and duplicate explanations will be omitted as necessary for clarity. Furthermore, unless otherwise specified, predetermined values ​​such as predetermined values ​​and threshold values ​​are pre-stored in a storage unit or the like in a manner that allows a device that uses the values ​​to acquire them. Furthermore, unless otherwise specified, the storage unit is configured with one or more storage devices.

[0011] Fig. 1 is a diagram illustrating an example of an outline of the operation of the genetic characteristic estimation device 2000 of embodiment 1. Here, Fig. 1 is a diagram for facilitating understanding of the outline of the genetic characteristic estimation device 2000, and the operation of the genetic characteristic estimation device 2000 is not limited to that shown in Fig. 1.

[0012] The genetic characteristic estimation device 2000 calculates an index value (hereinafter referred to as a genetic characteristic index value) related to the genetic characteristics of a target organism 10. The target organism 10 is any organism for which a genetic characteristic index value is to be calculated, and may be a human being or other animal, or may be a plant.

[0013] Genetic characteristics are characteristics that are manifested by the effects of genes. For example, genetic characteristics are characteristics related to diseases, such as the likelihood of developing a disease or the rate at which a disease progresses. Other examples of genetic characteristics include physical characteristics, such as height and weight. Other examples of genetic characteristics include the magnitude of drug effects, such as resistance or sensitivity to a drug.

[0014] The genetic characteristic index value is, for example, a polygenic risk score. However, the genetic characteristic index value is not limited to a polygenic risk score as long as it is an index value that represents the genetic characteristic of the subject organism 10.

[0015] To calculate the genetic characteristic index value of the target organism 10, the genetic characteristic estimation device 2000 acquires genetic mutation information 30 and position information 40. The genetic mutation information 30 indicates information about genetic mutations in the DNA (deoxyribonucleic acid) sequence of a cell (target cell 20) obtained from the target organism 10. Here, the genetic mutation information 30 indicates at least the position in the DNA sequence for each of one or more genetic mutations possessed by the target cell 20.

[0016] The position information 40 is information that associates the type of cell or organ with the position on the DNA sequence. For example, the position information 40 indicates, for each type of cell or organ, the position on the DNA sequence that should be particularly noted when calculating the genetic characteristic index value.

[0017] Examples of cell types include nerve cells, glial cells, blood cells, and skin cells. The granularity of the classification is arbitrary. For example, glial cells may be further classified into more specific types such as microglia and oligodendrocytes.

[0018] Examples of organ types include the brain, heart, and lungs. However, the granularity of organ classification is also arbitrary. For example, a group including multiple organ types, such as "respiratory system," may be used as an organ type.

[0019] The genetic characteristic estimation device 2000 identifies a genetic mutation at a position associated with the type of target cell 20 or the type of organ containing the target cell 20 in the position information 40, from among the genetic mutations indicated by the genetic mutation information 30. Hereinafter, the genetic mutation identified here will be referred to as a "mutation of interest." The genetic characteristic estimation device 2000 calculates a genetic characteristic index value for the target organism 10 based on the characteristics of the mutation of interest.

[0020] Here, among the genetic mutations possessed by the target cell 20, the characteristics of genetic mutations other than the mutation of interest may or may not be used in calculating the genetic characteristic index value. In the latter case, however, the characteristics of the mutation of interest are made to have a greater influence on the genetic characteristic index value (contribution to the genetic characteristic index value) than the characteristics of genetic mutations other than the mutation of interest. Specific methods for this will be described later.

[0021] <Example of effects> According to the genetic characteristic estimation device 2000 of this embodiment, a genetic mutation (mutation of interest) is identified from among the genetic mutations possessed by the target cell 20 of the target organism 10 at a position that is associated in the position information 40 with the type of the target cell 20 or the type of organ in which the target cell 20 is included. Then, an index value related to the genetic characteristic of the target organism 10 is calculated based on the characteristics of the mutation of interest. The characteristics of genetic mutations other than the mutation of interest are not used in calculating the genetic characteristic index value, or are used so as to have a smaller effect on the genetic characteristic index value than the characteristics of the mutation of interest.

[0022] According to this method, it is possible to calculate a genetic characteristic index value that more accurately represents the genetic characteristic of the target organism 10, compared to a case in which the target mutation is not distinguished from other genetic mutations. For example, in the position information 40, the type of target cell 20 or the type of organ containing the target cell 20 is associated with a position on the DNA sequence that is thought to have a large influence on the genetic characteristic. By doing so, in calculating the genetic characteristic index value, attention is paid to the characteristics of the genetic mutation at the position that is thought to have a large influence on the genetic characteristic. Therefore, it is possible to calculate a genetic characteristic index value that more accurately represents the genetic characteristic of the target organism 10, compared to a case in which such attention is not paid.

[0023] The genetic characteristic estimating device 2000 of this embodiment will be described in more detail below.

[0024] <Example of functional configuration> 2 is a block diagram illustrating an example of the functional configuration of a genetic characteristic estimation device 2000 according to the first embodiment. The genetic characteristic estimation device 2000 includes an acquisition unit 2020 and a calculation unit 2040. The acquisition unit 2020 acquires genetic mutation information 30 and position information 40 regarding a target cell 20 in a target organism 10. The calculation unit 2040 identifies, from the genetic mutations indicated by the genetic mutation information 30, a genetic mutation at a position associated in the position information 40 with the type of the target cell 20 or the type of organ containing the target cell 20. The calculation unit 2040 then calculates a genetic characteristic index value based on the characteristics of the identified genetic mutation.

[0025] <Example of hardware configuration> Each functional component of the genetic characteristic estimation device 2000 may be realized by hardware that realizes the functional component (e.g., a hardwired electronic circuit, etc.), or by a combination of hardware and software (e.g., a combination of an electronic circuit and a program that controls it). Below, a case where each functional component of the genetic characteristic estimation device 2000 is realized by a combination of hardware and software will be further described.

[0026] 3 is a block diagram illustrating an example of the hardware configuration of a computer 500 that realizes the genetic characteristic estimation device 2000. The computer 500 is any computer. For example, the computer 500 is a stationary computer such as a PC (Personal Computer) or a server machine. Alternatively, the computer 500 may be a portable computer such as a smartphone or a tablet terminal. The computer 500 may be a dedicated computer designed to realize the genetic characteristic estimation device 2000, or may be a general-purpose computer.

[0027] For example, by installing a predetermined application on the computer 500, each function of the genetic characteristic estimation device 2000 is realized on the computer 500. The application is configured with a program for realizing each functional component of the genetic characteristic estimation device 2000. The program can be acquired by any method. For example, the program can be acquired from a storage medium (such as a DVD disk or USB memory) on which the program is stored. Alternatively, the program can be acquired by downloading the program from a server device that manages the storage unit on which the program is stored.

[0028] The computer 500 includes a bus 502, a processor 504, a memory 506, a storage device 508, an input / output interface 510, and a network interface 512. The bus 502 is a data transmission path for the processor 504, the memory 506, the storage device 508, the input / output interface 510, and the network interface 512 to transmit and receive data to and from each other. However, the method for connecting the processor 504 and other components to each other is not limited to a bus connection.

[0029] The processor 504 is a variety of processors, such as a central processing unit (CPU), a graphics processing unit (GPU), or a field-programmable gate array (FPGA). The memory 506 is a main storage device realized using a random access memory (RAM) or the like. The storage device 508 is an auxiliary storage device realized using a hard disk, a solid state drive (SSD), a memory card, or a read-only memory (ROM) or the like.

[0030] The input / output interface 510 is an interface for connecting the computer 500 with input / output devices. For example, the input / output interface 510 is connected to an input device such as a keyboard and an output device such as a display device.

[0031] The network interface 512 is an interface for connecting the computer 500 to a network. This network may be a LAN (Local Area Network) or a WAN (Wide Area Network).

[0032] The storage device 508 stores programs (programs that realize the above-mentioned applications) that realize the various functional components of the genetic characteristic estimation device 2000. The processor 504 reads these programs into the memory 506 and executes them, thereby realizing the various functional components of the genetic characteristic estimation device 2000.

[0033] The genetic characteristic estimation device 2000 may be realized by one computer 500 or by multiple computers 500. In the latter case, the configurations of the computers 500 do not need to be the same, and can be different from each other.

[0034] <Processing flow> 4 is a flowchart illustrating the flow of processing executed by the genetic characteristic estimation device 2000 of embodiment 1. The acquisition unit 2020 acquires genetic variation information 30 (S102). The acquisition unit 2020 acquires location information 40 (S104). The calculation unit 2040 identifies a mutation of interest using the genetic variation information 30 and the location information 40 (S106). Specifically, the calculation unit 2040 identifies, from among the genetic variations indicated by the genetic variation information 30, a genetic variation that is associated in the location information 40 with the type of the target cell 20 or the type of organ containing the target cell 20, as the mutation of interest. The calculation unit 2040 then calculates a genetic characteristic index value based on the magnitude of the contribution of the mutation of interest to the genetic characteristic (S108).

[0035] Here, the processing flow shown in Fig. 4 is an example of the processing flow executed by the genetic characteristic estimation device 2000, and the processing flow executed by the genetic characteristic estimation device 2000 is not limited to that shown in Fig. 4. For example, the acquisition of genetic variation information 30 (S102) and the acquisition of position information 40 (S104) may be performed in the reverse order from the above-mentioned order, or may be performed in parallel with each other.

[0036] <Acquisition of gene mutation information 30: S102> The acquisition unit 2020 acquires genetic mutation information 30 (S102). As described above, the genetic mutation information 30 indicates information about genetic mutations in the DNA sequence of the target cell 20. The genetic mutation information 30 indicates, at least, the position of each genetic mutation in the DNA sequence of the target cell 20.

[0037] FIG. 5 is a diagram illustrating an example of gene mutation information 30 in table format. The gene mutation information 30 in FIG. 5 has two columns: position 32 and gene mutation 34. Position 32 indicates the position on the DNA sequence of the target cell 20. Gene mutation 34 indicates the gene mutation that the target cell 20 has at the position on the DNA sequence indicated by the corresponding position 32. For example, the record in the first row of FIG. 5 indicates that the target cell 20 has gene mutation V1 at position P1.

[0038] There are various methods for the acquiring unit 2020 to acquire the genetic mutation information 30. For example, the genetic mutation information 30 is pre-stored in a storage unit accessible from the genetic characteristic estimation device 2000. In this case, the acquiring unit 2020 acquires the genetic mutation information 30 by accessing this storage unit. As a more specific example, if the target organism 10 is a patient in a hospital, the genetic mutation information 30 may be included in data representing the medical record of the target organism 10 (so-called electronic medical record). In this case, the acquiring unit 2020 acquires the genetic mutation information 30 of the target organism 10 from the electronic medical record of the target organism 10. Note that existing technology can be used to acquire desired information from the electronic medical record of a specific person. Alternatively, for example, the genetic mutation information 30 may be transmitted to the genetic characteristic estimation device 2000 from another device.

[0039] <Acquisition of location information 40: S104> The acquiring unit 2020 acquires the position information 40 (S104). As described above, the position information 40 is information that associates the type of cell or organ with a position on a DNA sequence. FIG. 6 is a diagram illustrating an example of the position information 40 in table format. The position information 40 has two columns: type 42 and position 44. Type 42 indicates the type of cell or organ. Position 44 indicates one or more positions on a DNA sequence. When position 44 indicates multiple positions, position 44 may indicate a specific region on the DNA sequence. For example, the record on the first row in FIG. 6 associates cell type C1 with range R1 on a DNA sequence.

[0040] Specific examples of regions in a DNA sequence include promoters, enhancers, chemically modified regions (regions where DNA methylation has occurred), or specific genes. These regions directly or indirectly affect gene expression and protein structure. Therefore, genetic mutations in these regions are thought to have a greater impact on the genetic characteristics of an organism than genetic mutations in other regions. Therefore, by focusing particularly on genetic mutations in these regions among the genetic mutations in the target organism 10, the genetic characteristics of the target organism 10 can be more accurately understood.

[0041] There are various methods for the acquiring unit 2020 to acquire the position information 40. For example, the position information 40 is stored in advance in a storage unit accessible from the genetic characteristic estimation device 2000. In this case, the acquiring unit 2020 acquires the position information 40 by accessing this storage unit. Alternatively, for example, the position information 40 may be transmitted to the genetic characteristic estimation device 2000 from another device.

[0042] The location information 40 may be prepared for each type of genetic characteristic for which a genetic characteristic index value is to be calculated. In this case, for example, different location information 40 is used to calculate a genetic characteristic index value representing the risk of lung cancer and a genetic characteristic index value representing the risk of Alzheimer's disease.

[0043] Here, multiple pieces of position information 40 may be prepared for one type of genetic characteristic. In this case, the genetic characteristic estimation device 2000 may calculate one genetic characteristic index value using the multiple pieces of position information 40, or may calculate genetic characteristic index values ​​individually for each piece of position information 40. By calculating genetic characteristic index values ​​individually for the multiple pieces of position information 40, it is possible to evaluate risks, etc. for one genetic characteristic related to the target organism 10 for each type of organ or type of cell.

[0044] For example, for a patient with schizophrenia, genetic characteristic index values ​​that indicate the risk of developing schizophrenia are calculated for each of three organs: the brain, liver, and intestine. From these, it is possible to predict which of these organs is most closely related to the patient's schizophrenia. In this case, location information 40 for "type 42 = brain," location information 40 for "type 42 = liver," and location information 40 for "type 42 = intestine" are prepared. The genetic characteristic estimation device 2000 then calculates a genetic characteristic index value for each of these three pieces of location information 40.

[0045] If the genetic characteristic index values ​​for both the brain and the intestine indicate a high risk of schizophrenia, while the genetic characteristic index value for the liver indicates a low risk of schizophrenia, it can be seen that there is a high probability that the brain and the intestine are involved in this patient's schizophrenia.

[0046] When the position information 40 is defined for each type of genetic characteristic, for example, the genetic characteristic estimation device 2000 acquires information specifying the type of genetic characteristic for which the genetic characteristic index value is to be calculated (the type of genetic characteristic for which the genetic characteristic index value is to be calculated). For example, this information is input by the user. In this case, the acquisition unit 2020 acquires the position information 40 corresponding to the type of genetic characteristic specified by the user.

[0047] <Identification of target mutation: S106> The calculation unit 2040 identifies a position in the position information 40 that is associated with the type of the target cell 20 or the type of organ containing the target cell 20, and identifies the genetic mutation indicated for that position by the genetic mutation information 30 as a mutation of interest (S106). For example, the calculation unit 2040 identifies, from the position information 40, a record that indicates the type of the target cell 20 or the type of organ containing the target cell 20 as type 42.

[0048] The calculation unit 2040 identifies, from among the records of the genetic mutation information 30, the record in the identified position information 40 in which the position indicated by position 44 is indicated by position 32. Then, the calculation unit 2040 identifies the genetic mutation indicated by genetic mutation 34 in the identified record of the genetic mutation information 30 as the mutation of interest.

[0049] For example, suppose that a record of the position information 40 indicating the type of the target cell 20 as type 42 indicates two elements, "promoter" and "enhancer," at position 44. In this case, the calculation unit 2040 identifies a record from the genetic mutation information 30 that indicates a position included in the promoter or enhancer at position 32. Then, the calculation unit 2040 identifies the genetic mutation indicated in genetic mutation 34 of the identified record as the mutation of interest.

[0050] Here, whether to use the cell type or the organ type, or both, may be predetermined in the genetic characteristic estimation device 2000 or may be dynamically determined by the user. In the latter case, for example, the genetic characteristic estimation device 2000 provides the user of the genetic characteristic estimation device 2000 with an input interface (e.g., an input screen) that allows the user to select the cell type or the organ type to be used in identifying the mutation of interest. The genetic characteristic estimation device 2000 then identifies the mutation of interest based on the result of the user input. For example, suppose the user selects "cell type." In this case, the calculation unit 2040 identifies the position associated with the type of the target cell 20 in the position information 40.

[0051] <Calculation of genetic characteristic index value: S108> The calculation unit 2040 calculates a genetic characteristic index value based on the characteristics of the mutations of interest (S108). For example, the calculation unit 2040 calculates a score based on the characteristics of each mutation of interest. Then, the calculation unit 2040 calculates a genetic characteristic index value based on the score calculated for each mutation of interest.

[0052] For example, a formula for calculating a score from the characteristics of a genetic mutation and a formula for calculating a genetic characteristic index value based on the score calculated for each mutation of interest are determined in advance. For example, these formulas are expressed as the following formula (1).

number

[0053] For example, the genetic characteristic index value is calculated as a simple sum or a weighted sum of the scores calculated for each mutation of interest. In this case, formula (1) can be expressed as the following formula (2).

number

[0054] There are various ways to convert the characteristics of a target mutation into a score. For example, the number of specific alleles that the target mutation has can be used as a score. Other examples include the strength of the correlation between the target mutation and surrounding mutations on the DNA, as expressed as linkage disequilibrium, and the strength of promoter and enhancer activity.

[0055] Here, the magnitude of the influence of the characteristics of gene mutations on genetic characteristics may vary depending on the type of genetic characteristic. For example, for a certain genetic mutation, the magnitude of the influence on the risk of lung cancer, the magnitude of the influence on the risk of Alzheimer's disease, and the magnitude of the influence on the likelihood of height growth are likely to be different. Therefore, it is preferable to determine a calculation formula for calculating a score from the characteristics of gene mutations for each type of genetic characteristic.

[0056] In the case where a calculation formula for calculating a score from the characteristics of a genetic mutation is defined for each type of genetic feature, for example, the genetic feature estimation device 2000 acquires information specifying the type of genetic feature for which a genetic feature index value is to be calculated. As described above, for example, this information is input by a user. The calculation unit 2040 calculates the genetic feature index value using a calculation formula corresponding to the specified type of genetic feature from among the calculation formulas provided in advance. Note that, similarly, a calculation formula for calculating a genetic feature index value based on the score calculated for each mutation of interest may also be defined for each genetic feature for which a genetic feature index value is to be calculated.

[0057] <<Using genetic mutations other than the target mutation>> As described above, the genetic characteristic index value may be calculated using the characteristics of gene mutations other than the target mutation. In this case, the formula for calculating the genetic characteristic may be expressed as, for example, the following formula (3).

number

[0058] Furthermore, when the genetic characteristic index value is calculated as a simple sum or weighted sum of the scores calculated for each mutation of interest, Equation (3) can be expressed as Equation (4) below.

number

[0059] In formula (4), the constraint "α[i]>β[j]" is one way of realizing the constraint that "the magnitude of the influence of the characteristics of the mutation of interest on the genetic characteristic index value is greater than the influence of the characteristics of genetic mutations other than the mutation of interest on the genetic characteristic index value." However, the way of realizing this constraint is not limited to "α[i]>β[j]."

[0060] <<Consideration of contribution>> In calculating the genetic characteristic index value, the magnitude of contribution (degree of contribution) of each genetic mutation to the genetic characteristic may be taken into consideration. In this case, for example, the calculation unit 2040 selects a genetic mutation to be used in calculating the genetic characteristic index value based on the degree of contribution of each genetic mutation to the genetic characteristic. More specifically, the calculation unit 2040 selects, from the genetic mutations contained in the target cell 20, genetic mutations whose degree of contribution to the genetic characteristic is equal to or greater than a threshold, and calculates the genetic characteristic index value based on the characteristics of the selected genetic mutation. By selecting a genetic mutation to be used in calculating the genetic characteristic index value based on the magnitude of contribution to the genetic characteristic in this way, it is possible to calculate a genetic characteristic index value that more accurately represents the genetic characteristic of the target organism 10.

[0061] When only the mutation of interest is used and the contribution is taken into consideration, the formula for calculating the genetic characteristic index value can be expressed as, for example, the following formula (5).

number

[0062] When both the mutation of interest and other gene mutations are used and the contribution is taken into consideration, the formula for calculating the genetic characteristic index value can be expressed as, for example, the following formula (6).

number

[0063] Here, in equation (6), only those genetic mutations of interest and other genetic mutations whose contribution rate is equal to or greater than a threshold are selected. However, the calculation unit 2040 may select only genetic mutations other than the mutation of interest based on their contribution rate, without selecting the mutation of interest based on its contribution rate. In this case, the mutation of interest is used to calculate the genetic characteristic index value regardless of its contribution rate. On the other hand, only genetic mutations other than the mutation of interest whose contribution rate is equal to or greater than a threshold are used to calculate the genetic characteristic index value. In this case, the calculation formula for the genetic characteristic index value can be expressed as the following equation (7).

number

[0064] As described above, in order to take into account the contribution of each genetic variation to a genetic feature, the calculation unit 2040, for example, acquires information indicating the contribution of a genetic variation to a genetic feature (hereinafter, "contribution information"). The contribution information is stored in advance in an arbitrary storage unit in a form that can be acquired from the genetic feature estimation device 2000. The calculation unit 2040 acquires the contribution information for the genetic feature that is the target of calculation of the genetic feature index value, and uses the acquired contribution information to select a genetic variation to be used in calculating the genetic feature index value.

[0065] Fig. 7 is a diagram illustrating an example of contribution information in table format. The contribution information 50 in Fig. 7 has two columns: a genetic variation 52 and a contribution 54. The genetic variation 52 indicates identification information of the genetic variation. The contribution 54 indicates the contribution of the genetic variation indicated by the corresponding genetic variation 52 to the genetic feature.

[0066] Contribution information 50 is prepared for each type of genetic feature. For example, the contribution information 50 in FIG. 7 indicates the contribution of each genetic mutation to a genetic feature of type Fa. Therefore, for example, the record in the first row of the contribution information 50 in FIG. 7 indicates that the contribution of genetic mutation V1 to genetic feature Fa is Ka1.

[0067] Contribution information may be prepared for each indicator (such as blood glucose level or brain volume) of an organism that may affect a genetic characteristic. Specifically, contribution information 50 is prepared that indicates a higher contribution for a genetic mutation that has a stronger correlation with a specific indicator. When calculating a genetic characteristic index value for a genetic characteristic related to a specific index, the genetic characteristic estimation device 2000 uses contribution information 50 generated based on the strength of the correlation with the index.

[0068] For example, the strength of correlation with blood glucose level is examined for each genetic variation, and contribution information 50 is generated in advance, which indicates a higher contribution for a genetic variation with a stronger correlation with blood glucose level. The genetic characteristic estimation device 2000 uses this contribution information 50 when calculating a genetic characteristic index value for a disease related to blood glucose level (e.g., risk of developing diabetes).

[0069] When using contribution information 50 prepared for each index, information associating genetic features with indexes related to the genetic features is prepared in advance. For example, this information associates an index called "blood glucose level" with a genetic feature called "risk of developing diabetes." When calculating a genetic feature index value for a certain genetic feature, the genetic feature estimation device 2000 uses, in addition to a target mutation, genetic mutations whose contributions are equal to or greater than a threshold in the contribution information 50 for the index associated with the genetic feature.

[0070] <Output by genetic characteristic estimation device 2000> The genetic characteristic estimation device 2000 outputs information indicating a genetic characteristic index value (hereinafter, referred to as output information). For example, the output information includes a type of genetic characteristic and a genetic characteristic index value calculated for that type of genetic characteristic. In addition, for example, the output information may include various other information used in calculating the genetic characteristic index value. Examples of information used in calculating the genetic characteristic index value include a position (such as a promoter or enhancer) in the position information 40 associated with the type of target cell 20 or the type of organ containing the target cell 20, and a mutation of interest identified by the calculation unit 2040. Furthermore, when genetic variations are selected based on a threshold value for contribution, the output information may further include information such as the selected genetic variation and the threshold value for contribution.

[0071] The output information may be output in any manner. For example, the genetic characteristic estimation device 2000 stores the output information in any storage unit accessible from the genetic characteristic estimation device 2000. Alternatively, for example, the genetic characteristic estimation device 2000 displays the output information on any display device accessible from the genetic characteristic estimation device 2000. Alternatively, for example, the genetic characteristic estimation device 2000 transmits the output information to any device accessible from the genetic characteristic estimation device 2000.

[0072] Although the present invention has been described above with reference to the embodiments, the present invention is not limited to the above embodiments. Various modifications that can be understood by those skilled in the art can be made to the configuration and details of the present invention within the scope of the present invention.

[0073] In the above example, the program can be stored and provided to a computer using various types of non-transitory computer-readable media. Non-transitory computer-readable media include various types of tangible storage media. Examples of non-transitory computer-readable media include magnetic storage media (e.g., flexible disks, magnetic tapes, hard disk drives), magneto-optical storage media (e.g., magneto-optical disks), CD-ROMs, CD-Rs, CD-R / Ws, and semiconductor memories (e.g., mask ROMs, programmable ROMs (PROMs), erasable PROMs (EPROMs), flash ROMs, and RAMs). The program may also be provided to a computer by various types of transitory computer-readable media. Examples of transitory computer-readable media include electrical signals, optical signals, and electromagnetic waves. The transitory computer-readable media can provide the program to a computer via a wired communication path such as an electric wire or optical fiber, or via a wireless communication path.

[0074] A part or all of the above-described embodiments can be described as, but not limited to, the following supplementary notes. (Appendix 1) an acquisition unit that acquires gene mutation information regarding gene mutations in a DNA (deoxyribonucleic acid) sequence possessed by a target cell obtained from a target organism, and position information in which positions on the DNA sequence are associated with cell types or organ types; a calculation unit that identifies a mutation of interest, which is a genetic mutation at a position that corresponds in the position information to the type of the target cell or the type of organ containing the target cell, from the genetic mutations indicated by the genetic mutation information, and calculates a genetic characteristic index value that represents the genetic characteristic of the target organism based on the characteristics of the mutation of interest. (Appendix 2) A genetic characteristic estimation device as described in Appendix 1, wherein the location indicated by the location information represents a promoter, enhancer, chemically modified region, or region of a specific gene in DNA of the associated cell or a cell of the associated organ. (Appendix 3) The calculation unit Obtaining contribution information that indicates the degree of contribution of each gene mutation to the genetic characteristics; 3. The genetic characteristic estimation device according to claim 1, wherein the genetic characteristic index value is calculated based on the characteristics of the mutation of interest whose contribution is equal to or greater than a threshold. (Appendix 4) The calculation unit Calculating a first score based on characteristics of the mutation of interest and a second score based on characteristics of genetic mutations other than the mutation of interest; assigning different weights to the first score and the second score so that the influence of the first score on the genetic characteristic index value is greater than the influence of the second score on the genetic characteristic index value; 3. The genetic characteristic estimation device according to claim 1, wherein the genetic characteristic index value is calculated using the weighted first score and the weighted second score. (Appendix 5) The calculation unit Obtaining contribution information that indicates the degree of contribution of each gene mutation to the genetic characteristics; 5. The genetic characteristic estimation device according to claim 4, wherein the second score is calculated for only genetic mutations other than the target mutation whose contribution is equal to or greater than a threshold. (Appendix 6) 6. The genetic characteristic estimating device according to claim 5, wherein the calculation unit calculates the first score only for the mutation of interest whose contribution is equal to or greater than a threshold. (Appendix 7) 7. A genetic characteristic estimation device according to any one of appendices 1 to 6, wherein the genetic characteristic index value is a polygenic risk score. (Appendix 8) 1. A computer-implemented control method comprising: an acquisition step of acquiring gene mutation information regarding gene mutations in a DNA (deoxyribonucleic acid) sequence possessed by a target cell obtained from a target organism, and position information in which positions on the DNA sequence are associated with cell types or organ types; a calculation step of identifying a mutation of interest, which is a genetic mutation at a position that corresponds in the position information to the type of the target cell or the type of organ containing the target cell, from the genetic mutations indicated by the genetic mutation information, and calculating a genetic characteristic index value that represents the genetic characteristics of the target organism based on the characteristics of the mutation of interest. (Appendix 9) The control method described in Appendix 8, wherein the location indicated by the location information represents a promoter, enhancer, chemically modified region, or region of a specific gene in DNA of the associated cell or a cell of the associated organ. (Appendix 10) In the calculation step, Obtaining contribution information that indicates the degree of contribution of each gene mutation to the genetic characteristics; 10. The control method according to claim 8 or 9, wherein the genetic characteristic index value is calculated based on the characteristics of the mutation of interest whose contribution is equal to or greater than a threshold. (Appendix 11) In the calculation step, Calculating a first score based on characteristics of the mutation of interest and a second score based on characteristics of genetic mutations other than the mutation of interest; assigning different weights to the first score and the second score so that the influence of the first score on the genetic characteristic index value is greater than the influence of the second score on the genetic characteristic index value; 10. The control method according to claim 8 or 9, wherein the genetic characteristic index value is calculated using the weighted first score and the weighted second score. (Appendix 12) In the calculation step, Obtaining contribution information that indicates the degree of contribution of each gene mutation to the genetic characteristics; The control method according to claim 11, wherein the second score is calculated for only genetic mutations other than the target mutation whose contribution is equal to or greater than a threshold. (Appendix 13) 13. The control method according to claim 12, wherein in the calculation step, the first score is calculated for only the mutation of interest whose contribution is equal to or greater than a threshold. (Appendix 14) 14. The control method of any one of appendices 8 to 13, wherein the genetic characteristic index value is a polygenic risk score. (Appendix 15) On the computer, an acquisition step of acquiring gene mutation information regarding gene mutations in a DNA (deoxyribonucleic acid) sequence possessed by a target cell obtained from a target organism, and position information in which positions on the DNA sequence are associated with cell types or organ types; a calculation step of identifying a mutation of interest, which is a genetic mutation at a position that corresponds in the position information to the type of the target cell or the type of organ containing the target cell, from among the genetic mutations indicated by the genetic mutation information, and calculating a genetic characteristic index value that represents the genetic characteristics of the target organism based on the characteristics of the mutation of interest. (Appendix 16) The computer-readable medium of claim 15, wherein the location indicated by the location information represents a promoter, enhancer, chemically modified region, or region of a specific gene in DNA of the associated cell or cells of the associated organ. (Appendix 17) In the calculation step, Obtaining contribution information that indicates the degree of contribution of each gene mutation to the genetic characteristics; 17. The computer-readable medium of claim 15, wherein the genetic characteristic index value is calculated based on the characteristics of the mutation of interest whose contribution is equal to or greater than a threshold. (Appendix 18) In the calculation step, Calculating a first score based on characteristics of the mutation of interest and a second score based on characteristics of genetic mutations other than the mutation of interest; assigning different weights to the first score and the second score so that the influence of the first score on the genetic characteristic index value is greater than the influence of the second score on the genetic characteristic index value; 17. The computer-readable medium of claim 15, wherein the genetic characteristic index value is calculated using the weighted first score and the weighted second score. (Appendix 19) In the calculation step, Obtaining contribution information that indicates the degree of contribution of each gene mutation to the genetic characteristics; The computer-readable medium of claim 18, wherein the second score is calculated for only genetic mutations other than the target mutation whose contribution is equal to or greater than a threshold. (Appendix 20) 20. The computer-readable medium of claim 19, wherein in the calculation step, the first score is calculated for only the mutation of interest whose contribution is equal to or greater than a threshold. (Appendix 21) 21. The computer-readable medium of any one of claims 15 to 20, wherein the genetic characteristic index value is a polygenic risk score. [Explanation of symbols]

[0075] 10 Target organisms 20 Target cells 30 Gene mutation information 32 positions 34 Genetic Mutations 40 Location information 42 types 44 position 50 Contribution Information 52 Genetic Mutations 54 Contribution 500 computers 502 Bus 504 processor 506 memory 508 Storage Devices 510 Input / Output Interface 512 network interface 2000 Genetic Trait Estimation Device 2020 Acquisition Department 2040 Calculation Department

Claims

1. an acquisition unit that acquires gene mutation information regarding gene mutations in the DNA (deoxyribonucleic acid) sequence of a target cell obtained from a target organism, and position information that associates positions on the DNA sequence with cell types or organ types; a calculation unit that identifies a mutation of interest, which is a genetic mutation at a position associated with the type of the target cell or the type of organ containing the target cell in the position information, from the genetic mutations indicated by the genetic mutation information, and calculates a genetic characteristic index value that represents a genetic characteristic of the target organism using a value that represents a characteristic of the mutation of interest, A genetic characteristic estimation device, wherein the characteristics of the target mutation include at least the number of specific alleles that the target mutation has, the strength of correlation between the target mutation and surrounding mutations on DNA, or the strength of activity of the promoter or enhancer in which the target mutation is present.

2. The genetic characteristic estimation device of claim 1, wherein the location indicated by the location information represents a promoter, enhancer, chemically modified region, or region of a specific gene in DNA of the associated cell or a cell of the associated organ.

3. The calculation unit Obtaining contribution information that indicates the degree of contribution of each gene mutation to the genetic characteristics; The genetic characteristic estimation device according to claim 1 , wherein the genetic characteristic index value is calculated based on a characteristic of the mutation of interest whose contribution is equal to or greater than a threshold value.

4. The calculation unit calculating a first score based on characteristics of the mutation of interest and a second score based on characteristics of genetic mutations other than the mutation of interest; assigning different weights to the first score and the second score so that the influence of the first score on the genetic characteristic index value is greater than the influence of the second score on the genetic characteristic index value; The genetic characteristic estimating device according to claim 1 , wherein the genetic characteristic index value is calculated using the weighted first score and the weighted second score.

5. The calculation unit Obtaining contribution information that indicates the degree of contribution of each gene mutation to the genetic characteristics; The genetic characteristic estimation device according to claim 4 , wherein the second score is calculated for only genetic variations other than the target variation whose contribution is equal to or greater than a threshold value.

6. The genetic characteristic estimating device according to claim 5 , wherein the calculation unit calculates the first score only for the mutation of interest whose contribution is equal to or greater than a threshold value.

7. The genetic characteristic estimation device according to claim 1 , wherein the genetic characteristic index value is a polygenic risk score.

8. 1. A computer-implemented control method comprising: an acquisition step of acquiring gene mutation information regarding gene mutations in a DNA (deoxyribonucleic acid) sequence possessed by a target cell obtained from a target organism, and position information in which positions on the DNA sequence are associated with cell types or organ types; a calculation step of identifying a mutation of interest, which is a genetic mutation at the position associated with the type of the target cell or the type of organ containing the target cell in the position information, from the genetic mutations indicated by the genetic mutation information, and calculating a genetic characteristic index value representing the genetic characteristic of the target organism using a value representing the characteristic of the mutation of interest, A control method in which the characteristics of the target mutation include at least the number of specific alleles that the target mutation has, the strength of the correlation between the target mutation and surrounding mutations on the DNA, or the strength of activity of the promoter or enhancer in which the target mutation is present.

9. The control method described in claim 8, wherein the location indicated by the location information represents a promoter, enhancer, chemically modified region, or region of a specific gene in DNA of the associated cell or a cell of the associated organ.

10. On the computer, an acquisition step of acquiring gene mutation information regarding gene mutations in a DNA (deoxyribonucleic acid) sequence possessed by a target cell obtained from a target organism, and position information in which positions on the DNA sequence are associated with cell types or organ types; a calculation step of identifying a mutation of interest, which is a genetic mutation at the position associated with the type of the target cell or the type of organ containing the target cell in the position information, from among the genetic mutations indicated by the genetic mutation information, and calculating a genetic characteristic index value representing the genetic characteristic of the target organism using a value representing the characteristic of the mutation of interest; A program in which the characteristics of the target mutation include at least the number of specific alleles that the target mutation has, the strength of correlation between the target mutation and surrounding mutations on DNA, or the strength of activity of a promoter or enhancer in which the target mutation is present.

Citation Information

Patent Citations

  • Methods and Systems for Medical Sequencing Analysis

    US20130184161A1

  • Inter-rater and intra-rater reliability of physiological scan interpretation

    US20140296733A1

  • Genetic mutation assessment device, assessment method, program, and recording medium

    WO2019181022A1