IBD Fragment Prediction Methods and Related Equipment
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-23
- Publication Date
- 2026-08-14
AI Technical Summary
[0003]现有技术中大多数模型只能可靠地找出长IBD片段,却很少能可靠地找出短IBD片段
Smart Images

Figure CN116543836B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of detection technology, and in particular to an IBD fragment prediction method and related equipment. Background Technology
[0002] Although humans are genetically nearly identical, subtle differences in human DNA are what cause some visible variations between individuals. In fact, by comparing these subtle differences in individual DNA, long chromosome segments that suggest inheritance from a most recent common ancestor can be detected, and these detected segments can then be used to estimate the affinity between two individuals. In population genetics literature, the process of identifying segments that suggest recent common inheritance is called identity by descent (IBD) analysis. IBD analysis can be used to predict familial relationships (e.g., second-degree relatives) between any two individuals in a population. Understanding population structure from genetic polymorphism data is an important topic in genetics. For example, in IBD mapping, a pathogenic mutation that may be inherited in a particular family can be located by looking for IBD segments shared by affected individuals in that family. Another example is performing IBD segment analysis on every two individuals in a population; by looking at the overall proportion of IBD segments in the genome, demographic information can be obtained, helping researchers analyze the historical changes of that population.
[0003] Most existing models can reliably identify long IBD fragments, but rarely can they reliably identify short IBD fragments. Therefore, it is necessary to provide a method that can effectively identify short IBD fragments. Summary of the Invention
[0004] This invention provides an IBD fragment prediction method and related equipment, which can achieve more accurate and reliable prediction of short IBD fragments, thereby improving the prediction accuracy of IBD fragments.
[0005] In a first aspect, embodiments of the present invention provide an IBD fragment prediction method, comprising:
[0006] Obtain the predicted value of the first IBD fragment between the DNA sequences of at least two targets, wherein the length of the first IBD fragment is greater than a preset length;
[0007] The first probability of observing the DNA sequences of the at least two targets when the genotype corresponding to the same first variant is an IBD fragment, and the second probability of observing the DNA sequences of the at least two targets when the genotype corresponding to the same first variant is a non-IBD fragment, are calculated based on multiple first variants, wherein the occurrence frequency of the first variant is less than a preset frequency.
[0008] The predicted value of a second IBD fragment containing the first variant is calculated based on the first probability and the second probability, wherein the length of the second IBD fragment is less than the preset length;
[0009] The IBD fragment between the DNA sequences of the at least two targets is determined based on the predicted values of the first IBD fragment and the second IBD fragment between the DNA sequences of the at least two targets.
[0010] This invention involves obtaining a predicted value for a first IBD fragment (long IBD fragment) between the DNA sequences of at least two targets; then, calculating a predicted value for a second IBD fragment (short IBD fragment) containing low-frequency variations between the DNA sequences of the at least two targets based on multiple first variations; and finally determining the IBD fragment between the DNA sequences of the at least two targets based on the predicted values of the first and second IBD fragments. This method, by introducing low-frequency variations, makes the prediction of short IBD fragments more accurate and reliable, thereby improving the accuracy of IBD fragment prediction; and solves the problem of insufficient ability to predict short IBD fragments in existing technologies.
[0011] Secondly, embodiments of the present invention provide an IBD fragment prediction device, comprising:
[0012] The acquisition module is used to acquire the predicted value of a first IBD fragment between the DNA sequences of at least two targets, wherein the length of the first IBD fragment is greater than a preset length;
[0013] The prediction module is used to calculate, based on multiple first variants, a first probability of observing the DNA sequences of the at least two targets when the genotype corresponding to the same first variant is an IBD fragment, and a second probability of observing the DNA sequences of the at least two targets when the genotype corresponding to the same first variant is a non-IBD fragment, wherein the occurrence frequency of the first variant is less than a preset frequency.
[0014] The calculation module is used to calculate, based on the first probability and the second probability, a predicted value of a second IBD fragment containing the first mutation between the DNA sequences of the at least two targets, wherein the length of the second IBD fragment is less than the preset length;
[0015] A determination module is configured to determine the IBD fragment between the DNA sequences of the at least two targets based on the predicted values of a first IBD fragment and a second IBD fragment between the DNA sequences of the at least two targets.
[0016] Thirdly, this application provides an IBD fragment prediction device, including: a processor and a memory;
[0017] The processor is connected to a memory, wherein the memory is used to store program code, and the processor is used to invoke the program code to perform a method as provided in any possible implementation of the first aspect.
[0018] Fourthly, this application provides a computer storage medium including computer instructions that, when executed on an electronic device, cause the electronic device to perform a method as provided in any possible implementation of the first aspect.
[0019] Fifthly, embodiments of this application provide a computer program product that, when run on a computer, causes the computer to perform the method provided in any possible implementation of the first aspect.
[0020] It is understood that the apparatus described in the second aspect, the device described in the third aspect, the computer storage medium described in the fourth aspect, or the computer program product described in the fifth aspect are all used to perform the method provided in any possible embodiment of the first aspect. Therefore, the beneficial effects that can be achieved can be referred to the beneficial effects in the corresponding methods, and will not be repeated here. Attached Figure Description
[0021] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a flowchart illustrating an IBD fragment prediction method provided in an embodiment of the present invention;
[0023] Figure 2 This is a flowchart illustrating another IBD fragment prediction method provided in an embodiment of the present invention;
[0024] Figure 3a This is a schematic diagram of window division provided by an embodiment of the present invention;
[0025] Figure 3b This is a block division diagram provided in an embodiment of the present invention;
[0026] Figure 4 This is a schematic diagram of a family structure provided in an embodiment of the present invention;
[0027] Figure 5 This is a schematic diagram of the structure of an IBD fragment prediction device provided in an embodiment of the present invention;
[0028] Figure 6This is a schematic diagram of the structure of an IBD fragment prediction device provided in an embodiment of the present invention. Detailed Implementation
[0029] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.
[0030] It should be understood that the terms "first," "second," etc., in the specification, claims, and drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.
[0031] In this invention, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the invention. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described in this invention can be combined with other embodiments.
[0032] Please see Figure 1 , Figure 1 This is a flowchart illustrating an IBD fragment prediction method provided in an embodiment of the present invention; as shown below. Figure 1 As shown, the method may include steps 101-104, as detailed below:
[0033] 101. Obtain the predicted value of the first IBD fragment between the DNA sequences of at least two targets, wherein the length of the first IBD fragment is greater than a preset length;
[0034] The at least two targets can be a portion or all of the target population of interest (to be predicted). The preset length can be, for example, 2 centimorgan (cM) or other values; this scheme does not impose any restrictions on this. The first IBD segment can be understood as a long IBD segment.
[0035] The prediction of the first IBD fragment between at least two target DNA sequences can be based on length threshold models, frequency threshold models, or statistical models. The core idea behind these three models in determining whether two genotypes or haplotypes corresponding to two targets are long IBD fragments is based on the probability of random fragment similarity. A longer length, lower frequency, or higher probability of being a long IBD fragment means a lower probability of random similarity. Length threshold models compare two genotypes or haplotypes for similarity and determine if the length of the identical fragment exceeds a preset threshold. If it does, the fragment is considered a long IBD fragment; otherwise, it is considered a non-IBD fragment. Frequency threshold models are similar to length threshold models; they compare two haplotypes for similarity and determine if the frequency of this haplotype fragment is below a preset threshold. If it is below the threshold, the fragment is considered a long IBD fragment. Statistical models calculate the probability of a fragment being a long IBD fragment. Common models include Truffle, hap-IBD, and Gerline2.
[0036] Of course, other methods or models can be used to obtain the predicted value of the first IBD fragment between the DNA sequences of at least two targets, and this scheme does not impose any restrictions on this.
[0037] 102. Calculate a first probability of observing the DNA sequences of the at least two targets when the genotype corresponding to the same first variant is an IBD fragment, and a second probability of observing the DNA sequences of the at least two targets when the genotype corresponding to the same first variant is a non-IBD fragment, based on multiple first variants, wherein the occurrence frequency of the first variant is less than a preset frequency.
[0038] The occurrence frequency of this first mutation is less than the preset frequency, and it can be understood as a low-frequency mutation. These multiple first mutations can be obtained from a preset database or through other means; this solution does not impose any restrictions on this.
[0039] The genotype contains both IBD and non-IBD segments. By calculating, for example, the probability of observing the DNA sequences of two targets when their genotypes for each of the aforementioned multiple first variants are IBD segments, and simultaneously calculating the probability of observing the DNA sequences of these two targets when their genotypes for each of the aforementioned multiple first variants are non-IBD segments, short IBD segments between the two targets can be predicted. The advantage of the first variant over the second variant is that the probability of two people sharing a minor allele at the first variant site is much lower than the probability at the second variant site. In a short DNA segment, it is generally difficult to infer that it is an IBD segment from two people simply because they share bases at the second variant site, as the probability of them randomly sharing the same base is not low. However, if two people share a minor allele at the first variant site in this short DNA segment, they are more likely to have inherited this segment from the same ancestor, and share minor alleles at more first variants, making the probability of random similarity even smaller. Therefore, by selecting high-quality first variant data (i.e., sequencing errors of the first variant are very low), two individuals who both carry the second allele at that locus can provide more information about IBD.
[0040] In one possible implementation, step 102 may include steps 1021-1025, as follows:
[0041] 1021. Obtain multiple first mutations from the training and prediction sets.
[0042] The prediction set may include single nucleotide polymorphism (SNP) data for the at least two targets. The training set is obtained from the 1000 Genomes Project (1000G) or other datasets, based on the ethnicity of the target population (the at least two targets). The training set may include, for example, 500 individuals.
[0043] In one possible implementation, based on the aforementioned training and prediction sets, SNP data containing only two alleles are selected, while SNP data containing three or more alleles are removed. Furthermore, SNP data with a missing value ratio (such as missing base information) greater than 10% are removed. Then, based on the minor allele frequency (MAF) in the training set, the SNP data are categorized into common variants (e.g., MAF ≥ 5%), low-frequency variants (e.g., 0.5% ≤ MAF < 5%), and rare variants (e.g., MAF < 0.5%). Based on this data processing, multiple first variants (i.e., low-frequency variants) can be obtained.
[0044] Optionally, only the SNP data of the above common variations and some low-frequency variations (e.g., 0.5% < MAF ≤ 1%) can be retained for subsequent calculations.
[0045] 1022. Calculate the haplotype of each of the at least two targets based on the genotypes of the at least two targets in the plurality of windows and the training set.
[0046] For example, calculate the haplotype of each of the at least two targets based on the genotypes of the at least two targets in the plurality of windows and the haplotype information of each person in the preset population in the training set.
[0047] 1023. Determine the number of major alleles and minor alleles of the plurality of first variations in the training set and the prediction set respectively according to the SNP data of the training set and the SNP data of the prediction set.
[0048] Among them, the major allele and the minor allele can be distinguished according to the frequency of the allele. For example, the allele with a high frequency is the major allele, and the allele with a low frequency is the minor allele, etc.
[0049] 1024. Calculate the fifth probability according to the haplotype of each of the at least two targets and the number of major alleles and minor alleles of the plurality of first variations in the training set and the prediction set respectively.
[0050] In a possible implementation, assume that the first variation follows a Bernoulli distribution (B(θ)), where θ is the probability of the major allele appearing in this first variation. Then the fifth probability p(h1, h2, h3|D) can be expressed as:
[0051]
[0052] Among them, h1 and h2 are the haplotypes of one of the two targets, h2 and h3 are the haplotypes of the other of the two targets, h2 is the haplotype inherited by these two people from the same ancestor, and D represents the number of major alleles and minor alleles of the above low-frequency variations appearing in the training set.
[0053] Let p(θ|D) follow a Beta distribution, then p(θ|D) = (θ
[0049] *(1 - θ) β-1 * Γ(α + β)) / (Γ(α) * Γ(β)), where α is the number of major alleles in the training set and the prediction set, and β is the number of minor alleles in the training set and the prediction set.
[0054] Based on this, the fifth probability can be calculated.
[0055] 1025. Based on the genotypes of the plurality of first variants corresponding to each target and the fifth probability, calculate the first probability of observing the DNA sequences of at least two targets when the genotype corresponding to the same first variant is an IBD fragment, and the second probability of observing the DNA sequences of at least two targets when the genotype corresponding to the same first variant is a non-IBD fragment.
[0056] In one possible implementation, assuming that the genotypes of the two individuals within a low-frequency variant m are IBD fragments, then the first probability of observing these two genotypes, p(g(m), g′(m)|IBD), can be expressed as:
[0057]
[0058] Where m is any one of the multiple low-frequency variants mentioned above, g(m) is the genotype of one person in this low-frequency variant m, and g′(m) is the genotype of another person in this low-frequency variant m. If g(m) = h1 + h2, then p(g(m)|h1,h2) = 1 - ε′, otherwise p(g(m)|h1,h2) = ε′ / 2. Where ε′ is the probability of sequencing errors for low-frequency variants. Based on this calculation, p(g′(m)|h2,h3) can also be calculated.
[0059] Furthermore, based on the fifth probability p(h1,h2,h3|D) calculated above, the first probability p(g(m),g′(m)|IBD) can be calculated.
[0060] Accordingly, assuming that the genotypes of the two individuals within a low-frequency variant m are non-IBD segments, then the second probability of observing these two genotypes, p(g(m), g′(m)|non-IBD), can be expressed as:
[0061]
[0062] Where h1 and h2 are the individual types of one of the at least two targets, and h3 and h4 are the individual types of the other of the at least two targets.
[0063] Based on the above method of calculating p(g(m)|h1,h2), p(g′(m)|h3,h4) can be calculated. Based on the above method of calculating p(h1,h2,h3|D), p(h1,h2,h3,h4|D) can be calculated. Furthermore, the second probability p(g(m),g′(m)|non-IBD) can be calculated.
[0064] 103. Calculate the predicted value of a second IBD fragment containing the first variant between the DNA sequences of the at least two targets based on the first probability and the second probability, wherein the length of the second IBD fragment is less than the preset length;
[0065] The second IBD segment can be understood as a short IBD segment. The aforementioned IBD segments include the first IBD segment (long IBD segment) and the second IBD segment (short IBD segment).
[0066] In one possible implementation, for each low-frequency variant, the logarithm η of the ratio of the first probability to the second probability is taken. m , the η m It can be represented as:
[0067] η m =log(p(g(m),g′(m)|IBD) / p(g(m),g′(m)|non-IBD)).
[0068] Based on this, a predicted value can be obtained for the second IBD fragment containing the first variant between the DNA sequences of at least two targets.
[0069] 104. Determine the IBD fragment between the DNA sequences of the at least two targets based on the predicted values of the first IBD fragment and the second IBD fragment between the DNA sequences of the at least two targets.
[0070] In one possible implementation, the predicted value of the first IBD fragment is calculated by dividing the genotype into multiple blocks according to length; that is, the predicted value of the first IBD fragment is calculated on a block-by-block basis. Therefore, for a given first IBD fragment, if the predicted value of the first IBD fragment in a certain block is greater than a preset threshold, then that block is considered an IBD fragment.
[0071] In one possible implementation, for the second IBD segment, the window containing each of the plurality of first variants is determined based on the plurality of first variants and the plurality of windows of preset size.
[0072] Specifically, a DNA sequence is divided into multiple windows of a preset size, each containing the same number of second variants, the frequency of which is higher than the preset frequency. For example, this second variant can be understood as a common variant.
[0073] Based on the division of common mutation windows, the window in which low-frequency mutations are located is determined. If a low-frequency mutation is exactly between two windows, the window that is closer to it can be selected as its window.
[0074] Based on the predicted values of the first IBD fragments and the second IBD fragments corresponding to the windows where the plurality of first variants are located, the second IBD fragment between the DNA sequences of the at least two targets is determined.
[0075] It should be noted that the predicted value of the first IBD segment corresponding to the window where the first mutation is located in this scheme can be understood as the predicted value of the window where the first mutation is located.
[0076] For any first mutation, if the sum of the predicted value of the second IBD segment corresponding to the first mutation and the predicted value of the first IBD segment corresponding to the window where the first mutation is located is greater than a first preset value, the second IBD segment corresponding to the first mutation is expanded until one or more of the following conditions are met to obtain an updated second IBD segment:
[0077] 1. The predicted value of the first IBD segment of the window corresponding to the updated second IBD segment is less than a second preset value, and the second preset value is less than the first preset value;
[0078] 2. The sum of the predicted value of the updated second IBD segment corresponding to the first mutation and the predicted value of the first IBD segment corresponding to the window where the first mutation is located is less than the first preset value;
[0079] 3. The number and / or proportion of first windows in the window corresponding to the updated second IBD segment are greater than a third preset value, wherein the predicted value of the first IBD segment corresponding to the first window is less than 0;
[0080] 4. The length of the updated second IBD segment is greater than a fourth preset value. This fourth preset value may be, for example, 2 centimoles.
[0081] In other words, if the sum of the predicted value of the second IBD segment corresponding to the first mutation and the predicted value of the first IBD segment corresponding to the window where the first mutation is located is greater than the first preset value, the second IBD segment corresponding to the first mutation is expanded in units of each window. By continuously comparing whether the expanded second IBD segment meets any of the above conditions, if it does, the expansion is stopped, and the updated second IBD segment is determined as the predicted short IBD segment.
[0082] It should be noted that the updated second IBD segment in this scheme can be understood as the second IBD segment determined after expanding the predicted value based on the second IBD segment. Of course, other expressions can also be used, and this scheme does not restrict them.
[0083] Finally, the IBD fragment between the DNA sequences of the at least two targets is determined based on the first IBD fragment between the DNA sequences of the at least two targets and the updated second IBD fragment between the DNA sequences of the at least two targets.
[0084] Furthermore, for overlapping IBD fragments, they are merged to obtain an IBD fragment between the DNA sequences of the at least two targets.
[0085] In this embodiment, a predicted value of a first IBD fragment (long IBD fragment) between the DNA sequences of at least two targets is obtained. Then, a predicted value of a second IBD fragment (short IBD fragment) containing low-frequency variations between the DNA sequences of the at least two targets is calculated based on multiple first variations. The IBD fragment between the DNA sequences of the at least two targets is then determined based on the predicted values of the first and second IBD fragments. This method, by introducing low-frequency variations, makes the prediction of short IBD fragments more accurate and reliable, thereby improving the accuracy of IBD fragment prediction and solving the problem of insufficient ability to predict short IBD fragments in existing technologies.
[0086] Please see Figure 2 , Figure 2 This is a flowchart illustrating another IBD fragment prediction method provided in an embodiment of the present invention; as shown below. Figure 2 As shown, the method may include steps 201-209, as follows:
[0087] 201. Obtain multiple second mutations from the training set and the prediction set, wherein the prediction set includes SNP data of the at least two targets, the training set includes SNP data of a preset population and haplotype information of each person in the preset population, wherein the population type corresponding to the preset population is the same as the population type of the at least two targets, and the occurrence frequency of the second mutation is higher than the preset frequency.
[0088] This second variant can be understood as a common variant.
[0089] For details regarding the training and prediction sets, as well as the acquisition of multiple second mutations, please refer to [link to relevant documentation]. Figure 1 The description of step 102 in the illustrated embodiment will not be repeated here.
[0090] 202. Based on the multiple second variants in the training set and the prediction set, a DNA sequence is divided into multiple windows of a preset size, wherein each window of the multiple preset size contains the same number of second variants;
[0091] like Figure 3aAs shown, common variations in the entire DNA sequence are divided into adjacent windows of size ω, meaning each window contains ω common variations.
[0092] 203. Calculate the third probability of observing the DNA sequence of the at least two targets when the genotype of the at least two targets in the multiple windows is an IBD fragment, and the fourth probability of observing the DNA sequence of the at least two targets when the genotype of the at least two targets in the multiple windows is a non-IBD fragment, based on the genotype of the at least two targets in the multiple windows and the training set.
[0093] Wherein, assuming that the genotypes of the two targets within a window are IBD fragments, then the third probability (g(w), g′(w)|IBD) of observing the genotypes of these two targets can be expressed as:
[0094]
[0095] Where w represents any window, g(w) is the genotype of one target within this window, and g′(w) is the genotype of the other target within this window; h1 and h2 are haplotypes of one target, h2 and h3 are haplotypes of the other target, and h2 is the haplotype inherited from the same ancestor by both individuals; f w (h) represents the frequency of a certain monomer type in window w, which can be calculated based on the data in the training set.
[0096] The above
[0097] Where m is a common variation of window w, and It is one of the targets with two haplotypes on the common variant m, if but otherwise ε represents the probability of sequencing errors for common variants. Similarly, p′(g(w)|h2,h3) can be calculated.
[0098] Based on this, the third probability mentioned above can be obtained.
[0099] Assuming that the genotypes of two targets within a window are non-IBD fragments, then the fourth probability p(g(w), g′(w)|non-IBD) of observing the genotypes of these two targets can be expressed as:
[0100]
[0101] Among them, h1 and h2 are the singleton types of one target, h3 and h4 are the singleton types of the other target, and f w(h) represents the frequency of a certain monomer type in window w. Where m is a common variation of window w, and It is two haplotypes of a person on the common variant m, if but otherwise ε represents the probability of sequencing errors for common variants. Similarly, p′(g(w)|h3,h4) can be calculated. For a detailed explanation of this part, please refer to the above description of calculating the third probability; it will not be repeated here.
[0102] Based on this, the fourth probability can be obtained.
[0103] 204. Calculate the predicted value of the first IBD fragment between the DNA sequences of the at least two targets based on the third probability and the fourth probability; the length of the first IBD fragment is greater than a preset length;
[0104] The predicted value of the first IBD fragment between the DNA sequences of at least two targets can be understood as the predicted value of each window between the DNA sequences of at least two targets.
[0105] In the first possible implementation, the predicted value γ of the first IBD segment w (g(w),g′(w)) can be represented as:
[0106] γ w (g(w),g′(w))=log(p(g(w),g′(w)|IBD) / p(g(w),g′(w)|non-IBD)).
[0107] In the second possible implementation, a first value is calculated based on the third probability and the fourth probability. For example, the first value can be expressed as inner-LLR = log(p(g(w),g′(w)|IBD) / p(g(w),g′(w)|non-IBD)).
[0108] Optionally, if the absolute value of the inner-LLR value is greater than a preset value in certain windows (or the variance, etc., is greater than a preset value), then based on the training set and the preset window size, a second value is obtained indicating that at least two people in the training set share an IBD fragment within the same window, and a third value is obtained indicating that at least two people in the training set do not share an IBD fragment within the same window; based on the second value and the third value, the predicted value of the first IBD fragment between the DNA sequences of the at least two targets is calculated. It should be noted that the preset value can be any value, such as 0, or any other arbitrary value, and this solution does not impose any restrictions on this.
[0109] Specifically, the first value of window w, inner-LLR, will be treated as a random variable Γ. w Two empirical distributions were simulated using the training set. One empirical distribution consisted of 1000 pairs of inner-LLR values of IBD segments shared by each other in window w, and the other empirical distribution consisted of 1000 pairs of inner-LLR values of no IBD segments shared by each other in window w.
[0110] Define the model in the two empirical models above that contains IBD fragments in window w as Q. I The model without any IBD fragments is Thus, the predicted value of the first IBD fragment between the DNA sequences of at least two of the above targets can be expressed as:
[0111] In one possible implementation, based on the aforementioned window division, a sliding block of length L centimeters (cm) is established, where L is any positive number. Each block can contain a certain number of windows, such as... Figure 3b As shown. Assuming that each window is independent of the others, the predicted value λ of the first IBD segment within any block B can be obtained. B (g(B),g′(B))=∑ w∈B λ w (g(w),g′(w)).
[0112] 205. Obtain the plurality of first mutations of the training set and the prediction set;
[0113] For an introduction to this section, please refer to the aforementioned text. Figure 1 The description of step 102 in the illustrated embodiment will not be repeated here.
[0114] 206. Calculate the fifth probability based on the genotype of each of the at least two targets and the training set;
[0115] For an introduction to this section, please refer to the aforementioned text. Figure 1 The description of step 102 in the illustrated embodiment will not be repeated here.
[0116] 207. Based on the genotypes of the plurality of first variants corresponding to each target and the fifth probability, calculate the first probability of observing the DNA sequences of at least two targets when the genotype corresponding to the same first variant is an IBD fragment, and the second probability of observing the DNA sequences of at least two targets when the genotype corresponding to the same first variant is a non-IBD fragment.
[0117] For an introduction to this section, please refer to the aforementioned text. Figure 1 The description of step 102 in the illustrated embodiment will not be repeated here.
[0118] 208. Calculate the predicted value of a second IBD fragment containing the first variant between the DNA sequences of the at least two targets based on the first probability and the second probability, wherein the length of the second IBD fragment is less than the preset length;
[0119] For an introduction to this section, please refer to the aforementioned text. Figure 1 The description of step 103 in the illustrated embodiment will not be repeated here.
[0120] 209. Determine the IBD fragment between the DNA sequences of the at least two targets based on the predicted values of the first IBD fragment and the second IBD fragment between the DNA sequences of the at least two targets.
[0121] For an introduction to this section, please refer to the aforementioned text. Figure 1 The description of step 104 in the illustrated embodiment will not be repeated here.
[0122] In one possible implementation, the method further includes:
[0123] Based on the predicted values of the second IBD fragments corresponding to the plurality of first variants, non-IBD fragments between the DNA sequences of the at least two targets are determined.
[0124] For any first variant, if the predicted value of the second IBD fragment corresponding to the first variant is less than 0, the second IBD fragment is expanded until one or more of the following conditions are met to obtain a non-IBD fragment between the DNA sequences of the at least two targets:
[0125] The first IBD prediction value of the window corresponding to the non-IBD segment is greater than the fifth preset value; or...
[0126] The length of the non-IBD segment is greater than a sixth preset value. This sixth preset value could be, for example, 1 cM.
[0127] Then, the IBD fragment between the DNA sequences of the at least two targets is updated based on the non-IBD fragment between the DNA sequences of the at least two targets to obtain the updated IBD fragment between the DNA sequences of the at least two targets.
[0128] In other words, based on the predicted value of the second IBD fragment containing the first variant between the DNA sequences of the at least two targets obtained from the above calculation, if the predicted value is less than 0, the fragment will also be considered for extension until the above stopping condition is met, and the fragment will be predicted as a non-IBD fragment.
[0129] Based on the above-predicted long IBD segments, the overlapping IBD segments are merged, and based on the corresponding predicted non-IBD segments, the IBD segments are finally determined.
[0130] In this embodiment, a predicted value for a first IBD fragment (long IBD fragment) between the DNA sequences of at least two targets is obtained based on common variants. Then, based on multiple low-frequency variants and the predicted value of the first IBD fragment, a predicted value for a second IBD fragment (short IBD fragment) containing low-frequency variants between the DNA sequences of the at least two targets is calculated. Finally, the IBD fragment between the DNA sequences of the at least two targets is determined based on the predicted values of the first and second IBD fragments. This method, by introducing low-frequency variants, makes the prediction of short IBD fragments more accurate and reliable; simultaneously considering information from both common and low-frequency variants, the predicted IBD fragments are more accurate, solving the problem of insufficient ability to predict short IBD fragments in existing technologies.
[0131] The following example illustrates how to extract corresponding training data based on the ethnicity of a family of interest.
[0132] Optionally, single-type data can be obtained from the 1000 Genomes Project.
[0133] Then, based on this haplotype data, families and corresponding training sets were selected. The IDs of family members are NA19660, NA19661, NA19662, NA19684, NA19685, and NA19686. The family structure can be found by referring to [reference needed]. Figure 4 As shown. From the remaining people, those who are not related by blood and belong to the following groups (shown as codes) are selected, and the data corresponding to these people are used as the training set.
[0134] Then, SNPs are screened. Specifically, only SNPs containing two alleles are retained, and SNPs containing three or more alleles are removed; SNPs with a missing value greater than 10% are also removed; and SNPs with a frequency of 1 for a certain allele in the training set are also removed.
[0135] Identify common and low-frequency variants. Specifically, based on allele frequency, select SNPs with a frequency greater than or equal to 5% in the training set and define them as common variants. Based on allele frequency, select SNPs with a frequency greater than or equal to 0.5% and less than 1% in the training set and define them as low-frequency variants. Remove other SNPs based on these common and low-frequency variants.
[0136] Based on the above processing, the short and long IBD fragments between each pair of individuals in the family data are predicted using the IBD fragment prediction method provided in this scheme.
[0137] Then, optionally, the predictive ability of the above method is evaluated based on the IBD fragments predicted between NA19662 and NA19686, specifically as follows:
[0138] (1) Calculate the proportion of the predicted IBD fragment to the whole chromosome, take the average value in 22 autosomes, and then compare the average value with the expected IBD proportion.
[0139] (2) Calculate the predicted short IBD fragments inherited from NA19660 or NA19661 between NA19662 and NA19686, and compare them with IBD fragments identified by existing methods. Determine the proportion of short IBD fragments identified by this method that were found by existing methods or covered by long IBD fragments identified by existing methods. Further, it can be determined how many long IBD fragments identified using existing methods were found by this method. Here, a fragment being identified is defined as follows: in a fragment identified by method 1, method 2 finds at least 80% of that fragment; in this case, method 2 can be said to have found the fragment identified by method 1.
[0140] (3) Specifically, a short IBD fragment containing a low-frequency variant is given as an example to show the situation of haplotypes containing this low-frequency variant in families, thereby demonstrating the role of low-frequency variants in identifying short IBD fragments.
[0141] Based on the above analysis, it was found that the method proposed in this scheme can find many short IBD fragments that other methods cannot find, and the method proposed in this scheme can find most or all of the long IBD fragments that other methods can find, thus proving the superiority of the method proposed in this scheme.
[0142] Based on the description of the above-described IBD fragment prediction method embodiments, this invention also discloses an IBD fragment prediction device, with reference to... Figure 5 , Figure 5 This is a schematic diagram of an IBD fragment prediction device provided in an embodiment of the present invention. The IBD fragment prediction device includes an acquisition module 501, a prediction module 502, a calculation module 503, and a determination module 504; wherein:
[0143] The acquisition module 501 is used to acquire the predicted value of a first IBD fragment between the DNA sequences of at least two targets, wherein the length of the first IBD fragment is greater than a preset length;
[0144] The prediction module 502 is used to calculate, based on multiple first variants, a first probability of observing the DNA sequences of the at least two targets when the genotype corresponding to the same first variant is an IBD fragment, and a second probability of observing the DNA sequences of the at least two targets when the genotype corresponding to the same first variant is a non-IBD fragment, wherein the occurrence frequency of the first variant is less than a preset frequency.
[0145] The calculation module 503 is used to calculate, based on the first probability and the second probability, a predicted value of a second IBD fragment containing the first mutation between the DNA sequences of the at least two targets, wherein the length of the second IBD fragment is less than the preset length;
[0146] The determination module 504 is used to determine the IBD fragment between the DNA sequences of the at least two targets based on the predicted value of the first IBD fragment between the DNA sequences of the at least two targets and the predicted value of the second IBD fragment.
[0147] In one possible implementation, the acquisition module 501 is used for:
[0148] Obtain multiple second mutations from the training set and the prediction set. The prediction set includes SNP data of the at least two targets. The training set includes SNP data of a preset population and haplotype information of each person in the preset population. The population type corresponding to the preset population is the same as the population type of the at least two targets. The occurrence frequency of the second mutation is higher than the preset frequency.
[0149] Based on the multiple second variants in the training set and the prediction set, a DNA sequence is divided into multiple windows of a preset size, each of the multiple windows of the preset size containing the same number of second variants;
[0150] The third probability of observing the DNA sequence of the at least two targets when the genotype of the at least two targets in the multiple windows is an IBD fragment, and the fourth probability of observing the DNA sequence of the at least two targets when the genotype of the at least two targets in the multiple windows is a non-IBD fragment, are calculated based on the genotype of the at least two targets in the multiple windows and the training set, respectively.
[0151] The predicted value of the first IBD fragment between the DNA sequences of the at least two targets is calculated based on the third probability and the fourth probability.
[0152] In one possible implementation, if the first value is greater than a preset value, the first value is calculated based on the third probability and the fourth probability, and the acquisition module 501 is further configured to:
[0153] Based on the training set and the preset window size, a second value is obtained that at least two people in the training set share IBD segments within the same window, and a third value is obtained that at least two people in the training set do not have IBD segments within the same window;
[0154] The predicted value of the first IBD fragment between the DNA sequences of the at least two targets is calculated based on the second value and the third value.
[0155] In one possible implementation, the acquisition module 501 is further configured to:
[0156] Obtain the plurality of first mutations in the training set and the prediction set;
[0157] The prediction module 502 is also used for:
[0158] The haplotype of each of the at least two targets is calculated based on the genotypes of the at least two targets in the plurality of windows and the training set;
[0159] The number of major alleles and minor alleles in the plurality of first variants are determined in the training set and the prediction set, respectively, based on the SNP data of the training set and the SNP data of the prediction set.
[0160] The fifth probability is calculated based on the haplotype of each of the at least two targets and the number of major and minor alleles in the plurality of first variants in the training set and the prediction set, respectively.
[0161] Based on the genotypes of the plurality of first variants corresponding to each target and the fifth probability, calculate the first probability of observing the DNA sequences of at least two targets when the genotype corresponding to the same first variant is an IBD fragment, and the second probability of observing the DNA sequences of at least two targets when the genotype corresponding to the same first variant is a non-IBD fragment.
[0162] In one possible implementation, the determining module 504 is configured to:
[0163] Based on the plurality of first mutations and the plurality of windows of preset sizes, determine the window in which each of the plurality of first mutations is located;
[0164] Based on the predicted values of the first IBD fragments corresponding to the windows where the plurality of first variants are located and the predicted values of the second IBD fragments, a second IBD fragment between the DNA sequences of the at least two targets is determined; wherein, for any first variant, if the sum of the predicted value of the second IBD fragment corresponding to the first variant and the predicted value of the first IBD fragment corresponding to the window where the first variant is located is greater than a first preset value, the second IBD fragment corresponding to the first variant is expanded until one or more of the following conditions are met to obtain an updated second IBD fragment;
[0165] The predicted value of the first IBD segment in the window corresponding to the updated second IBD segment is less than a second preset value, and the second preset value is less than the first preset value; or...
[0166] The sum of the predicted value of the updated second IBD segment corresponding to the first mutation and the predicted value of the first IBD segment corresponding to the window where the first mutation is located is less than the first preset value; or,
[0167] The number and / or proportion of first windows in the window corresponding to the updated second IBD segment are greater than a third preset value, and the predicted value of the first IBD segment corresponding to the first window is less than 0; or,
[0168] The length of the updated second IBD segment is greater than the fourth preset value;
[0169] The IBD fragment between the DNA sequences of the at least two targets is determined based on a first IBD fragment between the DNA sequences of the at least two targets and an updated second IBD fragment between the DNA sequences of the at least two targets.
[0170] In one possible implementation, the device further includes an update module for:
[0171] Based on the predicted values of the second IBD fragments corresponding to the plurality of first variants, non-IBD fragments between the DNA sequences of the at least two targets are determined;
[0172] Wherein, for any first mutation, if the predicted value of the second IBD fragment corresponding to the first mutation is less than 0, the second IBD fragment is extended until one or more of the following conditions are met to obtain a non-IBD fragment between the DNA sequences of the at least two targets.
[0173] The first IBD prediction value of the window corresponding to the non-IBD segment is greater than the fifth preset value; or...
[0174] The length of the non-IBD segment is greater than the sixth preset value;
[0175] The IBD fragment between the DNA sequences of the at least two targets is updated based on the non-IBD fragment between the DNA sequences of the at least two targets to obtain the updated IBD fragment between the DNA sequences of the at least two targets.
[0176] It is worth noting that the specific functional implementation of the IBD fragment prediction device can be found in the description of the IBD fragment prediction method above, and will not be repeated here. The various units or modules in the IBD fragment prediction device can be individually or entirely merged into one or more other units or modules, or some of the units or modules can be further divided into multiple functionally smaller units or modules. This achieves the same operation without affecting the technical effects of the embodiments of the present invention. The aforementioned units or modules are based on logical functional division. In practical applications, the function of one unit (or module) can also be implemented by multiple units (or modules), or the function of multiple units (or modules) can be implemented by one unit (or module).
[0177] Based on the description of the above method and device embodiments, this invention also provides an IBD fragment prediction device.
[0178] Reference Figure 6 The diagram shown is a hardware structure schematic of another IBD fragment prediction device provided in an embodiment of this application. Figure 6 The IBD fragment prediction device 600 shown (which may specifically be a computer device) includes a memory 601, a processor 602, a communication interface 603, and a bus 604. The memory 601, processor 602, and communication interface 603 are interconnected via the bus 604.
[0179] The memory 601 may be a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM).
[0180] The memory 601 can store a program. When the program stored in the memory 601 is executed by the processor 602, the processor 602 and the communication interface 603 are used to execute the various steps of the IBD fragment prediction method of the present application embodiments.
[0181] Processor 602 is a circuit with signal processing capabilities. In one implementation, processor 602 can be a circuit with instruction reading and execution capabilities, such as a central processing unit (CPU), microprocessor, graphics processing unit (GPU) (which can be understood as a type of microprocessor), or digital signal processor (DSP). In another implementation, processor 602 can implement certain functions through the logical relationships of hardware circuits. These logical relationships of hardware circuits are fixed or reconfigurable. For example, processor 602 can be a hardware circuit implemented as an ASIC or a programmable logic device (PLD), such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document and configuring the hardware circuit can be understood as the process of the processor loading instructions to implement the functions of some or all of the above modules. Furthermore, it can also be a hardware circuit designed for artificial intelligence, which can be understood as a type of ASIC, such as a neural network processing unit (NPU), tensor processing unit (TPU), or deep learning processing unit (DPU). The processor 602 is used to execute related programs to implement the functions required by the units in the IBD fragment prediction device of the present application embodiments, or to execute the IBD fragment prediction method of the method embodiments of the present application.
[0182] As can be seen, each module in the above devices can be one or more processors (or processing circuits) configured to implement the above methods, such as: CPU, GPU, NPU, TPU, DPU, microprocessor, DSP, ASIC, FPGA, or a combination of at least two of these processor types.
[0183] Furthermore, the modules in the above devices can be integrated in whole or in part, or they can be implemented independently. In one implementation, these modules are integrated together as a system-on-a-chip (SOC). The SOC may include at least one processor for implementing any of the above methods or for implementing the functions of the modules in the device. The at least one processor can be of different types, such as CPU and FPGA, CPU and AI processor, CPU and GPU, etc.
[0184] The communication interface 603 uses transceiver devices, such as, but not limited to, transceivers, to enable communication between device 600 and other devices or communication networks. For example, data can be acquired through the communication interface 603.
[0185] Bus 604 may include a pathway for transmitting information between various components of device 600 (e.g., memory 601, processor 602, communication interface 603).
[0186] It should be noted that, although Figure 6 The illustrated device 600 only shows the memory, processor, and communication interface. However, those skilled in the art should understand that in specific implementations, device 600 may also include other components necessary for normal operation. Furthermore, depending on specific needs, those skilled in the art should understand that device 600 may also include hardware components for implementing other additional functions. Moreover, those skilled in the art should understand that device 600 may only include the components necessary for implementing the embodiments of this application, and may not necessarily include... Figure 6 All the devices shown.
[0187] This application also provides a computer-readable storage medium storing instructions that, when executed on a computer or processor, cause the computer or processor to perform one or more steps of any of the above methods.
[0188] This application also provides a computer program product containing instructions. When the computer program product is run on a computer or processor, it causes the computer or processor to perform one or more steps of any of the methods described above.
[0189] It should be understood that in the description of this application, unless otherwise stated, " / " indicates that the objects before and after it are in an "or" relationship. For example, A / B can represent A or B; where A and B can be singular or plural. Furthermore, in the description of this application, unless otherwise stated, "multiple" refers to two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple. Additionally, to facilitate a clear description of the technical solutions of the embodiments of this application, the terms "first" and "second" are used in the embodiments of this application to distinguish identical or similar items with substantially the same function and effect. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and the terms "first" and "second" do not necessarily imply difference. In this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" or "for example" in this application should not be construed as being better or more advantageous than other embodiments or designs. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner to facilitate understanding.
[0190] In the embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the division of units is merely a logical functional division, and in actual implementation, there may be other division methods. For instance, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. The coupling, direct coupling, or communication connection shown or discussed between each other may be through some interfaces, indirect coupling or communication connection of devices or units, and may be electrical, mechanical, or other forms.
[0191] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0192] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. This computer program product includes one or more computer instructions. When these computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in or transmitted through a computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available media can be read-only memory (ROM), random access memory (RAM), or magnetic media, such as floppy disks, hard disks, magnetic tapes, magnetic disks, or optical media, such as digital versatile discs (DVDs), or semiconductor media, such as solid-state disks (SSDs).
[0193] The above description is merely a specific implementation of the embodiments of this application, but the protection scope of the embodiments of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the embodiments of this application should be covered within the protection scope of the embodiments of this application. Therefore, the protection scope of the embodiments of this application should be determined by the protection scope of the claims.
Claims
1. A method for predicting IBD fragments, characterized in that, include: Obtain the predicted value of the first IBD fragment between the DNA sequences of at least two targets, wherein the length of the first IBD fragment is greater than a preset length; The step of obtaining the predicted value of the first IBD fragment between the DNA sequences of at least two targets includes: obtaining multiple second variants from a training set and a prediction set, wherein the prediction set includes SNP data of the at least two targets, and the training set includes SNP data of a preset population and haplotype information of each person in the preset population, wherein the population type corresponding to the preset population is the same as the population type of the at least two targets, and the occurrence frequency of the second variants is higher than a preset frequency; dividing the DNA sequence into multiple windows of a preset size based on the multiple second variants from the training set and the prediction set, wherein each window of the multiple preset size contains the same number of second variants; calculating a third probability of observing the DNA sequence of the at least two targets when the genotypes of the at least two targets in the multiple windows are IBD fragments, and a fourth probability of observing the DNA sequence of the at least two targets when the genotypes of the at least two targets in the multiple windows are non-IBD fragments, based on the genotypes of the at least two targets in the multiple windows and the training set; and calculating the predicted value of the first IBD fragment between the DNA sequences of the at least two targets based on the third probability and the fourth probability. The first probability of observing the DNA sequences of the at least two targets when the genotype corresponding to the same first variant is an IBD fragment, and the second probability of observing the DNA sequences of the at least two targets when the genotype corresponding to the same first variant is a non-IBD fragment, are calculated based on multiple first variants, wherein the occurrence frequency of the first variant is less than a preset frequency. The predicted value of a second IBD fragment containing the first variant is calculated based on the first probability and the second probability, wherein the length of the second IBD fragment is less than the preset length; The IBD fragment between the DNA sequences of the at least two targets is determined based on the predicted values of a first IBD fragment and a second IBD fragment between the DNA sequences of the at least two targets; the determination of the IBD fragment between the DNA sequences of the at least two targets based on the predicted values of a first IBD fragment and a second IBD fragment includes: determining the window in which each of the plurality of first variants is located based on the plurality of first variants and the plurality of windows of preset sizes; determining the second IBD fragment between the DNA sequences of the at least two targets based on the predicted values of the first IBD fragment and the second IBD fragment corresponding to the window in which the plurality of first variants are located; wherein, for any first variant, if the second IBD fragment corresponding to the first variant is... If the sum of the predicted value of the IBD fragment and the predicted value of the first IBD fragment corresponding to the window containing the first mutation is greater than a first preset value, the second IBD fragment corresponding to the first mutation is expanded until one or more of the following conditions are met to obtain an updated second IBD fragment: The predicted value of the first IBD fragment in the window corresponding to the updated second IBD fragment is less than a second preset value, and the second preset value is less than the first preset value; or, the sum of the predicted value of the updated second IBD fragment corresponding to the first mutation and the predicted value of the first IBD fragment corresponding to the window containing the first mutation is less than the first preset value; or, the number and / or proportion of first windows in the window corresponding to the updated second IBD fragment is greater than a third preset value, and the predicted value of the first IBD fragment corresponding to the first window is less than 0; or, the length of the updated second IBD fragment is greater than a fourth preset value; Based on the first IBD fragment between the DNA sequences of the at least two targets and the updated second IBD fragment between the DNA sequences of the at least two targets, the IBD fragment between the DNA sequences of the at least two targets is determined.
2. The method according to claim 1, characterized in that, If the first value is greater than a preset value, the first value is calculated based on the third probability and the fourth probability. The calculation of the predicted value of the first IBD fragment between the DNA sequences of the at least two targets based on the third probability and the fourth probability includes: Based on the training set and the preset window size, a second value is obtained that at least two people in the training set share IBD segments within the same window, and a third value is obtained that at least two people in the training set do not have IBD segments within the same window; The predicted value of the first IBD fragment between the DNA sequences of the at least two targets is calculated based on the second value and the third value.
3. The method according to claim 1 or 2, characterized in that, The method further includes: Obtain the plurality of first mutations in the training set and the prediction set; The calculation of a first probability of observing the DNA sequences of at least two targets when the genotype corresponding to the same first variant is an IBD fragment, and a second probability of observing the DNA sequences of at least two targets when the genotype corresponding to the same first variant is a non-IBD fragment, includes: The haplotype of each of the at least two targets is calculated based on the genotypes of the at least two targets in the plurality of windows and the training set; The number of major alleles and minor alleles in the plurality of first variants are determined in the training set and the prediction set, respectively, based on the SNP data of the training set and the SNP data of the prediction set. The fifth probability is calculated based on the haplotype of each of the at least two targets and the number of major and minor alleles in the plurality of first variants in the training set and the prediction set, respectively. Based on the genotypes of the plurality of first variants corresponding to each target and the fifth probability, calculate the first probability of observing the DNA sequences of at least two targets when the genotype corresponding to the same first variant is an IBD fragment, and the second probability of observing the DNA sequences of at least two targets when the genotype corresponding to the same first variant is a non-IBD fragment.
4. The method according to claim 1, characterized in that, The method further includes: Based on the predicted values of the second IBD fragments corresponding to the plurality of first variants, non-IBD fragments between the DNA sequences of the at least two targets are determined; Wherein, for any first mutation, if the predicted value of the second IBD fragment corresponding to the first mutation is less than 0, the second IBD fragment is extended until one or more of the following conditions are met to obtain a non-IBD fragment between the DNA sequences of the at least two targets. The first IBD prediction value of the window corresponding to the non-IBD segment is greater than the fifth preset value; or... The length of the non-IBD segment is greater than the sixth preset value; The IBD fragment between the DNA sequences of the at least two targets is updated based on the non-IBD fragment between the DNA sequences of the at least two targets to obtain the updated IBD fragment between the DNA sequences of the at least two targets.
5. An IBD fragment prediction device, characterized in that, include: The acquisition module is used to acquire the predicted value of a first IBD fragment between the DNA sequences of at least two targets, wherein the length of the first IBD fragment is greater than a preset length; The acquisition module is used to acquire multiple second variants from a training set and a prediction set. The prediction set includes SNP data of the at least two targets, and the training set includes SNP data of a preset population and haplotype information of each person in the preset population. The population type corresponding to the preset population is the same as the population type of the at least two targets, and the occurrence frequency of the second variants is higher than a preset frequency. Based on the multiple second variants from the training set and the prediction set, the DNA sequence is divided into multiple windows of a preset size. Each window contains the same number of second variants. Based on the genotypes of the at least two targets in the multiple windows and the training set, a third probability of observing the DNA sequence of the at least two targets is calculated when the genotype of the at least two targets in the corresponding windows is an IBD fragment, and a fourth probability of observing the DNA sequence of the at least two targets is calculated when the genotype of the at least two targets in the corresponding windows is a non-IBD fragment. Based on the third probability and the fourth probability, a predicted value of the first IBD fragment between the DNA sequences of the at least two targets is calculated. The prediction module is used to calculate, based on multiple first variants, a first probability of observing the DNA sequences of the at least two targets when the genotype corresponding to the same first variant is an IBD fragment, and a second probability of observing the DNA sequences of the at least two targets when the genotype corresponding to the same first variant is a non-IBD fragment, wherein the occurrence frequency of the first variant is less than a preset frequency. The calculation module is used to calculate, based on the first probability and the second probability, a predicted value of a second IBD fragment containing the first mutation between the DNA sequences of the at least two targets, wherein the length of the second IBD fragment is less than the preset length; A determining module is configured to determine an IBD fragment between the DNA sequences of the at least two targets based on the predicted values of a first IBD fragment and a second IBD fragment between the DNA sequences of the at least two targets. The determining module is configured to determine the window containing each of the plurality of first variants based on the plurality of first variants and the plurality of windows of preset sizes; and to determine the second IBD fragment between the DNA sequences of the at least two targets based on the predicted values of the first IBD fragment and the second IBD fragment corresponding to the windows containing the plurality of first variants. Wherein, for any first variant, if the predicted value of the second IBD fragment corresponding to the first variant and the predicted value of the first IBD fragment corresponding to the window containing the first variant are... If the sum of the predicted values of the D fragments is greater than a first preset value, the second IBD fragment corresponding to the first mutation is expanded until one or more of the following conditions are met to obtain an updated second IBD fragment: The predicted value of the first IBD fragment in the window corresponding to the updated second IBD fragment is less than a second preset value, and the second preset value is less than the first preset value; or, the sum of the predicted value of the updated second IBD fragment corresponding to the first mutation and the predicted value of the first IBD fragment corresponding to the window containing the first mutation is less than the first preset value; or, the number and / or proportion of first windows in the window corresponding to the updated second IBD fragment is greater than a third preset value, and the predicted value of the first IBD fragment corresponding to the first window is less than 0; or, the length of the updated second IBD fragment is greater than a fourth preset value; Based on the first IBD fragment between the DNA sequences of the at least two targets and the updated second IBD fragment between the DNA sequences of the at least two targets, the IBD fragment between the DNA sequences of the at least two targets is determined.
6. The apparatus according to claim 5, characterized in that, The acquisition module is also used for: Obtain the plurality of first mutations in the training set and the prediction set; The prediction module is also used for: The haplotype of each of the at least two targets is calculated based on the genotypes of the at least two targets in the plurality of windows and the training set; The number of major alleles and minor alleles in the plurality of first variants are determined in the training set and the prediction set, respectively, based on the SNP data of the training set and the SNP data of the prediction set. The fifth probability is calculated based on the haplotype of each of the at least two targets and the number of major and minor alleles in the plurality of first variants in the training set and the prediction set, respectively. Based on the genotypes of the plurality of first variants corresponding to each target and the fifth probability, calculate the first probability of observing the DNA sequences of at least two targets when the genotype corresponding to the same first variant is an IBD fragment, and the second probability of observing the DNA sequences of at least two targets when the genotype corresponding to the same first variant is a non-IBD fragment.
7. An IBD fragment prediction device, characterized in that, include: Processor and memory; The processor is connected to a memory, wherein the memory is used to store program code, and the processor is used to call the program code to execute the IBD fragment prediction method as described in any one of claims 1-4.
8. A computer storage medium, characterized in that, The computer storage medium stores a computer program, which includes program instructions that, when executed by a processor, perform the IBD fragment prediction method as described in any one of claims 1-4.
Citation Information
Patent Citations
Community assignments in identity by descent networks and genetic variant origination
CN112154508A
Group genealogy construction method and device, equipment and storage medium
CN114974414A