A method and device for locating the Alpha thalassemia SEA mutation chain
By using population-level data to construct the ancestral haplotype of the Alpha Thalassium SEA mutation chain and perform similarity analysis, the problem of the Alpha Thalassium SEA mutation chain in families with incomplete probands and family lines is solved, and the detection effect with high accuracy is achieved.
Patent Information
- Application Number
- CN202510243623.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-03-03
AI Technical Summary
It is difficult to identify the Alpha Thalassium SEA mutation chain of families with probands and incomplete families through family chain analysis.
The population-level data were used to obtain the ancestral haplotype of Alpha Thalassium SEA mutation chain. By conducting similarity analysis with the haplotype of the sample to be examined, the Alpha Thalassium SEA mutation chain and the normal chain were distinguished.
It realizes the accurate identification of Alpha Thalassium SEA mutation chain without relying on family chain analysis, which improves the accuracy and efficiency of detection.
Smart Images

Figure CN119719808B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of gene detection technology, and in particular to a method and device for detecting an Alpha thalassemia (SEA) mutation chain. Background Art
[0002] Depending on the type of globin gene defect, thalassemia is mainly divided into Alpha thalassemia and Beta thalassemia. Alpha thalassemia is mainly caused by point mutation or copy number variation of the alpha globin gene. 95% of alpha thalassemia patients are caused by large deletions of the alpha globin gene, that is, deletion-type alpha thalassemia. At present, at least 36 types of alpha thalassemia deletions have been identified, among which Alpha thalassemia SEA is a very common one.
[0003] Thalassemia is a single gene genetic disease. When both husband and wife are carriers of the same type of thalassemia gene, there is a 1 / 4 probability that their children will be patients with moderate to severe thalassemia, a 1 / 2 probability that they will be carriers of the thalassemia gene like their parents, and a 1 / 4 probability that the fetus will be a normal fetus.
[0004] Therefore, it is particularly important to discover the Alpha-thalassemia SEA mutation chain. The current conventional method requires samples from at least two generations, that is, samples from the proband and the couple, or samples from the couple and parents, a total of six people, if there is no proband. For families with incomplete probands and incomplete pedigrees, it is impossible to identify the Alpha-thalassemia SEA mutation chain of male or female carriers through family linkage analysis. Summary of the invention
[0005] In order to solve the above technical problems, the present application uses population-level data to obtain the ancestral haplotype linked to the Alpha thalassemia SEA mutation. The haplotype of the sample to be tested can distinguish the Alpha thalassemia SEA mutation chain from the normal chain by comparing the haplotype with the ancestral haplotype.
[0006] The specific technical solutions are as follows:
[0007] A method for locating the Alpha thalassemia SEA mutation chain, comprising the following steps:
[0008] S1, collect the whole genome chip test data of family members of multiple families with Alpha thalassemia SEA mutation, and determine the Alpha thalassemia SEA mutation chain and normal chain of each family member;
[0009] S2, calculate the P value of the difference in base usage frequency between all Alpha-thalassemia SEA mutant chains and all normal chains for each SNP site;
[0010] S3, screening SNP sites according to the P value in S2, and forming the ancestral haplotype of the Alpha thalassemia SEA mutation chain according to the screened SNP sites;
[0011] S4, forming a similarity matrix, forming a similarity matrix according to the proportion of multiple bases of the SNP sites of the ancestral haplotype in all Alpha thalassemia SEA mutation chains;
[0012] S5, after obtaining the sample to be tested, calculating the similarity value between the haplotype of each chain of the sample to be tested and the ancestral haplotype according to the similarity matrix;
[0013] S6, according to the similarity value corresponding to each chain, determine the Alpha thalassemia SEA mutation chain from the two chains of the sample to be tested; if the similarity value corresponding to one chain of the sample to be tested is greater than the first threshold, and the difference between the similarity value of the chain and the other chain of the sample to be tested is greater than the second threshold, determine that the chain is the Alpha thalassemia SEA mutation chain.
[0014] Furthermore, in S1, the haplotype phasing method is used to construct the haplotype of the chromosome where the Alpha thalassemia SEA mutation of any family is located, and the Alpha thalassemia SEA mutation chain and normal chain of each family member are determined by family linkage analysis.
[0015] Furthermore, the P value in S2 is obtained according to the following method:
[0016] S21, summarize all Alpha thalassemia SEA mutation chains to obtain the Alpha thalassemia SEA mutation chain set, and count the base usage frequency P1 of each SNP site in the Alpha thalassemia SEA mutation chain set;
[0017] S22, summing up all normal chains to obtain a normal chain set, and counting the base usage frequency P2 of each SNP site in the normal chain set;
[0018] S23, for each SNP site, calculate the P value of the difference in base usage frequency between the Alpha thalassemia SEA mutant chain set and the normal chain set, P=|P1-P2|.
[0019] Furthermore, the steps of screening SNP sites in S3 are as follows:
[0020] S31, arrange the SNP sites according to the P value from small to large;
[0021] S32, obtaining a first SNP site set formed by SNP sites whose P values are less than a first P value threshold; the first P value threshold is adjusted according to the number requirement of the screened SNP sites;
[0022] S33, deleting the SNP sites that are farther away from the HBA1 gene than the first distance from the first SNP site set to obtain a second SNP site set;
[0023] S34, deleting any one of the two SNP sites whose distance is less than the second distance from the second SNP site set to obtain a screened effective SNP site.
[0024] Furthermore, the similarity calculation method of the samples to be tested in S5 is as follows:
[0025] S51, obtaining all SNP sites of the haplotype of each chain of the sample to be tested according to the S3 method, wherein the number of effective SNP sites among all SNP sites is greater than or equal to 6;
[0026] S52, obtaining the ratio of the bases at the effective SNP site of the haplotype of the chain to the corresponding bases at the corresponding SNP site in the similarity matrix, and using the ratio as the similarity value of each SNP site;
[0027] S53, taking the ratio of the sum of the similarity values of the effective SNP sites of the haplotype of the chain to the number of all SNP sites as the similarity value between the haplotype of the chain and the ancestral haplotype.
[0028] Furthermore, the first threshold is the average value of the similarity values corresponding to all Alpha thalassemia SEA mutation chains; the second threshold is the difference between the average value of the similarity values corresponding to all normal chains and the first threshold.
[0029] A detection device for locating an Alpha thalassemia SEA mutation chain, the device comprising: a collection module, a P value calculation module, an ancestral haplotype formation module, a similarity matrix formation module, a similarity analysis module of a sample to be tested, and a mutation chain determination module;
[0030] The acquisition module includes a module for collecting whole genome chip detection data of family members of multiple families containing Alpha thalassemia SEA mutations, and determining the Alpha thalassemia SEA mutation chain and normal chain of each family member;
[0031] The P value calculation module is used to calculate the P value of the difference in base usage frequency between all Alpha thalassemia SEA mutant chains and all normal chains for each SNP site; the P value is obtained by the following steps:
[0032] (1) Summarize all Alpha thalassemia SEA mutation chains to obtain the Alpha thalassemia SEA mutation chain set, and count the base usage frequency P1 of each SNP site in the Alpha thalassemia SEA mutation chain set;
[0033] (2) Summarize all normal chains to obtain a normal chain set, and count the base usage frequency P2 of each SNP site in the normal chain set;
[0034] (3) For each SNP site, calculate the P value of the difference in base usage frequency between the Alpha thalassemia SEA mutant chain set and the normal chain set, P = | P1-P2|;
[0035] The ancestral haplotype forming module is used to screen the SNP sites according to the P value corresponding to each SNP site, and form the ancestral haplotype of the Alpha thalassemia SEA mutation chain according to the screened SNP sites; the specific rules are as follows:
[0036] 1) Arrange the SNP sites from small to large according to the P value;
[0037] 2) obtaining a first SNP site set formed by SNP sites having a P value less than a first P value threshold;
[0038] 3) deleting the SNP sites whose distance from the HBA1 gene exceeds the first distance from the first SNP site set to obtain a second SNP site set;
[0039] 4) deleting one of the two SNP sites whose distance is less than the second distance from the second SNP site set to obtain the screened SNP site;
[0040] The similarity matrix forming module is used to form a similarity matrix according to the proportion of multiple bases of the SNP sites of the ancestral haplotype in all Alpha thalassemia SEA mutation chains;
[0041] The similarity analysis module of the sample to be tested is used to calculate the similarity value between the haplotype of each chain of the sample to be tested and the ancestral haplotype according to the similarity matrix after obtaining the sample to be tested; the specific calculation rules are as follows:
[0042] i. Obtain all SNP sites of the haplotype of each chain of the sample to be tested according to the S3 method, wherein the number of effective SNP sites among all SNP sites is greater than or equal to 6;
[0043] ii. Obtain the ratio of the bases at the effective SNP sites of the haplotype of the chain to the corresponding bases at the corresponding SNP sites in the similarity matrix, and use the ratio as the similarity value of each SNP site;
[0044] iii. The ratio of the sum of the similarity values of the effective SNP sites of the haplotype of the chain to the number of all SNP sites is taken as the similarity value between the haplotype of the chain and the ancestral haplotype.
[0045] The mutant chain determination module is used to determine the Alpha thalassemia SEA mutant chain from the two chains of the sample to be tested according to the similarity value corresponding to each chain.
[0046] A computer device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements any one of the methods of claims 1 to 6 when executing the computer program.
[0047] A computer-readable storage medium stores a computer program, wherein the computer program implements any one of the methods of claims 1 to 6 when executed by a processor.
[0048] A computer program product, comprising a computer program, wherein the computer program implements any one of the methods of claims 1 to 6 when executed by a processor.
[0049] The working principle of the present invention is as follows: First, the acquisition module collects the whole genome chip detection data of multiple family members containing Alpha thalassemia SEA mutations, and uses the haplotype phasing method to construct the haplotype of the chromosome where the Alpha thalassemia SEA mutation is located in the family, and then determines the mutation chain and normal chain of each member through family linkage analysis. Next, the P value calculation module summarizes all mutation chains and normal chains, and respectively counts the base usage frequencies P1 and P2 of each SNP site therein, and calculates the P value of the difference in base usage frequencies between the two. The ancestral haplotype formation module arranges the SNP sites from small to large according to the P value, obtains the effective SNP sites through a series of screening conditions (such as the distance from the HBA1 gene, etc.), and then constructs the ancestral haplotype of the Alpha thalassemia SEA mutation chain. The similarity matrix formation module generates a similarity matrix based on the proportion of multiple bases in the SNP sites of the ancestral haplotype. After obtaining the sample to be tested, the similarity analysis module of the sample to be tested calculates the similarity value of each chain haplotype with the ancestral haplotype according to the similarity matrix and specific rules. Finally, the mutant chain determination module compares the similarity values of each chain. If the similarity value of a chain is greater than the average value of all mutant chains (the first threshold), and the difference in similarity with another chain is greater than the difference between the average value of the normal chain and the first threshold (the second threshold), then this chain is determined to be an Alpha thalassemia SEA mutant chain. The entire process is achieved through the collaboration of various modules of the supporting detection device, or by running corresponding programs in computer equipment, storage media, and program products to achieve precise positioning.
[0050] Compared with the prior art, the present invention has the beneficial effect that the method can be used for the detection of Alpha thalassemia SEA mutation chains when male or female carriers cannot be identified through family linkage analysis due to incomplete probands and incomplete families. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 This is a flow chart of the method for detecting the Alpha thalassemia SEA mutation chain in an embodiment of the present invention;
[0052] Figure 2 This is a schematic diagram of the Alpha thalassemia SEA mutation chain of family 1 in an embodiment of the present invention;
[0053] Figure 3 This is a schematic diagram of the Alpha thalassemia SEA mutation chain of family 2 in an embodiment of the present invention;
[0054] Figure 4 This is a schematic diagram of determining the Alpha thalassemia SEA mutation chain of the female through embryo mutual inference after adding genomic data of two embryos to family 2 in an embodiment of the present invention;
[0055] Figure 5 This is a schematic diagram of the structure of the Alpha thalassemia SEA mutation chain detection device in an embodiment of the present invention;
[0056] Figure 6 Schematic diagram of a computer device in an embodiment of the present invention. DETAILED DESCRIPTION
[0057] In order to make the purpose, technical scheme and advantages of the present invention clearer, the present invention is further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings. It should be understood that these descriptions are only exemplary and are not intended to limit the scope of the present invention. In addition, in the following invention, the description of well-known structures and technologies is omitted to avoid unnecessary confusion of the concept of the present invention.
[0058] The Alpha thalassemia SEA mutation may have a single origin. In addition, studies on Alpha thalassemia SEA mutation carriers have shown that there is a haplotype composed of 5 SNPs (rs3760053, rs1211375, rs3918352, rs1203974, rs11248914). The SEA mutations of most carriers are linked to this haplotype. Therefore, finding the ancestral haplotype linked to the Alpha thalassemia SEA mutation at the population level is an effective solution to assist in identifying the Alpha thalassemia SEA mutation site, especially when the proband is incomplete and the family is incomplete and the Alpha thalassemia SEA mutation chain cannot be determined through family linkage analysis.
[0059] Therefore, the embodiment of the present invention proposes to use population-level data to obtain the ancestral haplotype linked to the Alpha thalassemia SEA mutation, and the haplotype of the sample to be tested can distinguish the Alpha thalassemia SEA mutation chain and the normal chain by comparing the haplotype with the ancestral haplotype.
[0060] Figure 1This is a flow chart of the method for detecting the Alpha thalassemia SEA mutation chain in an embodiment of the present invention, comprising:
[0061] S1, collect the whole genome chip test data of family members of multiple families with Alpha thalassemia SEA mutation, and determine the Alpha thalassemia SEA mutation chain and normal chain of each family member;
[0062] S2, calculate the P value of the difference in base usage frequency between all Alpha thalassemia SEA mutant chains and all normal chains for each SNP site (variation of a single nucleotide in the genome);
[0063] S3, based on the P value corresponding to each SNP site, the SNP site is screened, and based on the screened SNP site, the ancestral haplotype of the Alpha thalassemia SEA mutation chain is formed;
[0064] S4, forming a similarity matrix based on the proportion of multiple bases of the SNP sites of the ancestral haplotype in all Alpha thalassemia SEA mutation chains;
[0065] S5, after obtaining the sample to be tested, calculating the similarity value between the haplotype of each chain of the sample to be tested and the ancestral haplotype according to the similarity matrix;
[0066] S6, determining the Alpha thalassemia SEA mutation chain from the two chains of the sample to be tested according to the similarity value corresponding to each chain.
[0067] In the embodiment of the present invention, there is no need to perform mutation site detection and SNP linkage typing on samples of at least two generations (the proband or the parents of a couple), and the whole genome chip detection data of family members of multiple families are directly collected to form the ancestral haplotype of the Alpha thalassemia SEA mutation chain. After the similarity matrix is formed, the similarity value between the haplotype of each chain of the sample to be tested and the ancestral haplotype is calculated according to the similarity matrix, and then the Alpha thalassemia SEA mutation chain is determined from the two chains of the sample to be tested, with high accuracy.
[0068] Each step is described in detail below.
[0069] In S1, the whole genome chip detection data of family members of multiple families with Alpha-thalassemia SEA mutations are collected, and the Alpha-thalassemia SEA mutation chain and normal chain of each family member are determined;
[0070] The whole genome chip detection data of family members of multiple families reflect population data. The embodiment of the present invention uses population-level data to obtain ancestral haplotypes linked to mutations, wherein the more population data there is, the higher the accuracy of the obtained ancestral haplotypes.
[0071] In one embodiment, determining the Alpha thalassemia SEA mutant chain and normal chain of each family member comprises:
[0072] The haplotype phasing method was used to construct the haplotype of the chromosome where the Alpha-thalassemia SEA mutation was located in each family, and the Alpha-thalassemia SEA mutation chain and normal chain of each family member were determined through family linkage analysis.
[0073] Specifically, the method of haplotype phasing is very effective in constructing haplotypes. The basic principle of the method of haplotype phasing is as follows:
[0074] (1) obtaining a first data set from a child object, a second data set from the father of the child object, a third data set from the mother of the child object, and at least one reference data set from a reference object;
[0075] (2) Select multiple target sites, analyze and detect molecular markers in their upstream and downstream regions, and determine at least one molecular marker for each target site in its upstream and downstream regions;
[0076] (3) In each data set, each molecular marker is annotated to obtain a first data set, a second data set, a third data set and a reference data set that are annotated with the molecular markers;
[0077] (4) constructing a binary genetic vector for each molecular marker site upstream and downstream of each target site in each data set;
[0078] (5) For each target site, the maximum likelihood estimate L is determined using the hidden Markov model, and the maximum possible composition is estimated through the Viterbi dynamic programming algorithm to determine the haplotype of the offspring object.
[0079] In S2, the P value of the difference in base usage frequency between all Alpha-thalassemia SEA mutant chains and all normal chains was calculated for each SNP site;
[0080] In one embodiment, the P value of the difference in base usage frequency between all Alpha thalassemia SEA mutant chains and all normal chains for each SNP site is calculated, including:
[0081] Summarize all Alpha thalassemia SEA mutation chains to obtain the Alpha thalassemia SEA mutation chain set, and count the base usage frequency of each SNP site in the Alpha thalassemia SEA mutation chain set;
[0082] Summarize all normal chains to obtain a normal chain set, and count the base usage frequency of each SNP site in the normal chain set;
[0083] For each SNP site, the P value of the difference in base usage frequency between the Alpha thalassemia SEA mutant chain set and the normal chain set was calculated.
[0084] Specifically, Fisher's exact experience can be used to calculate the P value of the difference in base usage frequency between the Alpha thalassemia SEA mutant chain set and the normal chain set. The smaller the P value, the greater the difference in the usage frequency of the SNP site between the Alpha thalassemia SEA mutant chain and the normal chain.
[0085] In S3, the SNP sites are screened according to the P value corresponding to each SNP site, and the ancestral haplotype of the Alpha thalassemia SEA mutation chain is formed according to the screened SNP sites;
[0086] In one embodiment, screening SNP sites according to the P value corresponding to each SNP site includes:
[0087] Arrange the SNP sites according to the P value from small to large;
[0088] Obtaining a first SNP site set formed by SNP sites having a P value less than a first P value threshold, wherein the first P value threshold is adjusted according to the number requirement of the screened SNP sites;
[0089] Deleting SNP sites that are farther away from the HBA1 gene than a first distance from the first SNP site set to obtain a second SNP site set;
[0090] One of the two SNP sites whose distance is less than the second distance is deleted from the second SNP site set to obtain the screened SNP site.
[0091] Specifically, the first P value threshold, the first distance and the second distance can be determined according to actual conditions. For example, after arranging the SNP sites according to the P value from small to large, it is determined that the first P value threshold is 1E-11, and there are 21 SNP sites with a P value less than 1E-11; among the 21 SNP sites, 6 SNP sites are more than the first distance of 300kb from the HBA1 gene, indicating that these 6 SNP sites are far away, and these 6 SNP sites are discarded; among the 21 SNP sites, one SNP site is very close to another site by 36bp (second distance), and this SNP site is discarded; finally, 14 SNP sites are screened to form the ancestral haplotype of the Alpha thalassemia SEA mutation chain. Table 1 is an example of 14 SNP sites of the ancestral haplotype.
[0092] SNP P-value Ancestral haplotype snp01 1.16E-62 A snp02 7.97E-58 T snp03 3.63E-39 C snp04 1.67E-34 A snp05 6.04E-25 A snp06 1.00E-23 T snp07 6.12E-21 A snp08 3.65E-20 T snp09 3.95E-17 C snp10 6.85E-17 C snp11 1.34E-14 C snp12 3.12E-13 G snp13 1.57E-12 C snp14 1.18E-11 A .
[0093] In S4, a similarity matrix is formed according to the proportion of multiple bases of the SNP sites of the ancestral haplotype in all Alpha thalassemia SEA mutation chains;
[0094] Specifically, for each SNP site, the proportion of each base in the Alpha thalassemia SEA mutation chain was counted. Taking the 14 SNP sites of the aforementioned ancestral haplotype as an example, the similarity matrix formed is shown in Table 2. For example, in snp07, the ratio of the chain containing base A to all Alpha thalassemia SEA mutation chains is 0.994, so the proportion of base A is 0.994, and the others are similar.
[0095] SNP A T C G snp07 0.994 0 0.006 0 snp08 0 0.994 0.006 0 snp02 0 0.959 0.041 0 snp01 0.859 0 0.141 0 snp11 0 0.029 0.971 0 snp12 0.075 0 0 0.925 snp14 0.938 0 0 0.062 snp13 0 0 1 0 snp04 0.858 0 0 0.142 snp03 0 0.054 0.946 0 snp05 0.854 0 0.146 0 snp06 0 0.861 0.139 0 snp10 0 0.078 0.921 0 snp09 0 0.013 0.987 0
[0096] In S5, after obtaining the sample to be tested, the similarity value between the haplotype of each chain of the sample to be tested and the ancestral haplotype is calculated according to the similarity matrix;
[0097] In one embodiment, according to the similarity matrix, calculating the similarity value between the haplotype of each chain of the sample to be tested and the ancestral haplotype comprises:
[0098] For each chain of the sample to be tested, all SNP sites of the haplotype of the chain are obtained;
[0099] Obtain the ratio of the bases at each SNP site of the haplotype of the chain to the corresponding bases at the corresponding SNP site in the similarity matrix, and use the ratio as the similarity value of each SNP site;
[0100] The ratio of the sum of the similarity values of all SNP sites of the haplotype of the chain to the number of all SNP sites is taken as the similarity value between the haplotype of the chain and the ancestral haplotype.
[0101] Specifically, multiple SNP sites of the haplotype of each chain of the sample to be tested need to be calculated according to the similarity matrix to evaluate the similarity with the ancestral haplotype. The method for obtaining the haplotype of each chain of the sample to be tested adopts the aforementioned haplotype phasing method.
[0102] If a certain SNP site of a certain chain of the sample to be tested is not detected, the SNP site is not a valid site, and the similarity value of the site is 0. Therefore, taking the above Table 2 as an example, the number of SNP sites of the haplotype of each chain of the sample to be tested may be less than 14, that is, the valid sites are less than 14, and then the sum of the similarity values of all SNP sites of the haplotype of the chain obtained is less than 14. In order to ensure the accuracy of the similarity score, the embodiment of the present invention requires that the number of valid SNP sites is greater than or equal to 6.
[0103] The similarity value between the haplotype of this chain and the ancestral haplotype = the sum of the similarity values of all effective SNP sites / the ratio of the number of all effective SNP sites.
[0104] For example, if snp07, one of the SNP sites of the haplotype of a chain of the sample to be tested, exists, then snp07 is a valid site, wherein, referring to Table 2, the proportion of base A in snp07 is 0.994, then the similarity value of snp07 is 0.994, and the similarity values of all other valid SNP sites are obtained in the same way. The sum of the similarity values of these valid SNP sites is divided by the number of all valid SNP sites to obtain the similarity value of the chain with the ancestral haplotype.
[0105] In S6, the Alpha thalassemia SEA mutation chain is determined from the two chains of the sample to be tested according to the similarity value corresponding to each chain.
[0106] In one embodiment, the method further comprises:
[0107] Calculate the similarity value between the haplotype of each Alpha thalassemia SEA mutation chain in the Alpha thalassemia SEA mutation chain set and the ancestral haplotype;
[0108] The average of the similarity values corresponding to all Alpha thalassemia SEA mutation chains is used as the first threshold;
[0109] Calculate the similarity value between the haplotype of each normal chain in the normal chain set and the ancestral haplotype;
[0110] The difference between the average value of the similarity values corresponding to all normal chains and the first threshold is calculated, and the difference is used as the second threshold.
[0111] Specifically, taking Table 2 as an example, the first threshold is calculated to be 0.5, and the second threshold is calculated to be 0.1.
[0112] In one embodiment, according to the similarity value corresponding to each chain, determining the Alpha thalassemia SEA mutation chain from the two chains of the sample to be tested includes:
[0113] If the similarity value corresponding to one chain of the sample to be tested is greater than the first threshold, and the difference between the similarity values of the chain and another chain of the sample to be tested is greater than the second threshold, the chain is determined to be the Alpha thalassemia SEA mutation chain.
[0114] The above gives the steps of detecting the Alpha thalassemia SEA mutation chain. In order to verify the effectiveness of the method of the present invention, group data can also be collected to verify the method of the present invention, that is, the whole genome chip detection data of family members of multiple families containing Alpha thalassemia SEA mutations are collected as the verification set, and the two chains (known Alpha thalassemia SEA mutation chain and normal chain) of the 120 SEA mutation carriers in the verification set are all analyzed for similarity with the ancestral haplotype. Referring to the two threshold conditions (first threshold and second threshold) for judging whether a chain of the sample to be tested is an Alpha thalassemia SEA mutation chain, see Table 3, there are 8 cases that do not meet the threshold conditions, and among the 112 cases that meet the threshold conditions, 111 cases have Alpha thalassemia SEA mutation chains identified by similarity analysis and the true Alpha thalassemia SEA mutation chain results are consistent.
[0115] Number of carrier samples Accuracy Error rate Threshold condition not met 8 \ \ Meet the threshold condition 112 99.11% 0.89% .
[0116] It can be seen that the present invention finds the ancestral haplotype linked to the Alpha thalassemia SEA mutation at the population level, and uses the similarity matrix method to accurately identify the Alpha thalassemia SEA mutation chain in the SEA carriers of the validation set without relying on family linkage analysis, with an accuracy rate of 99.11%.
[0117] Several specific embodiments are given below to illustrate the effectiveness of the method proposed in the embodiment of the present invention.
[0118] Example 1: Alpha-thalassemia SEA mutation carrier family 1
[0119] All members of family 1 underwent whole genome typing using the Illumina ASA chip. The male and female members of the family were SEA mutation carriers. Figure 2 This is a schematic diagram of the Alpha thalassemia SEA mutation chain of family 1 in an embodiment of the present invention, and a schematic diagram of the Alpha thalassemia SEA mutation chain of the male and the Alpha thalassemia SEA mutation chain of the female determined by family linkage analysis.
[0120] Using the ancestral haplotype similarity matrix, the results of family 1 are shown in Table 4. The similarity value between the haplotype of the male_P chain and the ancestral haplotype is 0.92, and the similarity value between the haplotype of the male_M chain and the ancestral haplotype is 0.48. The male_P chain is determined to be the Alpha thalassemia SEA mutation chain, which is consistent with the Alpha thalassemia SEA mutation chain determined by family linkage analysis.
[0121] As shown in Table 4, using the ancestral haplotype similarity calculation, the similarity value between the haplotype of the female_M chain and the ancestral haplotype is 0.92, and the similarity value between the haplotype of the female_P chain and the ancestral haplotype is 0.33. It is determined that the female_M chain is the Alpha thalassemia SEA mutation chain, which is consistent with the results of the Alpha thalassemia SEA mutation chain determined by the family linkage analysis in Table 4.
[0122] Ancestral haplotype Male_P Male_M Female_P Female_M snp12 G G A A G snp13 C C T T C snp07 A A A A A snp08 T T T T T snp10 C C C T C snp11 C C C C C snp03 C C T T C snp05 A A C C A snp04 A A G G A snp01 A A C C A snp06 T T T C T snp02 T - - - - snp14 A - - - - snp09 C - - - - Similarity value 0.92 0.48 0.33 0.92 .
[0123] Example 2: Alpha-thalassemia SEA mutation carrier family 2
[0124] All members of the family were tested for whole genome typing using the Illumina ASA chip. The male and female members of the family were SEA mutation carriers. Figure 3 This is a schematic diagram of the Alpha thalassemia SEA mutation chain of family 2 in the embodiment of the present invention. The Alpha thalassemia SEA mutation chain of the male was determined by family linkage analysis, but the mutation chain of the female could not be determined due to incomplete information about the female's father.
[0125] Figure 4 This is a schematic diagram of determining the Alpha thalassemia SEA mutation chain of the female through embryo mutual inference after adding genomic data of two embryos to Family 2 in an embodiment of the present invention, one of which is a homozygous proband embryo.
[0126] Using the ancestral haplotype similarity matrix, the results of family 2 are shown in Table 5. The similarity value between the haplotype of the female_P chain and the ancestral haplotype is 0.91, and the similarity value between the haplotype of the female_M chain and the ancestral haplotype is 0.19. It is determined that the female_P chain is the Alpha thalassemia SEA mutation chain, which is consistent with the Alpha thalassemia SEA mutation chain determined by embryo mutual inference.
[0127] Ancestral haplotype Female_P Female_M snp12 G - - snp14 A - - snp13 C - - snp07 A - - snp08 T - - snp02 T T C snp09 C C C snp10 C C T snp11 C C T snp03 C C T snp05 A A C snp04 A A G snp01 A A C snp06 T T C Similarity value 0.91 0.19 .
[0128] The embodiment of the present invention also proposes an Alpha thalassemia SEA mutant chain detection device, the principle of which is similar to the Alpha thalassemia SEA mutant chain detection method, and will not be repeated here.
[0129] Figure 5 This is a schematic diagram of the structure of the Alpha thalassemia SEA mutation chain detection device in an embodiment of the present invention. The Alpha thalassemia SEA mutation chain detection device in an embodiment of the present invention includes:
[0130] The acquisition module 501 is used to acquire whole genome chip detection data of family members of multiple families containing Alpha thalassemia SEA mutation, and determine the Alpha thalassemia SEA mutation chain and normal chain of each family member;
[0131] A P value calculation module 502 is used to calculate the P value of the difference in base usage frequency between all Alpha thalassemia SEA mutant chains and all normal chains for each SNP site;
[0132] An ancestral haplotype forming module 503 is used to screen the SNP sites according to the P value corresponding to each SNP site, and form the ancestral haplotype of the Alpha thalassemia SEA mutation chain according to the screened SNP sites;
[0133] A similarity matrix forming module 504 is used to form a similarity matrix according to the proportions of multiple bases of the SNP sites of the ancestral haplotypes in all Alpha thalassemia SEA mutation chains;
[0134] The similarity analysis module 505 of the sample to be tested is used to calculate the similarity value between the haplotype of each chain of the sample to be tested and the ancestor haplotype according to the similarity matrix after obtaining the sample to be tested;
[0135] The mutant chain determination module 506 is used to determine the Alpha thalassemia SEA mutant chain from the two chains of the sample to be tested according to the similarity value corresponding to each chain.
[0136] In one embodiment, the acquisition module is used to:
[0137] The haplotype phasing method was used to construct the haplotype of the chromosome where the Alpha-thalassemia SEA mutation was located in each family, and the Alpha-thalassemia SEA mutation chain and normal chain of each family member were determined through family linkage analysis.
[0138] In one embodiment, the P value calculation module is used to:
[0139] Summarize all Alpha thalassemia SEA mutation chains to obtain the Alpha thalassemia SEA mutation chain set, and count the base usage frequency of each SNP site in the Alpha thalassemia SEA mutation chain set;
[0140] Summarize all normal chains to obtain a normal chain set, and count the base usage frequency of each SNP site in the normal chain set;
[0141] For each SNP site, the P value of the difference in base usage frequency between the Alpha thalassemia SEA mutant chain set and the normal chain set was calculated.
[0142] In one embodiment, ancestral haplotype forming modules are used to:
[0143] Arrange the SNP sites according to the P value from small to large;
[0144] Obtaining a first SNP site set formed by SNP sites whose P values are less than a first P value threshold;
[0145] Deleting SNP sites that are farther away from the HBA1 gene than a first distance from the first SNP site set to obtain a second SNP site set;
[0146] One of the two SNP sites whose distance is less than the second distance is deleted from the second SNP site set to obtain the screened SNP site.
[0147] In one embodiment, the similarity matrix forming module is used to:
[0148] For each chain of the sample to be tested, all SNP sites of the haplotype of the chain are obtained;
[0149] Obtain the ratio of the bases at each SNP site of the haplotype of the chain to the corresponding bases at the corresponding SNP site in the similarity matrix, and use the ratio as the similarity value of each SNP site;
[0150] The ratio of the sum of the similarity values of all SNP sites of the haplotype of the chain to the number of all SNP sites is taken as the similarity value between the haplotype of the chain and the ancestral haplotype.
[0151] In one embodiment, the mutant chain determination module is used to:
[0152] If the similarity value corresponding to one chain of the sample to be tested is greater than the first threshold, and the difference between the similarity values of the chain and another chain of the sample to be tested is greater than the second threshold, the chain is determined to be the Alpha thalassemia SEA mutation chain.
[0153] In one embodiment, the similarity matrix forming module is further used for:
[0154] Calculate the similarity value between the haplotype of each Alpha thalassemia SEA mutation chain in the Alpha thalassemia SEA mutation chain set and the ancestral haplotype;
[0155] The average of the similarity values corresponding to all Alpha thalassemia SEA mutation chains is used as the first threshold;
[0156] Calculate the similarity value between the haplotype of each normal chain in the normal chain set and the ancestral haplotype;
[0157] The difference between the average value of the similarity values corresponding to all normal chains and the first threshold is calculated, and the difference is used as the second threshold.
[0158] In summary, in the method and device proposed in the embodiment of the present invention, the whole genome chip detection data of family members of multiple families containing Alpha thalassemia SEA mutations are collected, and the Alpha thalassemia SEA mutation chain and normal chain of each family member are determined; the P value of the difference in base usage frequency of each SNP site in all Alpha thalassemia SEA mutation chains and all normal chains is calculated; according to the P value corresponding to each SNP site, the SNP site is screened, and according to the screened SNP site, the ancestral haplotype of the Alpha thalassemia SEA mutation chain is formed; according to the proportion of multiple bases of the SNP site of the ancestral haplotype in all Alpha thalassemia SEA mutation chains, a similarity matrix is formed; after obtaining the sample to be tested, according to the similarity matrix, the similarity value of the haplotype of each chain of the sample to be tested and the ancestral haplotype is calculated; according to the similarity value corresponding to each chain, the Alpha thalassemia SEA mutation chain is determined from the two chains of the sample to be tested. Through the above steps, there is no need to perform mutation site detection and SNP linkage typing on samples of at least two generations (the proband or the parents of a couple), and the whole genome chip detection data of family members of multiple families can be directly collected to form the ancestral haplotype of the Alpha thalassemia SEA mutation chain. After the similarity matrix is formed, the similarity value between the haplotype of each chain of the sample to be tested and the ancestral haplotype is calculated according to the similarity matrix, and then the Alpha thalassemia SEA mutation chain is determined from the two chains of the sample to be tested, with high accuracy.
[0159] An embodiment of the present invention further provides a computer device, Figure 6 This is a schematic diagram of a computer device in an embodiment of the present invention, wherein the computer device 600 includes a memory 610, a processor 620, and a computer program 630 stored in the memory 610 and executable on the processor 620, and when the processor 620 executes the computer program 630, the above-mentioned Alpha thalassemia SEA mutation chain detection method is implemented.
[0160] An embodiment of the present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned Alpha thalassemia SEA mutation chain detection method.
[0161] An embodiment of the present invention also provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the above-mentioned Alpha thalassemia SEA mutation chain detection method.
[0162] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0163] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0164] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0165] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0166] It should be understood that the above specific embodiments of the present invention are only used for illustrative inventions or to explain the principles of the present invention, and do not constitute limitations on the present invention. Therefore, any modifications, equivalent substitutions, improvements, etc. made without departing from the spirit and scope of the present invention should be included in the protection scope of the present invention. In addition, the appended claims of the present invention are intended to cover all changes and modifications that fall within the scope and boundaries of the appended claims, or the equivalent forms of such scope and boundaries.
Claims
1. A method for locating the Alpha-thalassemia SEA mutation chain, characterized in that: S1, collect the whole genome chip test data of family members of multiple families with Alpha thalassemia SEA mutation, use the haplotype phasing method to construct the haplotype of the chromosome where the Alpha thalassemia SEA mutation of each family is located, and determine the Alpha thalassemia SEA mutation chain and normal chain of each family member through family linkage analysis; S2, calculate the P value of the difference in base usage frequency between all Alpha-thalassemia SEA mutant chains and all normal chains for each SNP site; The P value is obtained as follows: S21, summarize all Alpha thalassemia SEA mutation chains to obtain the Alpha thalassemia SEA mutation chain set, and count the base usage frequency P1 of each SNP site in the Alpha thalassemia SEA mutation chain set; S22, summing up all normal chains to obtain a normal chain set, and counting the base usage frequency P2 of each SNP site in the normal chain set; S23, for each SNP site, calculate the P value of the difference in base usage frequency between the Alpha thalassemia SEA mutant chain set and the normal chain set, P=| P1-P2| S3, screening SNP sites according to the P value in S2, and forming the ancestral haplotype of the Alpha thalassemia SEA mutation chain according to the screened SNP sites; S4, forming a similarity matrix, forming a similarity matrix according to the proportion of multiple bases of the SNP sites of the ancestral haplotype in all Alpha thalassemia SEA mutation chains; S5, after obtaining the sample to be tested, calculating the similarity value between the haplotype of each chain of the sample to be tested and the ancestral haplotype according to the similarity matrix; The similarity calculation method of the haplotype of the sample to be tested and the ancestral haplotype is as follows: S51, obtaining all SNP sites of the haplotype of each chain of the sample to be tested according to the S3 method, wherein the number of effective SNP sites among all SNP sites is greater than or equal to 6; S52, obtaining the ratio of the bases at the effective SNP site of the haplotype of the chain to the corresponding bases at the corresponding SNP site in the similarity matrix, and using the ratio as the similarity value of each SNP site; S53, taking the ratio of the sum of the similarity values of the effective SNP sites of the haplotype of the chain to the number of all SNP sites as the similarity value between the haplotype of the chain and the ancestral haplotype; S6, determining the Alpha thalassemia SEA mutation chain from the two chains of the sample to be tested according to the similarity value corresponding to each chain; if the similarity value corresponding to one chain of the sample to be tested is greater than the first threshold, and the difference between the similarity value of the chain and the other chain of the sample to be tested is greater than the second threshold, determining that the chain is the Alpha thalassemia SEA mutation chain; The first threshold is the average of the similarity values corresponding to all Alpha thalassemia SEA mutation chains; The second threshold is the difference between the average value of the similarity values corresponding to all normal links and the first threshold.
2. The method according to claim 1, characterized in that: The steps of screening SNP sites in S3 are as follows: S31, arrange the SNP sites according to the P value from small to large; S32, obtaining a first SNP site set formed by SNP sites whose P values are less than a first P value threshold; the first P value threshold S33, deleting the SNP sites that are farther away from the HBA1 gene than the first distance from the first SNP site set to obtain a second SNP site set; S34, deleting any one of the two SNP sites whose distance is less than the second distance from the second SNP site set to obtain a screened effective SNP site.
3. A detection device for locating the Alpha thalassemia SEA mutation chain, characterized in that: The device comprises: a collection module, a P value calculation module, an ancestral haplotype formation module, a similarity matrix formation module, a similarity analysis module of a sample to be tested and a mutation chain determination module; The acquisition module includes a module for collecting whole genome chip detection data of family members of multiple families containing Alpha thalassemia SEA mutations, and determining the Alpha thalassemia SEA mutation chain and normal chain of each family member; The P value calculation module is used to calculate the P value of the difference in base usage frequency between all Alpha thalassemia SEA mutant chains and all normal chains for each SNP site; the P value is obtained by the following steps: (1) Summarize all Alpha thalassemia SEA mutation chains to obtain the Alpha thalassemia SEA mutation chain set, and count the base usage frequency P1 of each SNP site in the Alpha thalassemia SEA mutation chain set; (2) Summarize all normal chains to obtain a normal chain set, and count the base usage frequency P2 of each SNP site in the normal chain set; (3) For each SNP site, calculate the P value of the difference in base usage frequency between the Alpha thalassemia SEA mutant chain set and the normal chain set, P = | P1-P2|; The ancestral haplotype forming module is used to screen the SNP sites according to the P value corresponding to each SNP site, and form the ancestral haplotype of the Alpha thalassemia SEA mutation chain according to the screened SNP sites; the specific rules are as follows: 1) Arrange the SNP sites from small to large according to the P value; 2) obtaining a first SNP site set formed by SNP sites having a P value less than a first P value threshold; 3) deleting the SNP sites whose distance from the HBA1 gene exceeds the first distance from the first SNP site set to obtain a second SNP site set; 4) deleting one of the two SNP sites whose distance is less than the second distance from the second SNP site set to obtain the screened SNP site; The similarity matrix forming module is used to form a similarity matrix according to the proportion of multiple bases of the SNP sites of the ancestral haplotype in all Alpha thalassemia SEA mutation chains; The similarity analysis module of the sample to be tested is used to calculate the similarity value between the haplotype of each chain of the sample to be tested and the ancestral haplotype according to the similarity matrix after obtaining the sample to be tested; the specific calculation rules are as follows: i. According to the method of S3 in claim 1, all SNP sites of the haplotype of each chain of the sample to be tested are obtained, and the number of effective SNP sites among all SNP sites is greater than or equal to 6; ii. Obtain the ratio of the bases at the effective SNP sites of the haplotype of the chain to the corresponding bases at the corresponding SNP sites in the similarity matrix, and use the ratio as the similarity value of each SNP site; iii. The ratio of the sum of the similarity values of the effective SNP sites of the haplotype of the chain to the number of all SNP sites is taken as the similarity value between the haplotype of the chain and the ancestral haplotype; The mutant chain determination module is used to determine the Alpha thalassemia SEA mutant chain from the two chains of the sample to be tested according to the similarity value corresponding to each chain; If the similarity value corresponding to one chain of the sample to be tested is greater than the first threshold, and the difference between the similarity values of the chain and another chain of the sample to be tested is greater than the second threshold, the chain is determined to be an Alpha thalassemia SEA mutation chain; The first threshold is the average of the similarity values corresponding to all Alpha thalassemia SEA mutation chains; The second threshold is the difference between the average value of the similarity values corresponding to all normal links and the first threshold.
4. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 2 is implemented.
5. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 2 is implemented.
6. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 2 is implemented.
Citation Information
Patent Citations
Method and system for determining fetus alpha thalassemia gene haplotype
CN108048541A
Kit for detecting whether individual to which sample to be detected belongs suffers from genetic disease
CN116004798A