Analysis method of alpha-thalassemia genotype
Through the amplicon sequencing technology and internal reference region screening methods, the accuracy and operational complexity of α-thalassemia gene detection in the prior art are solved, and efficient and accurate genotype analysis is achieved.
Patent Information
- Application Number
- CN202510429802.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-05-09
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing α-thalassemia gene detection methods are difficult to accurately detect various types of variants, and are cumbersome to operate and have high cost.
Using amplicon sequencing technology, amplicon sequencing data of multiple sequence regions in DNA samples were obtained, combined with genomic reference sequences for comparison, stable internal reference regions were screened out, and internal and inter-sample homogenization was performed, and the Ratio values of the target regions were calculated, which were used to analyze and output α-thalassemia genotype prompts.
Accurate analysis of α-thalassemia genotype was achieved, eliminating the sequencing depth differences caused by inconsistent sequencing data volume, reducing the deviation caused by PCR, and improving the accuracy and efficiency of detection.
Smart Images

Figure CN119955931A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of biological detection technology, and in particular to a method for analyzing α-thalassemia genotype. Background Art
[0002] α-thalassemia (abbreviated as "α-thalassemia") is a common hemolytic genetic disease. The molecular mechanism leading to α-thalassemia is: the mutation or deletion of the α-globin gene reduces or completely prevents the synthesis of α-globin chains, resulting in an imbalance in the proportion of globin chains that form hemoglobin, which in turn leads to heritable hemolytic anemia. Most areas south of the Yangtze River in my country are high-incidence areas for this disease, especially Guangdong, Guangxi and Hainan.
[0003] The α-globin gene is located on chromosome 16, and each chromosome 16 has two α-globin genes (abbreviated as "α-genes"). α-thalassemia includes deletion type and non-deletion type, among which deletion type mutations account for the majority. The severity of the α-thalassemia phenotype is directly related to the number of copies of the dysfunctional α gene. According to the number of α genes affected on a single chromosome, the deletion type can be divided into α + Thalassemia (single α gene deletion) and α 0 Thalassemia (two α genes on one chromosome are inactivated simultaneously). + Common variants in thalassemia are -α3.7 and -α4.2. 0 The common variant of thalassemia in the Chinese population is αSEA (Southeast Asian type). Non-deletional α-thalassemia is caused by point mutations, the most common of which include α CS , α QS and α WS The above six mutations account for about 95% to 96.3% of all mutations.
[0004] At present, the main methods for detecting α-thalassemia genes include the following: (1) Sanger sequencing: Primers are designed at both ends of the sequence to be tested, and after PCR amplification, sequencing is performed using the dideoxy chain termination method, and then compared with the wild-type sequence for analysis. Only point mutations within the amplified region can be detected.
[0005] (2) Gap-PCR: Primers are designed on both sides of the target deletion fragment. Since the upstream and downstream primers of the wild-type sequence are far apart, the product cannot be amplified. When a deletion occurs, a fragment of a specific length can be amplified. It can detect known deletion variants, such as -α3.7, -α4.2, --αSEA, and --αTHAI, but it cannot accurately detect multiple types of variants and is easily affected by homologous sequences.
[0006] (3) Multiplex ligation-dependent probe amplification (MLPA): It can detect known or unknown deletion / duplication variations within the probe coverage area, such as -α3.7, -α4.2, --αSEA, αααanti3.7, etc. However, the probe cost is high and the operation is cumbersome.
[0007] (4) Second-generation sequencing: Targeted enrichment is performed using target region probe capture technology. Through library preparation, high-throughput sequencing, and bioinformatics analysis to screen point mutations or bioinformatics analysis of deletion variants, known / unknown deletion variants and point mutations in the region can be detected. However, the probe cost is high and the operation is cumbersome. Summary of the invention
[0008] In view of the problems of the prior art, the present application provides a method for analyzing the α-thalassemia genotype, which can achieve accurate analysis of the α-thalassemia genotype based on amplicon sequencing technology.
[0009] Methods for analyzing α-thalassemia genotypes include: Step S10, obtaining amplicon sequencing data of multiple sequence regions in a DNA sample, wherein the DNA sample includes a sample to be analyzed and a control sample, and the sequence region includes a target region for α-thalassemia; Step S20, comparing the amplicon sequencing data with the genome reference sequence to obtain the number of reads and sequencing depth of each sequence region in each DNA sample; Step S30, obtaining the remaining sequence regions except the target region of α-thalassemia from all DNA samples, sorting the remaining sequence regions according to the stability of the sequencing depth of all DNA samples based on the same data amount, and selecting multiple stable remaining sequence regions as internal reference regions; Step S40, obtaining the ratio of the number of reads in each target region to the number of reads in multiple internal reference regions in each DNA sample, and calculating the average of the multiple ratios as the A value for each target region; Step S50, obtaining the ratio of the A value of each target area in the sample to be analyzed and the control sample, which is the Ratio value of each target area in the sample to be analyzed; Step S60: Analyze and output the α-thalassemia genotype hint of the sample to be tested according to the Ratio value of the target region.
[0010] Optionally, the amplicon sequencing data is obtained by performing a same-batch amplification reaction on all DNA samples using a polymerase chain reaction system containing multiple pairs of specific primers, obtaining amplicon mixtures of multiple sequence regions in each DNA sample, and then sequencing the amplicon mixtures on a machine; the specific primers include sequences shown in SEQ ID NOs: 1 to 120.
[0011] Optionally, the specific operation of the stability sorting is: counting the differences between all DNA samples for each remaining sequence region to obtain the coefficient of difference CV of each remaining sequence region, and sorting all the coefficients of difference CV from small to large; extracting the expected number of remaining sequence regions with the smallest coefficient of difference CV as internal reference regions.
[0012] Optionally, it includes: performing predictive screening of negative samples among all DNA samples, taking the DNA samples obtained from the predictive screening as candidate samples, or taking all DNA samples as a whole as candidate samples; and selecting control samples from the candidate samples.
[0013] Optionally, selecting a control sample from the candidate samples comprises the following steps: (1) Determine a target region and obtain the A value of the target region corresponding to the internal reference region in each candidate sample; (2) Calculate the median value of the target area in the set of A values of the target area of all candidate samples; (3) According to the degree of dispersion of the target area A value relative to the median value Median, select the samples with the lowest expected number of dispersion from all candidate samples as the control samples of the target area; (4) Execute steps (1) to (3) for each target area to obtain a control sample for each target area, and take the union of the control samples of all target areas to obtain the final control sample set.
[0014] Optionally, the DNA sample contains multiple control samples, and the ratios of the A values of each target region in the sample to be analyzed and the multiple control samples are calculated, and the average of the multiple ratios of each target region is obtained as the Ratio value of each target region in the sample to be analyzed.
[0015] Optionally, the analysis method comprises: Providing a DNA sample database with known α-thalassemia genotypes, including DNA samples of -α4.2, -α3.7, --SEA, --FIL, and --THAI, executing steps S10 to S50 to obtain the Ratio value of each target region in all DNA samples in the DNA sample database; According to the target region copies of all DNA samples in the DNA sample database and the Ratio values of the target region, a receiver operating characteristic curve analysis is performed to obtain a Ratio threshold for determining whether the copy of the target region is positive; The Ratio value of each target area in the sample to be analyzed is compared with the corresponding Ratio threshold value, and according to the comparison result, the α-thalassemia genotype prompt of the sample to be analyzed is analyzed and output.
[0016] Optionally, the analysis method includes: analyzing and outputting the α-thalassemia genotype hint of the sample to be analyzed based on the ratio of the target region Ratio values of -α4.2 and -α3.7.
[0017] Optionally, the specific primers include those shown in SEQ ID NO: 123 and SEQ ID NO: 124 for amplification. HBA1 and HBA2 primer pairs specific for the homologous region; The analysis method comprises: pre-shielding in the genome reference sequence HBA2 and HBA1 The homologous region of the genome is then compared with the reference genome sequence; for each pair of differential bases, the number of differential bases is counted, and the ratio of the number of one differential base to the total number of two differential bases is calculated to obtain a VAF value; and the α-thalassemia genotype hint of the sample to be analyzed is obtained according to the VAF value.
[0018] The present application also provides a computer device, comprising a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement any of the analysis methods described above.
[0019] Compared with the prior art, the present application uses the screened internal reference region to perform intra-sample homogenization on the target region of the sample to be analyzed, which can eliminate the sequencing depth difference caused by inconsistent sequencing data volume; and then uses the control sample to perform inter-sample homogenization on the target region of the sample to be analyzed, which can eliminate the deviation caused by the polymerase chain reaction (PCR), and the obtained Ratio value can be used to accurately indicate the copy status of the target region; the copy status of the target region can be comprehensively analyzed and the α-thalassemia genotype prompt of the sample to be tested can be accurately analyzed and output. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 is a flow chart of a method for analyzing α-thalassemia genotype in one embodiment; Figure 2 A flowchart of determining a control sample from a candidate sample in one embodiment; Figure 3 A flowchart of determining a control sample from all DNA samples in one embodiment; Figure 4 Two homologous regions HBA1 and HBA2 A comparison diagram, in which the red boxes represent two pairs of different bases in the comparison results, and the dots represent the difference matching relationship between the two bases at the corresponding positions; Figure 5This is the logic diagram for the classification of α-thalassemia; Figure 6 1 is a diagram of the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0021] The technical solution of the present application is further described below in conjunction with specific implementation methods, but the present application is not limited thereto.
[0022] See also Figure 1 The present application provides a method for analyzing α-thalassemia genotype, comprising: Step S10, obtaining amplicon sequencing data of multiple sequence regions in a DNA sample, wherein the DNA sample includes a sample to be analyzed and a control sample, and the sequence region includes a target region for α-thalassemia; Step S20, aligning the amplicon sequencing data with the genome reference sequence to obtain the number of reads and sequencing depth of each sequence region in each DNA sample; Step S30, obtaining the remaining sequence regions except the target region of α-thalassemia from all DNA samples, sorting the remaining sequence regions according to the stability of the sequencing depth of all DNA samples based on the same data amount, and selecting multiple stable remaining sequence regions as internal reference regions; Step S40, obtaining the ratio of the number of reads in each target region to the number of reads in multiple internal reference regions in each DNA sample, and calculating the average of the multiple ratios as the A value for each target region; Step S50, obtaining the ratio of the A value of each target area in the sample to be analyzed and the control sample, which is the Ratio value of each target area in the sample to be analyzed; Step S60: Analyze and output the α-thalassemia genotype hint of the sample to be tested according to the Ratio value of the target region.
[0023] Among them, the amplicon sequencing data is obtained by performing the same batch amplification reaction on all DNA samples using a polymerase chain reaction system containing multiple pairs of specific primers, obtaining amplicon mixtures of multiple sequence regions in each DNA sample, and then sequencing the amplicon mixtures on a machine.
[0024] Specifically, the specific primers include sequences shown in SEQ ID NOs: 1 to 120, wherein SEQ ID NOs: 1 to 12 are specific primers designed for common deletion types of α-thalassemia (-α4.2, -α3.7, --SEA, --FIL and --THAI), and the sequence regions corresponding to each deletion type are shown in Table 1. SEQ ID NOs: 13 to 120 are specific primers for amplifying the other 54 sequence regions.
[0025] The present application first uses the screened internal reference region to perform intra-sample homogenization on the target region of the sample to be analyzed, which can eliminate the difference in sequencing depth caused by the inconsistent sequencing data volume of different samples; then uses the control sample to perform inter-sample homogenization on the target region of the sample to be analyzed, which can eliminate the deviation caused by the polymerase chain reaction (PCR), such deviation includes the deviation caused by PCR capture efficiency and preference, and the obtained Ratio value can be used to accurately indicate the copy status of the target region; according to the copy status of the target region, the α-thalassemia genotype prompt of the sample to be analyzed can be accurately analyzed and output.
[0026] The same batch amplification reaction means that all DNA samples are subjected to amplification reaction using the same batch of polymerase chain reaction system and the same reaction conditions to ensure that the specific primers capture the consistency of the amplification region in all DNA samples. For multiple amplification regions, multiple sets of specific primers are usually required to be designed, that is, the polymerase chain reaction of the present application is a multiple amplification reaction (multiple PCR). The polymerase chain reaction system of the present application can ensure the stability of amplification, including the polymerase chain reaction system including a specific amplification reaction solution and an enrichment amplification reaction solution.
[0027] In one embodiment, the amplification reaction includes: (1) constructing a first amplification reaction system of a DNA sample, a specific primer, and a specific amplification reaction solution, performing a first round of amplification reaction, and obtaining a first amplification product mixture; (2) constructing a second amplification reaction system of the first amplification product mixture, a sequencing adapter primer, and an enriched amplification reaction solution, and performing a second round of amplification reaction to obtain a second amplification product mixture; (3) purifying the second amplification product mixture to obtain an amplicon mixture for sequencing.
[0028] The specific primers can be pre-mixed into a primer mixture, and when used, the final concentration of each primer in the amplification reaction system is 0.02-0.03 μM, preferably 0.025 μM. The product size obtained by amplification with each specific primer is 50-300 bp.
[0029] In one embodiment, step S30 includes performing a uniform calibration on the total sequencing data volume of each DNA sample so that the total sequencing data volume of each DNA sample is the same. For example, if the application requires the total sequencing data volume to be 1G, the total sequencing data volume of each DNA sample obtained by sequencing may fluctuate. The total sequencing data volume of each DNA sample is proportionally converted so that the total sequencing data volume of each DNA sample is close to 1G, i.e., uniform calibration.
[0030] The so-called stability sorting is: count the differences between all the remaining sequence regions in all DNA samples, obtain the coefficient of difference CV of each remaining sequence region, and sort all the coefficients of difference CV from small to large. After sorting, extract the expected number of remaining sequence regions with the smallest coefficient of difference CV as internal reference regions. The expected number can be 3 to 10.
[0031] The control sample is generally a negative sample, that is, a sample with a normal copy of the target region. In one embodiment, the control sample is a known negative sample, such as a known negative sample that can be detected by multiple ligation-dependent probe amplification (MLPA). Since the position of the control sample in the DNA sample in high-throughput sequencing is known, the amplicon sequencing data obtained by the detection channel corresponding to the control sample can be extracted after high-throughput sequencing.
[0032] Another embodiment is to determine a control sample from all DNA samples, including the following steps: extracting candidate samples from all DNA samples; and selecting a control sample from the candidate samples.
[0033] Specifically, candidate samples can be extracted from all DNA samples by: performing predictive screening of negative samples in all DNA samples, and using the DNA samples obtained from the predictive screening as candidate samples; or using all DNA samples as a whole as candidate samples.
[0034] The specific steps of prediction and screening include: using DNA samples with known target region copies and the A value of the target region obtained by the analysis method of the present application as a training set, training a random forest algorithm training model, using the random forest algorithm training model to predict and screen negative samples in all DNA samples, and using the DNA samples obtained by prediction and screening as candidate samples. Among them, DNA samples with known amplification region copy numbers include positive samples with abnormal target region copies and negative samples with normal target region copies.
[0035] In one embodiment, 150 DNA samples with known target region copies and A values of multiple internal reference regions corresponding to the target regions in the DNA samples are used as training sets to obtain a random forest algorithm training model.
[0036] See also Figure 2 , the specific steps for determining the control sample from the candidate samples are as follows: (1) Determine a target region and obtain the A value of the target region corresponding to the internal reference region in each candidate sample; (2) Calculate the median value of the target area in the set of A values of the target area of all candidate samples; (3) According to the degree of dispersion of the A value of the target area relative to the median value Median, select the expected number of samples with the lowest degree of dispersion from all candidate samples as the control samples of the target area; the expected number can be 5 to 15, for example, 10; (4) Execute steps (1) to (3) for each target area to obtain a control sample for each target area, and take the union of the control samples of all target areas to obtain the final control sample set.
[0037] See also Figure 3 In the specific embodiment shown, determining a control sample from all DNA samples comprises the following steps: The model was trained using the random forest algorithm to perform predictive screening on all DNA samples; If a DNA sample is predicted to be a negative sample, the DNA sample obtained by the prediction screening is used as a candidate sample, otherwise all samples are used as candidate samples; Execute steps (1) to (4) for all candidate samples to obtain a final control sample set. If the number of control samples in the final control sample set is ≥10, extract the top 10 control samples that appear in the process of obtaining the union as the control samples in step S50 and include them in the calculation. If the number of control samples in the final control sample set is <10, extract all control samples in the control sample set as the control samples in step S50 and include them in the calculation.
[0038] In one embodiment, if the DNA sample contains multiple control samples, the ratios of the target region A values of the sample to be analyzed and the multiple control samples are calculated, and the average of the multiple ratios is obtained as the Ratio value of the target region in the sample to be analyzed.
[0039] For the control samples obtained by the prediction screening, each control sample can be further analyzed and verified using the analysis method of the present application. Specifically, in the control sample set, the ratio of each control sample to the A value of each target area in all control samples is calculated, and the average of multiple ratios of each target area is obtained to obtain the Ratio value of each target area in each control sample.
[0040] The copying situation can be divided into copy missing or copy duplication. The target area has corresponding theoretical Ratio values corresponding to different copying situations. By comparing the Ratio value actually measured by the analysis method of the target area of the present application with the corresponding theoretical Ratio value, the copying situation of each target area in the sample to be analyzed can be determined.
[0041] Taking human genes as an example: (1) If the target region copy is normal, its theoretical Ratio value is 1; (2) If the target region copy is half missing, its theoretical Ratio value is 0.5; (3) If the target region copy is completely missing, its theoretical Ratio value is 0. Compare the measured Ratio value of the target region with the theoretical Ratio value to indicate the copy status of the target region of the sample to be tested.
[0042] However, the theoretical Ratio value will have a certain deviation in actual detection. In order to more accurately analyze the copy number changes in the target region, this application uses a large amount of data from clinical positive samples and negative samples to perform ROC analysis and obtain a suitable threshold. The specific method includes the following steps: A DNA sample database with known α-thalassemia genotypes is provided, including DNA samples of -α4.2, -α3.7, --SEA, --FIL, and --THAI, and steps S10 to S50 are performed to obtain the Ratio value of each target region (i.e., each specific primer amplification region) in all DNA samples in the DNA sample database; According to the target region copies and the Ratio values of the target region of all DNA samples in the DNA sample database, a receiver operating characteristic (ROC) curve analysis is performed to obtain a Ratio threshold for determining whether the target region copy is positive; The Ratio value of each target region in the sample to be analyzed is compared with the corresponding Ratio threshold, and based on the comparison result, the α-thalassemia genotype hint of the sample to be analyzed is analyzed and output.
[0043] In one embodiment, 142 samples with known genotypes were analyzed by ROC curve analysis, and the thresholds of the P7 amplification region were 0.65, the P9 amplification region was 0.66, the P10 amplification region was 0.66, the P13 amplification region was 0.66, the P16 amplification region was 0.65, and the P17 amplification region was 0.65. The following typing was performed based on these thresholds: (1) If the Ratio value of the P7 amplification region is ≤0.65, it indicates that -α4.2 occurs; (2) If the Ratio value of the P9 amplification region is ≤0.66 and the Ratio value of the P10 amplification region is ≤0.66, if either condition is met, it indicates that -α3.7 occurs; (3) If the Ratio value of the P13 amplification region is ≤0.66, it indicates that --SEA, --FIL, or --THAI has occurred; (4) If the Ratio value of the P16 amplification region is ≤0.65, it indicates that FIL or THAI has occurred; (5) If the Ratio value of the P17 amplification region is ≤0.65, it indicates that THAI has occurred.
[0044] See also Figure 5 In the illustrated embodiment, the output logic of the α-thalassemia genotype provided by the present application includes: distinguishing whether multiple deletion genotypes occur according to whether the Ratio value of the P7, P9 (or P10) amplification region is less than 0.1 (ie, close to 0).
[0045] First, determine whether the Ratio value of the amplified regions of P7 and P9 (or P10) is less than 0.1. If the Ratio value of either is <0.1, it indicates that a homozygous deletion has occurred; if the Ratio values of the amplified regions of P7 and P9 are both not less than 0.1, it indicates that a heterozygous deletion has occurred.
[0046] For the homozygous deletion genotype, the Ratio values of the amplified regions of P13, P16 and P17 are compared with the corresponding Ratio thresholds to determine whether the corresponding deletion type occurs, and the α-thalassemia genotype prompt I (a) is obtained.
[0047] For the heterozygous deletion genotype, the Ratio values of the amplified regions of P7, P9, P10, P13, P16, and P17 are compared with the corresponding Ratio thresholds to determine whether the corresponding deletion type occurs, and the α-thalassemia genotype prompt I (b) is obtained.
[0048] To improve the accuracy of the enhancement analysis, in one embodiment, the sequence region chr16:199351-199660 is selected as the background region, and the specific primer P15 is designed to capture and amplify the background region, as shown in SEQ ID NO: 121 and SEQ ID NO: 122. Accordingly, the analysis method includes: The threshold values of the ratio value of the P16 amplification region / the ratio value of the P15 amplification region and the ratio value of the P17 amplification region / the ratio value of the P15 amplification region were obtained from the DNA sample database with known α-thalassemia genotypes through ROC curve analysis.
[0049] Calculate the Ratio value of the P16 amplification region / the Ratio value of the P15 amplification region (P16 / P15 for short), and the Ratio value of the P17 amplification region / the Ratio value of the P15 amplification region (P17 / P15 for short) in the sample to be analyzed; compare P16 / P15 and P17 / P15 with the corresponding thresholds, and output the α-thalassemia genotype hint II of the sample to be analyzed based on the comparison results.
[0050] In another embodiment, the Ratio value of the P7 amplification region / the Ratio value of the P9 amplification region, or the Ratio value of the P7 amplification region / the Ratio value of the P10 amplification region, hereinafter referred to as P7 / P9, P7 / P10, is calculated to reflect the relative copy number changes in the region where -α4.2 and -α3.7 are located, and the α-thalassemia genotype indication III is obtained. The specific analysis is as follows: (1) If it approaches 0, it indicates that the genotype of -α4.2 homozygous deletion occurs, which may be a composite genotype of -α4.2 + (-α4.2) or -α4.2 + (--SEA / --FIL / --THAI) or a composite genotype of anti-HKαα + (--SEA / --FIL / --THAI); (2) If it is 0.33, it may be the anti-HKαα genotype; (3) If it is 0.5, it may be the -α4.2 genotype or the anti3.7+ (--SEA / FIL / THAI) composite genotype; (4) If it is 0.67, it may be the anti3.7 genotype; (5) If the ratio is 1, it may be normal or --SEA / --FIL / --THAI genotype; if the ratio is 1.5, it may be anti4.2 genotype; (6) If it is 2, it may be a composite genotype of -α3.7 or anti4.2+ (--SEA / --FIL / --THAI type); (7) If it is 3, it may be the HKαα genotype; (8) If the value is close to 10 or above, it indicates a -α3.7 homozygous deletion genotype, which may be a composite genotype of -α3.7+ (-α3.7) or -α3.7+ (--SEA / --FIL / --THAI type) or a composite genotype of HKαα+ (--SEA / --FIL / --THAI type).
[0051] Since the sequence chr16:222807-227565 contains HBA1, The sequence region chr16:219785-224767 contains HBA2 , and HBA1 There are homologous regions, and there are two pairs of different bases between them, see Figure 4 , affecting the result analysis.
[0052] One embodiment of the present application provides a method for amplifying HBA2 and HBA1 Specific primers for the homologous region are shown in SEQ ID NO: 123 and SEQ ID NO: 124. Accordingly, the analysis method includes: Pre-masked in genome reference sequence HBA2 and HBA1 The homologous regions of the DNA sequence were then compared with the reference genome sequence. For each pair of differential bases, the number of differential bases was counted, and the ratio of the number of one differential base to the total number of two differential bases was calculated as the VAF value. The two pairs of differential bases were marked as VAF1 and VAF2.
[0053] According to the VAF value, the α-thalassemia genotype of the sample to be analyzed is analyzed and outputted. The specific analysis is as follows: (1) If it is close to 1, it indicates that the genotype with homozygous deletion of -α4.2 may be a composite genotype of -α4.2+ (-α4.2) or -α4.2+ (--SEA / --FIL / --THAI type) or a composite genotype of anti-HKαα+ (--SEA / --FIL / --THAI type); (2) If it is 0.75, it may be the anti-HKαα genotype; (3) If it is 0.67, it may be the -α4.2 genotype or the anti3.7+ (--SEA / --FIL / --THAI) composite gene; (4) If it is 0.6, it may be the anti3.7 genotype; (5) If it is 0.5, it may be normal or --SEA / --FIL / --THAI genotype; (6) If it is 0.4, it may be the anti4.2 genotype; (7) If it is 0.33, it may be a composite genotype of -α3.7 or anti4.2+ (--SEA / --FIL / --THAI); (8) If it is 0.25, it may be the HKαα genotype; (9) If the value is close to 0, it indicates a genotype of homozygous deletion of -α3.7, which may be a composite genotype of -α3.7+ (-α3.7) or -α3.7+ (--SEA / --FIL / --THAI) or a composite genotype of HKαα+ (--SEA / --FIL / --THAI type).
[0054] Each α-thalassemia genotype hint can be used as evidence to determine and output the specific genotype of α-thalassemia. In practical applications, the specific genotype can be directly output based on a certain α-thalassemia genotype hint, or the specific genotype can be directly output based on multiple α-thalassemia genotype hints.
[0055] Furthermore, the results of the above-mentioned α-thalassemia genotypes I to IV are taken as a union, the number of occurrences of each typing result is counted, different weights are assigned to each typing evidence, and the typing with the largest weight is found, which is the α-thalassemia genotype of the sample to be analyzed.
[0056] In the process of analyzing heterozygous deletion genotypes, if the four typing results of Normal, SEA, THAI and TIL appear the same number of times (that is, the strength of evidence is the same), the following logic is used to determine: if SEA, P7 and P9 (or P10) are empty at the same time, the output is Normal; if SEA, P7 and P9 (or P10) are not empty at the same time, the Ratio value of the amplified regions of P16 and P17 is compared with the corresponding Ratio threshold to determine whether the corresponding deletion type occurs.
[0057] See also Figure 6 The present application also provides a computer device, including a memory, a processor and a computer program stored in the memory, and the processor executes the computer program to implement steps S10 to S60.
[0058] The present application also provides a computer program product, including computer instructions, which implement steps S10 to S60 when executed by a processor.
[0059] In the present application, the computer program product includes a program code portion for executing the method steps in each embodiment of the present application when the computer program product is executed by one or more computing devices. The computer program product may be stored on a computer-readable recording medium. The computer program product may also be provided for download via a data network (e.g., via a RAN, via the Internet and / or via an RBS). Alternatively or additionally, the method may be encoded in a field programmable gate array (FPGA) and / or an application specific integrated circuit (ASIC), or the functionality may be provided for download by means of a hardware description language.
[0060] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0061] Application Example 1 1. Provide 18 DNA samples, including 8 samples to be analyzed and 10 known negative samples. Obtain the amplicon mixture of each DNA sample through polymerase chain reaction of the same batch. The specific steps are as follows: (1) Extract human peripheral blood genomic DNA according to the instructions of nucleic acid extraction or purification reagents. The range of DNA sample detection usage is: 2ng≤DNA concentration≤4ng, purity: 260 / 280>1.8, 260 / 230>1.6.
[0062] (2) Table 1 shows the deletion regions of common α-thalassemia deletion types. Specific primer pairs designed for each α-thalassemia deletion type include P7, P9, P10, P13, P16 and P17 as shown in SEQ ID NOs: 1 to 12, and primer pairs for HBA1 and HBA2 The homologous regions to provide P045 shown in SEQ ID NO: 123 and SEQ ID NO: 124.
[0063] Table 1 Common deletion types of α-thalassemia and the corresponding deletion regions
[0064] (3) This application example also provides 54 regions for screening internal reference regions, as well as specific primer pairs for these regions, such as SEQ ID NO: 13~120; the 54 regions are chr1:45965892-45966144, chr1:45973019-45973276, chr1:45973821-45974091, chr6:49399032-49399297, chr6:49399357-49399632, chr6:49403128-49403415, chr6:49404191-49404468, chr6:49407785-49408112, chr6:49409493-49409707, chr6:494122 45-49412511, chr6:49415325-49415574, chr6:49416458-49416714, chr6:49421221-49421496, chr7:107301191-107301391, chr7:10730202 9-107302291、chr7:107303683-107303951、chr7:107312536-107312760、chr7:107315337-107315602、chr7:107329439-107329697、chr7:10 c hr7:107344693-107344904, chr7:107350465-107350676, chr7:107352933-107353211, chr7:107355810-107356033, chr7:107356455-10735 6725, chr12:103234102-103234345, chr12:103236767-103236990, chr12:103237359-103237626, chr12:103240535-103240802, chr12:1032 45345-103245624, chr12:103246518-103246791, chr12:103248230-10 3248461, chr12:103248466-103248738, chr12:103260227-103260493,chr12:103271099-103271378, chr12:103288476-103288741, chr12:103306515-103306762, chr12:103310791-1033 11022, chr13:52507387-52507604, chr13:52509695-52509868, chr13:52513113-52513362, chr13:52515172-52515 413. chr13:52516476-52516736, chr13:52523730-52524008, chr13:52524064-52524340, chr13:52524348-5252461 1. chr13:52531588-52531830, chr13:52534249-52534509, chr13:52535852-52536135, chr13:52542520-52542799. ,
[0065] (4) Construction of the first round of amplification reaction system The specific amplification reaction solution and the amplicon capture panel (containing SEQ ID NO: 1-124) were thawed on ice, vortexed for 3-5 seconds to mix, centrifuged instantaneously, and placed on ice for use. The amplification reaction system was prepared according to the following system dosage in Table 2, and 15 μL of reaction lubricant was dispensed into each well of the PCR plate, and then 5 μL of DNA sample was added to each reaction well. After sealing the film, vortexed and mixed for 5-10 seconds, and then centrifuged instantaneously to expel the bubbles at the bottom of the tube, and the droplets on the tube wall and the sealing film were collected at the bottom of the tube.
[0066] Table 2 Component amounts of the first round amplification reaction system
[0067] The specific amplification reaction solution was Equinox Multiplex PCR Master Mix (Watchmaker Genomics); the final concentration of each specific primer in the first round amplification reaction system was 0.025 μM.
[0068] (5) Perform the first round of amplification reaction The first round of amplification reaction was carried out according to Table 3. After amplification, each well of the PCR plate was filled up to 50 μL with water, and then purified by magnetic beads to obtain a first amplification product mixture.
[0069] Table 3 First round amplification reaction conditions
[0070] (6) Construction of the second round of amplification reaction system According to Table 4, the enrichment amplification reaction solution and sequencing adapter primer set were thawed on ice, vortexed for 3-5 seconds to mix, centrifuged instantaneously, and placed on ice for use. Take a new PCR plate, dispense 25 μL of the enrichment amplification reaction solution into each well of the PCR plate; add 1 μL of the corresponding forward and reverse index primers, i.e., the sequencing adapter primer set, to the reaction wells according to the required index position, and then add 23 μL of the first amplification product mixture, seal the PCR plate with a sealing film, vortexed for 5-10 seconds, and then centrifuged instantaneously to expel the bubbles at the bottom of the tube, and collect the droplets on the tube wall and the sealing film to the bottom of the tube.
[0071] Table 4 Second round amplification reaction system
[0072] The enrichment amplification reaction solution is Equinox Library Amplification Kit (Watchmaker Genomics). The sequencing adapter primer set is shown in SEQ ID NO: 125-126, where (nnnnnnnn) is an index sequence of a random base combination, and different samples use different index combinations as identity identifiers.
[0073] (7) Perform the second round of amplification reaction The second round of amplification reaction was carried out according to Table 5, and after the amplification was completed, magnetic bead purification was performed to obtain a second amplification product mixture, that is, an amplicon mixture of each DNA sample.
[0074] Table 5 Second round amplification reaction conditions
[0075] 2. The amplicon mixture is sequenced on a machine to obtain the amplicon sequencing data of each DNA sample.
[0076] 3. Pre-mask the region chr16:223942-224152 in the genome reference sequence, align the amplicon sequencing data with the reference genome sequence, and obtain the number of reads and sequencing depth of each amplified region.
[0077] 4. Extract the amplicon sequencing data of the remaining sequence regions except the α-thalassemia target region from all DNA samples, and count the differences in sequencing depth of each remaining sequence region among all DNA samples in the remaining sequence regions to obtain the difference coefficient CV of each remaining sequence region, and sort all the difference coefficients CV from small to large. After sorting, extract the first 5 remaining sequence regions as internal reference regions.
[0078] 5. In each DNA sample, calculate the ratio of the number of reads in each target region to the number of reads in the five internal reference regions, and take the average of the five ratios as the A value for each target region.
[0079] 6. Calculate the ratio of each sample to be analyzed to the A value of each target area in the 10 negative samples, and take the average of the 10 ratios as the Ratio value of each target area, that is, the Ratio values of the amplified areas of P7, P9, P10, P13, P15, P16, and P17 are obtained.
[0080] 7. For each sample to be analyzed, according to all Ratio values, the α-thalassemia genotypes are indicated as Ⅰ~Ⅲ; HBA1 and HBA2 The two pairs of different bases were obtained, VAF1 and VAF2, indicating the α-thalassemia genotype IV.
[0081] 8. Combined with α-thalassemia genotype prompts Ⅰ~Ⅳ, output the α-thalassemia genotype of each sample to be analyzed.
[0082] Analyze the results The analysis results of Application Example 1 are shown in Table 6. Each DNA sample was verified using the multiplex ligation-dependent probe amplification (MLPA) method.
[0083] Table 6 Analysis results of application example 1
[0084] Application Example 2 1. Provide 45 DNA samples and refer to steps (1) to (7) of Application 1 to obtain an amplicon mixture of each DNA sample.
[0085] 2. The amplification mixture is sequenced to obtain the amplicon sequencing data of each DNA sample.
[0086] 3. Pre-mask the region chr16:223942-224152 in the genome reference sequence, align the amplicon sequencing data with the reference genome sequence, and obtain the number of reads and sequencing depth of each amplified region.
[0087] 4. Extract the amplicon sequencing data of the remaining sequence regions except the α-thalassemia target region from all DNA samples, and count the differences in sequencing depth of each remaining sequence region among all DNA samples in the remaining sequence regions to obtain the difference coefficient CV of each remaining sequence region, and sort all the difference coefficients CV from small to large. After sorting, extract the first 5 remaining sequence regions as internal reference regions.
[0088] 5. In each DNA sample, calculate the ratio of the number of reads in each target region to the number of reads in the five internal reference regions, and take the average of the five ratios as the A value of each target region.
[0089] 6. Use the random forest algorithm model to predict and screen negative samples in all DNA samples, and use the predicted and screened DNA samples as candidate samples; in each candidate sample, obtain the A value obtained for the internal reference area corresponding to each target area; in the set of all candidate sample A values, calculate the median Median of each target area, and according to the degree of discreteness of each A value of each candidate sample relative to the median Median, select the samples with the lowest expected number of discreteness from all candidate samples as the control samples of each control area; take the union of the control samples of all target areas, and if the number of control samples in the union result exceeds 10, extract the top 10 control samples that appear in the union process to obtain the final control sample set.
[0090] 7. For each sample to be analyzed, calculate the ratio of the sample to the A value of each target area in the 10 control samples, and take the average of the 10 ratios as the Ratio value of each target area; in the control sample set, calculate the ratio of each control sample to the A value of each target area in the 10 control samples, and take the average of multiple ratios of each target area as the Ratio value of each target area in each control sample.
[0091] 8. Each DNA sample has α-thalassemia genotypes Ⅰ~Ⅲ according to all Ratio values; HBA1 and HBA2 The two pairs of different bases were obtained, VAF1 and VAF2, indicating the α-thalassemia genotype IV.
[0092] 9. Combined with α-thalassemia genotype hints Ⅰ~Ⅳ, output the α-thalassemia genotype of each DNA sample.
[0093] Analyze the results The analysis results of Application Example 2 are shown in Table 7. Each DNA sample was verified using the multiplex ligation-dependent probe amplification (MLPA) method.
[0094] Table 7 Analysis results of application example 2
[0095] Application Example 3 The specific analysis steps refer to Application Example 1, wherein all control samples are the 10 control samples predicted and screened in Application Example 2. The analysis results of Application Example 3 are shown in Table 8, and each DNA sample is verified by multiplex ligation-dependent probe amplification (MLPA) method.
[0096] Table 8 Analysis results of application example 3
[0097] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be construed as limiting the scope of the present application. It should be noted that, for a person of ordinary skill in the art, several modifications and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. A method for analyzing α-thalassemia genotype, characterized in that: include: Step S10, obtaining amplicon sequencing data of multiple sequence regions in a DNA sample, wherein the DNA sample includes a sample to be analyzed and a control sample, and the sequence region includes a target region for α-thalassemia; Step S20, comparing the amplicon sequencing data with the genome reference sequence to obtain the number of reads and sequencing depth of each sequence region in each DNA sample; Step S30, obtaining the remaining sequence regions except the target region of α-thalassemia from all DNA samples, sorting the remaining sequence regions according to the stability of the sequencing depth of all DNA samples based on the same data amount, and selecting multiple stable remaining sequence regions as internal reference regions; Step S40, obtaining the ratio of the number of reads in each target region to the number of reads in multiple internal reference regions in each DNA sample, and calculating the average of the multiple ratios as the A value for each target region; Step S50, obtaining the ratio of the A value of each target area in the sample to be analyzed and the control sample, which is the Ratio value of each target area in the sample to be analyzed; Step S60: Analyze and output the α-thalassemia genotype hint of the sample to be tested according to the Ratio value of the target region.
2. The analysis method according to claim 1, characterized in that The amplicon sequencing data is obtained by performing a same-batch amplification reaction on all DNA samples using a polymerase chain reaction system containing multiple pairs of specific primers to obtain a mixture of amplicon of multiple sequence regions in each DNA sample, and then sequencing the amplicon mixture on a machine; The specific primers include sequences shown in SEQ ID NOs: 1-120.
3. The analysis method according to claim 1, characterized in that The specific operation of the stability ranking is: Count the differences of the remaining sequence regions among all DNA samples to obtain the coefficient of difference CV of each remaining sequence region, and sort all the coefficients of difference CV from small to large; The expected number of remaining sequence regions with the smallest coefficient of variation (CV) were extracted as internal reference regions.
4. The analysis method according to claim 1, characterized in that include: Performing predictive screening of negative samples among all DNA samples, and using the DNA samples obtained through the predictive screening as candidate samples, or using all DNA samples as a whole as candidate samples; Select control samples from the candidate samples.
5. The analysis method according to claim 4, characterized in that The selection of control samples from candidate samples includes the following steps: (1) Determine a target region and obtain the A value of the target region corresponding to the internal reference region in each candidate sample; (2) Calculate the median value of the target area in the set of A values of the target area of all candidate samples; (3) According to the degree of dispersion of the target area A value relative to the median value Median, select the samples with the lowest expected number of dispersion from all candidate samples as the control samples of the target area; (4) Execute steps (1) to (3) for each target area to obtain a control sample for each target area, and take the union of the control samples of all target areas to obtain the final control sample set.
6. The analysis method according to claim 1, characterized in that The DNA sample contains multiple control samples. The ratios of the A values of each target region in the sample to be analyzed and the multiple control samples are calculated, and the average of the multiple ratios is obtained as the Ratio value of each target region in the sample to be analyzed.
7. The analysis method according to claim 1, characterized in that include: Providing a DNA sample database with known α-thalassemia genotypes, including DNA samples of -α3.7, -α4.2, --SEA, --FIL, and --THAI, executing steps S10 to S50 to obtain the Ratio value of each target region in all DNA samples in the DNA sample database; According to the target region copies of all DNA samples in the DNA sample database and the Ratio values of the target region, a receiver operating characteristic curve analysis is performed to obtain a Ratio threshold for determining whether the copy of the target region is positive; The Ratio value of each target area in the sample to be analyzed is compared with the corresponding Ratio threshold value, and according to the comparison result, the α-thalassemia genotype prompt of the sample to be analyzed is analyzed and output.
8. The analysis method according to claim 1, characterized in that include: According to the ratio of the Ratio values of the target regions of -α4.2 and -α3.7, the α-thalassemia genotype hint of the sample to be analyzed is analyzed and output.
9. The analysis method according to claim 2, characterized in that The specific primers include those shown in SEQ ID NO: 123 and SEQ ID NO: 124 for amplification HBA1 and HBA2 primer pairs specific for the homologous region; The analysis method comprises: Pre-masked in the genome reference sequence HBA2 and HBA1 homologous regions, and then aligning the amplicon sequencing data with the reference genome sequence; For each pair of differential bases, the number of differential bases was counted, and the ratio of the number of one differential base to the total number of the two differential bases was calculated to obtain the VAF value; According to the VAF value, the α-thalassemia genotype hint of the sample to be analyzed is analyzed and output.
10. A computer device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the analysis method described in any one of claims 1 to 9.
Citation Information
Patent Citations
Method and kit for identifying thalassemia alpha alpha alpha anti4.2 heterozygosis and HK alpha alpha heterozygosis
CN114277096A
Method and device for detecting alpha-globin genotype
CN117976059A
Method for detecting gene copy number variation based on amplicon sequencing technology
CN119307599A
Method for genotyping on basis of detecting copy number variation of α-globin gene, and device thereof
WO2024199195A1