Method for analyzing sex and sex chromosome abnormality of sample based on next-generation sequencing data
Through the bioinformatics method based on second-generation sequencing data, using technical means such as Y chromosome homogeneity processing and hierarchical clustering, the accuracy of sample gender and sex chromosome abnormality analysis was solved, and the accurate detection and analysis of sexual development abnormalities and sex chromosome copy number was achieved.
Patent Information
- Application Number
- CN202510223998.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-17
AI Technical Summary
It is difficult to accurately analyze the gender and sex chromosome abnormalities in the sample, especially in the case of sample confusion, data errors or sexual development abnormalities, resulting in the failure to detect gender misregistration and sex chromosome copy number abnormalities in time.
The bioinformatics method based on second-generation sequencing data was used to initially divide the gender through Y chromosome homogenization processing and k-means clustering of the training set; then the sequencing depth homogenization and hierarchical clustering of the samples to be tested were performed to further confirm the gender, and the copy number of the sex chromosomes was determined by CNV analysis.
Accurate analysis of the sample gender and sex chromosome abnormalities is achieved, reducing the occurrence of gender misregistration and the occurrence of sex chromosome copy number abnormalities, and providing important tips for sexual developmental abnormalities, chimeric sex chromosomes and transgender cross-contamination.
Smart Images

Figure CN120164526A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of biomedical gene detection, and particularly to an analysis method for analyzing the gender and sex chromosome abnormalities of a sample based on next-generation sequencing data analysis. Background Art
[0002] Normally, a normal female has two X chromosomes (one from the father and the other from the mother) and no Y chromosome; while a normal male has one X chromosome (inherited from the mother) and one Y chromosome (inherited from the father).
[0003] Next-generation sequencing (NGS) technologies, such as whole exome sequencing (WES), low-depth whole genome copy number variation sequencing technology (CNVseq), and whole genome sequencing (WGS), have been widely used in the screening and auxiliary diagnosis of genetic diseases. During the sample submission process, problems such as mislabeling of samples, confusion of samples in the laboratory, incorrect correspondence between samples and data during data analysis, and semen residue in aborted fetus samples may lead to incorrect determination of gender, resulting in a mismatch between the registered gender and the detected gender.
[0004] Among the normal population, there are quite a number of men with normal phenotypes or slightly abnormal phenotypes who have supermale syndrome (chromosome constitution is 47XYY) or women with superfemale syndrome (chromosome constitution is 47XXX). Such populations may have relatively high sex hormone levels. Superfemale syndrome may have premature ovarian failure, some have mild intellectual abnormalities, and are prone to giving birth to offspring with chromosomal abnormalities. Some people with supermale syndrome have violent tendencies or infertility. Currently, there are few screening technologies that can provide prompt information about the true chromosomal conditions of such adults. There is another group of people, such as women without a uterus, men with abnormal testicular development, hypospadias, cryptorchidism, etc., whose genders are not clear, and most of these people cannot find a clear cause.
[0005] There are also many important gene regions on the sex chromosomes, such as STS, PLP1, DMD, MEPC2. If copy number variations (CNVs) occur in these regions, serious genetic diseases will be caused. Female X-linked recessive genetic CNV (such as DMD) heterozygous carriers also have a relatively high risk of giving birth to offspring with DMD. If such populations can be screened before pregnancy, such situations can be detected and prompted, and the occurrence of birth defects in offspring can be avoided through the preimplantation genetic testing technology (PGT) of the third-generation test tube embryo. To correctly analyze the CNV on the sex chromosome, accurately distinguishing the gender of the sample is a prerequisite.
[0006] Currently, the conventional method for detecting sex chromosome abnormalities is the chromosome karyotype technique, but this technique is usually only used for culturing amniotic fluid cells of samples with abnormal prenatal ultrasound, and has a long cycle, low positive rate, and cannot detect CNVs less than 5M. Summary of the Invention
[0007] To solve the technical problems of inconsistent registered gender and NGS data analysis gender, sexual development disorders, abnormal sex chromosome copy numbers, transgender cross - contamination, and difficulty in timely and accurate detection due to sample confusion or other errors, the present invention provides an analysis method for analyzing the gender and sex chromosome abnormalities of a sample based on next - generation sequencing (NGS) data by using bioinformatics methods.
[0008] To achieve the above - mentioned technical purpose, the technical solution of the present invention is as follows.
[0009] An analysis method for analyzing the gender and sex chromosome abnormalities of a sample based on next - generation sequencing data, comprising the following steps:
[0010] Step 1: Select multiple NGS raw data of known genders with the same library type as the training set, where the number of NGS raw data of known female samples is not less than 100. Then align the raw data of each sample in the training set to the reference genome and calculate the background noise region: Extract the data of each female sample aligned to the Y chromosome separately, and use samtools merge to take the union. Calculate the region with a sequencing depth greater than 3 as the region where females can be aligned to the Y chromosome, that is, the background noise region.
[0011] Step 2: Calculate the number of reads aligned to the Y chromosome for the sequencing data of each sample in the training set of known genders. After deducting the reads in the background noise region, divide it by the sum of the reads of the autosomes of each sample to obtain the normalized reads of the Y chromosome for each sample. Perform kmeans clustering on the normalized reads of the Y chromosome of the sequencing data of all these training sets of known genders to cluster into two categories; Denote the maximum value of the category with a smaller average value of the normalized reads of the Y chromosome as a, and the minimum value of the category with a larger average value of the normalized reads of the Y chromosome in the sample as b. Use (a + b) / 2 as the threshold for subsequent preliminary gender division.
[0012] Step 3: Align the sequencing data of each test sample of unknown gender to the reference genome, calculate the number of reads aligned to the Y chromosome, deduct the reads in the background noise region, divide it by the sum of the reads of its autosomes to obtain the normalized reads of the Y chromosome for each sample, and compare this value with the threshold for preliminary gender division. Those greater than (a + b) / 2 are preliminarily determined to be male, and those less than (a + b) / 2 are preliminarily determined to be female.
[0013] Step 4: When the number of samples of the less represented gender obtained after the preliminary gender division of the test samples in Step 3 is less than 5 samples, or when the total number of test samples is less than 10, then all test samples are marked as cases. At the same time, another set of sequencing data samples of the same type is added and marked as ctrls, and then gender analysis is performed together with the cases; otherwise, directly proceed to Step 5. Among them, the number of male and female samples in ctrls is not less than 50 each, and there is no disease phenotype and all chromosome copy numbers are normal.
[0014] Step 5: For the captured regions of the sequencing results of each sample in Step 4, after excluding the background noise regions, calculate the sequencing depth, and then perform sequencing depth normalization processing.
[0015] Step 6: After adding 1 to the normalized value of each captured region on the X and Y chromosomes and the individual Y chromosome respectively, then take log10, and then perform the ward.D method in hclust on this value to achieve hierarchical clustering, so as to obtain two clustering results and perform gender marking.
[0016] Step 7: Compare the gender markings of the two clustering results. When they are inconsistent, the gender marking corresponding to the hierarchical clustering result of the individual Y chromosome is used as the final gender result.
[0017] Step 8: If ctrls are added in Step 4, each test sample uses all samples of the same gender including cases and ctrls except itself as the candidate control samples for CNV analysis; otherwise, all cases samples of the same gender except itself are used as the candidate control samples for CNV analysis.
[0018] Step 9: Take the logarithm base 2 of the ratio of the median of the sequencing depth after normalization and deviation removal of each captured region between each test sample in cases and the same-gender CNV final control sample, and the result is denoted as log2ratio. Then perform fragmentation processing on the continuous log2ratio on the same chromosome to obtain the log2ratio of continuous CNV segments, and then reverse-infer the copy number of each captured region according to log2ratio.
[0019] Step 10: According to the gender result, the copy number of the X chromosome, and the copy number of the Y chromosome, obtain the final copy number of the sex chromosomes. If ctrls are introduced during gender analysis in Step 4, only the results of the test samples cases are extracted.
[0020] For the method described above, in Step 5, the process of calculating the sequencing depth for the captured regions of the sequencing results of each sample includes:
[0021] Obtain the raw sequencing data of multiple next-generation sequencing within the same sequencing batch, then align it with the hg19 reference genome. After sorting by genomic coordinates and marking repetitive sequences, calculate the sequencing depth for each captured region of the captured sequencing.
[0022] For the method described above, when the obtained raw sequencing data of next-generation sequencing is high-depth whole-genome sequencing data or low-depth whole-genome sequencing data, use a fixed length formed by at least 36bp continuous uniquely mappable regions on the reference genome as the sliding window region, and use the sliding window region as the virtual captured region to calculate the sequencing depth.
[0023] For the method described above, in step five, the sequencing depth normalization process is as follows:
[0024] Divide the sequencing depth of each captured region by the total sequencing depth sum of all captured regions on autosomes, namely chromosomes 1 to 22.
[0025] For the method described above, in step six, the gender marking based on the two clustering results includes:
[0026] After performing hierarchical clustering on the normalized sequencing depths of each captured region on the X and Y chromosomes, mark the group with a smaller sum of the normalized sequencing depths of each region on the X chromosome obtained by clustering as male, and the other group with a larger sum as female;
[0027] After performing hierarchical clustering on the normalized sequencing depths of each captured region on the Y chromosome, mark the group with a larger sum of the normalized sequencing depths of each region on the Y chromosome obtained by clustering as male, and the other group with a smaller sum as female.
[0028] For the method described above, in step eight, the selection of the final CNV control sample for each sample from the candidate CNV control samples with the same gender marking includes:
[0029] Calculate the pairwise Pearson correlation coefficients of the normalized sequencing depths of each corresponding captured region between each sample, and select all samples with a Pearson correlation coefficient greater than 0.9 for all captured regions of a single sample and the same gender marking as the final CNV control sample set for the test sample. If there is no sample with a Pearson correlation coefficient greater than 0.9 for a certain sample, then sort the Pearson correlation coefficients with it from large to small, and select the 10 samples with the largest Pearson correlation coefficient and the closest overall GC content as its final CNV control sample set.
[0030] The method is characterized in that after completing step seven, it further includes the step of separately calculating the copy number of the SRY gene for samples with the final gender result of female for the test samples:
[0031] Take the logarithm base 2 of the ratio of the normalized sequencing depth of the SRY gene in the female sample to the median of the normalized sequencing depth of the SRY gene in the male control sample, and record the result as log2ratio s , and use the result of 1 * 2 ^ (log2ratio s ) as the copy number of the final SRY gene in the female sample.
[0032] In the method described above, in step nine, the method for back-calculating the copy number of each capture region based on log2ratio includes:
[0033] The copy number formula is:
[0034] baseCN * 2 ^ (log2ratio)
[0035] Among them, for ordinary autosomal regions, regions on the X chromosome of female samples, and pseudoautosomal (PAR) regions of male samples, baseCN in the formula is set to 2, and baseCN is the basic copy number; for regions on the Y chromosome of female samples, baseCN in the formula is set to 0.0015; for the X chromosome and Y chromosome regions of male samples except the PAR region, baseCN in the formula is set to 1.
[0036] In the method described above, after completing step four, it further includes a step of detecting whether there is cross-contamination between samples in the high-depth whole-genome sequencing data:
[0037] Filter out low-quality SNVs on all chromosomes of the high-depth whole-genome sequencing data, and then analyze the number of heterozygous SNVs on the X chromosome; then calculate the CHARR value and INCONSISTENT_AB_HET_RATE value for each sample based on the SNVs to detect whether there is cross-contamination between samples; where CHARR is the contamination of the sequencing reads infiltrating the reference genome in homozygous mutant variants, and INCONSISTENT_AB_HET_RATE is the proportion of inconsistent allele heterozygosity.
[0038] In the method described above, detecting whether there is cross-contamination between samples includes:
[0039] When CHARR > 0.02 | INCONSISTENT_AB_HET_RATE > 0.1, it indicates that the sample is contaminated with 1 - 5% of other samples;
[0040] When CHARR > 0.03 | INCONSISTENT_AB_HET_RATE > 0.15, it indicates that the sample is contaminated with more than 5% of other samples.
[0041] After obtaining the number of heterozygous SNVs on the X chromosome, it further includes a verification step based on the final gender result in step seven:
[0042] Samples with more than 95% of SNV heterozygous deletions on the X chromosome are marked as male, and then the final gender result of this sample in step seven is verified. If the final gender result is female, then the X chromosome copy number of this sample in step seven is further verified to see if it is 2. If so, it means that this sample has X chromosome uniparental isodisomy, that is, iUPD.
[0043] For the described method, after completing step ten, it further includes an analysis step for the deletion and chimerism of the sample chromosomes:
[0044] 1) If the sample gender is determined to be female, the X chromosome copy number is 1, and the Y chromosome copy number is less than 0.015, then the sample is entirely composed of 45X cells;
[0045] 2) If the sample gender is determined to be female, the X chromosome copy number is greater than 1 but less than 2, and the Y chromosome copy number is less than 0.015, then the sample is a 45X / 46XX chimera, that is, it has both some 45X cells and some 46XX cells; or a 45X / 47XXX chimera, that is, it has both 45X cells and some 47XXX cells, and the number of 45X karyotype cells is greater than that of 47XXX, and there are no 46XX karyotype cells;
[0046] 3) If the sample gender is determined to be female, the X chromosome copy number is greater than 2 but less than 3, and the Y chromosome copy number approaches 0, then the sample is a 46XX / 47XXX chimera, that is, it has both some 46XX cells and some 47XXX cells;
[0047] 4) If the sample gender is determined to be female, the X chromosome copy number is 2, the Y chromosome copy number approaches 0, but the SRY gene copy number is 1, then the sample has a normal 46XX karyotype and has one copy of the SRY gene;
[0048] 5) If the sample gender is determined to be female, the X chromosome copy number is 3, and the Y chromosome copy number approaches 0, then the sample is entirely composed of 47XXX cells;
[0049] 6) If the sample gender is determined to be male, the X chromosome copy number is 1, and the Y chromosome copy number is greater than 0.1 but less than 0.9, then the sample is a 45X / 46XY chimera, that is, it has both some 45X cells and some 46XY cells;
[0050] 7) If the sample sex is determined to be male, and the copy number of the X chromosome is greater than 1 but less than 2, and the copy number of the Y chromosome is greater than 0.1 but less than 0.9, then the sample is a 45X / 46XY / 47XXY chimera, that is, part of the cells are 45X, part of the cells are 46XY, and part of the cells are 47XXY;
[0051] 8) If the sample sex is determined to be male, the copy number of the X chromosome is greater than 1 but less than 2, and the copy number of the Y chromosome is 1, then the sample is a 46XY / 47XXY chimera, that is, part of the cells are 46XY, and part of the cells are 47XXY;
[0052] 9) If the sample sex is determined to be male, the copy number of the X chromosome is 2, and the copy number of the Y chromosome is greater than 0.1 but less than 0.9, then the sample is a 46XX / 47XXY chimera, that is, part of the cells are 46XX, and part of the cells are 47XXY;
[0053] 10) If the sample sex is determined to be male, the copy number of the X chromosome is 2, and the copy number of the Y chromosome is 1, then the sample is 47XXY, that is, it is composed of all 47XXY cells;
[0054] 11) If the sample sex is determined to be male, the copy number of the X chromosome is 1, and the copy number of the Y chromosome is 2, then the sample is 47XYY, that is, it is composed of all 47XYY cells;
[0055] 12) If the sample sex is determined to be male, the copy number of the X chromosome is 1, and the copy number of the Y chromosome is greater than 1 but less than 2, then the sample is a 46XY / 47XYY chimera, that is, part of the cells are 46XY, and part of the cells are 47XYY.
[0056] The technical effect of the present invention is that the present invention can further assist in determining the sex of the sample according to whether there is heterozygous deletion of the SNV on the X chromosome, can also indicate the presence of (chimeric) ROH on the X chromosome, and can also indicate a lower proportion of transgender cross-contamination. The present invention provides important clues for situations such as abnormal sexual development, abnormal sex chromosome copy number (including chimerism), transgender cross-contamination between samples, and wrong sample information recording of transgender. Brief Description of the Drawings
[0057] Figure 1 It is a hierarchical clustering diagram (based on the ward.D method in hclust) of the normalized sequencing depth of the X and Y chromosomes. Each row represents a sample, the upper half are the clustered female samples, the lower half are the clustered male samples, the right side is the hierarchy of the clustered sex, each column represents the normalized sequencing depth of a sequencing region, the left part is the Y chromosome, and the right part is the X chromosome (the X chromosome is longer than the Y chromosome).
[0058] Figure 2 It is a hierarchical clustering diagram of the sequencing depth after Y chromosome-based normalization (based on the ward.D method in hclust). Each row represents a sample. The upper half shows the clustered female samples, the lower half shows the clustered male samples, and the right side shows the hierarchy of the clustered genders. Each column represents the sequencing depth after normalization of a sequencing region.
[0059] Figure 3 It is a schematic diagram of the VAF distribution of 45X.
[0060] Figure 4 It is a schematic diagram of the VAF distribution of 45Xmos / 46XX.
[0061] Figure 5 It is a schematic diagram of the VAF distribution of 45Xmos / 46XY.
[0062] Figure 6 It is seq[GRCh37]del(X)(p22.33p11.21)chrX:g.200837_57936883del;
[0063] A schematic diagram of the VAF distribution of seq[GRCh37]dup(X)(q11.1q28)chrX:g.62569918_155240163dup.
[0064] Figure 7 It is a schematic diagram of the VAF distribution of 47XXX.
[0065] Figure 8 It is a schematic diagram of the VAF distribution of 47XXY.
[0066] Figure 9 It is a schematic diagram of the VAF distribution of 46XY / 47XYYmos.
[0067] Figure 10 It is a schematic diagram of the VAF distribution of 47XYY.
[0068] Figure 11 It is a schematic diagram of the VAF distribution of 46XX iUPD.
[0069] Figure 12 It is a schematic diagram of the VAF distribution of 45X(p22.11p22.33dup).
[0070] Figure 13 It is a schematic diagram of the VAF distribution of 45,X;seq[GRCh37]dup(X)(p11.3p11.23)chrX:g.46387744_46713604dup.
[0071] Figure 14It is a schematic diagram of the VAF distribution with a low proportion of cross - contamination in the sample. (In the VAF distribution diagram, the contamination of sequencing reads infiltrating the reference genome in homozygous mutant variations can be clearly seen.) Detailed implementation manners
[0072] See Figures 1 - 14 . To further illustrate each embodiment, the present invention provides accompanying drawings, which are a part of the disclosure of the present invention. They are mainly used to illustrate the embodiments and can be combined with the relevant descriptions in the specification to explain the operating principle of the embodiments. With reference to these contents, those of ordinary skill in the art should be able to understand other possible implementation manners and the advantages of the present invention. An analysis method for analyzing the gender and sex chromosome abnormalities of a sample based on second - generation sequencing data provided in this embodiment includes:
[0073] Step 1: Select multiple NGS raw data of known genders with the same library type as the training set. For example, they are all whole - exome sequencing from the same brand of kit or all whole - genome sequencing. The raw data is not less than 200, and the NGS raw data of known female samples is not less than 100. Then align the raw data of each sample in the training set to the reference genome and calculate the background noise region: Extract the data of each female sample aligned to the Y chromosome separately and use samtools merge to take the union. Calculate the region with a sequencing depth greater than 3 as the region where females can be aligned to the Y chromosome. Since these regions may interfere with gender analysis, these regions need to be subtracted when calculating the Y - chromosome region later, that is, these regions are the background noise regions.
[0074] Step 2: Calculate the number of reads aligned to the Y chromosome for the sequencing data of each sample in the training set of known genders in Step 1. After deducting the reads in the background noise region, that is, the reads in the region where female samples can be aligned to the Y chromosome calculated in Step 1, divide it by the sum of the reads of the autosomes of each sample to obtain the normalized reads of the Y chromosome for each sample. Perform kmeans clustering on the normalized reads of the Y chromosome of the sequencing data of all these training sets of known genders to cluster them into 2 categories. Denote the maximum value of the category with a smaller average value of the normalized reads of the Y chromosome as a, and the minimum value of the category with a larger average value of the normalized reads of the Y chromosome in the sample as b. Use (a + b) / 2 as the empirical value threshold for subsequent preliminary gender division.
[0075] Step 3: Align the sequencing data of each test sample of unknown gender (1 to n samples) to the reference genome, calculate the number of reads aligned to the Y chromosome, subtract the reads in the background noise region, and then divide by the sum of the reads on its autosomes (chromosomes 1 - 22). Obtain the normalized reads number of the Y chromosome for each sample, and compare this value with the threshold for preliminary gender classification. Among them, those greater than (a + b) / 2 are preliminarily determined to be male, and those less than (a + b) / 2 are preliminarily determined to be female.
[0076] Step 4: When the number of samples of the less common gender obtained after the preliminary gender classification of the test samples in Step 3 is less than 5 samples (for example, only 0 - 5 females and the rest are males; or only 0 - 5 males and the rest of the samples are females), or when the total number of test samples is less than 10 (i.e., 1 - 9), then mark all test samples as cases, and at the same time add another set of sequencing data sample sets of the same type and mark them as ctrls, and perform gender analysis together with the cases; otherwise, directly proceed to Step 5; among them, the number of male and female samples in ctrls is not less than 50, and there is no disease phenotype and all chromosome copy numbers are normal.
[0077] Then filter out low-quality SNVs on all chromosomes of the high-depth whole-genome sequencing data, and then analyze the number of heterozygous SNVs on the X chromosome; then calculate the CHARR value and INCONSISTENT_AB_HET_RATE value for each sample based on the SNVs to detect whether there is cross-contamination between samples; where CHARR is the contamination of the sequencing reads infiltrated into the reference genome in homozygous mutant variants, and INCONSISTENT_AB_HET_RATE is the proportion of inconsistent allelic heterozygosity.
[0078] Among them, detecting whether there is cross-contamination between samples includes:
[0079] When CHARR > 0.02 | INCONSISTENT_AB_HET_RATE > 0.1, it indicates that the sample is contaminated with 1 - 5% of other samples;
[0080] When CHARR > 0.03 | INCONSISTENT_AB_HET_RATE > 0.15, it indicates that the sample is contaminated with more than 5% of other samples.
[0081] Step 5: For the capture regions of the sequencing results of each sample in Step 4 (if there are ctrls, ctrls need to be included), calculate the sequencing depth after excluding the background noise region, and then perform sequencing depth normalization processing.
[0082] Specifically for this Step 5, first, obtain the original sequencing data of multiple cases of next-generation sequencing within the same sequencing batch, then align it with the hg19 reference genome. After sorting by genomic coordinates and marking repetitive sequences, calculate the sequencing depth for each captured region of the captured sequencing. If the obtained original sequencing data of next-generation sequencing is high-depth whole-genome sequencing data or low-depth whole-genome sequencing data, use the fixed length connected by at least 36bp continuous uniquely aligned regions on the reference genome as the sliding window region, and use the sliding window region as the virtual captured region to calculate the sequencing depth.
[0083] Then perform sequencing depth normalization: divide the sequencing depth of each captured region by the total sequencing depth sum of all captured regions on autosomes, i.e., chromosomes 1 to 22.
[0084] In Step 6, after adding 1 to the normalized value of each captured region on the X and Y chromosomes and the separate Y chromosome, then take log10, and then perform the ward.D method in hclust to achieve hierarchical clustering (hierarchical clustering uses the ward.D method in hclust), thus obtaining two clustering results and performing gender marking:
[0085] After performing hierarchical clustering on the normalized sequencing depths of each captured region on the X and Y chromosomes, mark the group with a smaller sum of the normalized sequencing depths of each region on the X chromosome obtained by clustering as male, and the other group with a larger sum as female;
[0086] After performing hierarchical clustering on the normalized sequencing depths of each captured region on the Y chromosome, mark the group with a larger sum of the normalized sequencing depths of each region on the Y chromosome obtained by clustering as male, and the other group with a smaller sum as female.
[0087] In Step 7, compare the gender markings of the two clustering results. When they are inconsistent, use the gender marking corresponding to the hierarchical clustering result of the separate Y chromosome as the final gender result. Then, separately calculate the copy number of the SRY gene for samples with the final gender result of female for the test samples:
[0088] Take the log2 of the ratio of the normalized sequencing depth of the SRY gene in female samples to the median of the normalized sequencing depth of the SRY gene in male control samples, record the result as log2ratios, and use the result of 1*2^(log2ratios) as the final copy number of the SRY gene in female samples.
[0089] Then, according to the results obtained by analyzing the number of heterozygous SNVs on the X chromosome in Step 4, samples with more than 95% SNV heterozygous deletions on the X chromosome are marked as male. Then, check the final gender result of this sample in Step 7. If the final gender result is female, then check whether the copy number of the X chromosome of this sample in Step 7 is 2. If so, it means that this sample has X chromosome uniparental isodisomy, i.e., iUPD.
[0090] Step 8: If ctrls are added in Step 4, for each sample to be tested, all samples of the same gender including cases and ctrls except itself are used as candidate control samples for CNV analysis. Otherwise, all cases samples of the same gender except itself are used as candidate control samples for CNV analysis. That is, calculate the pairwise Pearson correlation coefficients of the normalized sequencing depths of each corresponding capture region between each sample, and select all samples with Pearson correlation coefficients greater than 0.9 for all capture regions of a single sample and the same gender marker as the final CNV control sample set of the sample to be tested. If there is no sample with a Pearson correlation coefficient greater than 0.9 for a certain sample, then arrange the Pearson correlation coefficients with it from large to small, and select the 10 samples with the largest Pearson correlation coefficient and the closest overall GC content to it as its final CNV control sample set.
[0091] Step 9: Take the logarithm to the base 2 of the ratio of the median of the normalized and outlier-removed sequencing depths of each capture region between each sample to be tested in cases and the same-gender CNV final control samples, and record the result as log2ratio. Then, fragment the continuous log2ratio on the same chromosome to obtain the log2ratio of continuous CNV segments. Then, infer the copy number of each capture region based on log2ratio, where the copy number formula is:
[0092] baseCN*2^(log2ratio)
[0093] Among them, for ordinary autosomal regions, regions on the X chromosome of female samples, and pseudoautosomal (PAR) regions of male samples, baseCN in the formula is set to 2, and baseCN is the base copy number; for regions on the Y chromosome of female samples, baseCN in the formula is set to 0.0015; for the X chromosome and Y chromosome regions of male samples except the PAR region, baseCN in the formula is set to 1.
[0094] Step 10: Obtain the final copy number of the sex chromosomes based on the gender result, X chromosome copy number, and Y chromosome copy number. If ctrls are introduced when analyzing gender in Step 4, only extract the results of the samples to be tested in cases.
[0095] Finally, the analysis of the deletion and chimerism of the sample chromosomes is as follows:
[0096] 1) If the sample gender is determined to be female, the copy number of the X chromosome is 1, and the copy number of the Y chromosome is less than 0.015, then the sample is entirely composed of 45X cells;
[0097] 2) If the sample gender is determined to be female, the copy number of the X chromosome is greater than 1 but less than 2, and the copy number of the Y chromosome is less than 0.015, then the sample is a 45X / 46XX chimera, that is, there are both some 45X cells and some 46XX cells; or a 45X / 47XXX chimera, that is, there are 45X cells and some 47XXX cells, and the number of 45X karyotype cells is greater than that of 47XXX, and there are no 46XX karyotype cells;
[0098] 3) If the sample gender is determined to be female, the copy number of the X chromosome is greater than 2 but less than 3, and the copy number of the Y chromosome approaches 0, then the sample is a 46XX / 47XXX chimera, that is, there are both some 46XX cells and some 47XXX cells;
[0099] 4) If the sample gender is determined to be female, the copy number of the X chromosome is 2, the copy number of the Y chromosome approaches 0, but the copy number of the SRY gene is 1, then the sample has a normal 46XX karyotype and one copy of the SRY gene;
[0100] 5) If the sample gender is determined to be female, the copy number of the X chromosome is 3, and the copy number of the Y chromosome approaches 0, then the sample is entirely composed of 47XXX cells;
[0101] 6) If the sample gender is determined to be male, the copy number of the X chromosome is 1, and the copy number of the Y chromosome is greater than 0.1 but less than 0.9, then the sample is a 45X / 46XY chimera, that is, there are both some 45X cells and some 46XY cells;
[0102] 7) If the sample gender is determined to be male, the copy number of the X chromosome is greater than 1 but less than 2, and the copy number of the Y chromosome is greater than 0.1 but less than 0.9, then the sample is a 45X / 46XY / 47XXY chimera, that is, some cells are 45X, some cells are 46XY, and some cells are 47XXY;
[0103] 8) If the sample gender is determined to be male, the copy number of the X chromosome is greater than 1 but less than 2, and the copy number of the Y chromosome is 1, then the sample is a 46XY / 47XXY chimera, that is, some cells are 46XY and some cells are 47XXY;
[0104] 9) If the sample sex is determined to be male, the copy number of the X chromosome is 2, and the copy number of the Y chromosome is greater than 0.1 but less than 0.9, then the sample is 46XX / 47XXY mosaic, that is, part of the cells are 46XX and part of the cells are 47XXY;
[0105] 10) If the sample sex is determined to be male, the copy number of the X chromosome is 2, and the copy number of the Y chromosome is 1, then the sample is 47XXY, that is, it is composed of all 47XXY cells;
[0106] 11) If the sample sex is determined to be male, the copy number of the X chromosome is 1, and the copy number of the Y chromosome is 2, then the sample is 47XYY, that is, it is composed of all 47XYY cells;
[0107] 12) If the sample sex is determined to be male, the copy number of the X chromosome is 1, and the copy number of the Y chromosome is greater than 1 but less than 2, then the sample is 46XY / 47XYY mosaic, that is, part of the cells are 46XY and part of the cells are 47XYY.
[0108] Example
[0109] 1) Select the original sequencing data (in fastq format) of approximately 7,000 whole exome sequencing samples (the kit is KAPA HyperExome V2 from Roche) with a capture region size of 43.11M from the MGI T7 sequencing platform;
[0110] 2) Use fastp (PMID: 30423086) to remove adapters and low-quality reads, align the original data to the GRCh37 reference genome using bwa mem (PMID: 19451168), sort the reads using samtools sort (PMID: 19505943), and mark duplicates using GATK4.4 markduplicates (PMID: 20644199) to obtain a bam file, and build an index using samtools index;
[0111] 3) For each sample, use samtools bedcov to calculate the sequencing depth of each capture region, and divide the sequencing depth of each capture region by the sum of the total sequencing depths of all capture regions on autosomes to obtain the normalized depth;
[0112] 4) Use the ward.D method in the hclust function of the R language to perform hierarchical clustering on the normalized sequencing depths of each captured region on the X and Y chromosomes. Mark the group with a smaller sum of the normalized sequencing depths of each region on the X chromosome as male, and the rest as female samples. At the same time, perform hierarchical clustering on the normalized sequencing depths of each captured region on the Y chromosome alone, and mark the group with a larger sum of the normalized sequencing depths of each region on the Y chromosome as male, and the rest as female samples. If there is a gender inconsistency between the two methods, the gender calculated based on the Y chromosome alone is taken as the final gender.
[0113] 5) Perform PCA clustering analysis on the normalized sequencing depths of each region of all samples. Samples that are clustered into the same group and have the same gender are used as control samples for each other. The median of the normalized sequencing depths of each captured region of the control samples after removing outliers is used as the final control reference value. The ratio of the test sample to the final control reference value is taken, and then log2 is taken to obtain log2ratio. Then, use the DNAcopy in the R language to perform fragmentation processing on each chromosome to obtain the fragmented log2ratio.
[0114] 6) Calculate the copy numbers of different regions of the sex chromosomes according to the formula baseCN*2^(log2ratio) based on the different genders and different regions of the sex chromosomes of each sample.
[0115] 7) Detect SNVs for all samples according to the GATK best practice, and filter out low-quality SNVs (including cases such as low sequencing depth, VAF bias in multiple samples, and low quality values in multiple samples), and calculate the number of heterozygous SNVs on the X chromosome. In normal cases, samples with most of the X chromosome being heterozygous deletions are marked as male, otherwise as female. If the gender calculated in step 4) is a female sample, and the copy number of the X chromosome calculated in step 6) is 2, and the entire X chromosome shows a situation of missing heterozygous SNVs, it is judged that there is iUPD on the X chromosome.
[0116] 8) Use the SCE-VCF (https: / / github.com / HTGenomeAnalysisUnit / SCE-VCF) software to calculate the cross-contamination situation between samples of each sample.
[0117] 9) Calculate the copy number of the SRY gene for all female samples separately.
[0118] 10) Calculate the copy number (chimerism) situation of each sex chromosome.
[0119] Table 1 Detection of sex chromosomes in the samples of the examples
[0120]
[0121]
[0122]
[0123] It should be emphasized that for the CNV or SV results detected based on next-generation sequencing data, other detection techniques such as array CGH, MLPA, PCR, and FISH should be used in combination with the actual situation to verify the phenotype-associated variations. The laboratory and the clinic need to jointly explore the report disclosure and result verification plan. Therefore, the present invention is only for reference and cannot be directly used as the final result.
Claims
1. A method for analyzing the gender and sex chromosome abnormalities of a sample based on second-generation sequencing data, characterized in that: The following steps are involved: Step 1: Select multiple NGS raw data of known genders of the same library type as the training set, where there are no less than 100 known female sample NGS raw data. Then, align the raw data of each sample in the training set to the reference genome, and calculate the background noise area: extract the data of each female sample aligned to the Y chromosome separately, and use samtoolsmerge to take the union, and calculate the area with a sequencing depth greater than 3 as the area that can be aligned to the Y chromosome for females, that is, the background noise area; Step 2: Calculate the number of reads aligned to the Y chromosome for each sample sequencing data in the known gender training set in step 1, deduct the reads in the background noise area, and divide it by the sum of the reads of each sample autosome to obtain the normalized number of reads of the Y chromosome of each sample; perform kmeans clustering on the normalized number of reads of the Y chromosome of all the sequencing data of the known gender training sets to cluster them into 2 categories; the maximum value of the category with a smaller average value of the normalized number of reads of the Y chromosome in the sample is recorded as a, and the minimum value of the category with a larger average value of the normalized number of reads of the Y chromosome in the sample is recorded as b, and (a+b) / 2 is used as the threshold for subsequent preliminary gender division; Step 3: Align the sequencing data of each unknown sex sample to the reference genome, calculate the number of reads aligned to the Y chromosome, deduct the reads in the background noise area, and divide it by the sum of the reads of the autosomes to obtain the normalized number of reads of the Y chromosome of each sample. This value is compared with the threshold for preliminary gender division, where those greater than (a+b) / 2 are preliminarily judged as males, and those less than (a+b) / 2 are preliminarily judged as females; Step 4: When the number of samples with the least gender obtained after the preliminary gender division of the samples to be tested in step 3 is less than 5 samples, or the total number of samples to be tested is less than 10, all samples to be tested are marked as cases, and another set of sequencing data samples of the same type of library is added and marked as ctrls, and then gender analysis is performed together with cases; otherwise, go directly to step 5; the number of male and female samples in ctrls is not less than 50, and there is no disease phenotype, and the copy number of all chromosomes is normal; Step 5, for the captured region of the sequencing results of each sample in step 4, after excluding the background noise region, the sequencing depth is calculated, and then a sequencing depth normalization process is performed; Step 6: After adding 1 to the normalized value of each captured region on the X and Y chromosomes and the single Y chromosome, log10 is taken, and then the ward.D method in hclust is executed on this value to realize hierarchical clustering, thereby obtaining two clustering results and marking the gender; Step 7: Compare the sex markers of the two clustering results. If they are inconsistent, the sex marker corresponding to the hierarchical clustering result of the single Y chromosome is used as the final sex result. Step 8: If ctrls is added in step 4, all samples of the same sex including cases and ctrls except the sample to be tested are used as candidate control samples for CNV analysis; otherwise, all case samples of the same sex except the sample to be tested are used as candidate control samples for CNV analysis; Step 9: Take the log2 of the median ratio of the sequencing depth of each sample to be tested in cases and the final control sample of the same sex after normalization and removal of deviation values in each capture region. The result is recorded as log2ratio, and the continuous log2ratio on the same chromosome is fragmented to obtain the log2ratio of the continuous CNV fragments, and then the copy number of each capture region is inferred based on the log2ratio; Step 10, according to the gender result, X chromosome copy number and Y chromosome copy number, the final copy number of the sex chromosome is obtained. If ctrls is introduced when analyzing the gender in step 4, only the results of the sample cases to be tested are extracted.
2. The method according to claim 1, characterized in that In step 5, the process of calculating the sequencing depth of the capture region of the sequencing result of each sample includes: The raw sequencing data of multiple second-generation sequencing samples in the same sequencing batch were obtained, and then compared with the hg19 reference genome. After sorting by genomic coordinates and marking repeated sequences, the sequencing depth was calculated for each capture region of the capture sequencing.
3. The method according to claim 2, characterized in that When the raw sequencing data obtained by the second-generation sequencing is high-depth whole-genome sequencing data or low-depth whole-genome sequencing data, a fixed length consisting of at least 36bp continuous unique alignment regions on the reference genome is used as the sliding window area, and the sliding window area is used as the virtual capture area to calculate the sequencing depth.
4. The method according to claim 1, characterized in that In the step 5, the sequencing depth normalization process is as follows: The sequencing depth of each captured region was divided by the sum of the total sequencing depths of all captured regions on the autosomes, i.e. chromosomes 1 to 22.
5. The method according to claim 1, characterized in that In the step 6, gender marking according to the two clustering results includes: After hierarchical clustering of the normalized sequencing depth of each captured region on chromosomes X and Y, the group with the smaller sum of the normalized sequencing depth of each region on chromosome X obtained by clustering is marked as male, and the other group with the larger sum is marked as female; After hierarchical clustering of the normalized sequencing depth of each captured region on the Y chromosome, the group with a larger sum of the normalized sequencing depth of each region on the clustered Y chromosome is marked as male, and the other group with a smaller sum is marked as female.
6. The method according to claim 1, characterized in that In the step eight, selecting a final CNV control sample from candidate CNV control samples with the same sex marker for each sample includes: Calculate the pairwise Pearson correlation coefficients of the sequencing depth after normalization of each corresponding capture region between each sample, and select all samples with the same sex marker and Pearson correlation coefficients greater than 0.9 for all capture regions of a single sample as the final CNV control sample set for the sample to be tested. If there are no samples with Pearson correlation coefficients greater than 0.9 for a certain sample, arrange the Pearson correlation coefficients with it from large to small, and select the 10 samples with the largest Pearson correlation coefficient and the closest overall GC content as its final CNV control sample set.
7. The method according to claim 1, characterized in that After completing step 7, the step of calculating the copy number of the SRY gene for samples whose final gender result is female is also included: The log2ratio is the ratio of the normalized sequencing depth of the SRY gene in female samples to the median of the normalized sequencing depth of the SRY gene in male control samples. s , use 1*2^(log2ratio s ) was used as the final SRY gene copy number of the female sample.
8. The method according to claim 1, characterized in that In step nine, inferring the copy number of each capture region according to log2ratio includes: The copy number formula is: baseCN*2^(log2ratio) For the common autosomal regions, the regions on the X chromosome of female samples, and the pseudoautosomal (PAR) regions of male samples, the baseCN in the formula is set to 2, where baseCN is the basic copy number; for the regions on the Y chromosome of female samples, the baseCN in the formula is set to 0.0015; for the X chromosome and Y chromosome regions of male samples except the PAR region, the baseCN in the formula is set to 1.
9. The method according to claim 3, characterized in that: After completing step 4, the following steps are also included to detect whether there is cross-contamination between samples in the high-depth whole genome sequencing data: Low-quality SNVs were filtered out for all chromosomes of the high-depth whole-genome sequencing data, and then the number of heterozygous SNVs on chromosome X was analyzed. The CHARR value and INCONSISTENT_AB_HET_RATE value of each sample were calculated based on the SNV to detect whether there was cross-contamination between samples. CHARR is the sequencing read contamination of the homozygous mutant variant introgressed into the reference genome, and INCONSISTENT_AB_HET_RATE is the inconsistent allele heterozygous ratio.
10. The method according to claim 9, characterized in that Situations where cross contamination between samples is detected include: When CHARR>0.02|INCONSISTENT_AB_HET_RATE>0.1, it means that the sample is contaminated by 1-5% of other samples; When CHARR>0.03|INCONSISTENT_AB_HET_RATE>0.15, it means that the sample is contaminated by more than 5% of other samples.
11. The method according to claim 9, characterized in that After obtaining the number of heterozygous SNVs on the analyzed X chromosome, a step of checking based on the final gender result of step seven is also included: The sample with more than 95% SNV heterozygous deletion on the X chromosome is marked as male, and then the final gender result of this sample in step 7 is checked. If the final gender result is female, then check whether the X chromosome copy number of this sample in step 7 is 2. If so, it means that this sample has X chromosome uniparental disomy, iUPD.
12. The method according to claim 9, characterized in that After completing step 10, the following steps are also included to analyze the deletion and mosaicism of the sample chromosomes: 1) If the sample gender is determined to be female, the X chromosome copy number is 1, and the Y chromosome copy number is less than 0.015, then the sample is composed entirely of 45X cells; 2) If the gender of the sample is determined to be female, the number of copies of chromosome X is greater than 1 but less than 2, and the number of copies of chromosome Y is less than 0.015, then the sample is 45X / 46XX mosaicism, that is, there are both some 45X cells and some 46XX cells; or 45X / 47XXX mosaicism, that is, there are both 45X cells and some 47XXX cells, and the number of 45X karyotype cells is greater than 47XXX, and there are no 46XX karyotype cells; 3) If the gender of the sample is determined to be female, the X chromosome copy number is greater than 2 but less than 3, and the Y chromosome copy number is close to 0, then the sample is 46XX / 47XXX mosaic, that is, there are both some 46XX cells and some 47XXX cells; 4) If the sex of the sample is determined to be female, the copy number of chromosome X is 2, the copy number of chromosome Y is close to 0, but the copy number of SRY gene is 1, then the sample has a normal 46XX karyotype and has one copy of SRY gene; 5) If the sample gender is determined to be female, the X chromosome copy number is 3, and the Y chromosome copy number is close to 0, then the sample is composed entirely of 47XXX cells; 6) If the sample is determined to be male, the X chromosome copy number is 1, and the Y chromosome copy number is greater than 0.1 but less than 0.9, then the sample is 45X / 46XY mosaicism, that is, there are both some 45X cells and some 46XY cells; 7) If the sample gender is determined to be male, the X chromosome copy number is greater than 1 but less than 2, and the Y chromosome copy number is greater than 0.1 but less than 0.9, then the sample is 45X / 46XY / 47XXY mosaic, that is, some cells are 45X, some cells are 46XY, and some cells are 47XXY; 8) If the sample gender is determined to be male, the X chromosome copy number is greater than 1 but less than 2, and the Y chromosome copy number is 1, then the sample is 46XY / 47XXY mosaicism, that is, some cells are 46XY and some cells are 47XXY; 9) If the sample gender is determined to be male, the X chromosome copy number is 2, and the Y chromosome copy number is greater than 0.1 but less than 0.9, the sample is 46XX / 47XXY mosaic, that is, some cells are 46XX and some cells are 47XXY; 10) If the sample gender is determined to be male, the X chromosome copy number is 2, and the Y chromosome copy number is 1, then the sample is 47XXY, that is, it is composed entirely of 47XXY cells; 11) If the gender of the sample is determined to be male, the number of copies of the X chromosome is 1, and the number of copies of the Y chromosome is 2, then the sample is 47XYY, that is, it is composed entirely of 47XYY cells; 12) If the sex of the sample is determined to be male, the copy number of chromosome X is 1, and the copy number of chromosome Y is greater than 1 but less than 2, then the sample is 46XY / 47XYY mosaicism, that is, some cells are 46XY and some cells are 47XYY.