A method and device for identifying maternal contamination and detecting copy number abnormalities using peaks

By detecting the SNP site information of the abortion sample, using a mixed Gaussian model and sliding window to determine the parent source contamination and copy number abnormalities, the problem of the impact of parent cell contamination is solved, and rapid and accurate detection is achieved, reducing costs and time.

CN119673271BActive Publication Date: 2025-07-25巫昭祺
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411548646.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-01
Publication Date
2025-07-25
Estimated Expiration
2044-11-01

AI Technical Summary

Technical Problem

The prior art detects copy number variation in abortion samples, which is affected by maternal cell contamination, resulting in inaccurate detection results, requiring additional maternal samples and experimental steps, increasing costs and time.

Method used

By detecting the SNP site information of the detected sample, the peak number and peak value of the autosome are obtained using a mixed Gaussian model and sliding window, and combining the LRR characteristic data to determine parent contamination and copy number abnormalities, without additional parent samples.

Benefits of technology

Quickly judge the sample parent source pollution, improve detection efficiency, reduce costs, improve detection comprehensiveness, and reduce time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119673271B_ABST
    Figure CN119673271B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for identifying maternal contamination and detecting copy number abnormalities using peaks, comprising: S10: detecting SNP locus information of a sample to be tested, and extracting LRR feature data, BAF feature data, and genotype feature data of each autosome of the sample to be tested based on the genomic position order of SNP loci; S20: obtaining the number of peaks and peak values on each autosome according to the BAF feature data, and judging the chromosomal state of each autosome according to the number of peaks, peak values, and LRR feature data; S30: judging the sample state of the sample to be tested according to the chromosomal state with the highest proportion among all chromosomal states; S40: obtaining the CNV result and ROH result of the sample to be tested, and / or judging whether the sample to be tested is a whole-genome ROH according to the LRR feature data, BAF feature data, and genotype feature data. The present invention uses the detection data of aborted fetuses to judge whether there is maternal contamination and copy number abnormalities in the sample, without additional experiments and maternal samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of genetic testing, and particularly to a method and device for identifying maternal contamination and detecting copy number abnormalities by using peaks. Background Art

[0002] Chromosomal diseases are a large class of serious genetic diseases caused by chromosomal number or structural abnormalities, including clinical manifestations such as multiple malformations, mental retardation, and growth and development retardation. They are also important causes of developmental abnormalities and male and female infertility. Embryonic chromosomal abnormalities are one of the important causes of early spontaneous abortion. Clinically, about 10-15% of pregnant women experience spontaneous abortion, and 50%-60% of them are related to embryonic chromosomal abnormalities. Among the spontaneous abortions caused by chromosomal abnormalities, about 96% are numerical abnormalities (i.e., copy number abnormalities), and only 3%-5% are structural abnormalities.

[0003] Genomic copy number variation (CNV) generally refers to the decrease or increase in the copy number of large genomic fragments longer than 1 kb, mainly manifested as submicroscopic-level duplications and deletions, and is an important part of genomic structural variation. Detecting copy number variation (CNV detection) in abortuses can help doctors identify the causes of abortion, provide a scientific basis for subsequent treatment or intervention measures, provide guidance for the next pregnancy, and avoid unnecessary examinations and incorrect treatments.

[0004] Due to the limitations of sampling techniques, currently collected abortus tissues are often mixed with maternal peripheral blood, uterine fundus decidua tissues, etc., resulting in different degrees of maternal cell contamination (MCC) in the test samples, which ultimately affects the accuracy of CNV detection results. Therefore, before performing CNV detection, it is usually necessary to first identify the maternal contamination in the abortus sample through short tandem repeats (STR) detection. However, this method requires an additional blood sample from the mother and additional experimental steps, which not only increases the detection cost but also prolongs the detection cycle. Summary of the Invention

[0005] Based on this, the object of the present invention is to provide a method for identifying maternal contamination and detecting copy number abnormalities by using peaks, which can use the SNP locus information of abortuses to determine whether there is maternal contamination and copy number abnormalities in the sample, without additional experiments and additional maternal peripheral blood samples.

[0006] The present invention is achieved by the following technical solutions:

[0007] A method for identifying maternal contamination and detecting copy number abnormalities using peaks, comprising the following steps:

[0008] Step S10: Detect the SNP locus information of the sample to be tested, and extract the LRR feature data, BAF feature data, and genotype feature data of each autosome of the sample to be tested based on the genomic position order of the SNP loci according to the SNP locus information;

[0009] Step S20: Obtain the number of peaks and peak values on each autosome according to the BAF feature data, and judge the chromosomal state of each autosome according to the number of peaks, peak values, and LRR feature data;

[0010] Step S30: Judge whether the sample state of the sample to be tested is maternal contamination or no maternal contamination according to the chromosomal state with the highest proportion among all chromosomal states; if the sample state is no maternal contamination, further judge whether the sample state is haploid, triploid, or other, and perform step S40;

[0011] Step S40: Obtain the CNV result of the sample to be tested according to the LRR feature data and BAF feature data; judge whether the sample to be tested is whole-genome ROH and / or obtain the ROH result of the sample to be tested according to the LRR feature data and genotype feature data.

[0012] Compared with the prior art, by comprehensively evaluating the peak conditions of BAF data on autosomes, the present invention can not only quickly judge the maternal contamination of the sample, but also quickly judge whether the sample belongs to special types of samples such as haploid, triploid, and whole-genome ROH, which can improve the detection efficiency, enhance the detection comprehensiveness, and reduce the detection time and cost.

[0013] Further, in the step S20, a mixture Gaussian model and / or a sliding window is used to obtain the number of peaks and peak values on each autosome; wherein, using the mixture Gaussian model to obtain the number of peaks and peak values on each autosome includes:

[0014] Use the mixture Gaussian model to fit the distribution of the BAF feature data on each autosome, calculate the Bayesian information criterion BIC of the mixture Gaussian model corresponding to 1 Gaussian distribution to 10 Gaussian distributions, and the number and mean of the Gaussian distributions corresponding to the mixture Gaussian model with the lowest BIC are the number of peaks and peak values on each autosome.

[0015] Further, step S20 further includes obtaining the number and peak values of the common peaks of the Gaussian mixture model and the sliding window on each autosome, including: obtaining the intersection region between the peaks of the Gaussian mixture model and the peaks of the sliding window. If the proportion of the intersection region is greater than the first set threshold, the union region of the peaks of the Gaussian mixture model and the peaks of the sliding window that form the intersection region is the common peak region; using the sliding window to detect the peak values of the common peak region to obtain the peak values and the number of common peaks.

[0016] Further, in step S20, the determination of the chromosomal status of each autosome according to the number and peak values of the peaks includes:

[0017] When the number of peaks is 0 and the BAF feature data is mainly distributed near 0 and 1, if the mean value of the LRR feature data is less than -0.2, the chromosomal status of this autosome is chromosomal overall deletion; otherwise, the chromosomal status is unknown.

[0018] When the number of peaks is 1, if the peak value is within the interval [0.45, 0.55] and its peak proportion is greater than the second set threshold, the chromosomal status of this autosome is chromosomal normal; otherwise, the chromosomal status is maternal contamination.

[0019] When the number of peaks is 2, if the two peak values are respectively within the intervals [0.3, 0.4] and [0.6, 0.7], the chromosomal status of this autosome is chromosomal duplication; otherwise, the chromosomal status is unknown.

[0020] If the number of peaks is greater than or equal to 3, the chromosomal status of this autosome is maternal contamination.

[0021] Among them, the calculation formula for the peak proportion is:

[0022]

[0023] Further, step S40 of obtaining the CNV result of the test sample includes the following steps: according to the LRR feature data, BAF feature data, and HMM model, obtain the segment CNV results on each autosome in the test samples with the sample status of others, merge the adjacent segment CNV results, and filter the segment CNV results according to the length of the CNV and the SNP data included in the CNV to obtain the CNV result.

[0024] Further, step S40 of determining whether the test sample is a genome-wide ROH includes: calculating the total heterozygosity ratio of all autosomes of the test sample with the sample status of others. If the heterozygosity ratio is less than the third set threshold, the sample status is updated to genome-wide ROH; otherwise, the sample status is others. Among them, the calculation formula for the heterozygosity ratio is as follows:

[0025]

[0026] Further, the ROH result in step S40 is the ROH result of a tested sample with a sample state of haploid, triploid or other. The method for obtaining the ROH includes the following steps: According to the positions of SNP sites on the genome, n adjacent SNPs are divided into bins. If the homozygous proportion of SNPs in the bin > 95% and the mean value of LRR characteristic data is near 0, then the bin is a potential ROH region; adjacent potential ROH regions are merged to form a new potential ROH region. If the homozygosity rate of SNPs in the new potential ROH region > 99%, then the new potential ROH region is an ROH region. All ROH regions are summarized to obtain the ROH result. Wherein, the calculation formula of the homozygosity rate is as follows:

[0027]

[0028] The present invention also provides a device for identifying maternal contamination and detecting copy number abnormalities using peaks, including:

[0029] A data extraction unit extracts LRR characteristic data, BAF characteristic data, and genotype characteristic data of each autosome of the tested sample based on the genomic position order of SNP sites according to the SNP site information of the tested sample in the database;

[0030] A chromosome state acquisition unit obtains the number of peaks and peak values on each autosome according to the BAF characteristic data, and judges the chromosome state of each autosome according to the number of peaks, peak values, and LRR characteristic data;

[0031] A first sample state determination unit preliminarily judges whether the sample state of the tested sample is maternal contamination present or no maternal contamination according to the chromosome state with the highest proportion among all chromosome states. If the sample state is no maternal contamination, then further judges the sample state of the tested sample as haploid, triploid or other according to the chromosome state with the highest proportion among all chromosome states;

[0032] A second sample state determination unit, if the sample state is not maternal contamination present, obtains the sample CNV result of the tested sample according to the LRR characteristic data and BAF characteristic data; and judges whether the tested sample is a whole-genome ROH and / or obtains the sample ROH result of the tested sample according to the LRR characteristic data and genotype characteristic data.

[0033] The present invention also provides a computer device, which includes a processor and a memory:

[0034] The memory is used to store a computer program and transmit the computer program to the processor;

[0035] The processor is configured to execute the method for identifying maternal contamination and detecting copy number abnormalities using peaks according to the instructions in the computer program.

[0036] The present invention also provides a computer-readable storage medium for storing a computer program, which, when executed by a processor, causes the processor to execute the method for identifying maternal contamination and detecting copy number abnormalities using peaks.

[0037] For better understanding and implementation, the present invention will be described in detail below with reference to the accompanying drawings. Description of the Drawings

[0038] Figure 1 It is a schematic diagram of a method for identifying maternal contamination and detecting copy number abnormalities using peaks according to the present invention;

[0039] Figure 2 It is a flowchart of a method for identifying maternal contamination and detecting copy number abnormalities using peaks provided by the present invention;

[0040] Figure 3 It is a schematic structural diagram of a device for identifying maternal contamination and detecting copy number abnormalities using peaks provided by the present invention;

[0041] Figure 4 It is a graph of LRR results and BAF results of a maternally contaminated sample in Example 1;

[0042] Figure 5 It is a peak graph on autosomes of a maternally contaminated sample in Example 1;

[0043] Figure 6 It is a graph of LRR results and BAF results of a slightly maternally contaminated sample in Example 2;

[0044] Figure 7 It is a peak graph on autosomes of a slightly maternally contaminated sample in Example 2;

[0045] Figure 8 It is a peak graph of a triploid sample in Example 3;

[0046] Figure 9 It is a peak graph of a whole-genome ROH sample in Example 4;

[0047] Figure 10 It is a CNV result graph of chromosome 6 duplication of a whole-genome ROH sample in Example 4;

[0048] Figure 11 It is a CNV result graph of chromosome 9 duplication of a whole-genome ROH sample in Example 4;

[0049] Figure 12 The peak graph of the sample with fragment deletion in Example 5;

[0050] Figure 13 The CNV result graph of the sample with fragment deletion in Example 5. Specific implementation manners

[0051] Aiming at the problem that before performing CNV detection in the prior art, it is necessary to first collect maternal peripheral blood samples for maternal contamination detection, and only the samples confirmed to have no maternal contamination can be used for subsequent CNV detection-related experiments, the present invention aims to provide a method that can use the data obtained from the detection of convection product samples to judge the maternal contamination of the samples and can comprehensively detect the samples without additional maternal samples and additional experiments.

[0052] Refer to Figure 1 , the present invention provides a method for identifying maternal contamination and detecting copy number abnormalities using peaks, specifically including the following steps:

[0053] Step S10: Detect the SNP locus information of the sample to be tested, and extract the LRR feature data, BAF feature data, and genotype feature data of each autosome of the sample to be tested based on the genomic position order of the SNP loci;

[0054] Step S20: According to the BAF feature data, obtain the number of peaks and peak values on each autosome, and judge the chromosomal state of each autosome based on the number of peaks, peak values, and LRR feature data;

[0055] Step S30: Judge whether the sample state of the sample to be tested is with maternal contamination or without maternal contamination according to the chromosomal state with the highest proportion among all chromosomal states; if the sample state is without maternal contamination, further judge whether the sample state is haploid, triploid, or others, and perform step S40 processing;

[0056] Step S40: According to the LRR feature data and BAF feature data, obtain the sample CNV result of the sample to be tested; according to the LRR feature data and genotype feature data, judge whether the sample to be tested is whole-genome ROH and / or obtain the sample ROH result of the sample to be tested.

[0057] To more clearly understand the technical solutions provided by the embodiments of the present application, some key terms related to the embodiments of the present application are introduced here first:

[0058] Single nucleotide polymorphism (SNP) refers to the polymorphism caused by the change of a single nucleotide at a certain locus in the chromosomal DNA sequence. The frequency of SNP in the population is generally >1%. There is one SNP per 300 - 1000 bp on average in the human genome. Currently, SNP databases can be obtained from multiple public databases, including, for example, http: / / cgap.ncbi.nih.gov / GAI; http: / / www.ncbi.nlm.nih.gov / SNP; the human SNP database http: / / hgbas.cgr.ki.sei or http: / / hgbase.interactiva.de / .

[0059] Chromosomal microarray analysis (CMA) technology is a molecular karyotyping analysis technology for high-throughput detection of genomic DNA copy number variations, including two technologies: microarray comparative genomic hybridization chip (Array-CGH) and single nucleotide polymorphism array (SNP array). Among them, the single nucleotide polymorphism array technology can not only diagnose genetic syndromes and chromosomal diseases caused by genomic deletions or duplications based on CNV, but also diagnose homozygous regions, unbalanced translocations, etc. through SNP analysis.

[0060] Region of homozygosity (ROH) is a phenomenon of continuous loss of heterozygosity within a certain range in the genomic region. For most diploid cells such as human somatic cells, there are two copies of the genome, one from the father and the other from the mother. At a certain allele locus, if the bases from the paternal and maternal sides are different, then this locus is heterozygous. If due to some mechanism (such as consanguinity or inbreeding marriage or gene conversion), the consecutive allele sequences within a certain range are all homozygous without heterozygotes (the copy number is still 2), then this region is the genomic homozygous region ROH. The generation of ROH involves reasons such as identity by descent (IBD) or uniparental disomy (UPD).

[0061] Copy number variation (CNV) refers to the change in the copy number of at least a part of the nucleic acid molecule in the test sample compared to the corresponding nucleic acid sequence in the normal sample, where the part has a length greater than 1 kb. The situations and reasons for copy number variation can include: deletion, such as microdeletion; insertion, such as microinsertion, microduplication, duplication; inversion, transposition, and complex multi-locus variation.

[0062] The signal log ratio, that is, the log2R ratio (LRR, that is, the log2-transformed value of the normalized SNP intensity), reflects the copy number change of the sample to be tested relative to the normal sample at the measured SNP locus. BAF, that is, the B allele frequency (B Allele Frequency, that is, the ratio of the B allele signal intensity to the total SNP signal intensity), reflects the occurrence of different alleles of the sample to be tested at the measured SNP locus. In normal diploid samples, the BAF values are distributed around 0, 0.5, and 1 (that is, for each SNP, there are three possible forms AA, AB, and BB in the genome), and the LRR values are distributed around 0. Abnormalities in the tested sample (including sample contamination and copy number abnormalities) will cause an increase or decrease in the total SNP signal intensity and / or the B allele frequency, resulting in BAF and / or LRR values deviating from the above values. For example, when a chromosomal region undergoes copy number loss, the BAF in this region is only 0 and 1 (genotypes AA and BB), without values around 0.5 (genotype AB), and the LRR value in this region will decrease to a value less than 0; when a chromosomal region undergoes copy number gain, the BAF values in this region will be around 0, 0.33, 0.67, and 1.0 (genotypes AAA, AAB, ABB, and BBB), and the LRR value in this region will increase to a value greater than 0. Based on this change in BAF and LRR values, abnormal copy number variations in the sample can be estimated in principle.

[0063] The mixture Gaussian model is a probability model that assumes that the observed data is composed of a mixture of multiple Gaussian distributions, and each Gaussian distribution represents a potential class or group.

[0064] The Bayesian information criterion (BIC) is an evaluation index for the reliability of the result of correcting the occurrence probability using the Bayesian formula by subjectively estimating the state of partial unknowns under incomplete information.

[0065] The sliding window (Sliding Window) is a commonly used technique in data processing and analysis for performing local operations on sequence data or continuous data. It divides the input data into windows of a fixed size and calculates the corresponding results each time the window is moved. The sliding step (Sliding Step) is the step distance defined for each window movement. It determines the degree of overlap between windows and their relative positional relationship. For example, if the size of the sliding window is 10 and the sliding step is 5, the window will slide 5 units to the right each time it moves.

[0066] The Hidden Markov Model (HMM) is a statistical model used to describe a Markov process of a system with hidden states. In this model, the states of the system are not directly visible (they are "hidden"), but the changes in states can be inferred from a series of observed events or signals.

[0067] In the above method, the detection of SNP locus information described in step S10 includes using a chromosomal microarray chip to detect the nucleic acids of the tested sample, and obtaining the information of the SNP loci corresponding to the chip by analyzing the fluorescence signals after hybridization. Specifically, the chromosomal microarray chip includes the Infinium Asian Screening Array (ASA) chip of Illumina and the like.

[0068] Other steps of the above method can be carried out by a device for identifying maternal contamination and detecting copy number abnormalities. Hereinafter, in combination with the device for identifying maternal contamination and detecting copy number abnormalities, the method for identifying peak maternal contamination and detecting copy number abnormalities provided by the embodiments of the present invention will be described in detail.

[0069] Refer to Figure 2 and Figure 3 , Figure 2 is a schematic flowchart of the method for identifying maternal contamination and detecting copy number abnormalities in this embodiment, Figure 3 is a schematic structural diagram of the device for identifying maternal contamination and detecting copy number abnormalities in this embodiment. The device for identifying maternal contamination and detecting copy number abnormalities described in the embodiments of the present invention is provided with a database, and the SNP locus information of the tested sample is stored in the database. The device further includes a data extraction unit 10, a chromosomal state acquisition unit 20, a sample state determination unit 30, a second sample state determination unit 40, and a result output unit 50.

[0070] Among them, the data extraction unit 10 of the device for identifying maternal contamination and detecting copy number abnormalities is used to execute step S10: according to the SNP locus information of the tested sample, extract and obtain the LRR feature data, BAF feature data, and genotype feature data of each autosome of the tested sample based on the genomic position order of the SNP loci. Specifically, the data preprocessing unit 10 includes a BAF feature data extraction module 11a, an LRR feature data extraction module 11b, and a genotype feature data extraction module 11c.

[0071] The BAF feature data extraction module 11a is used to execute step S11a: obtain the BAF feature data of each autosome based on the genomic position order of the SNP loci.

[0072] Specifically, the BAF data corresponding to the SNP sites are extracted from the raw chip data through the annotation file corresponding to the chip, and the following sites in the BAF data are filtered out: 1) sites where the chromosome is not an autosome or a sex chromosome; 2) sites with incorrect and duplicate coordinate information; 3) sites where the genotype is not AA, AB, or BB; then the filtered BAF data are sorted according to the chromosome and coordinates to obtain the BAF feature data based on the genomic position order of the SNP sites for each autosome. In this embodiment, the BeadArrayFiles software is used to extract the data in the Illumina chip.

[0073] The LRR feature data extraction module 11b is used to execute step S11b: obtaining the LRR feature data based on the genomic position order of the SNP sites for each autosome.

[0074] Specifically, the LRR data corresponding to the SNP sites are extracted from the raw chip data through the annotation file corresponding to the chip, and the following sites in the LRR data are filtered out: 1) sites where the chromosome is not an autosome or a sex chromosome; 2) sites with incorrect and duplicate coordinate information; 3) sites where the genotype is not AA, AB, or BB; then the filtered LRR data are sorted according to the chromosome and coordinates to obtain the LRR feature data based on the genomic position order of the SNP sites for each autosome.

[0075] The genotype feature data extraction module 11c is used to execute step S11c: extracting the genotype feature data corresponding to the SNP sites from the raw chip data through the annotation file corresponding to the chip.

[0076] Specifically, the genotype data corresponding to the SNP sites are extracted from the raw chip data through the annotation file corresponding to the chip, and the following sites in the genotype data are filtered out: 1) sites where the chromosome is not an autosome or a sex chromosome; 2) sites with incorrect and duplicate coordinate information; 3) sites where the genotype is not AA, AB, or BB; then the filtered genotype data are sorted according to the chromosome and coordinates to obtain the genotype feature data based on the genomic position order of the SNP sites for each autosome.

[0077] Furthermore, the chromosome status acquisition unit 20 is used to execute step S20: obtaining the number of peaks and peak values on each autosome according to the BAF feature data, and judging the chromosome status of each autosome according to the number of peaks, peak values, and LRR feature data. Specifically, the chromosome status acquisition 20 includes a peak acquisition module 21 and a chromosome status judgment module 22.

[0078] The peak acquisition module 21 is used to execute step S21: obtaining the number of peaks and peak values on each autosome by using a Gaussian mixture model and / or a sliding window. In this embodiment, the peak acquisition module 21 includes a Gaussian mixture model module 21a, a sliding window module 21b, and a common peak acquisition module 21c.

[0079] The Gaussian mixture model module 21a is used to execute step S21a: fitting the distribution of BAF feature data on each autosome by using a Gaussian mixture model, and calculating the Bayesian information criterion BIC of the Gaussian mixture model corresponding to 1 Gaussian distribution to 10 Gaussian distributions. Among them, the number of Gaussian distributions and the mean value of the Gaussian mixture model corresponding to the lowest BIC are the number of peaks and peak values of peak F1.

[0080] Among them, the Gaussian mixture model is:

[0081] Among them, x is the observed data point, and π k is the mixing weight of k Gaussian distributions, and N(x|μ k , ∑ k ) represents the probability density function of the k-th Gaussian distribution with a mean of μ k and a covariance matrix of ∑ k .

[0082] The calculation formula of the Bayesian information criterion BIC is:

[0083] BIC = kln(n) - 2ln(L); where k is the number of Gaussian distributions, n is the number of BAF feature data, and L is the likelihood function.

[0084] In this embodiment, it is assumed that there is only 1 Gaussian distribution to at most 10 Gaussian distributions in the current chromosome BAF feature data. Each time an assumption is made, the BAF feature data is used for model training and the BIC value is calculated simultaneously. 10 assumptions generate 10 BIC values, and the model with the lowest BIC value is taken as the optimal model; the number of Gaussian distributions corresponding to the optimal model is the number of peaks of peak F1, and the mean value of each distribution of the optimal model is the peak value of each peak F1.

[0085] The sliding window module 21b is used to execute step S21b: fitting the distribution of BAF feature data on each autosome by using a sliding window, obtaining BAF distribution data, and using a peak detection method to obtain the peak value and distribution of peak F2 in the BAF distribution data, specifically including the following steps:

[0086] Process the BAF feature data by selecting appropriate window sizes and step sizes. Move the window one step in the direction from 0 to 1, and repeat this process until all regions are divided into windows. Calculate the BAF distribution data by computing the number of SNPs in each window. In this embodiment, the step size is less than or equal to the window size.

[0087] Then use the peak detection method to identify significant peaks from the BAF distribution data, that is, the local maxima in the BAF distribution data, and remove the peaks that may be caused by random noise with the number of SNPs below the threshold, to obtain the peak value and distribution of peak F2.

[0088] The common peak acquisition module 21c is used to execute step S21c: obtain the intersection region between all peaks in peak F1 identified by the mixture Gaussian model and peak F2 identified by the sliding window. If the proportion of the intersection region between peak F1 and peak F2 is greater than the first set threshold, then use the union region of peak F1 and peak F2 that form this intersection region as the common peak region; use the sliding window to detect the peak value of the common peak region to obtain the peak value and the number of peaks of the final common peak F.

[0089] Among them, the calculation formula for the proportion of the intersection region is as follows:

[0090] In the formula, F1 refers to the peak range of peak F1 identified by the mixture Gaussian model, and F2 refers to the peak range of peak F2 identified by the sliding window.

[0091] In this embodiment, the first set threshold is set to 60%.

[0092] The chromosome state type judgment module 22 is used to execute step S22: determine the chromosome state type of each autosome according to the number of peaks, peak values of the common peak F obtained in step S21c, and the LRR feature data obtained in step S11b.

[0093] (1) If the number of peaks of the common peak F is 0, and the BAF feature data is mainly distributed near 0 and 1. Further, if the mean value of the LRR feature data is less than -0.2, then the chromosome state of this autosome is chromosome whole deletion, otherwise, the chromosome state of this autosome is unknown.

[0094] (2) When the number of peaks of the common peak F is 1, if the peak value of the common peak F is near 0.5, that is, in the interval [0.45, 0.55], and the peak value proportion of the common peak F is greater than the second set threshold, then the chromosome state of this autosome is chromosome normal, otherwise, the chromosome state of this autosome is chromosome slightly contaminated, that is, there is maternal contamination in the chromosome and the contamination ratio > 5%; among them, the calculation formula for the peak value proportion is:

[0095] In this embodiment, the second set threshold is 0.2;

[0096] (3) When the number of peaks of the common peak F is 2, if the peak values of the double peaks satisfy that one peak is near 0.33, that is, within the interval [0.3, 0.4], and the other peak is near 0.67, that is, within the interval [0.6, 0.7], then the chromosomal state of this autosome is chromosomal duplication. Otherwise, the chromosomal state of this autosome is unknown;

[0097] (4) If the number of peaks of the common peak F is greater than or equal to 3, then the chromosomal state of this autosome is the existence of maternal contamination, and the contamination ratio > 10%.

[0098] Further, the first sample state determination unit 30 is used to execute step S30: Determine whether the sample state of the tested sample is the existence of maternal contamination or no maternal contamination is detected according to the chromosomal state with the highest proportion among all chromosomal states. If the sample state is no maternal contamination is detected, further determine whether the sample state is haploid, triploid or other.

[0099] Specifically, count the chromosomal state of each chromosome, and extract the chromosomal state type with the largest number and record it as the autosomal state; preliminarily determine the sample state based on the autosomal state. If the autosomal state is the existence of chromosomal maternal contamination, then the sample state is the existence of maternal contamination. Otherwise, the sample state is no maternal contamination is detected; if the sample state is no maternal contamination is detected, further judge: 1) If the autosomal state is chromosomal deletion, then the sample state is updated to haploid; 2) If the autosomal state is chromosomal duplication, then the sample state is updated to triploid; 3) If the autosomal state is other types such as normal chromosome or unknown, then the sample state is updated to other.

[0100] Further, the second sample state determination unit 40 is used to execute step S40: Obtain the CNV result and ROH result on each autosome according to the LRR feature data, BAF feature data and genotype feature data, and judge whether the tested sample is a genome-wide ROH according to the heterozygosity ratio of the tested sample. The second sample state determination unit 40 includes a sample CNV analysis module 40A, a genome-wide ROH analysis module 40B and a sample ROH analysis module 40C.

[0101] The sample CNV analysis module 40A is used to perform step S40A: obtaining the CNV results of the sample with the sample status of other according to the LRR feature data and the BAF feature data, and the method for obtaining the sample CNV results includes: using the PennCNV software, the BAF feature data of the sample, the LRR feature data of the sample and the pre-trained HMM model, obtaining the fragment CNV results on each chromosome in the sample, merging the adjacent small fragment CNVs, and filtering the fragment CNV results according to the length of the CNV and the SNP data contained in the CNV. The pre-trained HMM model described in this embodiment is an existing publicly available HMM model trained based on more than 100 research and development negative samples as a baseline.

[0102] The whole genome ROH analysis module 40B is used to perform step S40B: judging whether the sample status of the sample under test whose sample status is other is whole genome ROH according to the genotype feature data, specifically including: calculating the heterozygosity ratio of all autosomes of the sample under test whose sample status is other, if the heterozygosity ratio is less than the third set threshold, the sample status is updated to whole genome ROH; otherwise, the sample status is other; wherein the calculation formula of the heterozygosity ratio is as follows:

[0103] In this embodiment, the third set threshold is set to 0.1.

[0104] The sample ROH analysis module 40C is used to perform step S40C: according to the LRR characteristic data and the genotype characteristic data, the ROH result of the sample under test is obtained, and the sample state is haploid, triploid or other. The method for obtaining the sample ROH result includes the following steps: according to the position of the SNP on the genome, n adjacent SNPs are divided into bins. If the homozygosity rate of the SNP in the bin is greater than 95%, and the mean value of the LRR characteristic data is near 0, then the bin is a potential ROH region; merge adjacent potential ROH regions to form a new potential ROH region. If the homozygosity rate of the SNP in the new potential ROH region is greater than 99%, then the new potential ROH region is a ROH region. All ROH regions are aggregated, and regions with a small number of SNPs and a short length are filtered to obtain the final sample ROH result.

[0105] The calculation formula of the homozygosity rate of the SNP is as follows:

[0106] Wherein, k is the number of homozygous SNP sites in the bin, and n is the number of all SNP sites in the bin.

[0107] Further, the result output unit 50 is configured to execute step S50: output the final sample status of the sample to be tested, as well as the CNV result and / or ROH result of the sample to be tested:

[0108] 1) If the sample status obtained in step S30 is maternal contamination, then it is prompted that the sample has maternal contamination;

[0109] 2) If the sample status obtained in step S40B is genome-wide ROH, then it is prompted that the sample status is genome-wide ROH. Meanwhile, if the sample has CNV, the CNV result is output;

[0110] 3) If the sample status obtained in step S30 is haploid or triploid, then it is prompted that the sample status is haploid or triploid. Meanwhile, if the sample has ROH regions, the ROH result is output;

[0111] 4) If the sample status obtained in step S40B is other, then it is prompted that the sample status is other. Meanwhile, if the sample has CNV and / or ROH regions, the CNV result and / or ROH result is output.

[0112] This embodiment also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps in the method for identifying maternal contamination and detecting copy number abnormalities using peaks are implemented.

[0113] This embodiment provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps in the method for identifying maternal contamination and detecting copy number abnormalities using peaks as described above are implemented.

[0114] For better understanding and implementation, the present invention will be described in detail below with reference to specific embodiments and drawings.

[0115] Example 1

[0116] The gDNA of sample RD22112301 was detected using the Infinium Asian Screening Array chip of Illumina. The information of the SNP sites corresponding to the chip was obtained by analyzing the fluorescence signals after hybridization. Then, the method for identifying maternal contamination and detecting copy number abnormalities using peaks provided by the present invention was used to obtain the sample status, CNV result, and ROH result of the sample. The sample used in this embodiment is a research and development test sample. The specific steps of the method are as follows:

[0117] Step S11a: Use the BeadArrayFiles software of Illumina to extract the BAF data corresponding to all SNP loci from the chromosomal microarray data; filter out the SNP data where the chromosome is not autosome or sex chromosome, the coordinate information is incorrect or repeated, and the genotype information is not AA, AB, or BB; finally, sort according to the position of the SNP locus on the genome to obtain the BAF feature data of each autosome based on the genomic position order of the SNP locus.

[0118] Step S11b: Use the BeadArrayFiles software of Illumina to extract the LRR data corresponding to all SNP loci from the chromosomal microarray data; filter out the SNP data where the chromosome is not autosome or sex chromosome, the coordinate information is incorrect or repeated, and the genotype information is not AA, AB, or BB; finally, sort according to the position of the SNP locus on the genome to obtain the LRR feature data of each autosome based on the genomic position order of the SNP locus.

[0119] Step S11c: Use the BeadArrayFiles software of Illumina to extract the genotype data corresponding to all SNP loci from the chromosomal microarray data, filter out the SNP data where the chromosome is not autosome or sex chromosome, the coordinate information is incorrect or repeated, and the genotype information is not AA, AB, or BB, and finally sort according to the position of the SNP locus on the genome to obtain the genotype feature data.

[0120] Step S21a: Use the mixture Gaussian model to perform model training on the BAF feature data one by one from 1 distribution to 10 distributions and calculate the BIC values corresponding to K distributions at the same time. Take the model with the lowest BIC value as the optimal model; the number of distributions of the optimal model is the number of peaks of peak F1, and the mean of each distribution is the peak value of each peak F1.

[0121] Step S21b: Use a window of size 0.025 to slide at a step of 0.005, count the number of SNPs in each window one by one, and find all the peaks in the BAF feature data through the peak detection method. Remove the peaks with the number of SNPs less than 100 to obtain the number of peaks and peak values of peak F2.

[0122] Step S21c: Take the union region of peak F1 and peak F2 where the intersection region ratio of the peaks identified by the mixture Gaussian model (peak F1) and the sliding window method (peak F2) is greater than 60% as the common peak region. Use a 0.025 sliding window to perform peak detection at a step of 0.005 in the common peak region to determine the peak value and number of peaks of the final common peak F.

[0123] Referring to Table 1, in this embodiment, the peak detection results of sample RD22112301 show that the number of peaks of the common peak F on chromosome 1 is 5, and the peak values are 0.16, 0.36, 0.48, 0.62, and 0.88 respectively.

[0124] Table 1: Peak Detection Results of Chromosome 1 of Sample RD22112301

[0125]

[0126] Step S22: Determine the chromosomal status type of each chromosome based on the number of peaks and peak values of the common peak F obtained in step S21c and the LRR feature data obtained in step S11b.

[0127] Referring to Table 1 and Figures 4 - 5 , Figure 4 showing that the BAF map with severe maternal contamination has significant data separation, Figure 5 showing that multiple peaks appear on each chromosome, that is, in the sample RD22112301 of this embodiment, the number of peaks on chromosomes 1 to 22 is greater than 3, that is, the chromosomal status of each chromosome is maternal contamination, and the contamination ratio > 10%.

[0128] Step S30: Summarize the chromosomal status types of all autosomes obtained in step S22, and judge whether the sample status of the sample is maternal contamination or no maternal contamination. If the sample status is no maternal contamination, further judge whether the sample status is haploid, triploid or others.

[0129] In this embodiment, by summarizing the chromosomal status of 22 autosomes of sample RD22112301, it is found that the chromosomal status of the autosomes is maternal chromosomal contamination, that is, the sample status is maternal contamination, and it is severe maternal contamination with a contamination ratio greater than 10%.

[0130] Step S50: Output the sample status and / or CNV result and / or ROH result: Prompt that the sample has maternal contamination.

[0131] Example 2

[0132] The gDNA of sample RD23120601 was detected using the Infinium Asian Screening Array chip of Illumina. The information of the SNP sites corresponding to the chip was obtained by analyzing the fluorescence signals after hybridization, and then the following steps were used to obtain the sample status, CNV result, and ROH result of the sample:

[0133] Step S11a: Obtain the BAF feature data of each autosome based on the genomic position order of the SNP sites.

[0134] Step S11b: Obtain the LRR feature data of each autosome based on the genomic position order of SNP sites.

[0135] Step S11c: Obtain the genotype feature data of each autosome based on the genomic position order of SNP sites.

[0136] Step S21a: Use the Gaussian mixture model to obtain the peak value and the number of peaks of peak F1 on each autosome.

[0137] Step S21b: Use a sliding window to obtain the peak value and the number of peaks of peak F2 on each autosome.

[0138] Step S21c: Obtain the peak value and the number of peaks of the common peak F of peak F1 and peak F2 on each autosome.

[0139] Step S22: Judge the chromosomal state of each autosome according to the number of peaks, peak values of the common peak F, and the LRR feature data.

[0140] Refer to Figures 6 - 7 In this embodiment, in sample RD23120601, the number of peaks of chromosomes 1 to 22 is 1, and the peak value ratio is 0.158, which is less than the threshold 0.2, that is, the chromosomal state of each chromosome is chromosomal slight contamination.

[0141] Step S30: Summarize the chromosomal state types of all autosomes obtained in Step S22, judge whether the sample has maternal contamination or no maternal contamination, and if the sample state is no maternal contamination, further judge whether the sample state is haploid, triploid or others.

[0142] In this embodiment, by summarizing the chromosomal states of 22 autosomes of sample RD22112301, it is found that the chromosomal state is the existence of slight maternal contamination, that is, the sample state is maternal contamination.

[0143] Step S50: Output the sample state and / or CNV result and / or ROH result: Prompt that the sample has maternal contamination.

[0144] Embodiment 3

[0145] Use the Infinium Asian Screening Array chip of Illumina to detect the gDNA of sample RD24041801, obtain the information of the SNP sites corresponding to the chip by analyzing the fluorescence signals after hybridization, and then adopt the following steps to obtain the sample state, CNV result, and ROH result of the sample:

[0146] Step S11a: Obtain the BAF feature data of each autosome based on the genomic position order of SNP sites.

[0147] Step S11b: Obtain the LRR feature data of each autosome based on the genomic position order of SNP sites.

[0148] Step S11c: Obtain the genotype feature data of each autosome based on the genomic position order of SNP sites.

[0149] Step S21a: Use the mixture Gaussian model to obtain the peak value and the number of peaks of peak F1 on each autosome.

[0150] Step S21b: Use the sliding window to obtain the peak value and the number of peaks of peak F2 on each autosome.

[0151] Step S21c: Obtain the peak value and the number of peaks of the common peak F of peak F1 and peak F2 on each autosome.

[0152] Step S22: Judge the chromosomal state of each autosome according to the number of peaks, peak values of the common peak F, and the LRR feature data.

[0153] Refer to Figure 8 , in this embodiment, in sample RD24041801, chromosomes 1 to 22 are all bimodal, and the peak value of one peak is in the interval [0.3, 0.4], and the peak value of the other peak is in the interval [0.6, 0.7], that is, the chromosomal state is chromosomal duplication.

[0154] Step S30: Summarize the chromosomal state types of all autosomes obtained in Step S22, and judge whether the sample has maternal contamination or no maternal contamination. If the sample state is no maternal contamination, further judge whether the sample state is haploid or triploid or others.

[0155] In this embodiment, by summarizing the chromosomal states of 22 autosomes, it is found that there is no chromosomal maternal contamination in the chromosomal state, that is, the sample state is no maternal contamination; further judgment shows that the chromosomal states of all autosomes are chromosomal duplications, that is, the sample state is triploid.

[0156] S40C: Obtain the ROH results of the tested samples with the sample status of haploid, triploid or others: According to the positions of SNPs on the genome, 1000 adjacent SNPs are divided into one bin. When the homozygosity rate of SNPs in the bin > 95%, and the LRR in this region is near 0, then this bin is considered a potential ROH region; by merging adjacent potential ROH regions to form a new potential ROH region, if the homozygosity rate of SNPs in the new potential ROH region > 99%, then this region is considered an ROH region, and regions with the number of SNPs less than 500 and length less than 1M are filtered out to obtain the final ROH results of the sample.

[0157] In this embodiment, through analysis, it is found that there is no ROH region in this sample.

[0158] S50: Output the sample status and / or CNV results and / or ROH results: Prompt that the sample is triploid.

[0159] Example 4

[0160] Use the Infinium Asian Screening Array chip of Illumina to detect the gDNA of sample RD23103101, obtain the information of the SNP sites corresponding to the chip by analyzing the fluorescence signals after hybridization, and then adopt the following steps to obtain the sample status, CNV results and ROH results of the sample:

[0161] Step S11a: Obtain the BAF feature data of each autosome based on the genomic position order of SNP sites.

[0162] Step S11b: Obtain the LRR feature data of each autosome based on the genomic position order of SNP sites.

[0163] Step S11c: Obtain the genotype feature data of each autosome based on the genomic position order of SNP sites.

[0164] Step S21a: Use the mixture Gaussian model to obtain the peak value and the number of peaks of peak F1 on each autosome.

[0165] Step S21b: Use the sliding window to obtain the peak value and the number of peaks of peak F2 on each autosome.

[0166] Step S21c: Obtain the peak value and the number of peaks of the common peak F of peak F1 and peak F2 on each autosome.

[0167] Step S22: Judge the chromosomal status of each autosome according to the number of peaks, peak values of the common peak F and the LRR feature data.

[0168] Refer to Figure 9, in this embodiment, in sample RD23103101, chromosomes 6 and 9 show chromosomal duplications, chromosomes 11, 16, and 21 show unknown, and the remaining chromosomes show normal chromosomes.

[0169] Step S30: Summarize the chromosomal status types of all autosomes obtained in step S22, and determine whether the sample status of the test sample is maternal contamination or no maternal contamination based on the chromosomal status with the highest proportion. If the sample status is no maternal contamination, further determine whether the sample status is haploid, triploid, or other.

[0170] In this embodiment, by summarizing the chromosomal status of 22 autosomes, it is found that there is no maternal chromosomal contamination in the autosomal status, that is, the sample status is no maternal contamination; further judgment shows that most chromosomal statuses are normal, and finally the sample status is other.

[0171] S40A: Obtain the CNV results of the test sample with the sample status of other: Use the PennCNV software, the BAF feature data of the sample, the LRR feature data of the sample, and the pre-trained HMM model to obtain the segment CNV results on each chromosome in the sample, and filter out the CNVs with a CNV region length less than 100k and the number of SNPs in the CNV region less than 50.

[0172] Refer to Figures 10 - 11 , in this embodiment, through analysis, it is found that chromosomes 6 and 9 of this sample show chromosomal duplications.

[0173] S40B: Summarize the genotype data of SNPs in all chromosomes of the test sample with the sample status of other. If the heterozygosity ratio of the test sample is greater than the second set threshold, the sample status of the test sample is updated to genome-wide ROH.

[0174] In this embodiment, the heterozygosity ratio of sample RD23103101 is 0.0939, which is less than the threshold 0.1, that is, the sample status of this sample is updated to genome-wide ROH.

[0175] S50: Output the sample status and / or CNV results and / or ROH results: Prompt that the sample is genome-wide ROH, and the karyotype is arr[hg19](1-22)×2hmz, (X,N)×1hmz, (6,9)×3.

[0176] Example 5

[0177] The gDNA of sample RD23071601 was detected using the Illumina Infinium Asian Screening Array chip. Information on the SNP sites corresponding to the chip was obtained by analyzing the fluorescence signals after hybridization. Then, the following steps were used to obtain the sample status, CNV results, and ROH results of the sample:

[0178] Step S11a: Obtain the BAF feature data of each autosome based on the genomic position order of the SNP sites.

[0179] Step S11b: Obtain the LRR feature data of each autosome based on the genomic position order of the SNP sites.

[0180] Step S11c: Obtain the genotype feature data of each autosome based on the genomic position order of the SNP sites.

[0181] Step S21a: Use the mixture Gaussian model to obtain the peak value and the number of peaks of peak F1 on each autosome.

[0182] Step S21b: Use a sliding window to obtain the peak value and the number of peaks of peak F2 on each autosome.

[0183] Step S21c: Obtain the peak value and the number of peaks of the common peak F of peak F1 and peak F2 on each autosome.

[0184] Step S22: Judge the chromosomal status of each autosome according to the number of peaks, peak values of the common peak F, and the LRR feature data.

[0185] Refer to Figure 12 In this embodiment, in sample RD23071601, chromosomes 1 to 22 all showed normal chromosomes.

[0186] Step S30: Summarize the chromosomal status types of all autosomes obtained in Step S22. Judge whether the sample status of the tested sample is maternal contamination or no maternal contamination according to the chromosomal status with the highest proportion. If the sample status is no maternal contamination, further judge whether the sample status is haploid, triploid, or other.

[0187] In this embodiment, by summarizing the chromosomal status of 22 autosomes, it was found that there was no maternal chromosomal contamination in the chromosomal status, that is, the sample status was no maternal contamination; further judgment found that the chromosomal status of all autosomes was normal or unknown, that is, the sample status was other.

[0188] S40A: Obtain the CNV results of the tested samples with the sample status being other: Using the PennCNV software, the BAF feature data of the samples, the LRR feature data of the samples, and the pre-trained HMM model, obtain the segment CNV results on each chromosome in the samples, and filter out the CNVs with a CNV region length less than 100k and the number of SNPs in the CNV region less than 50.

[0189] See Figure 13 , in this embodiment, through analysis, it is found that there is a segment deletion in the region from 62546172 to 68494674 on chromosome 18 of this sample.

[0190] S40B: Aggregate the genotype data of SNPs in all chromosomes of the tested samples with the sample status being other. If the heterozygosity ratio of the tested sample is greater than the second set threshold, then update the sample status of the tested sample to genome-wide ROH.

[0191] In this embodiment, the heterozygosity ratio of sample RD23103101 is 0.1785, which is greater than the threshold 0.1, that is, the sample status remains unchanged and is still other.

[0192] S40C: Obtain the ROH results of the tested samples with the sample status being haploid, triploid or other: According to the positions of SNPs on the genome, divide 1000 adjacent SNPs into one bin. When the SNP homozygosity rate in the bin > 95%, and the LRR in this region is near 0, then consider this bin as a potential ROH region; form a new potential ROH region by merging adjacent potential ROH regions. If the homozygosity rate of SNPs in the new potential ROH region > 99%, then consider this region as an ROH region, and filter out the regions with the number of SNPs less than 500 and the length less than 1M to obtain the final sample ROH results.

[0193] In this embodiment, through analysis, it is found that there is no ROH region in this sample.

[0194] S50: Output the sample status and / or CNV results and / or ROH results: Prompt that there is a segment deletion in the sample, and the karyotype is arr[hg19]18q22.1q22.2(62546172_68494674)×1.

[0195] Comparative Example 1

[0196] This comparative example uses STR detection to obtain the maternal contamination results of samples RD22112301, RD22120801, RD23120601, RD24052801, and RD24050301: Collect the whole blood of the tested samples and the sample mothers, and perform traditional STR locus detection to obtain the maternal contamination results of the samples. At the same time, use the method provided by the present invention to obtain the maternal contamination results of samples RD22112301, RD22120801, RD23120601, RD24052801, and RD24050301.

[0197] Table 2 Results of maternal contamination of the same samples by the method provided by the present invention and STR locus detection

[0198]

[0199] The detection results are shown in Table 2. The method provided by the present invention can detect the maternal contamination results of samples, which is consistent with the results of traditional STR locus detection. Therefore, it can replace traditional STR locus detection.

[0200] Experimental Example 1

[0201] Use the method provided by the invention to calculate the heterozygosity ratios of different types of clinical R & D samples. Referring to Table 3, the heterozygosity ratio of whole-genome ROH is less than 0.1, while the heterozygosity ratios of the remaining samples are all greater than 0.1.

[0202] Table 3 Heterozygosity ratios of different types of samples

[0203]

[0204] In summary, compared with the prior art, the method provided by the present application has the following advantages:

[0205] (1) The method provided by the present invention does not require additional experiments to determine whether there is maternal contamination in the sample, and does not require additional provision of maternal peripheral blood, which can improve the detection efficiency of copy number abnormalities, improve the detection comprehensiveness, and reduce the detection time and cost.

[0206] (2) By comprehensively evaluating the peak conditions of BAF data on autosomes, the present invention can not only quickly judge the maternal contamination situation of the sample, but also quickly judge whether the sample belongs to special types of samples such as haploid, triploid, and whole-genome ROH.

[0207] (3) By using the method of determining the final peak number based on the common peak number, the present invention can effectively avoid the overfitting and underfitting situations that may occur in the results of the mixture Gaussian model or the sliding window result.

[0208] The above-described embodiments merely represent several implementation manners of the present invention. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and the present invention also intends to cover these modifications and variations.

Claims

1. A method for identifying maternal contamination and detecting copy number abnormalities using peaks, characterized in that, Including: Step S10: Detect the SNP locus information of the sample to be tested, and extract the LRR feature data, BAF feature data, and genotype feature data of each autosome of the sample to be tested based on the genomic position order of the SNP loci according to the SNP locus information; Step S20: According to the BAF feature data, use the mixture Gaussian model and / or the sliding window to obtain the number of peaks and peak values on each autosome, and judge the chromosomal state of each autosome according to the number of peaks, peak values, and LRR feature data: When the number of peaks is 0 and the BAF feature data is mainly distributed near 0 and 1, if the mean value of the LRR feature data is less than -0.2, the chromosomal state of this autosome is chromosomal overall deletion, otherwise, the chromosomal state is unknown; When the number of peaks is 1, if the peak value is within the interval [0.45, 0.55] and its peak ratio is greater than the second set threshold, the chromosomal state of this autosome is chromosomal normal, otherwise, the chromosomal state is maternal contamination; When the number of peaks is 2, if the two peak values are within the intervals [0.3, 0.4] and [0.6, 0.7] respectively, the chromosomal state of this autosome is chromosomal duplication, otherwise, the chromosomal state is unknown; If the number of peaks is greater than or equal to 3, the chromosomal state of this autosome is maternal contamination; Among them, the calculation formula of the peak ratio is: Among them, using the mixture Gaussian model to obtain the number of peaks and peak values on each autosome includes: using the mixture Gaussian model to fit the distribution of the BAF feature data on each autosome, calculating the Bayesian information criterion BIC of the mixture Gaussian model corresponding to 1 Gaussian distribution to 10 Gaussian distributions, and the number and mean value of the Gaussian distributions corresponding to the mixture Gaussian model with the lowest BIC are the number of peaks and peak values on each autosome; Step S30: Judge whether the sample state of the sample to be tested is maternal contamination or no maternal contamination according to the chromosomal state with the highest proportion among all chromosomal states; if the sample state is no maternal contamination, further judge whether the sample state is haploid, triploid or others, and perform step S40 processing; Step S40: According to the LRR feature data and BAF feature data, obtain the CNV result of the sample to be tested; according to the LRR feature data and genotype feature data, judge whether the sample to be tested is whole-genome ROH and / or obtain the ROH result of the sample to be tested; the ROH result is the ROH result of the sample to be tested with the sample state of haploid, triploid or others, and the acquisition method of the ROH includes the following steps: according to the position of the SNP locus on the genome, divide n adjacent SNPs into bins, if the homozygosity rate of the SNPs in the bin > 95%, and the mean value of the LRR feature data is near 0, then the bin is a potential ROH region; merge adjacent potential ROH regions to form a new potential ROH region, if the homozygosity rate of the SNPs in the new potential ROH region > 99%, then the new potential ROH region is an ROH region, summarize all ROH regions, and obtain the ROH result.

2. The method for identifying maternal contamination and detecting copy number abnormalities using peaks according to claim 1, characterized in that Step S20 further includes obtaining the number of peaks and peak values of the common peaks of the Gaussian mixture model and the sliding window on each autosome, including: obtaining the intersection region between the peaks of the Gaussian mixture model and the peaks of the sliding window method; if the proportion of the intersection region is greater than a first set threshold, then the union region of the peaks of the Gaussian mixture model and the peaks of the sliding window method that form the intersection region is the common peak region; using the sliding window to detect the peak values in the common peak region to obtain the peak values and the number of peaks of the common peaks.

3. The method for identifying maternal contamination and detecting copy number abnormalities using peaks according to claim 1, wherein The step of obtaining the CNV result of the sample to be tested in step S40 includes the following steps: according to the LRR feature data, BAF feature data, and HMM model, obtaining the segment CNV results on each autosome in the samples to be tested with the sample status of others, merging the adjacent segment CNV results, and filtering the segment CNV results according to the length of the CNV and the SNP data included in the CNV to obtain the CNV result.

4. The method for identifying maternal contamination and detecting copy number abnormalities using peaks according to claim 1, wherein The determination in step S40 of whether the sample to be tested is a genome-wide ROH includes: calculating the total heterozygosity ratio of all autosomes of the sample to be tested with the sample status of others; if the heterozygosity ratio is less than a third set threshold, then the sample status is updated to genome-wide ROH, otherwise, the sample status is others; wherein, the calculation formula of the heterozygosity ratio is as follows:

5. An apparatus for identifying maternal contamination and detecting copy number abnormalities using peaks, for performing the method according to any one of claims 1-4, characterized in that, The device includes: A data extraction unit that extracts and obtains the LRR feature data, BAF feature data, and genotype feature data based on the SNP locus genomic position order of each autosome of the sample to be tested according to the SNP locus information of the sample to be tested in the database; A chromosome status acquisition unit that, according to the BAF feature data, obtains the number of peaks and peak values on each autosome, and determines the chromosome status of each autosome according to the number of peaks, peak values, and LRR feature data; A first sample status determination unit that determines whether the sample status of the sample to be tested is the presence of maternal contamination or no maternal contamination according to the chromosome status with the highest proportion among all chromosome statuses; if the sample status is no maternal contamination, then further determines the sample status of the sample to be tested as haploid, triploid, or others according to the chromosome status with the highest proportion among all chromosome statuses; A second sample status determination unit that, if the sample status is not the presence of maternal contamination, obtains the sample CNV result of the sample to be tested according to the LRR feature data and BAF feature data; and determines whether the sample to be tested is a genome-wide ROH and / or obtains the ROH result of the sample to be tested according to the LRR feature data and genotype feature data.

6. A computer device, characterized in that, The computer device includes a processor and a memory: The memory is used to store a computer program and transmit the computer program to the processor; The processor is used to execute the method according to any one of claims 1 to 4 according to the instructions in the computer program.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program, and when the computer program is executed by the processor, the processor is caused to execute the method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Method or device for identifying parent tendency of nucleic acid sample

    CN114214425A

  • Method or device for identifying maternal source pollution and correcting copy number

    CN117558343A