A detection system, device and method for analyzing embryo chromosomal aneuploidy and parental contamination

By using gene linkage disequilibrium and PicoPLEX single-cell amplification technology, a simplified analysis of embryonic chromosomal aneuploidy and parental contamination in conventional IVF fertilized embryos was achieved through a single test. This solved the problems of complex testing, high cost, and high error rate in existing technologies, and achieved efficient and accurate embryo testing.

CN117238375BActive Publication Date: 2026-07-24SUZHOU BASECARE MEDICAL DEVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SUZHOU BASECARE MEDICAL DEVICE CO LTD
Filing Date
2023-07-20
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing technologies for detecting embryos fertilized in conventional IVF have problems such as high detection costs, long cycles, complex operations, and high error rates. They are difficult to effectively track the parental genetic information of the embryos and cannot accurately distinguish between parental contamination and triploidy.

Method used

An embryonic chromosome aneuploidy and parental contamination analysis and detection system was adopted to achieve embryonic aneuploidy detection and parental contamination analysis in a single test. The likelihood values ​​of reads in the sample originating from multiple strands were inferred by gene linkage disequilibrium. Combined with PicoPLEX single-cell amplification technology, whole genome sequencing and bioinformatics analysis were performed.

Benefits of technology

It achieves a simplified detection process, high success rate, low cost, and short cycle for analyzing embryonic chromosomal aneuploidy and parental contamination. It has high genome coverage, high detection accuracy, can effectively exclude paternal contamination introduced by sperm, and can analyze parental contamination without requiring blood samples from both parents of the embryo.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117238375B_ABST
    Figure CN117238375B_ABST
Patent Text Reader

Abstract

The application discloses a detection system, device and method for embryo chromosomal aneuploidy and parent contamination analysis. The system comprises a database module, an alignment module, an analysis and calculation module and a parent contamination detection module. The analysis and calculation module is used for performing the following steps: calculating the ratio of the number of effective sequences matched to each chromosome to the number of corresponding chromosome sequences in the reference database, performing statistical analysis, obtaining the number of target chromosomes of the sample to be detected, and obtaining the detection result of whether the frequency of alleles observed in the genetic linkage disequilibrium mode is abnormal. The parent contamination detection module comprises a contamination prediction model. The aneuploidy detection and parent contamination analysis of the embryo can be simultaneously realized by one detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of bioinformatics technology and relates to a detection system, device and method for analyzing embryonic chromosome aneuploidy and parental contamination. Background Technology

[0002] Conventional in-vitro fertilization (IVF) is primarily used to treat couples with infertility due to non-male factors. While increasing the number of oocytes can improve the fertilization rate of mature oocytes and thus the pregnancy rate to some extent in mild male-factor infertility, the risk of IVF failure is as high as 50% in patients with severe oligospermia, asthenospermia, or teratospermia compared to patients with normal sperm morphology. With the rapid development of molecular biology, preimplantation genetic diagnosis (PGD) has begun to develop and be applied clinically, building upon assisted reproductive technology and micromanipulation. Some reproductive centers both domestically and internationally recommend and have already implemented preimplantation genetic testing for aneuploidy (PGT-A) in addition to traditional IVF for couples with non-male-factor infertility.

[0003] There are two main concerns regarding PGT testing of conventional IVF fertilized embryos: i) paternal contamination caused by the introduction of sperm; ii) an increased proportion of chimeric embryos due to the introduction of sperm; and iii) maternal contamination from granulosa cells, with conventional IVF fertilized embryos exhibiting higher frequency and more severe maternal contamination. The qPCT detection method has been reported for routine IVF frozen embryo gene copy number analysis and tracking of parental genetic information. However, qPCT technology combines preimplantation genetic screening (PGS) with microarray analysis, requiring two experiments, resulting in high costs, long testing cycles, complex experimental procedures, and high error rates. In chromosomal number variation (CNV) analysis, each sample generates approximately 2Mb reads with a single-end read length of 55bp, resulting in low genome coverage, low sequencing depth, and high error rates. When tracking parental genetic information in embryo samples, based on the limited number of single nucleotide polymorphism (SNP) sites obtained, it is necessary to combine the whole genome information of both parents' gDNA to determine whether parental contamination exists in the sample. When performing parental contamination analysis, the parent embryo samples must be tested to determine the degree of contamination. In embryo ploidy analysis, based on the limited number of SNP sites, it is impossible to infer or distinguish whether the embryo is contaminated with parental genetic information or is triploid.

[0004] In conclusion, to demonstrate the clinical feasibility of PGT testing in routine IVF fertilized embryos, it is urgent to develop a simple and effective detection method for paternal contamination from sperm source and maternal contamination from granulosa cell source, which can effectively track parental genetic information in trophoblastic ectoderm cells of the embryo. Summary of the Invention

[0005] To address the shortcomings of existing technologies and practical needs, this invention provides a detection system, device, and method for analyzing embryonic chromosomal aneuploidy and parental contamination, which can simultaneously detect embryonic aneuploidy and analyze parental contamination in a single test.

[0006] To achieve the above objectives, the present invention adopts the following technical solution:

[0007] In a first aspect, the present invention provides an analysis and detection system for embryonic chromosome aneuploidy and parental contamination, the system comprising a database module, a comparison module, an analysis and calculation module, and a parental contamination detection module;

[0008] The database module stores human genomic DNA information;

[0009] The comparison module is used to perform the following:

[0010] The reads from the whole genome sequencing of the sample to be tested are compared with the human genomic DNA information to obtain the number of valid sequences matched on each chromosome;

[0011] The analysis and calculation module is used to perform the following:

[0012] The ratio of the number of valid sequences matched to each chromosome to the number of corresponding chromosome sequences in the reference database is calculated, and statistical analysis is performed to obtain the detection results of whether the number of target chromosomes in the test sample and the frequency of allele occurrence observed in the gene linkage disequilibrium mode are abnormal.

[0013] The parental contamination detection module includes a contamination prediction model. The parental contamination detection module establishes a contamination prediction model based on the principle of tracking parental genetic information in embryos to analyze the degree of parental contamination. The contamination prediction model uses gene linkage disequilibrium to infer the likelihood value of reads in the sample originating from multiple strands.

[0014] This invention develops a detection system for analyzing embryonic chromosomal aneuploidy and parental contamination. First, the whole-genome sequencing data of the sample to be tested undergoes quality control, including the removal of adapter sequences and low-quality sequences. Then, bioinformatics software is used to compare the filtered whole-genome sequencing data of the sample to be tested with the human genomic DNA information to obtain the number of valid sequences matching each chromosome. The ratio of the number of valid sequences to the number of corresponding chromosome sequences in the reference database is calculated, enabling the detection of chromosomal aneuploidy abnormalities. Simultaneously, a parental contamination prediction model based on the principle of gene linkage disequilibrium inferring the likelihood values ​​of reads originating from multiple strands in the sample is used to analyze the degree of parental contamination in the sample. Furthermore, the source of contamination is determined by comparing the genetic information of the embryonic parents' genomic DNA.

[0015] In this invention, the gene linkage disequilibrium refers to the situation where two genes are not completely independently inherited, but exhibit a certain degree of linkage; the likelihood value refers to the probability that the parameter θ = θ1 (relative to the other parameter value θ2) is the true value under a given sample X = x.

[0016] Preferably, the database module is located locally or in the cloud.

[0017] Preferably, the pollution prediction model is used to perform the following steps:

[0018] S1. Divide the chromosome into several small windows containing the number of first base pairs, and identify the single nucleotide polymorphism (SNP) sites covered by reads in each small window;

[0019] S2. Calculate the frequency of haplotypes formed by the SNPs in the reference population;

[0020] S3. Randomly select ≥4 reads, calculate the likelihood ratio of observed allele occurrences under the ploidy hypothesis (assuming haploid, triploid, or diploid), and perform logarithmic calculation on the ratio under the two assumed ploidy hypotheses.

[0021] S4. Repeat the steps in S3;

[0022] S5. After merging the several small windows containing the first number of base pairs into a large window containing the second number of base pairs, the average of all log values ​​is taken as LLR;

[0023] S6. If the LLR is greater than 0, it is likely that there are three completely different homologous chromosomes; if it is less than 0, it is likely that there are two identical homologs plus one different homolog.

[0024] S7. By constructing a reference library of normal diploid embryos, the Z-score of each chromosome is calculated based on the LLR value of the haplotype blocks formed by the first and second base numbers; the Z-score represents how many standard deviations above or below the mean, and the larger the absolute value, the more significant the difference.

[0025] S8. Based on the Z-score of samples with different pollution gradients, draw a fitted line to determine the pollution ratio and whether pollution exists.

[0026] Preferably, the number of the first base is ≥50 kilobase pairs.

[0027] Preferably, the number of the second base is ≥4 megabase pairs.

[0028] In this invention, the Z-score represents how many standard deviations above or below the mean; the larger the absolute value, the more significant the difference. For example, a Z-score of 3 means that the value is 3 standard deviations above the mean, and a Z-score of -3 means that the value is 3 standard deviations below the mean.

[0029] In this invention, parental genetic information is tracked by inferring the likelihood (log-likelihood ratio, LLR) of reads in a sample from multiple strands based on gene linkage disequilibrium. A single test can simultaneously detect aneuploidy in embryos and analyze parental contamination.

[0030] Preferably, the parental contamination detection module is further configured to perform contamination sample source tracing analysis, the contamination sample source tracing analysis including:

[0031] The raw sequencing data of the embryo samples to be tested were mixed with the raw sequencing data of the gDNA samples of the embryo's parents for analysis. The median of LLR>0 in MixF0 and the median of LLR>0 in MixM0 were taken, and the ratio between the two was set as the R value. If R<0.8, it indicates that the sample has paternal contamination; if R>1.4, it indicates that the sample has maternal contamination. MixF0 is the raw sequencing data of the embryo sample to be tested mixed with the raw sequencing data of the embryo's father's gDNA sample, and MixM0 is the raw sequencing data of the embryo sample to be tested mixed with the raw sequencing data of the embryo's mother's gDNA sample.

[0032] In this invention, an algorithm for identifying pollution sources is designed to further clarify the sources of pollution.

[0033] In a second aspect, the present invention provides an apparatus for analyzing embryonic chromosomal aneuploidy and parental contamination, the apparatus comprising the detection system for analyzing embryonic chromosomal aneuploidy and parental contamination as described in the first aspect.

[0034] Preferably, the device further includes a single-cell amplification system, a library construction system, a sample loading system, and a sequencing system.

[0035] Thirdly, the present invention provides a method for detecting embryonic chromosomal aneuploidy and parental contamination, the method comprising the following steps:

[0036] S1. Perform whole-genome amplification on the target biological sample;

[0037] S2. Fragmentation of whole-genome DNA;

[0038] S3. Construct a whole-genome DNA library;

[0039] S4. Sequencing of the whole-genome DNA library;

[0040] S5. The reads obtained from sequencing are completely matched and compared with the human genome to obtain the number of effective sequences on each chromosome, and the detection of chromosomal aneuploidy is realized based on the ratio of the number of effective sequences to the number of corresponding chromosome sequences in the reference database;

[0041] S6. Based on the effective amount of data, the likelihood values ​​of reads in the sample originating from multiple strands are inferred through gene linkage disequilibrium to perform parental contamination analysis on the sample to be tested.

[0042] Preferably, the parental contamination analysis includes the following steps:

[0043] S6.1 divides the chromosome into several small windows containing the number of first base pairs and identifies the SNP sites covered by reads in each small window;

[0044] S6.2 Calculate the frequency of haplotypes formed by the SNPs in the reference population;

[0045] S6.3 Randomly select ≥4 reads, calculate the likelihood ratio of observed allele occurrences under the ploidy hypothesis, and perform logarithmic calculation on the ratio under the two assumed ploidy hypotheses.

[0046] S6.4 Repeat step S6.3;

[0047] S6.5 After merging the several small windows containing the first number of base pairs into a large window containing the second number of base pairs, the average of all log values ​​is taken as LLR;

[0048] S6.6 If the LLR is greater than 0, it is highly likely that there are three completely different homologous chromosomes (three completely different ones). If it is less than 0, it is highly likely that there are two identical homologs plus one different homolog (two identical and one different).

[0049] S6.7 By constructing a reference library of normal diploid embryos, the Z-score of each chromosome is calculated based on the LLR value of the haplotype blocks formed by the first and second base numbers. The Z-score represents how many standard deviations above or below the mean; the larger the absolute value, the more significant the difference.

[0050] S6.8 uses the Z-score of samples with different pollution gradients to fit a line, determine the pollution ratio, and identify whether pollution exists.

[0051] Preferably, the target biological sample includes any one of the following: a small number of biopsy cells from conventional IVF fertilized embryos, a small number of biopsy cells from intracytoplasmic sperm injection (ICSI) fertilized embryos, or an embryo culture medium sample.

[0052] Preferably, the method for whole-genome amplification includes single-cell whole-genome amplification.

[0053] Preferably, the single-cell whole genome amplification is performed using the PicoPLEX single-cell amplification method.

[0054] In this invention, the PicoPLEX single-cell amplification method has a poor effect on sperm DNA amplification. Even 25 sperm cannot be effectively amplified. To a certain extent, it can effectively eliminate potential paternal contamination caused by the introduction of sperm into the biopsy sample.

[0055] In this invention, the parental contamination analysis specifically includes using sequencing data and genotype linkage disequilibrium analysis to infer the likelihood values ​​of reads in the sample originating from multiple strands, and then analyzing the degree of parental contamination through an established contamination prediction model.

[0056] Preferably, the method for constructing the pollution prediction model includes:

[0057] PicoPLEX monoblot products from clinically uncontaminated blastocyst TE biopsy samples (e.g., 28 cases) were used to mix with their corresponding maternal or paternal gDNA PicoPLEX monoblot products at different ratios (paternal / maternal:embryo DNA PicoPLEX monoblot product mixing ratios of 0:1, 1:9, 2:8, 3:7, 4:6, 5:5, 6:4, 7:3, 8:2, 9:1, 0:10), with each contamination ratio repeated at least three times. All samples were artificially simulated with different contamination ratios, and libraries were constructed through fragmentation, end repair, adapter ligation, and PCR amplification. All libraries were sequenced at the whole-genome level using the MGI 200 sequencing platform, with a data volume of 30Mb per sample. Reads, with a single-end length of 100bp; all raw data from samples undergo quality control, including the removal of adapter sequences and low-quality sequences; the processed data is compared with the human reference genome, sorted, and duplicate alignment sequences are removed to obtain the alignment results, and the alignment results are quality controlled based on indicators such as alignment rate, genome coverage, and unique alignment rate; the filtered data is optimized based on the likelihood values ​​of gene linkage disequilibrium originating from multiple strands; a reference library of normal diploid embryos is constructed, and the Z-score value of each chromosome is calculated based on the LLR value of haplotype blocks; based on the Z-scores of samples with different contamination gradients, a fitting line is plotted to determine the contamination ratio and establish a contamination prediction model.

[0058] Preferably, the number of the first base is ≥50 kilobase pairs.

[0059] Preferably, the number of the second base is ≥4 megabase pairs.

[0060] Preferably, the detection method for embryonic chromosomal aneuploidy and parental contamination analysis further includes contamination sample source tracing analysis, which includes:

[0061] The raw sequencing data of the embryo samples to be tested were mixed with the raw sequencing data of the gDNA samples of the embryo's parents for analysis. The median of LLR>0 in MixF0 and the median of LLR>0 in MixM0 were taken, and the ratio between the two was set as the R value. If R<0.8, it indicates that the sample has paternal contamination; if R>1.4, it indicates that the sample has maternal contamination. MixF0 is the raw sequencing data of the embryo sample to be tested mixed with the raw sequencing data of the embryo's father's gDNA sample, and MixM0 is the raw sequencing data of the embryo sample to be tested mixed with the raw sequencing data of the embryo's mother's gDNA sample.

[0062] Compared with the prior art, the present invention has the following beneficial effects:

[0063] (1) Embryo samples only need to undergo one IVF-PGTA test to simultaneously detect embryonic chromosomal aneuploidy and analyze parental contamination. The test procedure is simple, has a high success rate, low testing cost, and a short cycle.

[0064] (2) Each embryo sample has 30Mb reads, with a single-end read length of 100bp, high genome coverage, high sequencing depth, and low detection error rate.

[0065] (3) PicoPLEX single-cell amplification is used, which has a lower single-base amplification error rate and better repeatability, and can effectively eliminate potential paternal contamination caused by the introduction of sperm during the biopsy process.

[0066] (4) The LLR analysis method based on gene linkage disequilibrium analyzes PicoPLEX single-cell amplification products. It has low requirements for gene coverage and maintains high detection accuracy in low-coverage genomes. In addition, the LLR analysis method is not limited by the amplification method.

[0067] (5) A large number of research and development experiments were conducted using the PicoPLEX monoamplification products of known CNVs and their corresponding parental gDNA samples to establish a predictive contamination model. The correlation coefficient of the standard curve established by the contamination detection results was 0.999. In the established artificial contamination model, 10% of the simulated parental contamination samples can be effectively detected. IVF-PGTA can detect ≥10% of the parental contamination.

[0068] (6) When performing parental contamination analysis of embryo samples, the degree of contamination can be analyzed without the need for blood samples from both parents of the embryo, and it can be determined whether the embryo sample to be tested is contaminated with parental contamination. In embryo chromosome ploidy analysis, the LLR trend can be used to infer whether the embryo sample to be tested is contaminated with parental contamination or belongs to triploid embryos. This method can also detect uniparental disomy. Attached Figure Description

[0069] Figure 1 A flowchart for simulating the detection process of contaminated samples using IVF-PGTA;

[0070] Figure 2 A graph showing the correlation analysis results of the predicted pollution standard curve;

[0071] Figure 3 This is a schematic diagram illustrating the interpretation of the source of sample contamination.

[0072] Figure 4 The graph shows the LLR analysis results for uncontaminated samples.

[0073] Figure 5 Figures showing the LLR analysis results for samples with 20%, 30%, and 50% paternal contamination.

[0074] Figure 6 Figures showing the LLR analysis results for samples with 20%, 30%, and 50% parental contamination.

[0075] Figure 7 The results of simulating different proportions of paternal contamination by adding different numbers of frozen sperm into corresponding embryo biopsy trophoblastic ectoderm cells;

[0076] Figure 8 The results of simulating different proportions of maternal contamination by adding different numbers of frozen granulocytes to corresponding embryo biopsy trophoblastic ectoderm cells;

[0077] Figure 9 The LLR analysis results are shown to simulate samples with different proportions of parental contamination. NC represents uncontaminated samples.

[0078] Figure 10 The LLR analysis results are shown to simulate samples with different proportions of parental contamination. NC represents uncontaminated samples.

[0079] Figure 11 IVF-PGTA testing flowchart;

[0080] Figure 12 The image shows the results of SNPs and LLR analysis based on a clinical embryo biopsy sample.

[0081] Figure 13 The image shows the results of SNPs and LLR analysis based on a clinical embryo biopsy sample. Detailed Implementation

[0082] To further illustrate the technical means and effects of this invention, the following description, in conjunction with embodiments and accompanying drawings, provides a further explanation of the invention. It is understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it.

[0083] Where specific techniques or conditions are not specified in the examples, they shall be performed in accordance with the techniques or conditions described in the literature in this field, or in accordance with the product instructions. Reagents or instruments whose manufacturers are not specified are all conventional products that can be purchased through legitimate channels.

[0084] This invention develops a novel method for detecting parental contamination in conventional IVF fertilized embryos—IVF-PGTA. This method is based on linkage disequilibrium (LD) to infer the likelihood of reads originating from multiple strands in the sample. A single test can simultaneously detect embryonic aneuploidy and analyze parental contamination. Specifically, it includes: using PicoPLEX single-cell amplification, 5-10 trophectoderm cells (TEs) obtained from conventional IVF fertilized embryos at the blastocyst stage are amplified into a single cell genome, followed by high-throughput sequencing of the entire genome using the MGI 200 sequencing platform. Bioinformatics software is used to analyze the sequencing results, obtaining the number of valid sequences matching each chromosome. The ratio of the number of valid sequences to the number of corresponding chromosome sequences in a reference database is calculated, enabling the detection of chromosomal aneuploidy. Simultaneously, the genotype linkage disequilibrium analysis method is used to infer the log likelihood of reads originating from multiple strands in the sample. The parental contamination level was analyzed using a contamination prediction model (LLR). If further analysis of the source of contamination is needed, the source of contamination can be traced by comparing the gDNA genomic information of the embryonic parents.

[0085] Example 1

[0086] This embodiment establishes an artificially simulated pollution model.

[0087] To further improve the detection rate of contamination originating from parental cells, this embodiment utilized PicoPLEX single-cell amplification products from TE biopsies of 28 clinically identified, uncontaminated embryos with known CNVs. Parental gDNA samples from corresponding families were also collected for whole-genome amplification using PicoPLEX single-cell amplification. The PicoPLEX single-cell amplification products of the embryos and their parents were mixed in different ratios (paternal / maternal:embryo DNA PicoPLEX single-cell amplification product mixing ratios of 0:1, 1:9, 2:8, 3:7, 4:6, 5:5, 6:4, 7:3, 8:2, 9:1, 0:10). Library construction was performed through fragmentation, end repair, adapter ligation, and PCR amplification. High-throughput whole-genome sequencing was then conducted using the MGI 200 sequencing platform. Each sample yielded 30 Mb reads, with a single-end read length of 100 bp. The raw sequencing data for each sample undergoes quality control, including the removal of adapter sequences and low-quality sequences. A Q30 > 80% meets the requirements. The processed data is then compared with a human reference genome. Duplicate alignment sequences are sorted and removed to obtain the alignment results. Alignment results are quality controlled based on indicators such as alignment rate, genome coverage, and unique alignment rate. Bioinformatics software is then used to analyze the sequencing results, determining the number of valid sequences matching each chromosome. The ratio of the number of valid sequences to the corresponding chromosome sequences in the reference database is calculated, enabling the detection of chromosomal aneuploidy.

[0088] Meanwhile, the filtered data is used for algorithm optimization based on the likelihood values ​​of gene linkage disequilibrium originating from multiple strands, for analysis of parental contamination in embryos, and a corresponding contamination prediction model is established. Figure 1 The specific analytical principle based on the likelihood values ​​of gene linkage disequilibrium originating from multiple strands is as follows:

[0089] (1) Divide the chromosome into 50kb windows and identify common SNP sites covered by reads in each window;

[0090] (2) Calculate the frequency of haplotypes composed of these common SNPs in the reference population;

[0091] (3) Randomly select 4 reads and calculate the likelihood ratio of the observed alleles under the ploidy assumptions (BPH, SPH, dissimony, monosomy). The ratio under the two ploidy assumptions is calculated by log.

[0092] (4) Repeat steps (3);

[0093] (5) After merging the 50kb small window into a 4Mb large window, take the average of all log values.

[0094] To further quantify LLR abnormalities, a reference library of normal diploid embryos was constructed. Z-scores for each chromosome were calculated based on haplotype block detection results for data optimization. The Z-score represents how many standard deviations above or below the mean; the larger the absolute value, the more significant the difference. For example, a Z-score of 3 indicates that the value is 3 standard deviations above the mean, and a Z-score of -3 indicates that the value is 3 standard deviations below the mean. Based on the Z-score values ​​of samples with different contamination gradients, a fitting curve was constructed, and a contamination prediction model was established to fit and analyze the raw sample data, interpreting the degree of contamination and determining whether the sample was contaminated. Simultaneously, a standard curve for contaminated positive samples was established based on previous research data. Figure 2 The vertical axis represents the predicted pollution ratio, and the horizontal axis represents the proportion of artificial pollution.

[0095] For samples with parental contamination, further source tracing analysis was conducted. The original data from the embryo samples to be tested were mixed with the original data from the gDNA samples of the parents of the embryos for analysis. The median of LLR>0 in MixF0 and the median of LLR>0 in MixM0 were taken, and the ratio was set as the R value (R = Mix(father - embryo)(LLR>0)). median / Mix(maternal parent - embryo) (LLR > 0) median ).like Figure 3 As shown, the ratio of the median of LLR > 0 in MixF0 to the median of LLR > 0 in MixM0 is calculated and set as the R value. If R < 0.8, i.e., the region represented by PC, it indicates that the sample has paternal contamination; if R > 1.4, i.e., the region represented by MC, it indicates that the sample has maternal contamination. The LLR analysis results of uncontaminated samples are shown below. Figure 4 As shown, for uncontaminated embryonic TE biopsy cells, IVF-PGTA detection was performed. LLR analysis was conducted on the sequencing data using the genotype linkage disequilibrium analysis method to calculate the Z-score value for each chromosome. The vertical axis represents the Z-score value, and the horizontal axis represents the chromosome number. In the LLR analysis graph, the vertical axis represents the LLR value, and the horizontal axis represents the sample chromosome number. In the Z-score result graph, the vertical axis represents the Z-score value, and the horizontal axis represents the chromosome number. The results for samples with 20%, 30%, and 50% paternal contamination are shown below. Figure 5 As shown, PicoPLEX monoamplification products of embryos and their paternal gDNA were mixed at ratios of 8:2, 7:3, and 5:5 to obtain 20%, 30%, and 50% paternally contaminated samples, respectively. LLR analysis was performed based on IVF-PGTA detection, and the Z-score value for each chromosome was calculated. The vertical axis represents the Z-score value, and the horizontal axis represents the chromosome number. Results for 20%, 30%, and 50% maternally contaminated samples are shown below. Figure 6As shown, the PicoPLEX monoamplification products of embryos and their maternal gDNA were mixed at ratios of 8:2, 7:3, and 5:5 to obtain maternally contaminated samples of 20%, 30%, and 50%, respectively. LLR analysis was performed based on IVF-PGTA detection to calculate the Z-score value of each chromosome. The vertical axis represents the Z-score value, and the horizontal axis represents the chromosome number.

[0096] Example 2

[0097] This embodiment performs routine in vitro fertilization embryo contamination detection (IVF-PGTA method), and the specific steps are as follows:

[0098] (1) PicoPLEX single-cell amplification: A small number of embryonic TE biopsy cells were taken for PicoPLEX single-cell amplification. The concentration of the single-cell amplification product > 15 ng / μL was considered qualified.

[0099] (2) DNA fragmentation: The extracted whole genome is fragmented by enzyme digestion to meet the requirements for sequencing.

[0100] (3) Library construction: After the fragmented DNA is processed by end repair, adapter addition, PCR amplification, purification and other treatments, the library concentration >5ng / μL is considered qualified;

[0101] (4) High-throughput sequencing: After quality control of the constructed library, whole-genome high-throughput sequencing was performed based on the MGI200 sequencing platform. The sequencing data volume of each test sample was 30Mb reads, the average sequencing depth was 1×, and the read length was 100bp at one end.

[0102] (5) Data analysis: The obtained reads were fully matched and compared with the human genome using bioinformatics software to obtain the number of valid sequences matched on each chromosome. The ratio of the number of valid sequences to the number of corresponding chromosome sequences in the reference database was calculated, which can realize the detection of chromosomal aneuploidy. At the same time, the sequencing data were analyzed using the genotype linkage disequilibrium analysis method to infer the likelihood value of reads in the sample originating from multiple strands. The degree of parental contamination was analyzed by the established contamination prediction model. If the source of contamination is further analyzed, the source of contamination is traced by comparing the gDNA genome information of the embryonic parents.

[0103] To clarify the sensitivity and accuracy of IVF-PGTA testing for parental contamination, 10 clinically discarded, fresh-cycle routine IVF embryos were collected. 5–10 TE biopsy cells were extracted from each embryo, with a maximum of three biopsies per embryo. Different numbers of frozen sperm were also collected. Figure 7 ) or different quantities of frozen ( Figure 8Granulocytes were added to TE biopsy samples to establish artificial parental contamination samples. All samples underwent CNV and LLR analysis based on the above IVF-PGTA assay. Z-score values ​​for each chromosome were calculated based on LLR analysis. In the CNV analysis graph, the vertical axis represents chromosome copy number, and the horizontal axis represents the sample chromosome number. In the LLR analysis graph, the vertical axis represents the LLR value, and the horizontal axis represents the sample chromosome number. In the Z-score graph, the vertical axis represents the Z-score value, and the horizontal axis represents the chromosome number. The LLR analysis results for samples simulating different proportions of paternal contamination are shown in the figure below. Figure 9 As shown, NC represents an uncontaminated sample. S3, S7, S10, S15, and S23 represent artificially created paternal contamination samples established by adding 3, 7, 10, 15, and 23 frozen sperm cells to embryo biopsy samples, respectively. Even with the addition of 23 sperm cells (S23), paternal contamination could not be detected, and CNVs results were normal. NC represents an uncontaminated sample. The LLR analysis results simulating different proportions of maternal contamination samples are shown in the figure below. Figure 10 As shown, NC represents an uncontaminated sample; G1, G3, and G5 represent the addition of 1, 3, and 5 cryopreserved granulocytes to the embryo biopsy sample, respectively; even adding 1 (G1) cryopreserved granulocyte can accurately detect contamination; CO-G1, CO-G3, CO-G5, CO-G7, and CO-G10 represent the addition of 1, 3, 5, 7, and 10 co-cultured granulocytes to the embryo biopsy sample, respectively.

[0104] To further demonstrate the feasibility of IVF-PGTA clinical application, biopsy samples from 40 discarded, fresh, conventionally IVF fertilized embryos from 10 families were used for IVF-PGTA performance testing to prove its effectiveness in clinical application. TE biopsy cells and corresponding inner cell mass samples were taken from each embryo for contamination assessment and karyotype consistency analysis. (See schematic diagram below.) Figure 11As shown, TE biopsy cells and corresponding inner cell mass samples were collected from clinically discarded, fresh-cycle, conventionally IVF fertilized embryos. The previously reported qPCT detection method is based on the analysis of parental contamination of embryos using all obtained SNP loci. To clarify the accuracy and sensitivity of IVF-PGTA in detecting parental contamination, CNV analysis and parental contamination analysis were performed on all biopsy samples from 40 embryos using SNP analysis (PGT-Plus) and LLR analysis (IVF-PGTA technology). Inner cell mass (ICM) samples and TE biopsy samples underwent whole-genome amplification using PicoPLEX single-cell amplification. Simultaneously, peripheral blood samples from both parents were collected and gDNA extracted. The embryonic PicoPLEX amplification products and parental gDNA were analyzed using a PGT-Plus sequencer on an MGI 200. Each sample contained at least 80 Mb reads with a paired-end length of 100 bp. After data filtering, SNP contamination analysis was performed. Then, 30 Mb reads from the single-end were extracted for IVF-PGTA analysis, including LLR analysis based on gene linkage disequilibrium and CNV analysis. The results of SNP and LLR analysis of TE and ICM samples from one embryo (5#) are shown in the figure below. Figure 12 As shown, SNP analysis based on PGT-Plus detection showed no parental contamination in either the TE or ICM biopsy samples of embryo #5. However, LLR analysis based on IVF-PGTA detection showed 24% maternal contamination in the TE sample (5-1#) of embryo #5, while no parental contamination was found in the ICM biopsy sample. Furthermore, SNP and LLR analysis results for the TE and ICM biopsy samples of embryo #9 (F) are as follows... Figure 13 As shown, both analytical methods revealed paternal contamination in the TE and ICM biopsy samples of the embryo. However, further analysis of the LLR trend plot determined that the embryo was a paternally triploid embryo (3PN). These results indicate that the LLR contamination analysis method based on gene linkage disequilibrium is more sensitive to contamination and provides more accurate interpretation than the SNP-based method. Furthermore, the LLR analysis method can infer whether an embryo is contaminated or triploid from the trend plot.

[0105] In summary, this invention develops a novel method for detecting parental contamination specifically in routine IVF fertilized embryo biopsy cells. This detection technology traces parental genetic information based on the likelihood values ​​of reads originating from multiple strands in the sample due to gene linkage disequilibrium. It has low requirements for gene coverage and maintains high detection accuracy even in areas with low genomic coverage. The detection technology uses PicoPLEX single-cell amplification, which, compared to MALBAC single-cell amplification technology, exhibits a lower single-base amplification error rate and better repeatability, effectively eliminating potential paternal contamination caused by sperm introduction during biopsy. Embryo samples only require a single IVF-PGTA test to simultaneously detect embryonic chromosomal aneuploidy and analyze parental contamination. The experimental procedure is simple to operate, has a high success rate, low detection cost, and a short cycle.

[0106] The applicant declares that the detailed method of the present invention is illustrated by the above embodiments, but the present invention is not limited to the above detailed method, that is, it does not mean that the present invention must rely on the above detailed method to be implemented. Those skilled in the art should understand that any improvements to the present invention, equivalent substitutions of the raw materials of the product of the present invention, addition of auxiliary components, selection of specific methods, etc., all fall within the protection scope and disclosure scope of the present invention.

Claims

1. A detection system for analyzing embryonic chromosomal aneuploidy and parental contamination, characterized in that, The system includes a database module, a comparison module, an analysis and calculation module, and a parent contamination detection module; The database module stores human genomic DNA information; The comparison module is used to perform the following: The reads from the whole genome sequencing of the sample to be tested are compared with the human genomic DNA information to obtain the number of valid sequences matched on each chromosome; The analysis and calculation module is used to perform the following: The ratio of the number of valid sequences matched to each chromosome to the number of corresponding chromosome sequences in the reference database is calculated, and statistical analysis is performed to obtain the detection results of whether the number of target chromosomes in the test sample and the frequency of allele occurrence observed in the gene linkage disequilibrium mode are abnormal. The parental contamination detection module includes a contamination prediction model; The parental contamination detection module establishes a contamination prediction model based on the principle of tracking embryonic parental genetic information to analyze the degree of parental contamination. The contamination prediction model uses gene linkage disequilibrium to infer the likelihood that reads in a sample originate from multiple strands; the contamination prediction model is used to perform the following steps: S1. Divide the chromosome into several small windows containing the number of first base pairs, and identify the SNP sites covered by reads in each small window; S2. Calculate the frequency of haplotypes formed by the SNPs in the reference population; S3. Randomly select ≥4 reads, calculate the likelihood ratio of observed allele occurrences under the ploidy hypothesis, and perform logarithmic calculation on the ratio under the two assumed ploidy hypotheses. S4. Repeat the steps in S3; S5. After merging the several small windows containing the first number of base pairs into a large window containing the second number of base pairs, the average of all log values ​​is taken as LLR; S6. If the LLR is greater than 0, it is highly likely that there are three completely different homologous chromosomes; if the LLR is less than 0, it is highly likely that there are two identical homologs plus one different homolog. S7. By constructing a reference library of normal diploid embryos, the Z-score of each chromosome is calculated based on the LLR value of the haplotype blocks formed by the first and second base numbers. The Z-score represents how many standard deviations above or below the mean; the larger the absolute value, the more significant the difference. S8. Based on the Z-score of samples with different pollution gradients, draw a fitted line to determine the pollution ratio and whether pollution exists.

2. The detection system for analyzing embryonic chromosomal aneuploidy and parental contamination as described in claim 1, characterized in that, The database module is set up locally or in the cloud.

3. The detection system for analyzing embryonic chromosomal aneuploidy and parental contamination as described in claim 1, characterized in that, The number of the first base is ≥50 kilobase pairs; The number of the second base is ≥4 megabase pairs.

4. The detection system for analyzing embryonic chromosomal aneuploidy and parental contamination as described in claim 1, characterized in that, The parental contamination detection module is also used to perform contamination sample source tracing analysis, which includes: The raw sequencing data of the embryo samples to be tested were mixed with the raw sequencing data of the gDNA samples of the embryo's parents for analysis. The median of LLR>0 in MixF0 and the median of LLR>0 in MixM0 were taken, and the ratio between the two was set as the R value. If R<0.8, it indicates that the sample has paternal contamination; if R>1.4, it indicates that the sample has maternal contamination. MixF0 is the raw sequencing data of the embryo sample to be tested mixed with the raw sequencing data of the embryo's father's gDNA sample, and MixM0 is the raw sequencing data of the embryo sample to be tested mixed with the raw sequencing data of the embryo's mother's gDNA sample.

5. An apparatus for analyzing embryonic chromosomal aneuploidy and parental contamination, characterized in that, The apparatus for analyzing embryonic chromosomal aneuploidy and parental contamination includes the detection system for analyzing embryonic chromosomal aneuploidy and parental contamination as described in any one of claims 1-4.

6. A method for detecting embryonic chromosomal aneuploidy and parental contamination, characterized in that, The method includes the following steps: S1. Perform whole-genome amplification on the target biological sample; S2. Fragmentation of whole-genome DNA; S3. Construct a whole-genome DNA library; S4. Sequencing of the whole-genome DNA library; S5. The reads obtained from sequencing are completely matched and compared with the human genome to obtain the number of effective sequences on each chromosome, and the detection of chromosomal aneuploidy is realized based on the ratio of the number of effective sequences to the number of corresponding chromosome sequences in the reference database; S6. Based on the effective amount of data, infer the likelihood values ​​of reads in the sample originating from multiple strands under the gene linkage disequilibrium mode to perform parental contamination analysis on the sample to be tested; The parental contamination analysis includes the following steps: S6.1 divides the chromosome into several small windows containing the number of first base pairs and identifies the SNP sites covered by reads in each small window; S6.2 Calculate the frequency of haplotypes formed by the SNPs in the reference population; S6.3 Randomly select ≥4 reads, calculate the likelihood ratio of observed allele occurrences under the ploidy hypothesis, and perform logarithmic calculation on the ratio under the two assumed ploidy hypotheses. S6.4 Repeat step S6.3; S6.5 After merging the several small windows containing the first number of base pairs into a large window containing the second number of base pairs, the average of all log values ​​is taken as LLR; S6.6 If the LLR is greater than 0, it is likely that there are three completely different homologous chromosomes; if it is less than 0, it is likely that there are two identical homologs plus one different homolog. S6.7 By constructing a reference library of normal diploid embryos, the Z-score of each chromosome is calculated based on the LLR value of the haplotype blocks formed by the first and second base numbers. The Z-score represents how many standard deviations above or below the mean; the larger the absolute value, the more significant the difference. S6.8 uses the Z-score of samples with different pollution gradients to fit a line, determine the pollution ratio, and identify whether pollution exists.

7. The detection method for embryonic chromosomal aneuploidy and parental contamination analysis as described in claim 6, characterized in that, The target biological sample includes any one of the following: a small number of biopsy cells from conventional IVF fertilized embryos, a small number of biopsy cells from intracytoplasmic sperm injection (ICSI) fertilized embryos, or an embryo culture medium sample.

8. The detection method for embryonic chromosomal aneuploidy and parental contamination analysis as described in claim 6, characterized in that, The whole genome amplification method includes single-cell whole genome amplification.

9. The detection method for embryonic chromosomal aneuploidy and parental contamination analysis as described in claim 8, characterized in that, The single-cell whole genome amplification was performed using the PicoPLEX single-cell amplification method.

10. The method for analyzing and detecting embryonic chromosomal aneuploidy and parental contamination as described in claim 6, characterized in that, The number of the first base is ≥50 kilobase pairs.

11. The method for analyzing and detecting embryonic chromosomal aneuploidy and parental contamination as described in claim 6, characterized in that, The number of the second base is ≥4 megabase pairs.

12. The method for analyzing and detecting embryonic chromosomal aneuploidy and parental contamination as described in claim 6, characterized in that, The parental contamination analysis also includes a step of quantifying LLR, and the method for quantifying LLR includes: By constructing a reference library of normal diploid embryos, the Z-score of each chromosome is calculated based on the LLR value of the haplotype blocks formed by the first and second base numbers. The Z-score represents how many standard deviations above or below the mean; the larger the absolute value, the more significant the difference.

13. The method for analyzing and detecting embryonic chromosomal aneuploidy and parental contamination as described in claim 6, characterized in that, The detection method for embryonic chromosomal aneuploidy and parental contamination analysis also includes contamination sample source tracing analysis, which includes: The raw sequencing data of the embryo samples to be tested were mixed with the raw sequencing data of the gDNA samples of the embryo's parents for analysis. The median of LLR>0 in MixF0 and the median of LLR>0 in MixM0 were taken, and the ratio between the two was set as the R value. If R<0.8, it indicates that the sample has paternal contamination; if R>1.4, it indicates that the sample has maternal contamination. MixF0 is the raw sequencing data of the embryo sample to be tested mixed with the raw sequencing data of the embryo's father's gDNA sample, and MixM0 is the raw sequencing data of the embryo sample to be tested mixed with the raw sequencing data of the embryo's mother's gDNA sample.