Method and device for detecting chromosome uploidy

By using nanopore sequencing data and window division information when detecting chromosome euployness, combined with the correction data of negative samples, baseline values ​​are established to evaluate chromosome euployness, and the problem of insufficient detection speed and accuracy in the prior art is solved, and a fast and accurate chromosome euployness analysis is achieved.

CN120220797APending Publication Date: 2025-06-27SHANGHAI HORIZON MEDICAL SCI CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311821109.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-26
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art is difficult to improve speed and accuracy at the same time when detecting chromosomal euployness, especially the second-generation sequencing technology has problems with inaccurate detection results caused by long delivery time and GC preference.

Method used

By obtaining nanopore sequencing data and window partitioning information of the reference genome, the standard alignment sequence number of the detection samples in each window is determined, and by correcting the data of the negative samples, baseline values ​​are established to evaluate the chromosomal euplobilization of the samples to be analyzed.

Benefits of technology

Fast and accurate analysis of chromosomal euployness is achieved, reducing the noise of the detection results, improving the effective use of sequencing data, and supporting the completion of the entire analysis process within the reproductive center.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220797A_ABST
    Figure CN120220797A_ABST
Patent Text Reader

Abstract

The invention relates to a method and a device for detecting chromosome uploidy. The method comprises the following steps: acquiring nanopore sequencing data of a detection sample; obtaining window division information of the reference genome; for each detection sample, determining a standard comparison sequence number of the detection sample in each window based on the nanopore sequencing data of the detection sample and the window division information; for each negative sample, eliminating discrete values in the standard alignment sequence number of each negative sample in each window, and correcting the standard alignment sequence number of each negative sample in each window to obtain a standard alignment sequence number baseline value of each window; and according to the standard comparison sequence number of the to-be-analyzed sample in each window and the corresponding standard comparison sequence number baseline value of each window, evaluating the chromosome uploidy condition of the to-be-analyzed sample. According to the method, the third-generation sequencing data is creatively applied to chromosome aneuploidy analysis through a new baseline data construction mode.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of bioinformatics technology, and in particular, to a method and device for detecting chromosomal euploidy. Background Art

[0002] Aneuploid chromosomes are closely related to some human genetic diseases. The most common one is Down syndrome, with an incidence of about 1 / 800, which is caused by an extra copy of chromosome 21. Another example is trisomy 13 and trisomy 18 syndromes, which are caused by an extra copy of chromosome 13 and chromosome 18 respectively. Autosomal aneuploidy is a major cause of pregnancy failure and miscarriage.

[0003] Abnormalities in the number of sex chromosomes can cause abnormal sexual development. An individual with an extra X chromosome (47, XXY) in males has Klinefelter syndrome, a congenital testicular hypoplasia. In females, the absence of one X chromosome results in Turner syndrome, also known as congenital ovarian hypoplasia syndrome, with a karyotype of 45, X.

[0004] Traditional methods for detecting aneuploidy include fluorescence in situ hybridization (FISH), realtime-PCR, MLPA, biochips, etc. Biochips are divided into comparative genomic hybridization chips and SNP chips, which are the main means for aneuploidy detection. They have low throughput, can only detect a limited number of samples at a time, are costly, and the operation is relatively complex. FISH and realtime-PCR have been applied to more than 80% of aneuploidy detections. They are faster in detection speed, but limited by the number of probes in the method itself, they cannot comprehensively detect all 23 pairs of chromosomes simultaneously, and the throughput is very low.

[0005] Currently, the most commonly used method for aneuploidy detection is to use next-generation sequencing technology to perform low-depth whole-genome sequencing, and at the same time use normal euploid data to build a detection baseline to detect chromosomal microdeletions or microduplications. This method improves the detection sensitivity and can detect microdeletions or microduplications greater than 4 Mb. However, the delivery time of this method is relatively long, and it generally takes 24 hours to complete the delivery even under the most optimized conditions. Secondly, next-generation sequencing also requires bridge PCR amplification during the sequencing process, which will introduce GC bias, that is, regions with a GC content of about 50% on the genome will obtain more reads during sequencing, while in high-GC or low-GC regions, fewer reads are generated and are not easily detected, resulting in inaccurate detection results. Among them, the number of reads refers to the number of DNA fragment sequences read during the sequencing process, which is one of the important indicators for measuring the sequencing depth and is usually used to evaluate the coverage and data quality of the sample.

[0006] Therefore, how to improve the detection speed and accuracy simultaneously is a difficult point in chromosomal euploidy analysis. Summary of the Invention

[0007] To solve the above problems and improve the speed and accuracy of chromosomal euploidy analysis at the same time, the first object of the present application is to provide a method for detecting chromosomal euploidy, including:

[0008] Obtain nanopore sequencing data of a test sample, where the test sample includes a sample to be analyzed and a number of negative samples;

[0009] Obtain window partitioning information of a reference genome, where the window partitioning information includes the partitioning positions of each window and the GC content of each window, and the partitioning length of each window is selected from any value between 100 kb and 1000 kb;

[0010] For each test sample, determine the standard alignment sequence number of the test sample in each window based on the target nanopore sequencing data of the test sample and the window partitioning information;

[0011] For each negative sample, after removing the discrete values in the standard alignment sequence number of each negative sample in each window, correct the standard alignment sequence number of each negative sample in each window to obtain the baseline value of the standard alignment sequence number of each window;

[0012] Evaluate the chromosomal euploidy situation of the sample to be analyzed according to the standard alignment sequence number of the sample to be analyzed in each window and the corresponding baseline value of the standard alignment sequence number of each window.

[0013] Obtain nanopore sequencing data of a test sample, where the test sample includes a sample to be analyzed and a number of negative samples;

[0014] Obtain window partitioning information of a reference genome, where the window partitioning information includes the partitioning positions of each window and the GC content of each window, and the partitioning length of each window is selected from any value between 100 kb and 1000 kb;

[0015] For each test sample, determine the standard alignment sequence number of the test sample in each window based on the target nanopore sequencing data of the test sample and the window partitioning information;

[0016] For each negative sample, after removing the discrete values in the standard alignment sequence number of each negative sample in each window, correct the standard alignment sequence number of each negative sample in each window to obtain the baseline value of the standard alignment sequence number of each window;

[0017] Evaluate the chromosomal euploidy situation of the sample to be analyzed according to the standard alignment sequence number of the sample to be analyzed in each window and the corresponding baseline value of the standard alignment sequence number of each window

[0018] This application creatively uses third-generation sequencing data for chromosomal aneuploidy analysis through new GC correction means and a new baseline value construction method, and can perform chromosomal ploidy analysis on the third-generation sequencing data of single-cell amplification products.

[0019] In one embodiment, after removing the discrete values in the standard alignment sequence numbers of each negative sample in each window, the standard alignment sequence numbers of each negative sample in each window are corrected to obtain the baseline values of the standard alignment sequence numbers of each window, including:

[0020] For each window, according to the GC content of the window, remove the discrete values in the standard alignment sequence numbers of each negative sample in the window;

[0021] After removing the discrete values, select the median of the standard alignment sequence numbers of each negative sample in the window to obtain the baseline value of the standard alignment sequence numbers of the window.

[0022] In one embodiment, removing the discrete values in the standard alignment sequence numbers of each negative sample in the window according to the GC content of the window includes:

[0023] According to the GC content of the window and the standard alignment sequence numbers of each negative sample in the window, calculate the predicted standard alignment sequence numbers of each negative sample in the window;

[0024] According to the predicted standard alignment sequence numbers and the standard alignment sequence numbers of each negative sample in the window, determine the residuals of each negative sample in the window;

[0025] According to the residuals of each negative sample in the window, remove the discrete values in the standard alignment sequence numbers of each negative sample in the window.

[0026] In one embodiment of the baseline value, evaluating the chromosomal ploidy of the sample to be analyzed according to the standard alignment sequence numbers of the sample to be analyzed in each window and the corresponding baseline values of the standard alignment sequence numbers of each window includes:

[0027] According to the standard alignment sequence numbers of the sample to be analyzed in each window and the corresponding baseline values of the standard alignment sequence numbers of each window, calculate the logarithm values of the standard alignment sequence numbers of the sample to be analyzed in each window;

[0028] Evaluate the chromosomal ploidy of the sample to be analyzed according to the logarithm values of the standard alignment sequence numbers of the sample to be analyzed in each window.

[0029] In one embodiment, evaluating the chromosomal ploidy of the sample to be analyzed according to the logarithm values of the standard alignment sequence numbers of the sample to be analyzed in each window includes at least one of the following analyses (1) to (3):

[0030] (1) Obtain the CBS merging and splitting results for each window based on the logarithmic values of the standard alignment sequence numbers of the sample to be analyzed in each window, and determine the ploidy of the chromosomes of the sample to be analyzed according to the CBS merging and splitting results for each window;

[0031] (2) Calculate the chromosome fraction of the target chromosome based on the logarithmic values of the standard alignment sequence numbers of the sample to be analyzed in each window, and determine the aneuploidy of the target chromosome of the sample to be analyzed according to the chromosome fraction of the target chromosome;

[0032] (3) Denoise the logarithmic values of the standard alignment sequence numbers of the sample to be analyzed in each window, and visually characterize the chromosome ploidy of the sample to be analyzed according to the denoised logarithmic values of the standard alignment sequence numbers of the sample to be analyzed in each window.

[0033] In one implementation, determining the standard alignment sequence number of the detection sample in each window based on the nanopore sequencing data of the detection sample and the window division information includes:

[0034] Obtain the unique alignment sequence number of the detection sample in each window based on the nanopore sequencing data of the detection sample and the division positions of each window;

[0035] Correct the unique alignment sequence number of the detection sample in each window according to the GC content of each window to obtain the standard alignment sequence number of the detection sample in each window.

[0036] In one implementation, correcting the unique alignment sequence number of the detection sample in each window according to the GC content of each window to obtain the standard alignment sequence number of each detection sample in each window includes:

[0037] Perform GC correction on the unique alignment sequence number of the detection sample in each window according to the GC content of each window;

[0038] Use the scale function to perform sequencing depth correction on the uniquely aligned sequence numbers of the detection sample in each window after GC correction to obtain the standard alignment sequence numbers of the detection sample in each window.

[0039] In one implementation, for each detection sample, the method further includes the following quality control steps:

[0040] Obtain the first quality control parameter of the detection sample according to the nanopore sequencing alignment data of the detection sample, and use the first quality control parameter to determine whether the detection sample is qualified. The first quality control parameter includes at least one of Q10, Q12, Q15, Mean length 、Mean quality 、Median length 、Median quality and N50;

[0041] and / or

[0042] Obtain a second quality control parameter according to the number of unique alignment sequences and the number of standard alignment sequences of the detection sample in each window, and determine whether the detection sample is qualified according to the second quality control parameter. The second quality control parameter includes raw sd , cr sd , raw CV , cr CV , cov 20 , cov 50 and cov 80 and at least one of them.

[0043] The second object of the present application is to provide a device for detecting chromosomal aneuploidy, including:

[0044] Sequencing data acquisition module: used to acquire nanopore sequencing data of the detection sample, and the detection sample includes a sample to be analyzed and several negative samples;

[0045] Window division information acquisition module: used to acquire window division information of the reference genome, and the window division information includes the division positions of each window and the GC content of each window. The division length of each window is selected from any value between 100kb and 1000kb;

[0046] Standard alignment sequence number determination module: for each detection sample, determine the number of standard alignment sequences of the detection sample in each window based on the nanopore sequencing data of the detection sample and the window division information;

[0047] Baseline correction module: for each negative sample, after removing the discrete values in the number of standard alignment sequences of each negative sample in each window, correct the number of standard alignment sequences of each negative sample in each window to obtain the baseline value of the number of standard alignment sequences of each window;

[0048] Chromosomal aneuploidy evaluation module: used to evaluate the chromosomal aneuploidy situation of the sample to be analyzed according to the number of standard alignment sequences of the sample to be analyzed in each window and the corresponding baseline value of the number of standard alignment sequences of each window.

[0049] The present application also relates to a computer device, including a memory and a processor. The memory stores a computer program. The processor is characterized in that when the processor executes the computer program, the steps of the above method are implemented.

[0050] The present application also relates to a computer-readable storage medium, on which a computer program is stored. The computer program is characterized in that when the computer program is executed by the processor, the steps of the above method are implemented.

[0051] The present application also relates to a computer program product, including a computer program, characterized in that when the computer program is executed by a processor, the steps of the above method are implemented. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0053] Figure 1 It is a flowchart of a method for detecting chromosomal euploidy provided by an embodiment of the present application;

[0054] Figure 2 It is a graph showing the relationship between sequencing depth and GC content of sequencing sequences in second-generation sequencing and ONT sequencing provided by an embodiment of the present application;

[0055] Figure 3 It is a structural block diagram of a device for detecting chromosomal euploidy provided by an embodiment of the present application;

[0056] Figure 4 It is an internal structure diagram of a computer device provided by an embodiment of the present application;

[0057] Figure 5 It is a schematic flowchart of detecting chromosomal euploidy in Embodiment 1 of the present application;

[0058] Figure 6 It is the chromosomal euploidy analysis result of Sample 1 in Embodiment 2 of the present application;

[0059] Figure 7 It is the chromosomal euploidy analysis result of Sample 2 in Embodiment 2 of the present application;

[0060] Figure 8 It is the chromosomal euploidy analysis result of Sample 3 in Embodiment 2 of the present application;

[0061] Figure 9 It is the chromosomal euploidy analysis result of Sample 4 in Embodiment 2 of the present application;

[0062] Figure 10 It is the chromosomal euploidy analysis result of Sample 5 in Embodiment 2 of the present application;

[0063] Figure 11 It is the chromosomal euploidy analysis result of Sample 6 in Embodiment 2 of the present application;

[0064] Figure 12This is the result of chromosome ploidy analysis of sample 6 using traditional second-generation sequencing data for GC correction and sequencing depth correction in Example 3 of this application;

[0065] Figure 13 This is the result of chromosome ploidy analysis of sample 1 using traditional second-generation sequencing data for baseline construction in Example 4 of this application;

[0066] Figure 14 This is the result of chromosome ploidy analysis of the sample using a noise reduction algorithm in Example 5 of this application;

[0067] Figure 15 This is the result of chromosome ploidy analysis of the sample without using a noise reduction algorithm in Example 5 of this application. Detailed Description of the Invention

[0068] Reference will now be made in detail to embodiments of the present application, one or more examples of which are described below. Each example is provided by way of explanation and not limitation of the present application. In fact, it will be apparent to those skilled in the art that various modifications and variations can be made to the present application without departing from the scope or spirit of the present application. For example, features illustrated or described as part of one embodiment can be used in another embodiment to yield a still further embodiment.

[0069] Accordingly, it is intended that the present application cover such modifications and variations that fall within the scope of the appended claims and their equivalents. Other objects, features, and aspects of the present application are disclosed in or are apparent from the following detailed description. Those of ordinary skill in the art should understand that this discussion is only a description of exemplary embodiments and is not intended to limit the broader aspects of the present application.

[0070] As described above, in actual detection, changes in GC content and different sequencing depths can both affect the detection results. In addition, in the conventional detection process, often only one calculation model is used to calculate the chromosome state, which is likely to result in low detection accuracy of embryonic aneuploidy and other abnormal chromosomes.

[0071] Hereinafter, the "sample to be analyzed" can be an embryo sample before implantation in the uterus in assisted reproductive technology, or a fetal cell sample obtained by invasive methods during pregnancy, such as chorionic villus sampling, amniocentesis, and percutaneous umbilical blood sampling.

[0072] The "negative sample" refers to an embryo sample that has the same source as the sample to be analyzed and has a normal chromosome multiple, that is, a diploid embryo sample, which can also be called a euploid embryo sample.

[0073] Both the sample to be analyzed and the negative sample are derived from mammals. As used herein, the term "mammal" includes any diploid animal among human, tiger, wolf, mouse, deer, mink, monkey, tapir, sloth, zebra, dog, fox, bear, elephant, leopard, musk, cattle, lion, panda, wart hog, pig, antelope, reindeer, koala, rhinoceros, lynx, pangolin, giraffe, panda, anteater, orangutan, manatee, otter, civet, dolphin, walrus, platypus, hedgehog, arctic fox, polar bear, kangaroo, armadillo, hippopotamus, seal, whale, weasel, rabbit.

[0074] When the sample to be analyzed is human, the reference genome is the human reference genome. Specifically, it can be the human genome hg38.

[0075] The "test sample" refers to the collective term for the sample to be analyzed and the negative sample.

[0076] To at least partially solve the above technical problems, a first aspect of the present application provides a method for detecting chromosomal ploidy, as Figure 1 shown, including:

[0077] S10: Obtain nanopore sequencing data of the test sample, where the test sample includes the sample to be analyzed and a plurality of negative samples;

[0078] In the present application, nanopore sequencing data refers to sequencing data obtained based on the nanopore sequencing principle by the third-generation sequencing technology. The nanopore sequencing principle is as follows: In a container filled with electrolyte solution, a molecular membrane embedded with nanopore protein is placed, and only the nanopore on the membrane allows ions to pass through. A stable potential difference is provided on both sides of the nanopore by using an external power supply, so that a stable current passes through the nanopore. With the assistance of relevant proteins and enzymes, the DNA molecule passes through the nanopore in steps of single nucleotides. When a specific nucleotide occupies the nanopore, it will cause a perturbation to the current passing through the pore.

[0079] Specifically, the DNA molecule passes through the nanopore from the negative electrode to the positive electrode. Due to the different volumes and charged properties of different nucleotides, it causes changes in the current on both sides of the molecular membrane, thereby perturbing the stable current generated by the external power supply. When different bases on the DNA pass through, the perturbation signals are different. Based on this, DNA base sequencing can be achieved by recording the current changes and recognizing the current pattern. Finally, the sequencing instrument records the current signals generated during the DNA molecule passing through the pore, and then translates this specific electrical signal sequence into a nucleotide sequence using algorithm software.

[0080] It should be noted that obtaining the nanopore sequencing data of the test sample above does not mean detecting the sample to be analyzed and the negative sample simultaneously. For example, data acquisition and data processing can be first performed on multiple negative samples, and then data acquisition and data processing can be performed on the sample to be analyzed.

[0081] In some embodiments, to ensure the quality of sequencing data, after obtaining the nanopore sequencing data of the test sample, the following steps are further included:

[0082] Obtaining the first quality control parameter of the test sample based on the nanopore sequencing data of the test sample, and determining whether the test sample is qualified by using the first quality control parameter. The first quality control parameter includes at least one of Q10, Q12, Q15, Mean length 、Mean quality 、Median length 、Median quality and N50.

[0083] Specifically, after obtaining the bam file of the sequenced data, the nanopore supporting software GridION or MinIONMk1Csoftware or MinION Software for Mk1b can be used to directly process the downloaded bam file. Subsequently, nanoplot and / or NanoStat are used to count the sequencing profile to obtain the first quality control parameters Q10, Q12, Q15, Mean length 、Mean quality 、Median length 、Median quality and N50. Among them, the Q value represents the error rate, Q = 10 ╳ log 10 (P), where P is the probability of incorrect base calling. For example, Q10 refers to a base error probability of 1 / 10, or a precision of 90%; Q30 refers to a base error probability of 1 / 1000, or a precision of 99.9%. Mean length is the average length of sequencing, Mean quality is the average sequencing quality, Median length is the median sequencing length, Median quality is the median sequencing quality. N50 represents the assembly quality, and the larger the value, the better the assembly quality.

[0084] Furthermore, the thresholds of the first quality control parameters are specifically shown in Table 1 below.

[0085] Table 1

[0086] Q10 95.00% Q12 85.00% Q15 45.00% Mean_length 550 Mean_quality 13 Median_length 500 N50 600

[0087] S20: Obtaining the window division information of the reference genome, where the window division information includes the division positions of each window and the GC content of each window, and the division length of each window is selected from any value between 100 kb and 1000 kb;

[0088] Specifically, using the window partitioning tool of the bedtools software, the chromosomes can be partitioned into windows according to a preset length, thereby determining the partitioning position information of each window. Depending on the sequencing volume of the test sample, the selected window can be 100 kb, 200 kb, 500 kb, or 1000 kb. Using the nuc tool of the bedtools software, the GC content of each window of the reference genome can be counted, thereby generating the GC content information of the window. The partitioning position information and GC content information of each window can be obtained from the file with the reference genome partitioned into windows for subsequent chromosomal aneuploidy analysis. Among them, the GC content refers to the proportion of guanine (G) and cytosine (C) in the DNA sequence, and variation in the GC content may lead to bias in the sequencing results.

[0089] S30: For each test sample, determine the standard alignment sequence count of the test sample in each window based on the target nanopore sequencing data of the test sample and the window partitioning information;

[0090] Specifically, according to the window partitioning information, the unique alignment sequence count of the nanopore sequencing data of the test sample in each window can be counted.

[0091] In this application, the term "unique alignment sequence" refers to the read sequence that is uniquely aligned to the reference genome in the sequencing data, and the unique alignment sequence count refers to the number of reads that are uniquely aligned to the reference genome in the sequencing data.

[0092] To obtain the unique alignment sequence count of the nanopore sequencing data of the test sample in each window, after obtaining the nanopore sequencing data, the following steps are included: exclude the multi-aligned sequences in the nanopore sequencing data to obtain the uniquely aligned sequences in the sequencing data, sort and remove duplicates from all the uniquely aligned sequences, count the number of reads uniquely aligned to each window region, and generate the unique alignment sequence count of the nanopore sequencing data of the test sample in each window.

[0093] For example, after obtaining the bam file of the nanopore sequencing data, using the view function of the samtools software, in conjunction with the -e parameter, the original bam file with split mapped can be excluded to obtain the unique.bam file; use the sort and markdup tools of the sambamba software to sort and remove duplicates from the unique.bam file to obtain the unique.sort.rmdup.bam file; use the coverage tool of the bedtools software to count the number of reads of the unique.sort.rmdup.bam file in each window region to generate the unique alignment sequence count (raw reads )

[0094] In some embodiments, determining the standard alignment sequence count of a test sample in each window based on the nanopore sequencing data of the test sample and the window division information includes:

[0095] Obtaining the unique alignment sequence count of the test sample in each window based on the nanopore sequencing data of the test sample and the division positions of each window;

[0096] Correcting the unique alignment sequence count of the test sample in each window according to the GC content of each window to obtain the standard alignment sequence count of the test sample in each window.

[0097] Specifically, the standard alignment sequence count refers to the number of reads after GC content correction and sequencing depth correction of the unique alignment sequence count.

[0098] In some embodiments, in order to obtain the standard alignment sequence count of a test sample in each window, correcting the unique alignment sequence count of the test sample in each window to obtain the standard alignment sequence count of each test sample in each window includes:

[0099] Performing GC correction on the unique alignment sequence count of the test sample in each window according to the GC content of each window;

[0100] Performing sequencing depth correction on the uniquely aligned sequence counts of the test samples in each window after GC correction using the scale function to obtain the standard aligned sequence counts of the test samples in each window.

[0101] Specifically, use locally weighted scatterplot smoothing (lowess) to relate the GC content to the raw reads Establish a regression model, such as Figure 2 shown, Figure 2 where the abscissa represents the GC content and the ordinate represents the raw reads of each window. Use the regression model and the GC content of each window to output the predicted value of the unique alignment sequence count of each window. The predicted value of the unique alignment sequence count is represented by . Further, use Formula 1 to generate the uniquely aligned sequence count after GC correction for each window. The uniquely aligned sequence count after GC correction is represented by .

[0102] Formula 1:

[0103]

[0104] In some embodiments, for the single-cell amplification process involved in the PGT-A detection scenario, the above GC correction process can effectively reduce the GC bias caused by the amplification process.

[0105] Furthermore, since the sequencing amounts of different detection samples cannot be made exactly the same, the number of uniquely mapped sequences after GC correction for each window is normalized for sequencing depth using the scale function in Python and / or R. Specifically, calculate the average number of uniquely mapped sequences after GC correction for each window in the detection sample:

[0106] According to complete the sequencing depth correction of the number of uniquely mapped sequences after GC correction for each window to obtain the standard mapped sequence number cr for each window in the detection sample That is, according to finally generate the cr depth file. depth depth file.

[0107] In some embodiments, to ensure the accuracy of the detection results, after performing GC content correction and sequencing depth correction on the number of uniquely mapped sequences for each window of the detection sample to obtain the standard mapped sequence number for each window of the detection sample, it includes:

[0108] Obtain a second quality control parameter based on the number of uniquely mapped sequences and the standard mapped sequence number for each window of the detection sample, and determine whether the detection sample is qualified according to the second quality control parameter. The second quality control parameter includes at least one of raw sd 、cr sd 、raw CV 、cr CV 、cov 20 、cov 50 and cov 80 in it.

[0109] Among them, raw sd is the standard deviation of the raw reads values for each window. cr sd is the standard deviation of the cr depth for each window. raw CV is raw sd divided by the average value of the raw reads for each window. cr CV is cr sd divided by the average value of the cr depth for each window. cov 20 、cov 50 and cov 80 are respectively the proportions of the cr depth in the sample for each window that can reach 0.2, 0.5, and 0.8 times the average value of the cr depth for each window. For example, cov 20 represents the crdepth The cr of each window can be achieved depth The proportion of the number of windows whose average is multiplied by 0.2 in the total number of windows in the sample.

[0110] Among them, the threshold of the second quality control parameter is specifically shown in Table 2.

[0111] Table 2

[0112] Quality control parameter Threshold <![CDATA[raw sd > 40 <![CDATA[cr sd > 35 <![CDATA[raw CV > 1.9 <![CDATA[cr CV > 1.6 <![CDATA[cov 20 > 90 <![CDATA[cov 50 > 85 <![CDATA[cov 80 > 50

[0113] S40: For each negative sample, after removing the discrete values in the standard alignment sequence numbers of each negative sample at each window, correct the standard alignment sequence numbers of each negative sample at each window to obtain the baseline value of the standard alignment sequence numbers at each window;

[0114] Specifically, the baseline value of the standard alignment sequence number of each window refers to the median determined after removing the discrete values from the standard alignment sequence numbers of each negative sample at this window. Those skilled in the art can understand that the median of the standard alignment sequence numbers of each negative sample at each window after removing the discrete values is the baseline value of the standard alignment sequence number of each window.

[0115] In some embodiments, in order to obtain the baseline value of the standard alignment sequence numbers at each window, after removing the discrete values in the standard alignment sequence numbers of each negative sample at each window, correcting the standard alignment sequence numbers of each negative sample at each window to obtain the baseline value of the standard alignment sequence numbers at each window includes:

[0116] For each window, remove the discrete values in the standard alignment sequence numbers of each negative sample at this window according to the GC content of this window, and select the median of the standard alignment sequence numbers of each negative sample at this window after removing the discrete values to obtain the baseline value of the standard alignment sequence numbers at this window.

[0117] Specifically, remove the discrete values from the standard alignment sequence numbers of each negative sample at this window, so as to reduce the influence of noise on the baseline correction result and facilitate more accurate analysis of chromosomal aneuploidy in the subsequent process.

[0118] In some embodiments, in order to remove the discrete values in the normalized alignment numbers of each negative sample at each window, predicting the standard alignment sequence numbers by removing the discrete values in the standard alignment sequence numbers of each negative sample at this window according to the GC content of this window includes:

[0119] Calculate the predicted standard alignment sequence number of each negative sample at this window according to the GC content of this window and the standard alignment sequence number of each negative sample at this window;

[0120] Determine the residuals of each negative sample at this window according to the predicted standard alignment sequence numbers and the standard alignment sequence numbers of each negative sample at this window;

[0121] Remove the discrete values in the standard alignment sequence numbers of each negative sample in this window based on the residuals of each negative sample in this window.

[0122] Specifically, for negative samples, the predicted standard alignment sequence numbers for each window are the results calculated by a regression model constructed using the linear relationship fitted based on the GC content of each window and the standard alignment sequence numbers of each negative sample in each window.

[0123] The specific steps for the baseline value to correct the standard alignment sequence numbers of each negative sample are as follows: 1) Use the cr depth file that contains the standard alignment sequence numbers of negative samples after GC correction and sequencing depth correction, and summarize the standard alignment sequence numbers cr of each negative sample in each window depth as well as the GC content of each window, and use lowess to fit the linear relationship between the cr of all negative samples in each window depth and the GC content of each window, establish a regression model, and use the GC content of each window according to this model to output the predicted standard alignment sequence numbers predict of each negative sample in each window depth ; 2) Use the standard alignment sequence number cr of each negative sample in each window depth to subtract the predicted standard alignment sequence number predict of this negative sample in this window depth to obtain the residual of this negative sample in this window: residual; 3) Calculate the standard deviation sd of the residuals of all negative samples in each window residual , and remove the standard alignment sequence numbers corresponding to the negative samples with residuals exceeding 3 times sd in each window residual ; 4) After performing removal correction on the standard alignment sequence numbers of each window, use the median of the corrected standard alignment sequence numbers of this window as the baseline value of the standard alignment sequence number of this window, denoted by ref median for calculating the log value.

[0124] S50: Evaluate the chromosomal ploidy of the sample to be analyzed according to the standard alignment sequence numbers of the sample to be analyzed in each window and the corresponding baseline values of the standard alignment sequence numbers of each window.

[0125] Specifically, the chromosomal ploidy includes chromosomal euploidy and chromosomal aneuploidy. For diploid organisms, chromosomal aneuploidy includes monosomy, trisomy, tetrasomy, etc., which are used to represent chromosomal abnormalities.

[0126] In some embodiments, evaluating the chromosomal ploidy of the sample to be analyzed according to the baseline values of the standard alignment sequence numbers of the sample to be analyzed in each window and the corresponding baseline values of the standard alignment sequence numbers of each window includes:

[0127] Calculate the logarithmic value of the standard alignment sequence number for each window of the sample to be analyzed based on the standard alignment sequence number for each window of the sample to be analyzed and the corresponding baseline value of the standard alignment sequence number for each window.

[0128] Evaluate the chromosomal euploidy of the sample to be analyzed based on the logarithmic value of the standard alignment sequence number for each window of the sample to be analyzed.

[0129] Specifically, for each sample to be analyzed, use the standard alignment sequence number cr depth after GC correction and sequencing depth correction for each window median divided by the baseline value ref of the standard alignment sequence number for that window, and calculate the logarithmic value of the standard alignment sequence number for each window through Formula 2, that is, the log value.

[0130] Formula 2:

[0131]

[0132] In some embodiments, since there will be some individual discrete points after calculating the logarithmic value of the standard alignment sequence number, the interquartile range (IQR) is used to remove the discrete logarithmic values.

[0133] In some embodiments, evaluating the chromosomal euploidy of the sample to be analyzed based on the logarithmic value of the standard alignment sequence number for each window of the sample to be analyzed includes:

[0134] Adopt the CBS algorithm to merge and segment the logarithmic values of the standard alignment sequence numbers for each window of the sample to be analyzed, and determine the euploidy of the chromosomes of the sample to be analyzed according to the merge and segmentation results.

[0135] Specifically, the average value (log mean ) of the logarithmic values of the standard alignment sequence numbers for each window is used as a threshold to judge the euploidy of local chromosomes after window merging and segmentation. The threshold and the judgment results are shown in Table 3 specifically.

[0136] Table 3

[0137] <![CDATA[log mean numerical value]]> Microdeletion / duplication status <![CDATA[-1<log mean > Monosomy <![CDATA[-1 < log mean <0.58]]> Normal <![CDATA[0.58 < log mean <1]]> Trisomy <![CDATA[log mean >1]]> Tetrasomy

[0138] In some embodiments, evaluating the chromosomal euploidy of the sample to be analyzed based on the logarithmic value of the standard alignment sequence number for each window of the sample to be analyzed includes:

[0139] Calculate the chromosomal fraction of the target chromosome based on the logarithmic value of the standard alignment sequence number for each window of the sample to be analyzed, and determine the euploidy of the target chromosome of the sample to be analyzed according to the chromosomal fraction of the target chromosome.

[0140] Specifically, the chromosomal fraction Z scoreTo evaluate the ploidy of the entire chromosome, the calculation process is as follows:

[0141] For the target chromosome, calculate the average log value log mean and the standard deviation log sd of the log values of each window of the target chromosome in the sample to be detected, and calculate the median of the log values of each window of the target chromosome Obtain Z using Formula 3 score :

[0142] Formula 3:

[0143]

[0144] where

[0145] Specifically, the average log mean is the arithmetic mean of the log values of each window of the target chromosome in the sample to be detected, the standard deviation log sd is the standard deviation of the log values of each window of the target chromosome in the sample to be detected, n is the number of windows, and log represents the log value of the window.

[0146] In some specific embodiments, the chromosomal ploidy status corresponding to Z score is shown in Table 4.

[0147] Table 4

[0148] Z_score Chromosome status -3 < Z_score < -2 Deletion -2 < Z_score < -1 Monosomy -1 < Z_score < 1 Normal 1 < Z_score < 2 Trisomy 2 < Z_score < 3 Tetrasomy

[0149] In some embodiments, evaluating the chromosomal ploidy of the sample to be analyzed based on the logarithmic values of the standard alignment sequence numbers in each window of the sample to be analyzed includes:

[0150] Denoise the logarithmic values of the standard alignment sequence numbers in each window of the sample to be analyzed, and visually characterize the chromosomal ploidy of the sample to be analyzed according to the denoised logarithmic values of the standard alignment sequence numbers in each window of the sample to be analyzed.

[0151] Specifically, using the savgol denoising algorithm to denoise the logarithmic values of the standard alignment sequence numbers can optimize the visualization effect of the log value plot, make the scatter points closer, and thus make it easier to identify the chromosomal ploidy situation.

[0152] It should be noted that the traditional method for analyzing chromosomal aneuploidy based on second-generation sequencing data cannot achieve fresh embryo transfer even under the most optimized process. Currently, the throughput of third-generation sequencing is still relatively low. At the same time, in order to achieve fresh embryo transfer, the sequencing time cannot be directly increased. At the same time, in order to balance the cost, individual libraries cannot be constructed and sequenced separately for individual samples.

[0153] In addition, third-generation sequencing directly sequences DNA fragments of the original length. Under the same sequencing volume, the obtained reads data will be lower than that of second-generation sequencing, resulting in a smaller number of reads counted for each window. Compared with about 100 reads that can be obtained for each window in second-generation sequencing, only about 20 reads can be obtained in third-generation sequencing. When 5 fewer reads are counted in a window, it has a 5% impact on second-generation sequencing, while it will have a 25% impact on third-generation sequencing. Therefore, the results obtained by third-generation sequencing will have more noise compared with second-generation sequencing.

[0154] Through a new GC correction method and a new baseline data construction method, the present application creatively uses third-generation sequencing data for chromosomal aneuploidy analysis, and can realize chromosomal ploidy analysis of third-generation sequencing data of single-cell amplification products, and has the following beneficial effects:

[0155] 1. In cooperation with the rapid library construction process, for 5-7 embryonic cells obtained within an assisted reproductive cycle, through single-cell amplification for 1.5 h, library construction for 30 min, and sequencing for 1.5 h, the PGT-A analysis can be completed within 1-2 minutes after the sequencing ends, and the detection and delivery can be completed within 4 h, thus providing the possibility for fresh embryo transfer;

[0156] 2. Based on the nanopore sequencer, the present application is extremely small in volume and low in cost. Sequencing can be completed by cooperating with a laptop computer with a GPU, and the whole set of analysis processes can be completed in the reproductive center without outsourcing sequencing analysis, truly realizing zero-loss fresh embryo transfer;

[0157] 3. Based on the nanopore sequencing technology, there is no GC bias in the sequencing process, and the sequencing can be carried out evenly and randomly;

[0158] 4. The present application can perform complete sequencing on the original fragments after single-cell amplification (average about 500 bp, median about 400 bp, N50 about 700 bp), effectively improving the success rate, accuracy rate, and the proportion of uniquely mapped sequences of subsequent alignment to the genome, thereby effectively improving the effective utilization rate of sequencing reads;

[0159] 5. The present application effectively reduces the interference of data fluctuations of individual baseline samples and / or individual windows on the overall baseline data, and improves the accuracy of euploid detection.

[0160] The second aspect of the present application lies in providing a device for detecting chromosomal aneuploidy, as Figure 3 shown, including:

[0161] A sequencing data acquisition module: used to acquire nanopore sequencing data of a test sample, and the test sample includes a sample to be analyzed and several negative samples;

[0162] Window division information acquisition module: used to acquire the window division information of the reference genome, where the window division information includes the division positions of each window and the GC content of each window, and the division length of each window is selected from any value between 100 kb and 1000 kb;

[0163] Standard alignment sequence number determination module: for each test sample, based on the nanopore sequencing data of the test sample and the window division information, determine the standard alignment sequence number of the test sample in each window;

[0164] Baseline correction module: for each negative sample, after removing the discrete values in the standard alignment sequence number of each negative sample in each window, correct the standard alignment sequence number of each negative sample in each window to obtain the baseline value of the standard alignment sequence number of each window;

[0165] Chromosome ploidy evaluation module: used to evaluate the chromosome ploidy of the sample to be analyzed according to the standard alignment sequence number of the sample to be analyzed in each window and the corresponding baseline value of the standard alignment sequence number of each window.

[0166] In some embodiments, the baseline correction module includes:

[0167] Discrete value removal unit: for each window, remove the discrete values in the standard alignment sequence number of each negative sample in this window according to the GC content of this window;

[0168] Baseline value selection unit: used to select the median of the standard alignment sequence number of each negative sample in this window after removing the discrete values to obtain the baseline value of the standard alignment sequence number of this window.

[0169] In some embodiments, the discrete value removal unit includes: predicted standard alignment sequence number calculation sub-unit; used to calculate the predicted standard alignment sequence number of each negative sample in this window according to the GC content of this window and the standard alignment sequence number of each negative sample in this window;

[0170] Residual calculation sub-unit: used to determine the residual of each negative sample in this window according to the predicted standard alignment sequence number and the standard alignment sequence number of each negative sample in this window;

[0171] Discrete value removal sub-unit: used to remove the discrete values in the standard alignment sequence number of each negative sample in this window according to the residual of each negative sample in this window.

[0172] In some embodiments, the chromosome ploidy evaluation module includes:

[0173] Logarithmic value calculation unit: configured to calculate the logarithmic value of the number of standard alignment sequences of a sample to be analyzed in each window according to the number of standard alignment sequences of the sample to be analyzed in each window and the baseline value of the number of standard alignment sequences of each corresponding window;

[0174] Chromosome euploidy evaluation unit: configured to evaluate the chromosome euploidy of a sample to be analyzed according to the logarithmic value of the number of standard alignment sequences of the sample to be analyzed in each window.

[0175] In some embodiments, the chromosome euploidy evaluation unit includes:

[0176] First chromosome euploidy evaluation subunit: configured to obtain the CBS combined segmentation result of each window according to the logarithmic value of the number of standard alignment sequences of the sample to be analyzed in each window, and determine the euploidy of the chromosomes of the sample to be analyzed according to the CBS combined segmentation result of each window;

[0177] Second chromosome euploidy evaluation subunit: configured to calculate the chromosome fraction of a target chromosome according to the logarithmic value of the number of standard alignment sequences of the sample to be analyzed in each window, and determine the euploidy of the target chromosome of the sample to be analyzed according to the chromosome fraction of the target chromosome;

[0178] Third chromosome euploidy evaluation subunit: configured to denoise the logarithmic value of the number of standard alignment sequences of the sample to be analyzed in each window, and visually characterize the chromosome euploidy of the sample to be analyzed according to the denoised logarithmic value of the number of standard alignment sequences of the sample to be analyzed in each window.

[0179] In some embodiments, the standard alignment sequence number determination module includes:

[0180] Unique alignment sequence number acquisition unit: configured to obtain the number of unique alignment sequences of a detection sample in each window based on the nanopore sequencing data of the detection sample and the division positions of each window;

[0181] Standard alignment sequence number determination unit: configured to correct the number of unique alignment sequences of the detection sample in each window according to the GC content of each window to obtain the number of standard alignment sequences of the detection sample in each window.

[0182] In some embodiments, the standard alignment sequence number determination unit includes:

[0183] GC correction subunit: configured to perform GC correction on the number of unique alignment sequences of the detection sample in each window according to the GC content of each window;

[0184] Sequencing depth correction subunit: configured to perform sequencing depth correction on the number of unique alignment sequences of the detection sample in each window after GC correction by using the scale function to obtain the number of standard alignment sequences of the detection sample in each window.

[0185] In some embodiments, the device further includes a first quality control module:

[0186] configured to obtain first quality control parameters of the test sample according to the nanopore sequencing alignment data of the test sample, and determine whether the test sample is qualified by using the first quality control parameters, where the first quality control parameters include at least one of Q10, Q12, Q15, Mean length 、Mean quality 、Median length 、Median quality and N50.

[0187] In some embodiments, the device further includes a second quality control module:

[0188] configured to, for each test sample, obtain second quality control parameters according to the number of unique alignment sequences and the standard alignment sequences of the test sample in each window, and determine whether the test sample is qualified according to the second quality control parameters, where the second quality control parameters include at least one of raw sd 、cr sd 、raw CV 、cr CV 、cov 20 、cov 50 and cov 80 .

[0189] For the specific limitations of the device for detecting chromosomal aneuploidy, reference can be made to the limitations of the method for detecting chromosomal aneuploidy in the foregoing text, which will not be elaborated here. Each module in the above device for detecting chromosomal aneuploidy can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in the form of hardware or be independent of it, or be stored in the memory of the computer device in the form of software, so as to facilitate the processor to call and execute the operations corresponding to the above respective modules.

[0190] In some embodiments, a computer device is provided. The computer device can be the server 104 or the terminal 102, and its internal structure diagram can be as Figure 4As shown in the figure. The computer device includes a processor, a memory, and a communication interface connected through a system bus. When the computer device is a terminal, it further includes a display screen and an input device connected to the system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a method for detecting chromosomal aneuploidy. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, trackball, or touchpad provided on the computer device housing, or an external keyboard, touchpad, or mouse, etc.

[0191] Those skilled in the art can understand that Figure 4 the structure shown in the figure is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0192] The present application also provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, it implements the steps of the above method for detecting chromosomal aneuploidy.

[0193] The present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program. When the computer program is executed by the processor, it implements the steps of the above method for detecting chromosomal aneuploidy.

[0194] The present application also provides a computer program product. The computer program product includes a computer program. When the computer program is executed by the processor, it implements the steps of the above method for detecting chromosomal aneuploidy.

[0195] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties.

[0196] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the various embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the various embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the various embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.

[0197] The implementation solutions of the present application will be described in detail below in conjunction with the embodiments.

[0198] Example 1

[0199] In this example, 310 samples were collected and third-generation sequencing was performed based on the ONT platform. According to the Figure 5 process shown, the method for detecting chromosomal aneuploidy of the present application was used to perform chromosomal aneuploidy analysis on the samples. It was found that there were 45 cases with microdeletions / duplications greater than 4M in the chromosomes, 87 cases of various trisomy cases, 73 cases of various monosomy cases, and 105 cases of normal euploid embryos. After comparing with the real karyotype analysis of the samples, the accuracy was 100%, and the false positive rate was 0%.

[0200] Among them, sample 1 is the cell of a normal female with a karyotype of 46,XX, and the test results are as Figure 6 shown.

[0201] Sample 2 is the cell of a normal male with a karyotype of 46,XY, and the test results are as Figure 7 shown.

[0202] Sample 3 is the cell of a patient with Klinefelter syndrome with a karyotype of 47,XXY, and the test results are as Figure 8 shown;

[0203] The Zscore score of sample 3 is shown in Table 5. ChrX is within the threshold range of chrX of normal females, but there is also a value of chrY at the same time, so 47, XXY is judged.

[0204] Table 5

[0205]

[0206]

[0207] Sample 4 is a female patient with trisomy 22 of chr22 with a karyotype of 47,XX,+22, and the test results are as Figure 9 shown;

[0208] The Zscore scores of different chromosomes of sample 4 are shown in Table 6. Chr22 is higher than the threshold of normal chr22.

[0209] Table 6

[0210] chr Z_score chr Z_score chr Z_score chr1 0.124 chr11 0.007 chr21 0.044 chr2 -0.061 chr12 0.097 chr22 1.386 chr3 0.096 chr13 -0.026 chrX 0.166 chr4 -0.173 chr14 -0.033 chrY chr5 -0.078 chr15 -0.015 chr6 -0.023 chr16 -0.289 chr7 0.156 chr17 -0.108 chr8 -0.142 chr18 0.148 chr9 0.035 chr19 0.027 chr10 -0.046 chr20 -0.137

[0211] Sample 5 is a female patient with monosomy 8 of chr8 with a karyotype of 45,XX,-8, and the test results are as Figure 10 shown;

[0212] The Zscore scores of different chromosomes of sample 5 are shown in Table 7. Chr8 is lower than the threshold of normal chr8.

[0213] Table 7

[0214] chr Z_score chr Z_score chr Z_score chr1 0.004 chr11 0.099 chr21 0.272 chr2 0.104 chr12 -0.022 chr22 0.129 chr3 -0.048 chr13 -0.061 chrX 0.368 chr4 0.164 chr14 0.098 chrY chr5 0.081 chr15 0.198 chr6 0.343 chr16 0.206 chr7 0.191 chr17 0.256 chr8 -1.799 chr18 0.352 chr9 0.068 chr19 0.233 chr10 0.203 chr20 0.086

[0215] Sample 6 is a female patient with a partial deletion of chr2 of 46,XX,del(2)(q35-q37.3)(24.91Mb), and the test results are as Figure 11 shown, and the CBS segmentation results of chr2 are shown in Table 8.

[0216] Table 8

[0217]

[0218]

[0219] Example 2

[0220] In this example, a sample 6 with a chromosomal karyotype of 46,XX,del(2)(q35-q37.3)(24.91Mb) was selected for third-generation sequencing based on the ONT platform. According to the Figure 5 procedure in Example 1, third-generation sequencing data was used to perform chromosomal ploidy analysis on sample 6. At the same time, as a control, sample 6 with a chromosomal karyotype of 46,XX,del(2)(q35-q37.3)(24.91Mb) was selected for NGS sequencing. According to the Figure 5 procedure in Example 1, traditional second-generation sequencing data was used to replace the Figure 5 third-generation sequencing data in the procedure to perform chromosomal ploidy analysis on sample 6.

[0221] The analysis results of using third-generation sequencing data to perform chromosomal ploidy analysis on sample 6 are as Figure 11 shown. The analysis results of using traditional second-generation sequencing data to perform chromosomal ploidy analysis on sample 6 are as Figure 12 shown. According to Figure 11 and Figure 12 , it can be seen that when using second-generation sequencing data for chromosomal aneuploidy analysis, the calculation results of the log values of each window will be overall on the high side.

[0222] Example 3

[0223] In this example, a normal female sample 1 with a chromosomal karyotype of 46,XX was selected for third-generation sequencing based on the ONT platform. According to the Figure 5 procedure in Example 1, third-generation sequencing data was used to perform chromosomal ploidy analysis on sample 1. At the same time, as a control, a normal female sample 1 with a chromosomal karyotype of 46,XX was selected for NGS sequencing. According to the Figure 5 procedure in Example 1, traditional second-generation sequencing data was used to replace the Figure 5 third-generation sequencing data in the procedure to perform chromosomal ploidy analysis on sample 1.

[0224] The analysis results of using third-generation sequencing data to perform chromosomal ploidy analysis on sample 1 are as Figure 6 shown. The analysis results obtained by using traditional second-generation sequencing data for chromosomal ploidy analysis are as Figure 13 shown. According to the comparison between Figure 13 and Figure 6 , it can be seen that when using second-generation sequencing data for chromosomal ploidy analysis, the scatter plot obtained after visualizing the log values is very discrete.

[0225] Example 4

[0226] Samples from a female patient with monosomy of chr12 with a karyotype of 45,xx,-12 were selected for third-generation sequencing based on the ONT platform, and according to the procedure in Example 1 Figure 5 The samples were analyzed for chromosomal aneuploidy by comparing the addition and non-addition of a noise reduction algorithm.

[0227] The results of chromosomal euploidy analysis with the addition of the noise reduction algorithm are as Figure 14 shown; the results of chromosomal euploidy analysis without the addition of the noise reduction algorithm are as Figure 15 shown. According to Figure 14 it can be seen that compared with Figure 15 the results of chromosomal euploidy analysis without the addition of the noise reduction algorithm in , the log value data obtained using the noise reduction algorithm are more closely concentrated, so that the scatter points of normal chromosomes are more significant than those of abnormal chromosomes and are easier to distinguish.

[0228] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0229] The above embodiments only represent several implementation manners of the present application, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.

Claims

1. A method for detecting chromosomal euploidy, characterized in that, Including: Obtaining nanopore sequencing data of a detection sample, where the detection sample includes a sample to be analyzed and a number of negative samples; Obtaining window partitioning information of a reference genome, where the window partitioning information includes the partitioning positions of each window and the GC content of each window, and the partitioning length of each window is selected from any value between 100 kb and 1000 kb; For each detection sample, determining the standard alignment sequence count of the detection sample in each window based on the nanopore sequencing data of the detection sample and the window partitioning information; For each negative sample, after removing discrete values in the standard alignment sequence count of each negative sample in each window, correcting the standard alignment sequence count of each negative sample in each window to obtain the baseline value of the standard alignment sequence count for each window; Evaluating the chromosomal ploidy of the sample to be analyzed according to the standard alignment sequence count of the sample to be analyzed in each window and the corresponding baseline value of the standard alignment sequence count for each window.

2. The method according to claim 1, wherein "After removing discrete values in the standard alignment sequence count of each negative sample in each window, correcting the standard alignment sequence count of each negative sample in each window to obtain the baseline value of the standard alignment sequence count for each window" includes the following steps: For each window, removing discrete values in the standard alignment sequence count of each negative sample in this window according to the GC content of this window; After removing discrete values, selecting the median of the standard alignment sequence count of each negative sample in this window to obtain the baseline value of the standard alignment sequence count for this window.

3. The method according to claim 2, wherein "Removing discrete values in the standard alignment sequence count of each negative sample in this window according to the GC content of this window" includes the following steps: Calculating the predicted standard alignment sequence count of each negative sample in this window according to the GC content of this window and the standard alignment sequence count of each negative sample in this window; Determining the residual of each negative sample in this window according to the predicted standard alignment sequence count and the standard alignment sequence count of each negative sample in this window; Removing discrete values in the standard alignment sequence count of each negative sample in this window according to the residual of each negative sample in this window.

4. The method according to any one of claims 1 to 3, characterized in that The "evaluating the chromosomal ploidy of the sample to be analyzed according to the standard alignment sequence count of the sample to be analyzed in each window and the corresponding baseline value of the standard alignment sequence count for each window" includes the following steps: Calculating the logarithm value of the standard alignment sequence count of the sample to be analyzed in each window according to the standard alignment sequence count of the sample to be analyzed in each window and the corresponding baseline value of the standard alignment sequence count for each window; Evaluating the chromosomal ploidy of the sample to be analyzed according to the logarithm values of the standard alignment sequence count of the sample to be analyzed in each window.

5. The method according to claim 4, wherein The "evaluating the chromosomal ploidy of the sample to be analyzed according to the logarithm values of the standard alignment sequence count of the sample to be analyzed in each window" is at least one of the following methods: (1) Obtaining the CBS combined segmentation result of each window according to the logarithm values of the standard alignment sequence count of the sample to be analyzed in each window, and determining the ploidy of the chromosomes of the sample to be analyzed according to the CBS combined segmentation result of each window; (2) Calculate the chromosomal fraction of the target chromosome based on the numerical values of the standard alignment sequence counts of the sample to be analyzed in each window, and determine the ploidy of the target chromosome of the sample to be analyzed according to the chromosomal fraction of the target chromosome; (3) Denoise the numerical values of the standard alignment sequence counts of the sample to be analyzed in each window, and visually characterize the chromosomal ploidy of the sample to be analyzed according to the denoised numerical values of the standard alignment sequence counts of the sample to be analyzed in each window.

6. The method according to claim 1, characterized in that "Determining the standard alignment sequence count of the test sample in each window based on the nanopore sequencing data of the test sample and the window division information" includes the following steps: Obtain the unique alignment sequence count of the test sample in each window based on the nanopore sequencing data of the test sample and the division positions of each window; Correct the unique alignment sequence count of the test sample in each window according to the GC content of each window to obtain the standard alignment sequence count of the test sample in each window.

7. The method according to claim 6, wherein "Correcting the unique alignment sequence count of the test sample in each window according to the GC content of each window to obtain the standard alignment sequence count of each test sample in each window" includes: Perform GC correction on the unique alignment sequence count of the test sample in each window according to the GC content of each window; Use the scale function to perform sequencing depth correction on the unique alignment sequence count of the test sample in each window after GC correction to obtain the standard alignment sequence count of the test sample in each window.

8. The method according to claim 1, wherein For each test sample, the method further includes the following quality control steps: Obtain the first quality control parameter of the test sample according to the nanopore sequencing alignment data of the test sample, and determine whether the test sample is qualified by using the first quality control parameter. The first quality control parameter includes at least one of Q10, Q12, Q15, Mean length , Mean quality , Median length , Median quality and N50; and / or Obtain a second quality control parameter based on the number of unique alignment sequences and the number of standard alignment sequences of the detection sample in each window, and determine whether the detection sample is qualified according to the second quality control parameter. The second quality control parameter includes raw sd , cr sd , raw CV , cr CV , cov 20 , cov 50 and cov 80 at least one of them.

9. An apparatus for detecting chromosomal euploidy, characterized in that Including: Sequencing data acquisition module: used to acquire the nanopore sequencing data of the test sample, and the test sample includes the sample to be analyzed and several negative samples; Window division information acquisition module: used to acquire the window division information of the reference genome, and the window division information includes the division positions of each window and the GC content of each window, and the division length of each window is selected from any value between 100 kb and 1000 kb; Standard alignment sequence count determination module: used to determine the standard alignment sequence count of the test sample in each window for each test sample based on the nanopore sequencing data of the test sample and the window division information; Baseline correction module: used to remove the discrete values in the standard alignment sequence count of each negative sample in each window for each negative sample, and correct the standard alignment sequence count of each negative sample in each window after removing the discrete values to obtain the baseline value of the standard alignment sequence count of each window; Chromosomal ploidy evaluation module: used to evaluate the chromosomal ploidy of the sample to be analyzed according to the standard alignment sequence count of the sample to be analyzed in each window and the corresponding baseline value of the standard alignment sequence count of each window.

10. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method described in any one of claims 1 to 8.

Citation Information

Cited By

  • System, device and medium for analyzing chromosomal abnormalities based on nanopore sequencing

    CN121148472A