Analysis method and application of chromosome aneuploid
Patent Information
- Application Number
- CN202480049407.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-22
- Publication Date
- 2026-02-24
AI Technical Summary
Existing non-invasive prenatal genetic testing methods face challenges in eliminating batch-to-batch variability and window feature dispersion in the detection of chromosomal aneuploidy, such as large result fluctuations, high sample quantity requirements, and the curse of dimensionality, resulting in insufficient accuracy and stability.
The method of selecting a reference set to calculate the corrected baseline is adopted. The sample with the lowest difference is selected from the samples with known fetal chromosome information and maternal peripheral blood cfDNA sequencing data as the reference set. The PCA model is used for dimensionality reduction, window depth correction is calculated, and GC correction and Z-score judgment are combined to improve the stability and accuracy of detection.
In the absence of a sample reference from the same batch, it significantly improved the accuracy of single-sample NIPT detection, reduced batch fluctuations, enhanced the differentiation between true positive and false positive samples, and improved the stability and accuracy of chromosomal aneuploidy abnormality detection.
Smart Images

Figure CN121569342A_ABST
Abstract
Description
A method for analyzing chromosomal aneuploidy and its application Technical Field
[0001] This application relates to the field of chromosomal aneuploidy detection technology, and in particular to an analytical method and application for chromosomal aneuploidy. Background Technology
[0002] Chromosomal aneuploidy refers to abnormal changes in the ploidy of chromosomes in cells. Common chromosomal aneuploidy disorders include Down syndrome (trisomy 21), Edwards syndrome (trisomy 18), and Patau syndrome (trisomy 13), all of which lead to severe developmental abnormalities in children. Research has found that cell-free fetal DNA exists in maternal peripheral blood, making it an important material for non-invasive prenatal genetic testing (NIPT). Quantitative detection of cfDNA in maternal peripheral blood can detect whether the fetus has chromosomal structural or numerical abnormalities.
[0003] Existing methods can detect aneuploidy in fetal chromosomes by performing whole-genome sequencing on cfDNA and determine whether the chromosome copy number is abnormal by statistically analyzing the sequencing depth. However, due to differences in sequencing depth across different chromosomal regions, methods are needed to eliminate biases introduced by baseline variations in different regions. Furthermore, the baseline of sequencing data from different batches of samples varies significantly due to differences in reagents and experimental procedures. Therefore, it is crucial to eliminate both baseline biases introduced by different regions and baseline biases between batches to obtain more stable results.
[0004] To address the above biases, some methods employ thresholding for the proportion of URs on the chromosome to be tested, eliminating the influence of different sequencing preferences in different segments; some use information from a fixed additional reference sample set for data correction; some use samples from the same batch as a reference, calculating the chromosome baseline for various GC content ranges within that batch using the reference set, and then performing data correction for each sample in that batch; some use intra-batch correction methods; and some use the Manhattan distance of the window features between the test samples in the entire batch and the samples in the reference database to screen a dynamic reference database for that batch of samples, using a weighted linear fitting method to correct the chromosome baseline data.
[0005] However, existing deviation correction methods still have the following defects and shortcomings:
[0006] (1) Due to the large differences in experimental conditions and batches, using the same reference set for correction or setting the same fixed threshold for all samples to be tested cannot effectively eliminate the batch effect, and the results will fluctuate greatly.
[0007] (2) Data correction is performed using samples from the same batch. There are certain requirements for the number of samples in the same batch. When the number of samples in the same batch is small, such as less than 10 samples, it is impossible to construct an effective reference set.
[0008] (3) Because the number of windows in the detection is large, even far more than the number of reference set samples, the dimensionality curse is likely to occur during the screening process using features such as Manhattan distance, resulting in the problem that the appropriate reference set samples cannot be selected due to the features being too scattered.
[0009] Therefore, how to correct NIPT test data from a single sample remains a key focus and challenge in the research of non-invasive prenatal chromosomal aneuploidy detection.
[0010] Summary of the Invention
[0011] The purpose of this application is to provide an improved method for analyzing chromosomal aneuploidy and its application.
[0012] To achieve the above objectives, this application adopts the following technical solution:
[0013] The first aspect of this application discloses a method for analyzing chromosomal aneuploidy, including using sequencing data of the sample to be tested and its alignment results to calculate the window depth of the sample to be tested within a set window; calculating the average relative depth of samples in a selected reference set within the set window as a correction baseline; correcting the window depth of the sample to be tested using the correction baseline; calculating the Z-value of the sample to be tested based on the corrected window depth; determining whether the fetal chromosome of the sample to be tested has aneuploidy abnormality based on the Z-value; and selecting the reference set as a set of samples with the lowest difference from the sample to be tested, obtained by screening from a complete reference set consisting of samples with known fetal chromosome information and corresponding maternal peripheral blood cfDNA sequencing data.
[0014] The Z-score is used to determine whether aneuploidy occurs in the fetal chromosomes. The threshold used is determined by combining theoretical probability and the distribution of actual experimental results. Unless otherwise specified, the sequencing data in this application generally refers to the sequencing data of cfDNA from the peripheral blood of pregnant women, thereby achieving non-invasive prenatal genetic testing. The complete reference set consists of all samples used to generate the selection reference set, specifically clinical samples from BGI Genomics that reported no abnormalities between 2020 and 2021. The proportions of various types are kept roughly the same, and as many batches of experimental reagents as possible are included. The specific number of samples constructed is approximately 15,000.
[0015] It should be noted that the chromosome aneuploidy analysis method of this application creatively selects the samples with the lowest difference from the sample to be tested from past sample data as a selection reference set. The selection reference set is used to calculate the correction baseline and correct the sample to be tested. In the absence of samples from the same batch as a reference, it can more stably and accurately detect chromosome aneuploidy abnormalities in the sample to be tested and reduce batch fluctuations. It is especially suitable for single-sample NIPT detection, which can improve the accuracy of single-sample NIPT detection, has a good distinction between true positive and false positive samples, and reduces false positives.
[0016] In one implementation of this application, the method for analyzing chromosomal aneuploidy further includes calculating the degree of chimerism of the sample to be tested based on the corrected window depth, and determining whether the fetal chromosomes of the sample to be tested have aneuploidy abnormalities based on the degree of chimerism and the Z-value.
[0017] In one implementation of this application, the degree of chimerism is the ratio of abnormal fetal cells to all fetal cells.
[0018] In one implementation of this application, the method for calculating the degree of chimerism is as follows:
[0019] FF i The following calculation method was used to obtain it:
[0020] In the formula, q i That is, chimerism, FF i represents the relative fetal concentration of chromosome i, and FF represents the fetal concentration. This represents the average depth after chromosome i correction. The average depth of all corrected autosomes is represented by the value of i, which ranges from 1 to 22.
[0021] In one implementation of this application, the window is set as several windows divided according to the human reference genome, and there is no overlap between adjacent windows.
[0022] In one implementation of this application, the window depth is the number of unique alignment sequences of the sample sequencing data within the set window that are aligned to the human reference genome.
[0023] In one implementation of this application, the relative depth is the quotient of the sample's window depth in a set window divided by the average window depth of each autosome of the sample in the corresponding set window.
[0024] In one implementation of this application, the method for selecting the reference set includes: calculating the relative depth of the test sample and each sample in the complete reference set within a set window; using the complete reference set to learn and obtain a dimensionality reduction model to reduce the relative depth; using the dimensionality reduction features to calculate the difference between the test sample and each sample in the complete reference set; and selecting several samples with the lowest difference as the selection reference set for the test sample.
[0025] In one implementation of this application, the dimensionality reduction model is a PCA model obtained by PCA learning using a complete reference set; the dimensionality reduction features are the dimensionality reduction features output after the relative depth is input into the PCA model.
[0026] In one implementation of this application, the dissimilarity is the Euclidean distance between the first n principal components, where n is a positive integer. The first n principal components cover more than 70% of the variance, preferably 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the variance.
[0027] In one implementation of this application, the dissimilarity is the Euclidean distance of the first ten principal components, and the formula for calculating the dissimilarity is as follows.
[0028] In the formula, diversityreference_j represents the degree of difference, and PC_i sample PC_ireference_j represents the features after relative depth dimensionality reduction of the sample to be tested, and PC_ireference_j represents the features after relative depth dimensionality reduction of the samples in the complete reference set. PC only represents the features after relative depth dimensionality reduction and does not specifically refer to the features after dimensionality reduction by a certain dimensionality reduction model.
[0029] In one implementation of this application, at least 50 samples with the lowest difference are selected as the reference set for the selection of samples to be tested.
[0030] In one implementation of this application, the relative depth of the set window for autosomes other than chromosomes 13, 18, and 21 is used to select the reference set.
[0031] In one implementation of this application, the correction baseline consists of an autosomal correction baseline and an X chromosome correction baseline. The autosomal correction baseline is used for window depth correction of autosomes or female fetuses' X chromosomes, and the X chromosome correction baseline is used for window depth correction of male fetuses' X chromosomes. The autosomal correction baseline is the average of the relative depths of the autosomes in the set window of all samples in the selected reference set. The X chromosome correction baseline is the average of the relative depths of the X chromosomes in the set window of all female fetuses in the selected reference set.
[0032] In one implementation of this application, the window depth of the sample to be tested is corrected using a correction baseline. Specifically, this includes: (1) correcting for autosomes or the X chromosome of a female fetus using the following formula.
[0033] (2) The following formula is used to correct the X chromosome in male fetuses.
[0034] In the above formula, UR n UR represents the window depth after baseline correction. a The window depth before baseline correction. 常 This is the baseline for autosomal correction. X The baseline for X chromosome correction is shown, and FF represents the fetal concentration. In this application, the Y chromosome of the male fetus does not require correction.
[0035] In one implementation of this application, the autosomal correction baseline is,
[0036] The baseline for X chromosome correction is:
[0037] Among them, baseline 常 This is the baseline for autosomal correction. X The baseline for X chromosome correction is defined as follows: reference refers to all samples selected from the reference set; female-reference refers to all female fetal samples selected from the reference set; n reference The number of windows used to select all samples in the reference set is nfemale-reference, where n is the number of windows used to select all female fetuses in the reference set, and d is the relative depth of the window.
[0038] In one implementation of this application, the method for analyzing chromosomal aneuploidy further includes GC correction of the window depth and relative depth calculation using the corrected window depth.
[0039] In one implementation of this application, a GC correction step is performed before correcting the window depth of the sample to be tested in the baseline correction step.
[0040] In one implementation of this application, GC correction includes using the number of uniquely aligned sequences of the sample sequencing data to the human reference genome within a set window as the initial window depth, labeled as UR. o ; Calculate the GC content of the human reference genome within a specified window and label it as GC. r According to UR o With GC r The fitted values and UR values for all windows of the chromosomeo The average value of UR o Correction is required.
[0041] In one implementation of this application, GC correction further includes selecting all set windows on the chromosome and utilizing their UR... o With GC r The relationship is obtained by performing cubic spline interpolation fitting: Obtain the fitted values for each set window of the chromosome. Calculate the UR of all set windows of the chromosome. o The average value, denoted as The initial window depth is corrected using the following formula to obtain the GC-corrected window depth UR. a ;
[0042] In the formula, UR a This represents the window depth after GC correction, which is the window depth before baseline correction.
[0043] In one implementation of this application, the Z-value calculation method includes: for each chromosome i to be detected, calculating the Z-value for the window UR of that chromosome and the UR of other autosomal windows of the same chromosome, sorting the obtained Z-values, removing the maximum and minimum statistical values (i.e., removing one maximum statistical value and one minimum statistical value), and calculating the mean of the remaining Z-values to obtain the final Z-value, as detailed below.
[0044] (j = 1...22, j ≠ i, after sorting, remove the first and last two values)
[0045] in: Mean value of UR on chromosome i; Mean value of UR on chromosome j; SD i : Represents the standard deviation of the UR on chromosome i; SD j : Represents the standard deviation of the UR of chromosome j; L i : Indicates the number of windows used to divide chromosome i; l j : Indicates the number of windows used to divide chromosome j; Z i : Indicates the significance of aneuploidy on chromosome i, reflecting the difference from euploidy.
[0046] The second aspect of this application discloses a baseline correction method, which includes calculating the average relative depth of samples in a selected reference set within a set window as a correction baseline, and using the correction baseline to correct the window depth of the sample to be tested; wherein, the selected reference set is a number of samples with the lowest difference from the sample to be tested, selected from a complete reference set consisting of samples with known fetal chromosomal information and corresponding maternal peripheral blood cfDNA sequencing data.
[0047] In one implementation of this application, the baseline correction method further includes performing GC correction on the window depth and using the corrected window depth for relative depth calculation.
[0048] It should be noted that the baseline correction method in this application is actually the same as the correction baseline used in the analysis method for chromosomal aneuploidy in this application, which corrects the window depth of the sample to be tested. Therefore, the baseline correction method in this application, including setting the window, window depth, relative depth, selection of the reference set, dimensionality reduction model, difference degree and its calculation method, autosomal correction baseline, X chromosome correction baseline, correction method, GC correction, etc., all refer to the analysis method for chromosomal aneuploidy in the first aspect of this application, and will not be repeated here.
[0049] The baseline correction method provided in this application can also be applied to CNV detection. Before the specific location of the CNV is determined, it can be used to correct the UR and eliminate the influence of batch sequencing.
[0050] The third aspect of this application discloses a method for selecting a reference set for screening test samples. This reference set is used for chromosomal aneuploidy analysis. The method for selecting the reference set for test samples includes: calculating the relative depth of the test sample and each sample in the complete reference set within a set window; using the complete reference set to learn and obtain a dimensionality reduction model to reduce the relative depth; using the dimensionality reduction features to calculate the difference between the test sample and each sample in the complete reference set; and selecting several samples with the lowest difference as the selection reference set for the test sample. The complete reference set consists of several samples with known fetal chromosomal information and their corresponding maternal peripheral blood cfDNA sequencing data.
[0051] It should be noted that the method for selecting the reference set from the test samples in this application is actually the specific screening method for selecting the reference set from the complete reference set in the analysis method for chromosome aneuploidy in this application. Therefore, the specific settings of the window, relative depth, window depth, difference degree and its calculation method, GC correction, etc. in the method for selecting the reference set from the test samples are all the same as the analysis method for chromosome aneuploidy in the first aspect of this application, and will not be repeated here.
[0052] The fourth aspect of this application discloses an analytical apparatus for chromosome aneuploidy, comprising:
[0053] The window depth calculation module is used to calculate the window depth of the sample within a set window using the sequencing data and alignment results of the sample to be tested.
[0054] The baseline correction module is used to calculate the average relative depth of the samples in the selected reference set within a set window, which is used as the correction baseline. The window depth of the sample to be tested is corrected using the correction baseline. The selected reference set consists of several samples with the lowest difference from the sample to be tested, which are selected from a complete reference set composed of samples with known fetal chromosome information and corresponding maternal peripheral blood cfDNA sequencing data.
[0055] The Z-value calculation module is used to calculate the Z-value of the sample under test based on the corrected window depth.
[0056] The result judgment module is used to determine whether the fetal chromosomes of the sample under test have aneuploidy based on the Z value.
[0057] It should be noted that the analysis device for chromosome aneuploidy in this application is actually implemented by combining various modules to realize the analysis method for chromosome aneuploidy in this application. Therefore, the limitations and specific implementation methods of each module, such as setting the window, window depth, relative depth, selection method of the reference set, dimensionality reduction model, difference degree and its calculation method, autosomal correction baseline, X chromosome correction baseline, specific correction method, GC correction, etc., all refer to the analysis method for chromosome aneuploidy in the first aspect of this application, and will not be elaborated here.
[0058] It should also be noted that the key to this application lies in selecting a reference set from the complete reference set. In principle, a reference set selection process is required for each new sample to be tested. As for the complete reference set, it can be continuously supplemented by samples that have been tested, thereby continuously improving the complete reference set.
[0059] The fifth aspect of this application discloses a baseline correction apparatus, the apparatus comprising:
[0060] The baseline correction module is used to calculate the average relative depth of the samples in the selected reference set within a set window, which is used as the correction baseline. The window depth of the sample to be tested is corrected using the correction baseline. The selected reference set consists of several samples with the lowest difference from the sample to be tested, which are selected from a complete reference set composed of samples with known fetal chromosome information and corresponding maternal peripheral blood cfDNA sequencing data.
[0061] It should be noted that the baseline correction apparatus of this application is actually the baseline correction method of this application implemented through the baseline correction module; therefore, the limitations and specific implementation of the baseline correction module, such as setting the window, window depth, relative depth, selection method of the reference set, dimensionality reduction model, difference degree and its calculation method, autosomal baseline correction, X chromosome baseline correction, specific correction method, GC correction, etc., can all refer to the baseline correction method of the second aspect of this application, and will not be elaborated here.
[0062] The sixth aspect of this application discloses an apparatus for selecting a reference set by screening samples to be tested, the apparatus comprising:
[0063] The reference set selection module is used to calculate the relative depth of the test sample and each sample in the complete reference set within the window; a dimensionality reduction model is obtained by learning from the complete reference set to reduce the dimensionality of the relative depth; the difference between the test sample and each sample in the complete reference set is calculated using the dimensionality-reduced features, and several samples with the lowest difference are selected as the reference set for the test sample; the complete reference set consists of several samples with known fetal chromosome information and their corresponding maternal peripheral blood cfDNA sequencing data.
[0064] It should be noted that the apparatus for selecting a reference set for screening test samples in this application is actually the method for selecting a reference set for screening test samples in this application implemented by the reference set selection module. Therefore, the limitations and specific implementation methods of the reference set selection module, such as setting the window, relative depth, window depth, difference degree and its calculation method, GC correction, etc., can all refer to the method for selecting a reference set for screening test samples in the third aspect of this application, and will not be elaborated here.
[0065] The seventh aspect of this application discloses an analytical apparatus for chromosome aneuploidy, the apparatus comprising a memory and a processor; the memory for storing programs; and the processor for executing the programs stored in the memory to implement the chromosome aneuploidy analytical method, the baseline correction method, or the method for screening test samples and selecting a reference set of this application.
[0066] It is understood that when the apparatus of this application executes the program stored in memory to implement the method of selecting a reference set for screening test samples, the apparatus of this application is actually a device for selecting and constructing the reference set. The selection reference set constructed by this device can be used for non-invasive prenatal chromosomal aneuploidy analysis of single samples, improving the stability and accuracy of single-sample chromosomal aneuploidy detection and reducing batch fluctuations in detection. Similarly, when the apparatus of this application executes the program stored in memory to implement the baseline correction method of this application, the apparatus of this application is actually a baseline acquisition and correction device, which can also improve the stability and accuracy of single-sample chromosomal aneuploidy detection.
[0067] The eighth aspect of this application discloses a computer-readable storage medium storing a program that can be executed by a processor to implement the method for analyzing chromosomal aneuploidy, the method for baseline correction, or the method for selecting a reference set of samples to be tested.
[0068] It is understood that when the program stored in the computer-readable storage medium of this application can be executed by a processor to implement the method of screening test samples and selecting a reference set of this application, the computer-readable storage medium of this application is actually a computer-readable storage medium for screening the test sample selection reference set. This computer-readable storage medium can be directly used to implement the screening and construction of the selection reference set. The selection reference set obtained thereby can be used to detect fetal chromosomal aneuploidy abnormalities according to the method of this application. Similarly, when the program stored in the computer-readable storage medium of this application can be executed by a processor to implement the baseline correction method of this application, the storage medium of this application is actually a storage medium for baseline acquisition and correction to obtain the corresponding corrected baseline.
[0069] Due to the adoption of the above technical solutions, the beneficial effects of this application are as follows:
[0070] The method and apparatus for analyzing chromosomal aneuploidy in this application utilize a selected reference set to calculate a corrected baseline, which is then used to correct the sample to be tested. This allows for more stable and accurate detection of chromosomal aneuploidy abnormalities in the sample even without samples from the same batch as a reference, reducing batch-to-batch fluctuations. The method and apparatus of this application improve the accuracy of single-sample NIPT detection, exhibiting excellent differentiation between true positive and false positive samples, and reducing false positives. Attached Figure Description
[0071] Figure 1 is a flowchart of the chromosome aneuploidy analysis method in the embodiments of this application;
[0072] Figure 2 is a structural block diagram of the chromosome aneuploidy analysis device in an embodiment of this application;
[0073] Figure 3 shows the statistical results of the positive rate and gray area rate of different detection methods in the embodiments of this application. Detailed Implementation
[0074] The present application will now be described in further detail with reference to specific embodiments and accompanying drawings. In the following embodiments, many details are described to facilitate a better understanding of the present application. However, those skilled in the art will readily recognize that some features may be omitted in different situations, or may be replaced by other devices, materials, or methods. In some cases, certain operations related to the present application are not shown or described in the specification to avoid obscuring the core parts of the application with excessive description. For those skilled in the art, detailed description of these related operations is not necessary; a complete understanding of the related operations can be obtained from the description in the specification and general technical knowledge in the art.
[0075] This application creatively utilizes existing known databases to screen and select a reference set, and uses the selected reference set to calculate a correction baseline for the test sample to correct it, thereby improving the accuracy and stability of single-sample chromosomal aneuploidy abnormality detection and reducing batch fluctuations in detection.
[0076] Therefore, this application proposes a method for analyzing chromosomal aneuploidy, including using sequencing data of the sample to be tested and its alignment results to calculate the window depth of the sample to be tested within a set window; calculating the average relative depth of the selected reference set samples within the set window as a correction baseline; using the correction baseline to correct the window depth of the sample to be tested; calculating the Z-value of the sample to be tested based on the corrected window depth; and determining whether the fetal chromosomes of the sample to be tested have aneuploidy abnormalities based on the Z-value.
[0077] The selection reference set consists of several samples with the lowest difference from the test sample, selected from a complete reference set composed of samples with known fetal chromosomal information and corresponding maternal peripheral blood cfDNA sequencing data. Specifically, for example, the relative depth between the test sample and each sample in the complete reference set within a defined window is calculated; a PCA model is obtained through PCA learning using the complete reference set; the relative depth is then reduced in dimensionality using the PCA model obtained through PCA learning, with the relative depth as input and the reduced features as output; the difference between the test sample and each sample in the complete reference set is calculated using the reduced features, and the samples with the lowest difference are selected as the selection reference set for the test sample. The defined window consists of several windows divided according to the human reference genome, with no overlap between adjacent windows; the window depth is the number of unique alignment sequences of the sample sequencing data to the human reference genome within the defined window; the relative depth is the quotient of the sample's window depth within the defined window divided by the average window depth of each autosome of the sample within the corresponding defined window; and the chimerism is the ratio of abnormal fetal cells to all fetal cells.
[0078] In one implementation of this application, the method for analyzing chromosomal aneuploidy, as shown in Figure 1, specifically includes a sequence alignment step 11, a window depth calculation and GC correction step 12, a reference set selection and screening step 13, a fetal concentration calculation step 14, a baseline correction step 15, a Z-value calculation step 16, a chimerism calculation step 17, and a result judgment step 18.
[0079] The sequence alignment step 11 includes aligning the peripheral blood cfDNA sequencing data of the pregnant woman to be tested to a human reference genome, obtaining the coordinates of each sequence on the human reference genome, filtering out sequences with poor alignment quality and sequences with multiple alignments, and retaining uniquely aligned sequences. In one implementation of this application, the sequence information contained in the fastq format file generated by the sequencer is aligned to the human reference genome using alignment software BWA, such as GRCh37 / hg19. Finally, the coordinates and other information of each uniquely aligned sequence are stored in a bam format file. Sequences with poor quality refer to sequences containing mismatched bases; that is, only sequences without mismatched bases are retained.
[0080] Step 12, which involves calculating the window depth and performing GC correction, includes dividing the human reference genome into several windows with no overlap between adjacent windows, calculating the number of unique aligned sequences in each window as the initial window depth, and labeling it as UR. o The GC content of the human reference genome in each window was statistically analyzed and labeled as GC. r Select all windows on the chromosome, based on UR o With GC rThe fitted values and UR values for all windows of the chromosome o The average value of UR o Correction is required.
[0081] In one implementation of this application, for each sample, all windows on the chromosome are selected, and UR is used. o With GC r The relationship is obtained by performing cubic spline interpolation fitting: The fitted value for each window of the chromosome is obtained from this relationship. Calculate UR for all windows of the chromosome o The average value, denoted as UR is calculated according to the following formula. o Perform correction to obtain the window depth after GC correction;
[0082] In the formula, UR a This represents the window depth after GC correction.
[0083] Step 13 of the reference set selection process includes calculating the relative depth of the test sample and each sample in the complete reference set within the window; using the complete reference set for PCA learning to obtain a PCA model, and performing dimensionality reduction on the relative depth, i.e., using the PCA model obtained through PCA learning, inputting the relative depth, and outputting the dimensionality-reduced features; using the dimensionality-reduced features to calculate the difference between the test sample and each sample in the complete reference set, and selecting several samples with the lowest difference as the selection reference set for the test sample; the complete reference set consists of several samples with known fetal chromosomal conditions and their corresponding maternal peripheral blood cfDNA sequencing data, wherein the several fetal chromosomal conditions are preferably several cases without fetal chromosomal aneuploidy abnormalities; the relative depth is the quotient of the sample's window depth in the window divided by the average window depth of each autosome in the window of that sample.
[0084] The selection of the reference set is one of the key aspects of this application. In one implementation of this application, a complete reference set containing a large number of samples is established, and a method for selecting a specific reference set from the complete reference set for each sample to be tested is established.
[0085] The complete reference set contains samples from as many experimental conditions and batches as possible. For example, in one implementation of this application, clinical samples from BGI Genomics that did not report fetal chromosomal aneuploidy abnormalities between 2020 and 2021 are selected, the proportion of various peripheral blood sampling tubes is controlled to be roughly the same, and the experimental reagents from as many batches as possible are included. Specifically, a complete reference set with approximately 15,000 samples is constructed.
[0086] In one implementation of this application, the relative depth of the windows of autosomes other than chromosomes 13, 18, and 21 is selected and labeled as d, as a feature for selecting the reference set. The relative depth is the quotient of the window depth of the sample in the set window divided by the average window depth of each autosome in the corresponding set window. Usually, the window depth after GC correction is divided by the average window depth of each autosome in the corresponding set window, which is used as the result after the window depth is standardized.
[0087] In the above formula, d represents the relative depth, and UR... a This represents the window depth after GC correction. This represents the average window depth of each autosome in the sample within its corresponding window.
[0088] Because the window features of the samples have high dimensionality, feature dimensionality reduction must be considered when selecting the reference set. PCA is a commonly used unsupervised machine learning dimensionality reduction method. This invention preferably uses a PCA model for dimensionality reduction. The PCA model is obtained by learning using the complete reference set. For the test sample, the relative depth information of the sample's window is output to the PCA model, thus obtaining the dimensionality-reduced features (PC) of that sample. The dimensionality-reduced features are used to calculate the difference between the test sample and each sample in the complete reference set. The difference is defined as the Euclidean distance of the first ten principal components. The 50 samples with the lowest difference from the test sample are selected as the specific reference set for subsequent analysis.
[0089] The formula for calculating the degree of difference is as follows:
[0090] In the formula, diversityreference_j represents the degree of difference, and PC_i sample PC_ireference_j represents the features of the samples under test after relative depth dimensionality reduction, which is the feature of the samples in the complete reference set after relative depth dimensionality reduction.
[0091] It should be noted that PC_i in the formula sample The PC in PC_ireference_j only represents the features after relative depth dimensionality reduction, and does not specifically refer to any particular dimensionality reduction model.
[0092] Fetal concentration calculation step 14 includes determining the fetal concentration of male fetuses based on the proportion of the Y chromosome, or estimating the fetal concentration of female fetuses by establishing a high-dimensional regression model based on the non-uniform distribution of fetal cell-free DNA on the genome.
[0093] Fetal fraction (FF) refers to the proportion of fetal-derived cell-free nucleic acids present in the maternal peripheral blood circulation.
[0094] Specifically, the concentration of male fetuses is determined by the proportion of Y chromosomes, and the Y chromosome window UR o Mean divided by the UR of autosomes o The mean value, multiplied by 2, gives the fetal DNA concentration (FF) of a male fetus; therefore, the fetal DNA concentration of a male fetus is calculated as follows:
[0095] The concentration of fetal DNA in female fetuses was estimated using a high-dimensional regression model based on the non-uniform distribution of fetal cell-free DNA across the genome. The underlying assumption was that the distribution characteristics of fetal cfDNA and maternal cfDNA across the genome differed between male and female fetuses. Therefore, the fetal concentration estimated using the Y chromosome method for male fetuses was used as input to train the model. A regression model was constructed using neural network machine learning, as detailed below:
[0096] Where l is the layer number of the network, the first layer is the input layer, the last layer is the output layer (with only one neuron), and the middle layers are hidden layers. Let be the value of the j-th neuron in the l-th layer. This represents the value of the k-th neuron in the (l-1)-th layer. The connection weights are the connection weights from the k-th neuron in layer (l-1) to the j-th neuron in layer l. This represents the input bias of the j-th neuron in the l-th layer. The most common form of the function f is the rectified linear unit, i.e., f(x) = max(0,x). w and b are obtained during model training. When applying the model, the neuron values are calculated layer by layer according to the above formula, and the neuron values in the last layer are the predicted values of the fetal concentration model.
[0097] The baseline correction step 15 includes calculating the average relative depth of all samples in the selected reference set within the window as the correction baseline, and using the correction baseline to correct the window depth of the sample to be tested.
[0098] In one implementation of this application, the mean relative depth of the selected samples in the reference set within the window is the correction baseline. The autosomal correction baseline is the average relative depth of the autosomes in the set window for all samples in the reference set, and the X chromosome correction baseline is the average relative depth of the X chromosomes in the set window for all female fetuses in the reference set. The corrected window depth, UR, of the test sample divided by the correction baseline of that window, is denoted as UR. nThe UR (uric acid) of the X chromosome in male fetal samples needs to be corrected for fetal concentration. Specifically, the correction baseline includes an autosomal correction baseline and an X chromosome correction baseline. The autosomal correction baseline is used to correct the window depth for autosomes or the X chromosome in female fetuses, while the X chromosome correction baseline is used to correct the window depth for the X chromosome in male fetuses. The autosomal correction baseline is the average relative depth of the autosomes within a set window for all samples in the selected reference set; the X chromosome correction baseline is the average relative depth of the X chromosome within a set window for all female fetal samples in the selected reference set. Details are as follows:
[0099] Autosomal correction baseline:
[0100] X chromosome correction baseline:
[0101] baseline 常 This is the baseline for autosomal correction. X The baseline for X chromosome correction is defined as follows: reference refers to all samples selected from the reference set; female-reference refers to all female fetal samples selected from the reference set; n reference The number of windows used to select all samples in the reference set is nfemale-reference, where n is the number of windows used to select all female fetuses in the reference set, and d is the relative depth of the window.
[0102] The window depth of the sample to be tested is corrected using a correction baseline, as follows:
[0103] (1) For autosomal chromosomes or the X chromosome in female fetuses, the following formula is used for correction.
[0104] (2) The following formula is used to correct the X chromosome in male fetuses.
[0105] In the above formula, UR n UR represents the window depth after baseline correction. a The window depth after GC correction and before baseline correction. 常 This is the baseline for autosomal correction. X The baseline for X chromosome correction is FF, where FF is the fetal concentration.
[0106] Z-value calculation step 16 includes calculating the Z-value of the sample to be tested based on the window depth after baseline correction.
[0107] Autosomal chromosomal URs follow a Poisson distribution, and with a large number of windows, they follow a normal distribution. For normal samples, the distribution of each window is indistinguishable, while for abnormal samples, there are slight differences, influenced by fetal concentration; the higher the fetal concentration, the greater the difference. The Z-test can be used to determine fetal chromosomal aneuploidy at a certain fetal concentration. For each chromosome i to be tested, the Z-value is calculated by comparing the UR of that chromosome's window with the URs of other autosomal windows in the same sample. The calculation method is as follows: The obtained Z-values are sorted, and the largest and smallest statistical values are removed. The mean of the remaining Z-values is the final Z-value. Details are as follows:
[0108] (j = l...22, j ≠ i, after sorting, remove the first and last two values)
[0109] in: Mean value of UR on chromosome i; Mean value of UR on chromosome j; SD i : Represents the standard deviation of the UR on chromosome i; SD j : Represents the standard deviation of the UR of chromosome j; L i : Indicates the number of windows used to divide chromosome i; L j : Indicates the number of windows used to divide chromosome j; Z i : Indicates the significance of aneuploidy on chromosome i, reflecting the difference from euploidy.
[0110] The above formula compares the 22 autosomes within the same sample. This is based on the assumption that the vast majority of chromosomes in a sample should be normal diploid. Therefore, the target chromosome is compared 21 times with the remaining 21 chromosomes. If the target chromosome is normal diploid, the vast majority of the 21 Z-test values should be close to 0, and averaging them yields a negative Z-value. Conversely, if the target chromosome is trisomic, the vast majority of the 21 Z-test values will be much greater than 0, and averaging them yields a positive Z-value.
[0111] Step 17, which calculates the degree of chimerism, includes calculating the relative fetal concentration of each chromosome in the sample to be tested based on the window depth after baseline correction, and calculating the degree of chimerism of each chromosome based on the relative fetal concentration and the fetal DNA concentration.
[0112] For both mosaic and non-mosaic embryonic samples, adding the degree of mosaicism to the Z-score to jointly determine the aneuploidy of the fetal chromosomes in the sample under test can better distinguish true positive samples.
[0113] Calculate the relative fetal concentration FF of chromosome i iThe relative fetal concentration of a chromosome is defined as the fetal concentration estimated using that chromosome assuming a trisomy 2000 in the fetus. The calculation method is as follows:
[0114] In the formula, FF i represents the relative fetal concentration of chromosome i, and FF represents the fetal concentration. This represents the average depth after chromosome i correction. The average depth of all corrected autosomes is represented by the value of i, which ranges from 1 to 22.
[0115] Calculate the mosaicism of chromosome i, denoted as q. i Mosaicity is the ratio of abnormal cells to all fetal cells when a trisomy of a certain chromosome is assumed to exist in the fetus. The calculation method is as follows:
[0116] The result judgment step 18 includes determining whether the fetal chromosomes of the sample to be tested have aneuploidy based on the Z-value and the degree of mosaicism.
[0117] In this application, the final detection result for each chromosome is a comprehensive judgment based on the Z-score and the degree of mosaicism. The higher the degree of mosaicism and the larger the absolute value of the Z-score, the more likely the sample is to have an abnormality. Different thresholds for mosaicism and Z-scores are set for different chromosomes, and the final detection result is obtained by comparing the detection value with the threshold. An example of trisomy detection on chromosome 21 is shown in Table 1.
[0118] Table 1. Trisomy Detection of Chromosome 21
[0119] In Table 1, Z1, Z2, and Z3 are the threshold values for the Z-value, which are normal distribution thresholds with special significance according to probability theory. For example, they represent ±1.96 for a p-value of 0.05, and 3 for a p-value of approximately 0.01, etc., selected based on the distribution of actual results. The gray area represents samples for which a definite conclusion cannot be obtained at this time.
[0120] Based on the method for analyzing chromosomal aneuploidy in this application, this application proposes a baseline correction method, which includes calculating the average relative depth of samples in a selected reference set within a set window as the correction baseline, and using the correction baseline to correct the window depth of the sample to be tested; the selected reference set consists of several samples with the lowest difference from the sample to be tested, obtained by screening from a complete reference set composed of samples with known fetal chromosomal information and their corresponding maternal peripheral blood cfDNA sequencing data.
[0121] Based on the method for analyzing chromosomal aneuploidy in this application, this application proposes a method for selecting a reference set for screening test samples. This reference set is used for single-sample non-invasive prenatal chromosomal aneuploidy abnormality analysis. The method for selecting the reference set for test samples includes: calculating the relative depth of the test sample and each sample in the complete reference set within a set window; using the complete reference set to learn and obtain a model, and then reducing the dimensionality of the relative depth, i.e., using the learned dimensionality-reduced model, inputting the relative depth, and outputting the dimensionality-reduced features; using the dimensionality-reduced features to calculate the difference between the test sample and each sample in the complete reference set, and selecting several samples with the lowest difference as the selection reference set for the test sample; the complete reference set consists of several samples with known fetal chromosomal information and their corresponding maternal peripheral blood cfDNA sequencing data.
[0122] The defined window is a set of windows divided according to the human reference genome, with no overlap between adjacent windows; the relative depth is the quotient of the sample's window depth within the defined window divided by the average window depth of each autosome of the sample within the defined window; the window depth is the number of unique alignment sequences of the sample sequencing data within the defined window that match the human reference genome. The difference calculation and GC correction are the same as the analysis method for chromosomal aneuploidy in this application, and will not be repeated here.
[0123] Those skilled in the art will understand that all or part of the functions of the above methods can be implemented in hardware or by computer programs. When all or part of the functions of the above methods are implemented by computer programs, the program can be stored in a computer-readable storage medium, which may include: read-only memory, random access memory, disk, optical disk, hard disk, etc., and the program is executed by a computer to achieve the above functions. For example, the program can be stored in the memory of a device, and when the program in the memory is executed by the processor, all or part of the above functions can be achieved. In addition, when all or part of the functions in the above embodiments are implemented by computer programs, the program can also be stored in storage media such as a server, another computer, disk, optical disk, flash drive, or portable hard drive, and can be downloaded or copied to the memory of a local device, or the system of the local device can be updated. When the program in the memory is executed by the processor, all or part of the functions of the above methods can be achieved.
[0124] Therefore, based on the method for analyzing chromosomal aneuploidy in this application, this application proposes an analytical device for chromosomal aneuploidy, comprising: a window depth calculation module, used to calculate the window depth of the sample to be tested within a set window using the sequencing data of the sample to be tested and its alignment results; a baseline correction module, used to calculate the average relative depth of samples in the selected reference set within the set window as the correction baseline, and to correct the window depth of the sample to be tested using the correction baseline; the selected reference set is a set of samples with the lowest difference from the sample to be tested, obtained by screening samples with known fetal chromosomal conditions and their corresponding maternal peripheral blood cfDNA sequencing data from a complete reference set; a Z-value calculation module, used to calculate the Z-value of the sample to be tested based on the corrected window depth; and a result judgment module, used to judge whether the fetal chromosomes of the sample to be tested have aneuploidy abnormalities based on the Z-value.
[0125] In one implementation of this application, the specific analysis device for chromosome aneuploidy, as shown in Figure 2, includes a sequence alignment module 21, a window depth calculation and GC correction module 22, a reference set selection and screening module 23, a fetal concentration calculation module 24, a baseline correction module 25, a Z-value calculation module 26, a chimerism calculation module 27, and a result judgment module 28.
[0126] The sequence alignment module 21 includes components for aligning the peripheral blood cfDNA sequencing data of the pregnant woman to be tested onto the human reference genome, obtaining the coordinates of each sequence on the human reference genome, filtering out sequences with poor alignment quality and sequences with multiple alignments, and retaining uniquely aligned sequences. For example, the sequence information contained in the fastq format file generated by the sequencer is aligned to the human reference genome GRCh37 / hg19 using the alignment software BWA, leaving uniquely aligned sequences, and storing the coordinates and other information of each uniquely aligned sequence in a bam format file.
[0127] Window depth calculation and GC correction module 22 includes a module for dividing the human reference genome into several windows with no overlap between adjacent windows, calculating the number of unique aligned sequences in each window as the initial window depth, denoted as UR. o The GC content of the human reference genome in each window was statistically analyzed and labeled as GC. r Select all windows on the chromosome, based on UR o With GC r The fitted values and UR values for all windows of the chromosome o The average value of UR o Correction is required.
[0128] For example, using UR o With GC r The relationship is obtained by performing cubic spline interpolation fitting: The fitted value for each window of the chromosome is obtained from this relationship. Calculate UR for all windows of the chromosome o The average value, denoted as UR is calculated according to the following formula. o Perform correction to obtain the window depth after GC correction;
[0129] In the formula, UR a This represents the window depth after GC correction.
[0130] The reference set selection module 23 includes methods for calculating the relative depth of the test sample and each sample in the complete reference set within a window; using the complete reference set for PCA learning to obtain a PCA model, and then reducing the dimensionality of the relative depth, i.e., using the PCA model obtained through PCA learning, inputting the relative depth, and outputting the dimensionality-reduced features; using the dimensionality-reduced features to calculate the difference between the test sample and each sample in the complete reference set, and selecting several samples with the lowest difference as the reference set for the test sample; the complete reference set consists of several samples with known fetal chromosomal information and their corresponding maternal peripheral blood cfDNA sequencing data; the relative depth is the quotient of the sample's window depth divided by the average window depth of each autosome in the sample within the window. The calculation methods for relative depth and difference are the same as those used in the single-sample non-invasive prenatal chromosomal aneuploidy analysis of this application.
[0131] The fetal DNA concentration calculation module 24 includes methods for determining the fetal DNA concentration of male fetuses based on the proportion of the Y chromosome, or for estimating the fetal DNA concentration of female fetuses by establishing a high-dimensional regression model based on the non-uniform distribution of cell-free fetal DNA on the genome. The specific formulas for calculating the fetal DNA concentration of male fetuses and the methods for calculating the fetal DNA concentration of female fetuses are the same as the analytical methods for chromosomal aneuploidy in this application.
[0132] The baseline correction module 25 includes a method for calculating the average relative depth of the samples in the selected reference set within a window, which serves as the correction baseline. This correction baseline is used to correct the window depth of the sample to be tested. The correction baseline includes an autosomal correction baseline and an X-chromosome correction baseline. The autosomal correction baseline is used to correct the window depth of autosomes or the X chromosome of female fetuses, while the X-chromosome correction baseline is used to correct the window depth of the X chromosome of male fetuses. The autosomal correction baseline is the average relative depth of the autosomes within a set window for all samples in the selected reference set; the X-chromosome correction baseline is the average relative depth of the X chromosome within a set window for all female fetuses in the selected reference set. The specific calculation formula and the method for correcting the window depth of the sample to be tested using the correction baseline are the same as the method for single-sample non-invasive prenatal chromosomal aneuploidy analysis in this application.
[0133] Z-value calculation module 26 includes a method for calculating the Z-value of the sample to be tested based on the baseline-corrected window depth. The Z-value calculation formula is the same as the method for analyzing chromosomal aneuploidy in this application.
[0134] The chimerism calculation module 27 includes methods for calculating the relative fetal concentration of each chromosome in the sample based on the baseline-corrected window depth, and for calculating the chimerism of each chromosome based on the relative fetal concentration and fetal DNA concentration. The calculation of relative fetal concentration, fetal DNA concentration, and chimerism is the same as the method for single-sample non-invasive prenatal chromosomal aneuploidy analysis in this application.
[0135] The result determination module 28 includes a method for determining whether the fetal chromosomes in the test sample have aneuploidy based on the Z-score and mosaicism. The specific determination method is the same as the method for single-sample non-invasive prenatal chromosomal aneuploidy analysis in this application.
[0136] Another implementation of this application provides an analysis device for chromosomal aneuploidy, the device including a memory and a processor; the memory includes a program for storing programs; the processor includes a program for executing the program stored in the memory to implement the following method: using sequencing data of the sample to be tested and its alignment results, calculating the window depth of the sample to be tested within a set window; calculating the average relative depth of samples in a selected reference set within the set window as a correction baseline; correcting the window depth of the sample to be tested using the correction baseline; calculating the Z-value of the sample to be tested based on the corrected window depth; and determining whether the fetal chromosomes of the sample to be tested have aneuploidy abnormalities based on the Z-value.
[0137] Alternatively, the device includes a memory and a processor; the memory includes a program for storing a program; the processor includes a program for executing the program stored in the memory to implement the following method: including calculating the average relative depth of samples in a selected reference set at a set window as a correction baseline, correcting the window depth of the sample to be tested using the correction baseline; the selected reference set is a set of samples with the lowest difference from the sample to be tested, obtained by screening from a complete reference set consisting of samples with known fetal chromosomal information and their corresponding maternal peripheral blood cfDNA sequencing data.
[0138] Alternatively, the device includes a memory and a processor; the memory includes a program for storing a program; the processor includes a program for executing the program stored in the memory to implement the following method: including calculating the relative depth of the test sample and each sample in the complete reference set within a set window; learning from the complete reference set to obtain a dimensionality reduction model, reducing the relative depth dimensionality, i.e., using the learned dimensionality reduction model, inputting the relative depth, and outputting the dimensionality reduction features; using the dimensionality reduction features to calculate the difference between the test sample and each sample in the complete reference set, and selecting several samples with the lowest difference as the selection reference set for the test sample.
[0139] Another implementation of this application also provides a computer-readable storage medium, which includes a program that can be executed by a processor to implement the following method: using sequencing data of the sample to be tested and its alignment results, calculating the window depth of the sample to be tested within a set window; calculating the average relative depth of samples in a selected reference set within the set window as a correction baseline; correcting the window depth of the sample to be tested using the correction baseline; calculating the Z-value of the sample to be tested based on the corrected window depth; and determining whether the fetal chromosome of the sample to be tested has aneuploidy based on the Z-value.
[0140] Alternatively, the storage medium may include a program that can be executed by a processor to implement the following method: including calculating the average relative depth of samples in a selected reference set at a set window as a correction baseline, using the correction baseline to correct the window depth of the sample to be tested; the selected reference set is a set of samples with the lowest difference from the sample to be tested, obtained by screening from a complete reference set consisting of samples with known fetal chromosomal information and their corresponding maternal peripheral blood cfDNA sequencing data.
[0141] Alternatively, the storage medium may include a program that can be executed by a processor to implement the following method: calculating the relative depth of the test sample and each sample in the complete reference set within a set window; learning from the complete reference set to obtain a dimensionality reduction model, and reducing the relative depth by using the learned dimensionality reduction model, inputting the relative depth, and outputting the dimensionality-reduced features; using the dimensionality-reduced features to calculate the difference between the test sample and each sample in the complete reference set, and selecting several samples with the lowest difference as the selection reference set for the test sample.
[0142] This application innovatively selects samples with the lowest difference from the test sample from samples under different experimental conditions and batches as a selection reference set. A corrected baseline is calculated using this selection reference set and applied to the test sample. In the absence of samples from the same batch or with insufficient samples as a reference, this method can more stably and accurately detect chromosomal aneuploidy abnormalities in the test sample, reducing batch-specific fluctuations. This method improves the accuracy of single-sample NIPT detection, exhibits excellent differentiation between true positive and false positive samples, and reduces false positives.
[0143] Example 1
[0144] A total of 973 samples were selected from those tested in actual clinical applications by our company, and those that underwent prenatal diagnosis and postnatal follow-up were confirmed to have positive karyotypes for chromosomal number abnormalities. First, these samples were standardized using traditional in-batch correction methods in actual clinical testing to calculate Z-scores and provide the results. Second, the Z-score calculation method based on the screening and selection reference set established in this invention was used to evaluate the results.
[0145] The specific steps are as follows:
[0146] 1) Sequence alignment steps: The sequence information of the sequencing data of 973 samples was compared to the human reference genome GRCh37 / hg19 using BWA, and the coordinates and other information of the first unique aligned sequence were stored in a bam format file.
[0147] 2) Window depth calculation and GC correction steps: Divide the human reference genome into windows of approximately 60kb each, with no overlap between adjacent windows. Calculate the number of uniquely aligned sequences (URs) within each window. o and sequence GC content GC o The GC content (GCr) of the human reference genome in each window was statistically analyzed. All windows on the chromosome were selected, and the UR was used to... o With GC r The relationship is obtained by performing cubic spline interpolation fitting: The fitted value for each window of the chromosome is obtained from this relationship. Calculate UR for all windows of the chromosome o The average value, denoted as UR is calculated according to the following formula. o Perform correction to obtain the window depth after GC correction.
[0148] In the formula, UR o UR is the initial window depth. a This represents the window depth after GC correction.
[0149] 3) Selection of reference set screening steps: In this embodiment, clinical samples from BGI Genomics that did not report fetal chromosomal aneuploidy abnormalities between 2020 and 2021 were selected as a complete reference set, with the proportion of each type being roughly the same and including as many batches of experimental reagents as possible.
[0150] The relative depth of the windows of autosomes other than chromosomes 13, 18, and 21 is selected and labeled as d. This is used as a feature for selecting the reference set. The relative depth is the quotient of the window depth of the sample in the set window divided by the average window depth of each autosome in the corresponding set window. In this example, the window depth UR after GC correction is divided by the average window depth of each autosome in the sample, which is used as the result after standardization of the window depth.
[0151] PCA model is obtained by performing PCA learning using the complete reference set. For the test sample, the relative depth information of the window is output to the PCA model to obtain the dimensionality-reduced features of the sample. The dimensionality-reduced features are then used to calculate the diversity between the test sample and each sample in the complete reference set. The diversity is defined as the Euclidean distance of the first ten principal components. The 50 samples with the lowest diversity with the test sample are selected as the sample-specific selection reference set for subsequent analysis.
[0152] The formula for calculating the degree of difference is as follows:
[0153] In the formula, diversityreference_j represents the degree of difference, and PC_i sample PC_ireference_j represents the features of the samples under test after relative depth dimensionality reduction, which is the feature of the samples in the complete reference set after relative depth dimensionality reduction.
[0154] 4) Fetal concentration calculation steps: The fetal concentration of a male fetus is determined by the proportion of the Y chromosome. The mean UR value of the Y chromosome window is divided by the mean UR value of the autosomes, and then multiplied by 2 to obtain the fetal concentration FF of the male fetus. Therefore, the fetal DNA concentration of a male fetus is calculated as follows:
[0155] Fetal DNA concentration in female fetuses was estimated using a high-dimensional regression model built upon the non-uniform distribution of cell-free fetal DNA across the genome. Fetal concentration estimated using the Y chromosome method in male fetuses was used as input to train the model, and a regression model was constructed using neural network machine learning, as detailed below:
[0156] Where l is the layer number of the network, the first layer is the input layer, the last layer is the output layer (with only one neuron), and the middle layers are hidden layers. Let be the value of the j-th neuron in the l-th layer. This represents the value of the k-th neuron in the (l-1)-th layer. The connection weights are the connection weights from the k-th neuron in layer (l-1) to the j-th neuron in layer l. This represents the input bias of the j-th neuron in the l-th layer. The most common form of the function f is the rectified linear unit, i.e., f(x) = max(0,x). w and b are obtained during model training. When applying the model, the neuron values are calculated layer by layer according to the above formula, and the neuron values in the last layer are the predicted values of the fetal concentration model.
[0157] 5) Baseline correction steps: The autosomal correction baseline is the average relative depth of the autosomes in the set window of all samples in the selected reference set, and the X chromosome correction baseline is the average relative depth of the X chromosome in the set window of all female fetuses in the selected reference set.
[0158] Autosomal correction baseline:
[0159] X chromosome correction baseline:
[0160] Among them, baseline 常 This is the baseline for autosomal correction. X The baseline for X chromosome correction is defined as follows: reference refers to all samples selected from the reference set, and female-reference refers to all female fetuses selected from the reference set. n reference The number of windows used to select all samples in the reference set is nfemale-reference, where n is the number of windows used to select all female fetuses in the reference set, and d is the relative depth of the window.
[0161] (1) For autosomal chromosomes or the X chromosome in female fetuses, the following formula is used for correction.
[0162] (2) The following formula is used to correct the X chromosome in male fetuses.
[0163] In the above formula, UR n UR represents the window depth after baseline correction. a The window depth before baseline correction. 常 This is the baseline for autosomal correction. X The baseline for X chromosome correction is FF, where FF is the fetal concentration.
[0164] 6) Z-score calculation steps: For each chromosome i to be tested, calculate the Z-score for the window UR of that chromosome and the UR of other autosomal windows of the same sample. The calculation method is as follows: Sort the obtained Z-scores, remove the largest and smallest statistical values, and calculate the mean of the remaining Z-scores to obtain the final Z-score. Details are as follows:
[0165] (j = 1...22, j ≠ i, after sorting, remove the first and last two values)
[0166] in: Mean value of UR on chromosome i; Mean value of UR on chromosome j; SD i : Represents the standard deviation of the UR on chromosome i; SD j : Represents the standard deviation of the UR of chromosome j; L i : Indicates the number of windows used to divide chromosome i; L j : Indicates the number of windows used to divide chromosome j; Z i : Indicates the significance of aneuploidy on chromosome i, reflecting the difference from euploidy.
[0167] 7) Steps for calculating chimerism: Calculate the relative concentration FF of chromosome i. i The relative concentration of a chromosome is defined as the fetal concentration estimated using that chromosome assuming a trisomy 2000 in the fetus. The calculation method is as follows:
[0168] In the formula, FF i This represents the relative fetal concentration of chromosome i. This represents the average depth after chromosome i correction. The average depth of all autosomes after correction is given, where i ranges from 1 to 22.
[0169] Calculate the mosaicism of chromosome i, denoted as q. i Mosaicity is the ratio of abnormal cells to all fetal cells when a trisomy of a certain chromosome is assumed to exist in the fetus. The calculation method is as follows:
[0170] 8) Result judgment steps: Determine whether the fetal chromosome of the sample to be tested has aneuploidy based on the Z value and mosaicism. The results are shown in Table 2 and Table 3.
[0171] It should be noted that the experimental sample size in this embodiment is relatively large, and the calculation results of each sample according to the above steps are not shown one by one.
[0172] The results of the method for selecting the reference set, the batch correction method, and the fixed reference set method in this embodiment are compared and the statistical results are shown in Table 2.
[0173] The other two algorithms are:
[0174] 1) The intra-batch correction method is as follows: In the baseline correction step, the baseline is calculated using samples from the same batch as the test sample with similar total GC content (GC content difference less than 0.02). The only difference between the intra-batch correction method test and the scheme of this embodiment is that steps 3) and 5) are omitted, and the GC correction method in step 2) is replaced by the intra-batch correction method.
[0175] 2) The method using a fixed reference set is as follows: In the baseline correction step, a fixed baseline data is used for all test samples. All samples (approximately 1000 cases) in a fixed complete reference set are used to calculate the baseline data, and there is no step of selecting a reference set. The difference between the fixed reference set correction experiment and the scheme in this embodiment is only that steps 3) and 5) are omitted, and the GC correction method in step 2) is replaced by the fixed reference set correction method.
[0176] In the method for selecting the reference set, the correction window size was specifically chosen to be 60k. The threshold selections in the Z-value method are shown in Table 3. For chromosome 13, Z>=3 and q>=0.5 is considered positive, Z>=3 and q<0.5 is considered gray, 1.96<=Z<3 and q>=0.25 is considered gray, and 1.96<=Z<3 and q<0.25 is considered negative. For chromosome 18, Z>=3 and q>=0.5 is considered positive, and Z>=3 and q<0.5 is considered gray. 1.96 <= Z < 3 and q >= 0.2 is a gray area, and 1.96 <= Z < 3 and q < 0.2 is negative; for chromosome 21, Z >= 4 is positive, 3 <= Z < 4 and q >= 0.18 is positive, 3 <= Z < 4 and q < 0.18 is a gray area, 1.96 <= Z < 3 and q >= 0.18 is a gray area, and 1.96 <= Z < 3 and q < 0.18 is negative.
[0177] Table 2 Statistical analysis of test results
[0178] The results in Table 2 show that the method using a selected reference set provided correct detection results for 973 samples, with a significantly lower number of false negatives compared to the method using a fixed reference set. Therefore, the algorithm using a selected reference set has better sensitivity and is superior to the method using a fixed reference set in the detection of sex chromosomes.
[0179] Table 3. Z-scores and mosaicism thresholds for trisomy detection of chromosomes 13, 18, and 21
[0180] Example 2
[0181] This example uses the established model to detect continuous samples from the production line.
[0182] The karyotype-based sample data distribution characteristics cannot reflect the true distribution characteristics of the population, nor can the specificity of the method be evaluated. Therefore, to evaluate the performance of the method in practical applications, such as outlier reporting rate, retest rate, and gray zone rate, a series of tests were conducted using clinical samples from the production line. Several complete batches from our company's production line between June and December 2021 were used as the test set. The outlier reporting rates of various anomalies in the first experiment using the reference set selection method established in this invention for samples that passed quality control and were not GC outliers were statistically analyzed, and compared with the results of in-batch correction and fixed reference set. A total of 5572 samples were counted, and the results are shown in Table 4 and Figure 3. The in-batch correction method and the fixed reference set method are the same as in Example 1.
[0183] Table 4 shows the statistical results of continuous sample testing on the production line.
[0184] The results in Table 4 and Figure 3 show that, assuming no false negatives, the selected reference set method significantly reduces the number of abnormal reports compared to the single-sample method with a fixed reference set. Therefore, the selected reference set method has fewer false positives and better specificity than the fixed reference set method, resulting in a higher positive predictive value (PPV). Furthermore, compared to the in-batch reference set, the selected reference set method has a lower gray area rate, reducing the cost associated with retesting gray area samples.
[0185] The above description, in conjunction with specific embodiments, provides a further detailed explanation of this application and should not be construed as limiting the specific implementation of this application to these descriptions. For those skilled in the art, several simple deductions or substitutions can be made without departing from the concept of this application.
Claims
1. A method for analyzing chromosomal aneuploidy, characterized in that: This includes using the sequencing data of the sample to be tested and its alignment results to calculate the window depth of the sample to be tested within a set window; The average relative depth of the selected reference set samples within a set window is calculated as the correction baseline. The window depth of the test sample is corrected using the correction baseline. The Z value of the test sample is calculated based on the corrected window depth. The Z value is used to determine whether the fetal chromosome of the test sample has aneuploidy. The selection reference set consists of several samples with the lowest difference from the sample to be tested, selected from a complete reference set composed of samples with known fetal chromosome information and corresponding maternal peripheral blood cfDNA sequencing data.
2. The method according to claim 1, characterized in that: It also includes calculating the chimerism of the sample under test based on the corrected window depth, and determining whether the fetal chromosome of the sample under test has aneuploidy based on the chimerism and the Z value.
3. The method according to claim 2, characterized in that: The degree of chimerism is the ratio of abnormal fetal cells to all fetal cells.
4. The method according to claim 3, characterized in that: The method for calculating the degree of fitting is as follows: FF i The following calculation method was used to obtain it: In the formula, q i That is, chimerism, FF i represents the relative fetal concentration of chromosome i, and FF represents the fetal concentration. This represents the average depth after chromosome i correction. The average depth of all corrected autosomes is represented by the value of i, which ranges from 1 to 22.
5. The method according to claim 1, characterized in that: The defined window is a number of windows divided according to the human reference genome, with no overlap between adjacent windows; The relative depth is the quotient of the sample's window depth in a set window divided by the average window depth of each autosome in the corresponding set window. The window depth is the number of unique alignment sequences in the human reference genome that the sample sequencing data in the set window are compared to.
6. The method according to claim 1, characterized in that: The method for selecting the reference set includes: calculating the relative depth of the test sample and each sample in the complete reference set within a set window; using the complete reference set to learn and obtain a dimensionality reduction model to reduce the relative depth; using the dimensionality reduction features to calculate the difference between the test sample and each sample in the complete reference set, and selecting several samples with the lowest difference as the selection reference set for the test sample.
7. The method according to claim 6, characterized in that: The dimensionality reduction model is a PCA model obtained by PCA learning using a complete reference set; the dimensionality reduction features are the dimensionality reduction features output after the relative depth is input into the PCA model.
8. The method according to claim 1 or 6, characterized in that: The dissimilarity is the Euclidean distance between the first n principal components, where n is a positive integer.
9. The method according to claim 8, characterized in that: The dissimilarity is the Euclidean distance between the first ten principal components, and the formula for calculating the dissimilarity is as follows. In the formula, diversityreference_j represents the degree of difference, and PC_i sample PC_ireference_j represents the features of the samples under test after relative depth dimensionality reduction, which is the feature of the samples in the complete reference set after relative depth dimensionality reduction.
10. The method according to claim 6, characterized in that: At least 50 samples with the lowest difference were selected as the reference set for the test samples.
11. The method according to claim 10, characterized in that: The reference set is selected by setting the relative depth of the window for autosomes other than chromosomes 13, 18, and 21.
12. The method according to claim 1, characterized in that: The correction baseline consists of an autosomal correction baseline and an X chromosome correction baseline. The autosomal correction baseline is used for window depth correction of autosomes or female fetuses' X chromosomes, and the X chromosome correction baseline is used for window depth correction of male fetuses' X chromosomes. The autosomal correction baseline is the average relative depth of the autosomes in the set window for all samples in the selected reference set; the X chromosome correction baseline is the average relative depth of the X chromosome in the set window for all female fetuses in the selected reference set.
13. The method according to claim 12, characterized in that: The window depth of the sample to be tested is corrected using a corrected baseline, specifically including: (1) Correction for autosomes or the X chromosome of female fetuses using the following formula, (2) The following formula is used to correct the X chromosome in male fetuses. In the above formula, UR n UR represents the window depth after baseline correction. a The window depth before baseline correction. 常 This is the baseline for autosomal correction. X The baseline for X chromosome correction is FF, where FF is the fetal concentration.
14. The method according to claim 12, characterized in that: The autosomal correction baseline is as follows: The X chromosome correction baseline is, Among them, baseline 常 This is the baseline for autosomal correction. X The baseline for X chromosome correction is defined as follows: reference refers to all samples selected from the reference set; female-reference refers to all female fetal samples selected from the reference set; n reference The number of windows used to select all samples in the reference set is nfemale-reference, where n is the number of windows used to select all female fetuses in the reference set, and d is the relative depth of the window.
15. The method according to claim 1, characterized in that: It also includes GC correction of the window depth and the use of the corrected window depth for relative depth calculation.
16. The method according to any one of claims 1-15, characterized in that: The method for calculating the Z-value includes: for each chromosome i to be detected, calculating the Z-value for the window UR of that chromosome and the UR of other autosomal windows of the same sample, sorting the obtained Z-values, removing the largest and smallest statistical values, and calculating the mean of the remaining Z-values to obtain the final Z-value, as detailed below. (j = 1...22, j ≠ i, after sorting, remove the first and last two values) in: Mean value of UR on chromosome i; Mean value of UR on chromosome j; SD i : Represents the standard deviation of the UR on chromosome i; SD j : Represents the standard deviation of the UR of chromosome j; L i : Indicates the number of windows used to divide chromosome i; L j : Indicates the number of windows used to divide chromosome j; Z i : Indicates the significance of aneuploidy on chromosome i, reflecting the difference from euploidy.
17. A method for baseline correction, characterized in that: This includes calculating the average relative depth of the samples in the selected reference set within a set window, using this as the correction baseline, and then using the correction baseline to correct the window depth of the sample to be tested. The selection reference set consists of several samples with the lowest difference from the sample to be tested, selected from a complete reference set composed of samples with known fetal chromosome information and corresponding maternal peripheral blood cfDNA sequencing data.
18. A method for screening a reference set for a sample to be tested, wherein the reference set is used for chromosome aneuploidy analysis, characterized in that: This includes calculating the relative depth of the test sample and each sample in the complete reference set within a set window; using the complete reference set for learning to obtain a dimensionality reduction model, and reducing the dimensionality of the relative depth; using the dimensionality reduction features to calculate the difference between the test sample and each sample in the complete reference set, and selecting several samples with the lowest difference as the selection reference set for the test sample. The complete reference set consists of several samples with known fetal chromosomal information and corresponding maternal peripheral blood cfDNA sequencing data.
19. An analytical device for chromosome aneuploidy, characterized in that: It includes a window depth calculation module, a baseline correction module, a Z-value calculation module, and a result judgment module; The window depth calculation module is used to calculate the window depth of the sample within a set window using the sequencing data and alignment results of the sample to be tested. The baseline correction module is used to calculate the average relative depth of the samples in the selected reference set within a set window, which is used as the correction baseline. The window depth of the sample to be tested is corrected using the correction baseline. The selected reference set consists of several samples with the lowest difference from the sample to be tested, which are selected from a complete reference set composed of samples with known fetal chromosome information and corresponding maternal peripheral blood cfDNA sequencing data. The Z-value calculation module is used to calculate the Z-value of the sample under test based on the corrected window depth. The result judgment module is used to determine whether the fetal chromosomes of the sample under test have aneuploidy based on the Z value.
20. A baseline correction device, characterized in that: Includes a baseline correction module; The baseline correction module is used to calculate the average relative depth of the samples in the selected reference set within a set window, which is used as the correction baseline. The window depth of the sample to be tested is corrected using the correction baseline. The selected reference set consists of several samples with the lowest difference from the sample to be tested, which are selected from a complete reference set composed of samples with known fetal chromosome information and corresponding maternal peripheral blood cfDNA sequencing data.
21. An apparatus for selecting a reference set from samples to be tested, characterized in that: Includes a reference set selection and filtering module; The reference set selection module is used to calculate the relative depth of the test sample and each sample in the complete reference set within the window; the complete reference set is used to learn and obtain a dimensionality reduction model to reduce the relative depth; the dimensionality reduction features are used to calculate the difference between the test sample and each sample in the complete reference set, and the samples with the lowest difference are selected as the reference set for the test sample. The complete reference set consists of several samples with known fetal chromosomal information and corresponding maternal peripheral blood cfDNA sequencing data.
22. An analytical device for chromosome aneuploidy, characterized in that: The device includes a memory and a processor; The memory is used to store programs; The processor is configured to implement the method for analyzing chromosomal aneuploidy as described in any one of claims 1-16, the method for baseline correction as described in claim 17, or the method for selecting a reference set by screening test samples as described in claim 18 by executing a program stored in the memory.
23. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a program that can be executed by a processor to implement the method for analyzing chromosomal aneuploidy as described in claims 1-16, the baseline correction method as described in claim 17, or the method for screening test samples and selecting a reference set as described in claim 18.