Method and device for detecting copy number variation of nuclear genome, equipment and storage medium

By identifying the gender and sequencing depth value of the sample to be tested in the nuclear genome copy number variation detection, combined with confidence interval comparison and spatial clustering, the problems of poor resolution and high false positive of CNV detection in the existing technology are solved, and efficient and accurate CNV detection is achieved.

CN115331730BActive Publication Date: 2025-10-17ZHENGZHOU JINYU CLINICAL TESTING CENT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210862798.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-20
Publication Date
2025-10-17
Estimated Expiration
2042-07-20

AI Technical Summary

Technical Problem

Existing nuclear genome copy number variation detection methods have problems in metagenomic detection, such as poor resolution, high false positive and false negative rates, complex reliance on auxiliary information, and low detection efficiency. In particular, it is difficult to accurately identify CNVs in low sequencing depth data.

Method used

By identifying the sex of the sample to be tested and comparing the sequencing depth values ​​and confidence intervals of the X chromosome and autosomes, the copy abnormality sites are automatically identified, spatial clustering and classification are performed, and adjacent copy variation sites are merged. Unsupervised machine learning methods are used to improve the accuracy and sensitivity of detection.

Benefits of technology

It realizes automated and highly accurate CNV detection in nuclear genome chromosomes, reduces the misjudgment rate, improves the accuracy of sex chromosome CNV detection and the stability of CNV fragments, and enhances the sensitivity and stability of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115331730B_ABST
    Figure CN115331730B_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of biological information detection, and discloses a nuclear genome copy number variation detection method, which can identify whether the sample to be detected is a female sample. If yes, the half value of the sequencing depth value of the X chromosome site and the sequencing depth value of the autosomal site are compared with the respective confidence interval to determine the copy abnormal site, the copy number value of the copy abnormal site is calculated for spatial clustering classification, and the normal copy class and the copy variation class are obtained. The copy abnormal site belonging to the copy variation class is determined as the copy variation site. Then, the copy variation sites with adjacent positions and the same variation type are combined to obtain CNV fragments. Therefore, the present application can automatically and accurately perform CNV detection of all nuclear genome chromosomes, improve the accuracy of sex chromosome CNV detection, and improve the resolution of RD sites while ensuring the stability, accuracy and sensitivity of the CNV fragments.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of biological information detection, and particularly relates to a nuclear genome copy number variation detection method, device, equipment and storage medium. BACKGROUND

[0002] The macrogenomic clinical detection technology is a new type of clinical technology for identifying and diagnosing microbial infection from the perspective of genetic material by using second-generation high-throughput sequencing. The microbial sequences (i.e. non-human sequences) contained in the high-throughput sequencing data are relatively limited, and more than 90% of the sequence content is human sequence. However, in the current clinical application, only a small amount of microbial sequence is used for infectious identification, and the large amount of human sequence still lacks application and analysis. The macrogenomic technology uses free nucleic acid for microbial identification, and tumor cells are easily detected in free nucleic acid due to their high metabolic rate. Free nucleic acid has become an important experimental sample for tumor screening and examination. If the macrogenomic technology is used for infectious detection at the same time, early screening and detection of tumors or genetic diseases can be performed, which not only solves the current disease problem, but also enables early intervention and treatment of positive results of tumor or genetic disease early screening, which has great significance for human health. However, so far, there are few algorithms and software that combine macrogenomic and tumor gene copy number variation (CNV) detection. Therefore, the simultaneous tumor or genetic disease screening work through the macrogenomic infection detection means is limited by technology and cannot be performed.

[0003] Currently, the mainstream theories for CNV detection include Read-Pair (RP) method, Split-read (SR) method, Read-Depth (RD) method and Assembly (AS) method. Among them, RP is the earliest algorithm, which uses the length distribution of double-end sequencing inserts to detect CNV, also known as pair end mapping (PEM) method. When the insert length is too long or too short, it represents that the genome has undergone structural variation. The SR method uses reads that can be aligned on one end and cannot be aligned on the other end to identify CNV. The other end cannot be aligned, which may be due to CNV. By splitting the individual reads, they can be correctly aligned to the reference genome, and the splitting point is the CNV breakpoint. The RD method uses the correlation between copy number and sequencing depth of the corresponding region for analysis, and the basic model is that the sequencing depth of the deletion region is relatively low, and the sequencing depth of the insertion region is relatively high. The AS method uses short sequences obtained by sequencing to assemble, compares the assembled contig with the reference genome, and determines the region where structural variation has occurred.

[0004] Because the macro gene detection is a 50-75 bp second generation single end sequencing technology, it does not conform to the current mainstream software, and the site difference introduced by the regional population site and different experimental conditions cannot be well corrected, so that its resolution and results are not good in the current software, and there is still no algorithm that can be used for macro gene CNV detection. According to the mainstream theory of CNV detection, the RP and SR methods depend on the double end sequencing technology, which is not suitable for macro gene data, and the relative algorithm is not accurate enough. The AS method depends on the flux and sequencing coverage technology, and the coverage of the macro gene is quite different from the genome assembly technology. This method cannot be applied to the macro gene data. The three theoretical methods do not conform to the macro gene sequencing data, and only a small amount of software under the RD method can be applied.

[0005] However, the RD method depends on the flux and sequencing depth technology, and needs a higher and more stable depth change to identify CNV, so that the application of this method in the macro gene data with low sequencing depth will introduce many false positive sites. And the analysis model of the traditional RD method uses the same RD resolution and CNV resolution, and the too small resolution will lead to too strong data dispersion, too high false positive, and too large resolution will lead to CNV averaging, resulting in false negative results, and the edge position of CNV may also form a transition type due to the interval coverage in RD calculation, affecting the recognition and judgment of CNV, so that the accuracy and sensitivity of CNV detection are not enough.

[0006] In addition, the current CNV detection tool needs to input a large amount of auxiliary information for genetic disease early screening, including but not limited to variant group and normal group information, patient gender, step, chromosome segment, reference index and the like. Among them, for the recognition strategy of sex chromosome CNV variation, the first kind is based on the direct depth comparison with autosomes, and the second kind is dependent on manual input parameters for gender grouping analysis. The first strategy can only detect single sample, and has more false positives, which is not accurate enough. The second strategy needs to manually input parameters, which is relatively complex. However, the clinical sequencing data is relatively complex, the auxiliary information is not clear, and the detection efficiency is low. SUMMARY

[0007] The purpose of the present application is to provide a nuclear genome copy number variation detection method and device, equipment and storage medium, which can automatically and accurately detect CNV of all nuclear genome chromosomes, and improve the accuracy and sensitivity of CNV detection.

[0008] The first aspect of the present application discloses a nuclear genome copy number variation detection method, comprising:

[0009] The sequencing depth value of a plurality of specified sites is determined from the sequencing data of the sample to be tested; wherein each of the specified sites corresponds to a confidence interval and an RD set mean value;

[0010] determining whether the sample to be tested is a female sample according to the sequencing depth values of the multiple specified sites;

[0011] identifying X chromosome sites and autosomal sites in the multiple specified sites if the sample to be tested is a female sample;

[0012] determining the X chromosome site as a low-copy abnormal site if the half value of the sequencing depth value of any X chromosome site is less than the confidence interval corresponding to the X chromosome site, and determining the X chromosome site as a high-copy abnormal site if the half value of the sequencing depth value of any X chromosome site is greater than the confidence interval corresponding to the X chromosome site;

[0013] determining the autosomal site as a low-copy abnormal site if the sequencing depth value of any autosomal site is less than the confidence interval corresponding to the autosomal site, and determining the autosomal site as a high-copy abnormal site if the sequencing depth value of any autosomal site is greater than the confidence interval corresponding to the autosomal site;

[0014] taking the low-copy abnormal sites and the high-copy abnormal sites as copy abnormal sites;

[0015] calculating the copy number values of the copy abnormal sites according to the RD setting mean values and the sequencing depth values corresponding to the copy abnormal sites, wherein the copy number values and the sequencing depth values are in a positive correlation;

[0016] performing spatial clustering classification on the copy number values of all the copy abnormal sites to obtain a normal copy class and a copy variation class, and the copy variation class includes two copy variation subclasses, which are a high-copy variation subclass and a low-copy variation subclass;

[0017] determining the high-copy abnormal sites belonging to the high-copy variation subclass and the low-copy abnormal sites belonging to the low-copy variation subclass as copy variation sites, respectively;

[0018] merging the copy variation sites that are adjacent in position and belong to the same copy variation subclass to obtain copy variation fragments.

[0019] The second aspect of the present application discloses a nuclear genome copy number variation detection device, comprising:

[0020] a depth determination unit configured to determine sequencing depth values of multiple specified sites from sequencing data of a sample to be tested, wherein each specified site corresponds to a confidence interval and an RD setting mean value;

[0021] a gender determination unit configured to determine whether the sample to be tested is a female sample according to the sequencing depth values of the multiple specified sites;

[0022] The site recognition unit is configured to, when the gender judgment unit judges that the sample to be tested is a female sample, recognize an X chromosome site and an autosome site in the plurality of specified sites;

[0023] The abnormality detection unit is configured to, when the half value of the sequencing depth value of any X chromosome site is less than the confidence interval corresponding to the X chromosome site, determine that the X chromosome site is a low-copy abnormal site; and when the half value of the sequencing depth value of any X chromosome site is greater than the confidence interval corresponding to the X chromosome site, determine that the X chromosome site is a high-copy abnormal site.

[0024] The abnormality detection unit is further configured to, when the sequencing depth value of any autosome site is less than the confidence interval corresponding to the autosome site, determine that the autosome site is a low-copy abnormal site; and when the sequencing depth value of any autosome site is greater than the confidence interval corresponding to the autosome site, determine that the autosome site is a high-copy abnormal site; and take both the low-copy abnormal site and the high-copy abnormal site as copy abnormal sites.

[0025] The calculation unit is configured to calculate the copy number value of each copy abnormal site according to the RD mean value and the sequencing depth value corresponding to each copy abnormal site; wherein the copy number value and the sequencing depth value are in a positive correlation relationship.

[0026] The clustering unit is configured to perform spatial clustering classification on the copy number values of all the copy abnormal sites to obtain a normal copy class and a copy variation class, and the copy variation class includes two copy variation sub-classes, which are a high-copy variation sub-class and a low-copy variation sub-class.

[0027] The variation determination unit is configured to determine the high-copy abnormal site belonging to the high-copy variation sub-class and the low-copy abnormal site belonging to the low-copy variation sub-class as copy variation sites, respectively.

[0028] The merging unit is configured to merge the copy variation sites that are adjacent in position and belong to the same copy variation sub-class to obtain a copy variation fragment.

[0029] The third aspect of the present application discloses an electronic device, comprising a memory storing executable program codes and a processor coupled with the memory; the processor invokes the executable program codes stored in the memory, and is configured to execute the nuclear genome copy number variation detection method disclosed in the first aspect.

[0030] The fourth aspect of the present application discloses a computer readable storage medium storing a computer program, wherein the computer program causes a computer to execute the nuclear genome copy number variation detection method disclosed in the first aspect.

[0031] The present application has the beneficial effect that the provided nucleic genome copy number variation detection method and device, equipment and storage medium mainly comprise: identifying whether the sample to be detected is a female sample, if so, comparing half of the sequencing depth value of the X chromosome site and the sequencing depth value of the autosomal site with the respective confidence interval to determine the copy abnormal site, calculating the copy number value of the copy abnormal site for spatial clustering classification to obtain a normal copy class and a copy variation class, and then filtering out the copy abnormal site belonging to the normal copy class to reduce the misjudgment rate; then, the copy abnormal site belonging to the copy variation class is determined as a copy variation site; and then, the copy variation sites with adjacent positions and the same variation type are combined to obtain a CNV fragment.

[0032] Therefore, the present application can automatically and accurately perform CNV detection of all nucleic genome chromosomes, has high detection efficiency, and can improve the accuracy of sex chromosome CNV detection; meanwhile, by separating the interval size of the RD site and the interval size of the CNV fragment, different resolutions are used for the RD site and the CNV fragment, the stability of the CNV fragment is improved by improving the stability of the RD site, so that the resolution of the RD site can be improved while the stability of the CNV fragment is ensured, the copy variation site is double-identified by using the unsupervised machine learning method to cluster and classify the copy abnormal site and compare with the confidence interval, which can further improve the identification accuracy of the copy variation site, and then improve the accuracy, sensitivity and stability of the CNV detection. BRIEF DESCRIPTION OF DRAWINGS

[0033] The drawings herein show specific examples of the technical solutions of the present application, and constitute part of the specification together with the specific embodiments, for explaining the technical solutions, principles and effects of the present application.

[0034] Unless specifically stated or defined otherwise, the same reference signs in different drawings represent the same or similar technical features, and different reference signs may also be used to represent the same or similar technical features.

[0035] Figure 1 is a flowchart of a nucleic genome copy number variation detection method disclosed by an embodiment of the present application;

[0036] Figure 2 is a basic structure diagram of metagenomic sequencing data disclosed by an embodiment of the present application;

[0037] Figure 3 is a structural schematic diagram of a nucleic genome copy number variation detection device disclosed by an embodiment of the present application;

[0038] Figure 4 is a structural schematic diagram of an electronic device disclosed by an embodiment of the present application.

[0039] Reference Signs List:

[0040] 301, depth determination unit; 302, gender determination unit; 303, site identification unit; 304, anomaly detection unit; 305, calculation unit; 306, clustering unit; 307, variation determination unit; 308, merging unit; 401, memory; 402, processor. DETAILED DESCRIPTION

[0041] Unless otherwise defined or specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. In the context of the technical solutions of the present application in a real scenario, all technical and scientific terms used herein can also have meanings corresponding to the purposes of implementing the technical solutions of the present application. The terms "first, second, …" used herein are only used for distinguishing names and do not represent specific quantities or sequences. The term "and / or" used herein includes any and all combinations of one or more related listed items.

[0042] Unless otherwise specified or defined, "the", "this" used herein refers to the technical features or technical contents mentioned or described before the corresponding position, which can be the same as or similar to the technical features or technical contents mentioned.

[0043] Unless otherwise specified or defined, "the", "this" used herein refers to the technical features or technical contents mentioned or described before the corresponding position, which can be the same as or similar to the technical features or technical contents mentioned.

[0044] It should be noted that the nucleic genome copy number variation detection method and device disclosed in the present application can be applied to low sequencing depth data such as metagenomic data, and can also be applied to normal sequencing high sequencing depth data, such as whole exon sequencing data, whole gene sequencing data, and chip sequencing data. In the embodiments of the present application, metagenomic sequencing data is taken as an example for illustration, which should not be considered as a limitation of the present application.

[0045] (I) Baseline establishment process

[0046] Preferably, in the embodiments of the present application, before the nucleic genome copy number variation detection of the sequencing data of the sample to be tested, a large number of clinical sample data of the same batch or condition and source can be used in combination with the machine learning method and variance index to filter unstable and unreasonable noise sites, the K-mean algorithm can be used to identify centromere, mitochondrial and repetitive sequence sites in the sites, and the variance can be used to remove fluctuating sites, and the stable sites that can be used for CNV detection are retained, and then the baseline of each stable site is established.

[0047] Specifically, before performing the nucleic genome copy number variation detection on the sequencing data of the to-be-tested sample, the following steps S101-S109 can be performed:

[0048] S101, obtain sequencing data of N1 training samples; the sequencing data of each training sample includes sequencing depth values of M1 candidate sites.

[0049] In the present application, the site actually refers to a site interval, which refers to a chromosome segment, including X chromosome, Y chromosome and autosome, and the sequencing depth value corresponds to the Read-Depth (RD) data of the chromosome segment. Wherein, the name of each candidate site carries identification information, which can identify the chromosome category to which the candidate site belongs, that is, identify whether the candidate site is located on the X chromosome, Y chromosome or autosome.

[0050] Since the unsupervised clustering method is performed, the obtained sequencing data can only include raw fastq format sequencing data, and does not require clinical verification and follow-up data, for example, the training sample is the result data of CNV detection positive or negative. After obtaining the raw fastq format sequencing data of N1 training samples, further filtering, de-adapter, alignment, deduplication and other preprocessing work are performed on the raw fastq format sequencing data according to the pre-designed comparison and quality control process.

[0051] Then, according to the preset block interval gradient parameter, the size of the site interval (i.e. the chromosome segment) is determined, and the chromosomes in the sequencing data are divided into M1 chromosome segments according to the size of the site interval, and each chromosome segment corresponds to a candidate site. For example, assuming that the preset block interval gradient parameter is 10KB, the stable site interval size under this process is 10KB, based on which the chromosomes are divided into M1 chromosome segments, and the size of each chromosome segment is 10KB.

[0052] After obtaining the M1 candidate sites of each training sample, the sequencing depth values of the M1 candidate sites of the N1 training samples can be preferably subjected to data standardization processing, so that the sequencing depth values of the N1 training samples under the same candidate site with large differences fall within the set range [0, 1], thereby eliminating the batch differences of different genome batches and sample sequencing depth, and reducing the adverse effects caused by the singular characteristic data with too large or too small values in the sequencing data.

[0053] After step S101, it is necessary to identify whether each training sample is a female training sample or a male training sample. In some other possible embodiments, it can be determined that the training sample is a female training sample if there is at least one candidate site with a sequencing depth value of 0, and it is determined that the training sample is a male training sample if there is no candidate site with a sequencing depth value of 0.

[0054] Preferably, in the embodiment, the Y chromosome score and the X chromosome score of each training sample are calculated by the following formulas (1) and (2) respectively:

[0055] Yscore = INT (2 * RDy / RDmean) (1)

[0056] Xscore = INT (2 * RDx / RDmean) (2)

[0057] wherein Yscore is the Y chromosome score of the sample, Xscore is the X chromosome score of the sample, RDy is the average RD value of all candidate sites on the Y chromosome of the sample, RDx is the average RD value of all candidate sites on the X chromosome of the sample, RDmean is the average RD value of all candidate sites of the sample, and INT is an integer function for rounding off the final calculation value in the bracket.

[0058] Then, according to the Y chromosome score and the X chromosome score of each training sample, the gender information of each training sample is identified, and N1 training samples are divided into Y1 female training samples and Y2 male training samples according to the gender information of each training sample, so as to obtain a female sample set and a male sample set.

[0059] If the Yscore of any training sample is 0 and / or the Xscore of any training sample is 2, it is determined that the training sample is a 2n normal female sample; if the Yscore of any training sample is 1 and / or the Xscore of any training sample is 1, it is determined that the training sample is a 2n normal male sample.

[0060] If there is any other condition, it is determined that the training sample is an abnormal sample with an important chromosome aneuploidy mutation, for example, when Yscore = 1 and Xscore = 2, it is determined that the training sample has X chromosome aberration, which may be related to diseases such as Edwards syndrome. Optionally, when it is determined that the training sample is an abnormal sample, the training sample can be removed.

[0061] S102, unsupervised clustering classification is performed on the sequencing depth values of all candidate sites to obtain a plurality of classification categories.

[0062] In step S102, interval quality control is performed on the sequencing depth values of the N1XM1 candidate sites, mainly by using a K-mean algorithm to identify a classification tree, so as to unsupervisedly cluster and classify the sequencing depth values of the N1XM1 candidate sites into multiple classification categories.

[0063] In S103, the first sequencing depth average of each classification category is calculated, and a noise category is identified from the multiple classification categories according to the first sequencing depth average.

[0064] In step S103, the sequencing depth values of all candidate sites included in each classification category are averaged to obtain the first sequencing depth average of each classification category, and then based on the first sequencing depth average, the classification category with abnormal mean value can be identified, so that the classification category with abnormal mean value is determined as a noise category and is excluded. Thus, background noise sites such as mitochondrial sites, repetitive sequence sites, and centromere sites can be identified by machine learning clustering algorithm, and efficient and detailed noise reduction can be achieved.

[0065] Specifically, the classification category with a first sequencing depth average less than a first specified threshold is determined as a centromere category; and the classification category with a first sequencing depth average greater than a second specified threshold is determined as a repetitive sequence category; and the classification category with a first sequencing depth average greater than a third specified threshold is determined as a mitochondrial category; and the centromere category, the repetitive sequence category, and the mitochondrial category are all regarded as noise categories.

[0066] The first specified threshold is less than the second specified threshold, and the second specified threshold is less than the third specified threshold. The first specified threshold, the second specified threshold, and the third specified threshold can be pre-set by the developer according to actual needs, and in some preferred embodiments, the first specified threshold, the second specified threshold, and the third specified threshold can also be determined according to the overall sequencing depth average of the multiple classification categories.

[0067] For example, after calculating the first sequencing depth average of each classification category, the first sequencing depth averages of all classification categories can be further averaged to obtain a second sequencing depth average of all classification categories, and then one tenth of the second sequencing depth average is taken as the first specified threshold; and five times the second sequencing depth average is taken as the second specified threshold; and twenty times the second sequencing depth average is taken as the third specified threshold. That is, assuming that the second sequencing depth average of all classification categories is C U , the first specified threshold is C U / 10, the second specified threshold is C U *5, and the third specified threshold is C U *20.

[0068] S104, eliminate the candidate sites included in the noise category to obtain M2 set sites. M2 is less than M1.

[0069] In the clustering results, the candidate sites included in the centromere category are typical noise sites due to high sequencing difficulty and are generally irrelevant to tumor CNV variation. The candidate sites included in the repeat sequence category have high copy number and are unstable in genomic location, and have no clear diagnostic significance, so they are also removed as noise sites; the candidate sites included in the mitochondria category differ greatly in different cells and become important noise points in low sequencing depth data, so after eliminating the candidate sites of these noise categories, the remaining M2 set sites include but are not limited to set sites located on the Y chromosome, set sites located on the X chromosome, and set sites located on the autosome, which are used as stable CNV detection targets.

[0070] S105, identify the X chromosome set sites, Y chromosome set sites, and autosome set sites in the M2 set sites.

[0071] In step S105, the X chromosome set site refers to the set site located on the X chromosome, the Y chromosome set site refers to the set site located on the Y chromosome, and the autosome set site refers to the set site located on the autosome. After identifying each X chromosome set site, Y chromosome set site, and autosome set site in the M2 set sites, the baseline of each X chromosome set site and Y chromosome set site can be optimized using the female sample set and the male sample set, i.e., calculating the haploid baseline, and steps S106 and S107 are performed respectively.

[0072] In calculating the haploid baseline, for the X chromosome set site located on the X chromosome, the half value of the sequencing depth value of the female training sample at the X chromosome set site and the original value of the sequencing depth value of the male training sample at the X chromosome set site are used for calculation; for the Y chromosome set site located on the Y chromosome, only the original value of the sequencing depth value of the male training sample at the Y chromosome set site is used for calculation.

[0073] S106, calculate the RD set mean and variance of each X chromosome set site according to the half value of the sequencing data of Y1 female training samples and the sequencing data of Y2 male training samples.

[0074] The RD set mean and variance of each X chromosome set site are calculated by the following formula:

[0075]

[0076]

[0077] wherein, RDsetmeanis the RD set mean value of the wthX chromosome set position, Y1 is the number of female training samples, Y2 is the number of male training samples, RD wy Sis the sequencing depth value of the ythtraining sample on the mthY chromosome set position; S sw 2 RDsetvaris the variance of the mthY chromosome set position.

[0078] S107, according to the sequencing data of N1 training samples, the RD set mean value and the variance of each Y chromosome set position and the RD set mean value and the variance of each autosomal set position are calculated.

[0079] wherein, the RD set mean value and the variance of each Y chromosome set position are calculated by the following formula:

[0080]

[0081]

[0082] wherein, RDsetmeanis the RD set mean value of the mthY chromosome set position, RD my Sis the sequencing depth value of the ythtraining sample on the mthY chromosome set position; S ym 2 RDsetvaris the variance of the mthY chromosome set position.

[0083] As for the RD set mean value and the variance of the autosomal set position, the average value of the sequencing depth value (preferably the normalized RD data) corresponding to each autosomal set position of N1 training samples is calculated as the RD set mean value corresponding to the autosomal set position, and the variance of the sequencing depth value corresponding to each autosomal set position of N1 training samples is calculated. After calculating the RD set mean value and the variance of each X chromosome set position, Y chromosome set position and autosomal set position, the RD set mean value and the variance of M2 set positions are obtained, and the RD set mean value between each set position is different.

[0084] S108, all or part of the set positions in M2 set positions are determined as designated positions.

[0085] Among them, all the set positions can be designated as designated positions, that is, stable positions for reference when detecting the test sample subsequently. M2 set positions can also be filtered again, and a small number of set positions with large variance in the sample dimension are removed. For example, after calculating the variance of each set position under N1 training samples, the set positions with variance less than the specified variance threshold can be determined as designated positions to obtain M3 designated positions, and the set positions with variance greater than or equal to the specified variance threshold are removed.

[0086] It should be noted that M3 is less than or equal to M2. When the variances of the M2 set positions are all less than the specified variance threshold, the M3 designated positions include all the M2 set positions, and when there are some set positions in the M2 set positions whose variances are greater than or equal to the specified variance threshold, the M3 designated positions include some set positions.

[0087] S109, according to the RD set mean and variance, calculate the confidence interval of each designated position.

[0088] In calculating the confidence interval of each designated position, the following formula model (7) can be referred to:

[0089]

[0090] In the formula, S represents the RD set mean corresponding to the i-th designated position, and S i 2 S represents the variance corresponding to the i-th designated position, since the data conforms to the normal distribution, is 1.96.

[0091] Therefore, in combination with the RD set mean and variance of each X chromosome set position, Y chromosome set position and autosomal set position in the above, the confidence interval of each designated position can be determined as:

[0092] If the designated position is an X chromosome set position, the confidence interval is

[0093] If the designated position is a Y chromosome set position, the confidence interval is

[0094] If the designated position is an autosomal set position, the confidence interval is

[0095] In the formula, S represents the RD set mean corresponding to the i-th autosomal set position, and S i 2 S represents the variance corresponding to the i-th autosomal set position.

[0096] That is to say, after the calculation with N1 training samples is completed, each designated site corresponds to a baseline, the content of the baseline includes a confidence interval and a RD set mean value, the confidence interval is used to define the normal value range of the sequencing depth value of the designated site, and the points falling outside the confidence interval will be identified as abnormal values. By setting ci=95% as the confidence interval of the designated site, a dedicated residual range is established for each designated site, and as the sample size accumulates, the residual range is more accurate and does not depend on known outcome samples.

[0097] The steps S101-S109 are implemented, a baseline is established by using the identified relatively stable set sites, the baseline is used to remove the differences between different sites, remove the noise points of uneven site depth introduced by GC content, population genome source, experimental method and quality control method, and the like, the background sites can be well processed for heterogeneity, the noise brought by batch effect and local human characteristics can be better removed, and it is ensured that the identified copy number variation is a real change. In addition, gender recognition can be performed on each training sample to distinguish the designated sites located on different sex chromosomes, and secondary correction is performed on the baseline of the designated sites, so that when early screening of genetic diseases is performed by using sex chromosomes, false positive signals caused by euploidy changes of sex chromosomes can be avoided, and the accuracy of CNV detection of sex chromosomes can be further improved.

[0098] (II) Detection process of the to-be-tested sample / training sample

[0099] As shown in Figure 1 The embodiment of the present application discloses a nuclear genome copy number variation detection method, which comprises the following steps S201-S210:

[0100] S201, determining the sequencing depth values of a plurality of designated sites from the sequencing data of the to-be-tested sample.

[0101] In the present application, the designated site refers to a pre-specified chromosome segment, which can be specified by the developer according to actual needs in some other possible embodiments, and in the present embodiment, the M3 designated sites determined in the establishment process of the baseline are determined. In actual sequencing, the to-be-tested sample may not measure all the designated M3 sites, therefore, in step S201, the plurality of designated sites of the to-be-tested sample determined includes all or part of the M3 designated sites. Then, the baseline content corresponding to each designated site, i.e., the corresponding confidence interval and RD set mean value, can be retrieved according to the determination of each designated site of the to-be-tested sample.

[0102] S202, judging whether the sample to be tested is a female sample according to the sequencing depth values of the multiple specified sites. If the sample to be tested is a female sample, steps S203-S204 are executed, and then step S206 is turned to; if the sample to be tested is not a female sample, it is determined to be a male sample, and then step S205 is turned to.

[0103] wherein, it can be determined that the sample to be tested is a female sample if the sequencing depth value of at least one specified site in the sample to be tested is 0, otherwise, it is determined to be a male sample. Preferably, in the embodiment, the Y chromosome score and the X chromosome score of the sample to be tested are calculated by the above formulas (1) and (2) respectively, if the Y chromosome score of the sample to be tested is 0 and / or the X chromosome score of the sample to be tested is 2, it is determined that the sample to be tested is a female sample, if the Y chromosome score of the sample to be tested is 1 and / or the X chromosome score of the sample to be tested is 1, it is determined that the sample to be tested is a male sample.

[0104] S203, identifying the X chromosome sites and the autosomal sites in the multiple specified sites.

[0105] wherein, after it is determined that the sample to be tested is a female sample, the Y chromosome sites in the multiple specified sites of the female sample are removed, and the X chromosome sites and the autosomal sites are left.

[0106] S204, if the half value of the sequencing depth value of any X chromosome site is less than the confidence interval corresponding to the X chromosome site, it is determined that the X chromosome site is a low-copy abnormal site; if the half value of the sequencing depth value of any X chromosome site is greater than the confidence interval corresponding to the X chromosome site, it is determined that the X chromosome site is a high-copy abnormal site; if the sequencing depth value of any autosomal site is less than the confidence interval corresponding to the autosomal site, it is determined that the autosomal site is a low-copy abnormal site; if the sequencing depth value of any autosomal site is greater than the confidence interval corresponding to the autosomal site, it is determined that the autosomal site is a high-copy abnormal site.

[0107] When the sequencing depth value of the female sample at each X chromosome site is compared with the confidence interval corresponding thereto, the sequencing depth value of each X chromosome site is divided by 2 before comparison. For example, the sequencing depth value of one of the X chromosome sites w1 of the female sample is RD w1 , then the X chromosome site w1 is identified as a low-copy abnormal site when , and the specified site w1 is identified as a high-copy abnormal site when . Wherein, is the mean value of the RD corresponding to the X chromosome site w1, S xw1 2 is the variance of the RD corresponding to the X chromosome site w1.

[0108] S205, if the sequencing depth value of any specified site is less than the confidence interval corresponding to the specified site, it is determined that the specified site is a low-copy abnormal site; when the sequencing depth value of any specified site is greater than the confidence interval corresponding to the specified site, it is determined that the specified site is a high-copy abnormal site.

[0109] For each specified site of the male sample, or the sequencing depth value on the autosomal site of the female sample, the sequencing depth value is compared with the corresponding confidence interval. The original value can be used for comparison. Specifically, if the sequencing depth value is less than the corresponding confidence interval, it is determined that the specified site is a low-copy abnormal site; if the sequencing depth value is greater than the corresponding confidence interval, it is determined that the specified site is a high-copy abnormal site.

[0110] For example, assuming that the sequencing depth value of the first specified site of the male sample is RD1, if the first specified site is identified as a low-copy abnormal site; if the first specified site is identified as a high-copy abnormal site.

[0111] S206, the low-copy abnormal site and the high-copy abnormal site are both regarded as copy abnormal sites.

[0112] The sequencing depth value of each specified site is compared with the corresponding confidence interval. As long as the sequencing depth value is not located in the confidence interval corresponding to the specified site, it is determined that the specified site is a copy abnormal site. After traversing multiple specified sites, multiple low-copy abnormal sites and high-copy abnormal sites can be determined, thereby obtaining a low-copy abnormal site set and a high-copy abnormal site set.

[0113] The above high and low copy abnormal sites may be copy number variation sites or error points under ecological analysis (α=0.05). A single abnormal site is not very clear. In order to further improve the accuracy, the copy number value of the above identified abnormal copy abnormal site can be calculated for further analysis, that is, step S207 is executed.

[0114] S207, according to the mean value and the sequencing depth value corresponding to each copy abnormal site, the copy number value of each copy abnormal site is calculated. The copy number value is positively correlated with the sequencing depth value.

[0115] In this embodiment, if the copy abnormal site is an X chromosome setting site, half of the sequencing depth value is used to calculate the copy number value, and if the copy abnormal site is a Y chromosome setting site or an autosomal setting site, the original value of the sequencing depth value is used to calculate the copy number value. For example:

[0116] If the copy abnormal site is an X chromosome setting site, its copy number value is calculated by the following formula (8):

[0117]

[0118] CPj represents the copy number value of the jth copy abnormal site located on the X chromosome, RD j RDj represents the sequencing depth value of the jth copy abnormal site located on the X chromosome, j RDj represents the sequencing depth value of the jth copy abnormal site located on the X chromosome, RDj represents the RD set mean value of the jth copy abnormal site located on the X chromosome.

[0119] If the copy abnormal site is a Y chromosome set site or an autosome set site, the copy number value thereof is calculated by the following formula (9):

[0120]

[0121] CPi represents the copy number value of the ith copy abnormal site located on the Y chromosome or autosome, RD i RDi represents the sequencing depth value of the ith copy abnormal site located on the Y chromosome or autosome, i RDi represents the sequencing depth value of the ith copy abnormal site located on the Y chromosome or autosome, RDi represents the RD set mean value of the ith copy abnormal site located on the Y chromosome or autosome.

[0122] It should be noted that the copy number value of each copy abnormal site is not limited to the calculation method of directly taking the ratio between half of the sequencing depth value of the copy abnormal site or the original value and the RD set mean value as the copy number value of the copy abnormal site as described in the above formula (8) or (9). In some other possible embodiments, other alternative formulas or obvious variant formulas equivalent to the above formula (8) or (9) can also be used to calculate the copy number value, which can also ensure that the copy number value is positively correlated with the sequencing depth value. For example, when the copy abnormal site is an autosome set site, the calculation formula of the copy number value thereof can also be and the like.

[0123] S208, spatially clustering and classifying the copy number values of all copy abnormal sites to obtain a normal copy class and a copy variation class, wherein the copy variation class includes two copy variation subclasses, i.e., a high copy variation subclass and a low copy variation subclass.

[0124] Specifically, in the present embodiment, the sample conclusion is not pre-known, and unsupervised clustering is used for identification. The preselected range K value is from 1 to 10, and the optimal K value is searched by cross-validation method. After a large number of sample verifications, the best K value of the prior clustering is 4. Therefore, in the present embodiment, K = 4 is set, and the K nearest neighbor (KNN) method is used to naturally cluster the copy number values of all copy abnormal sites to obtain 4 classification categories. Then, the average value of the copy number (CP) of each classification category is calculated, and the 4 classification categories are sorted according to the CP average value. According to the order from high to low of the CP average value, the high copy variation subcategory, the high copy discrete category, the low copy discrete category, and the low copy variation subcategory are identified, respectively. Among them, the high copy discrete category and the low copy discrete category can be regarded as normal copy categories, which are used to filter out copy abnormal sites belonging to normal copy categories and reduce the misjudgment rate.

[0125] S209, the high copy abnormal sites belonging to the high copy variation subcategory and the low copy abnormal sites belonging to the low copy variation subcategory are determined as copy variation sites, respectively.

[0126] Then, the copy abnormal sites included in the identified high copy variation subcategory are taken as the first intersection with the high copy abnormal site set obtained by comparing with the confidence interval, and the copy abnormal sites in the first intersection are determined as high copy variation sites. Similarly, the copy abnormal sites included in the identified low copy variation subcategory are taken as the second intersection with the low copy abnormal site set obtained by comparing with the confidence interval, and the copy abnormal sites in the second intersection are determined as low copy variation sites. That is, if any high copy abnormal site belongs to the high copy variation subcategory, the high copy abnormal site is determined as a high copy variation site, and if any low copy abnormal site belongs to the low copy variation subcategory, the low copy abnormal site is determined as a low copy variation site. Among them, the high copy variation sites and the low copy variation sites are collectively referred to as copy variation sites.

[0127] By using the K nearest neighbor data clustering method to cluster and classify the copy abnormal sites and compare with the baseline confidence interval, the copy variation sites can be identified twice, which can further improve the identification accuracy of the copy variation sites. At the same time, the copy number variation sites can be well clustered by using the spatial clustering, thereby realizing the automatic identification of mutation types such as copy number increase and decrease, more accurately identifying high copy variation sites and low copy variation sites, and further improving the accuracy, sensitivity and stability of CNV detection.

[0128] S210, the copy variation sites with adjacent positions and belonging to the same copy variation subcategory are merged to obtain copy variation fragments.

[0129] wherein high copy variation sites belonging to the same high copy variant class and being adjacent in position are merged to obtain a high copy variation fragment; and low copy variation sites belonging to the same low copy variant class and being adjacent in position are merged to obtain a low copy variation fragment. The high copy variation fragment and the low copy variation fragment are collectively referred to as a copy variation fragment.

[0130] Specifically, copy variation sites of the same variation type (i.e., belonging to the same high copy variant class or low copy variant class) are continuously merged within a chromosome, i.e., chromosome segments of the same variation type and being adjacent in position are merged. In this embodiment, a sliding window is used to scan the multiple copy variation sites. The length of the sliding window should cover at least 2 copy variation sites. In this embodiment, the interval of each site is 10 KB, and thus the length of the sliding window can be set to 20 KB. The moving step of the sliding window is 10 KB, i.e., the sliding window moves by a distance of 1 copy variation site each time. There are always 2 copy variation sites in the window. After each movement of the sliding window, it is determined whether the latter copy variation site in the sliding window belongs to the same copy variant class as the former copy variation site adjacent thereto. If yes, the latter copy variation site and the former copy variation site adjacent thereto are merged.

[0131] For example, assuming that the multiple copy variation sites are A1, A2, A3, …, A8, A9, …, A j , the sliding window first selects {A1, A2}. If A2 belongs to the same copy variant class as A1, the two are merged. Then the sliding window moves by a step of 10 KB to {A2, A3}. If A3 belongs to the same copy variant class as A2, the two are merged. This process is repeated until all copy variation sites are traversed.

[0132] In this process, it is assumed that A1 to A8 are merged, then the sliding window moves to {A8, A9}. If it is determined that A9 belongs to a different copy variant class from A8, and the total step of the sliding window moving from A1 does not exceed 100 KB, A9 can be regarded as the first non-copy variation site appearing in the 100 KB step, and A9 is still regarded as a linked site and linked with A8. Then, if it is determined that A 10 belongs to a different copy variant class from A9, it is determined that the current merged copy variation fragment is interrupted, and A 10 will not be merged. If it is determined that A 10 belongs to the same copy variant class as A9, A 10Linkage merge with A9. Based on this example, if continuously merged to A j-1 , and it is determined that A j and A j-1 belong to different copy variant classes, it is necessary to determine the distance between A j and the last-appeared non-copy variant site A9, if the distance between A j and A9 exceeds 100 KB, A j may still be considered as a linkage site and be merged with A j-1 ; otherwise, if the distance between A j and A9 does not exceed 100 KB, it is determined that the current-merged copy variant fragment is interrupted.

[0133] It can be seen that, when the positions of multiple copy variant sites of the same variant type are adjacent, the linkage algorithm is used to extend the sites, the copy variant sites of the same variant type and adjacent in position are merged, more accurate CNV fragments are obtained, the interval size of the RD site and the interval size of the CNV fragment are separated, different resolutions are used for the RD site and the CNV fragment, the stability of the RD site is improved to improve the stability of the CNV fragment, so that the resolution of the RD site is improved while the stability of the CNV fragment is ensured, data structure problems such as large fragment depth instability and unit point dispersion are prevented, the copy variant site is double-recognized by using the unsupervised machine learning method to cluster and classify the copy variant site and compare the copy variant site with the baseline confidence interval, the recognition accuracy of the copy variant site is further improved, and the accuracy, sensitivity and stability of CNV detection are further improved.

[0134] Further, all the calculated copy variant fragments can be screened by setting a copy variant fragment threshold, so as to further improve the accuracy and further plot. Specifically, in the embodiment of the present application, after step S210 is performed, the following steps S211-S213 can also be performed:

[0135] S211, if the size of the copy variant fragment reaches a first fragment threshold, the copy variant fragment is determined as a target copy variant fragment; if the size of the copy variant fragment reaches a second fragment threshold and is less than the first fragment threshold, the copy variant fragment is determined as a suspected copy variant fragment.

[0136] For example, the first fragment threshold is set to 10 MB, and the second fragment threshold is set to 1 MB, that is, the copy variant fragment with a size of 1 MB and less than 10 MB is a suspected copy variant fragment, and the copy variant fragment with a size of 10 MB is a target copy variant fragment.

[0137] S212, obtaining the copy number values of each target copy abnormal site included in the target copy variant fragment, and performing log conversion on the copy number values of the target copy abnormal site by the following formula:

[0138] CP' z = log(CP z + 0.001) (10)

[0139] In the formula, CP' z represents the copy parameter for plotting, CP z represents the copy number value of the zth target copy abnormal site.

[0140] S213, plotting according to the order of chromosome distribution of the copy parameter for plotting of each target copy abnormal site.

[0141] By performing steps S211-S213, the copy number values of the target copy abnormal site are log converted to obtain stable and high readability copy parameters for plotting, and the identified target copy variant fragment is added with auxiliary lines for prompting, and the log transformed site copy number is plotted and visualized by means of the plotting tool.

[0142] After performing steps S101-S109, the above steps S201-S210 are detection steps for unknown single test samples, and in actual application, unknown batch test samples can also be implemented, and the sample and sample quantity targeted in step S021 are different. In the detection process of the test sample, first, the sequencing data of N2 test samples is obtained; the sequencing depth values of a plurality of specified sites are determined from the sequencing data of each test sample, and then the above steps S202-S210 are implemented, which can realize the test function, and the each specified site and its baseline determined in steps S101-S109 are optimized and updated.

[0143] For example, after implementing the N2 test samples according to step S208, the copy abnormal sites in the copy variant class obtained will be retained and used as new specified sites, and each new specified site will use the historical sample data (at least including N1 training samples) and the N2 test data of the batch to calculate a new baseline again, and the new baseline corresponding to each new specified site and the test data of the batch are stored in the database. In order to detect the unknown single test sample next time or optimize the specified site and its baseline for the next batch of test sample data.

[0144] The test samples involved in the above can be derived in large quantities from clinical test data, without the need for additional scientific research training sets and clinical information. The advantages of a large number of clinical test samples and large data can be converted into the stability of designated sites and their baselines. By continuously submitting for testing to optimize the designated sites and their baselines, better compatibility and stability can be achieved for the testing methods, test populations, analysis methods, and sequencing methods of the testing units.

[0145] like Figure 2 As shown, Figure 2 The GC content and depth of the autosomes, X / Y sex chromosomes, and mitochondrial genomes are displayed. Chromosomes exhibit unstable GC content, which can introduce false signals during analysis. The depth data indicates that the overall genome sequencing depth is very low, ranging from 1 to 2, which does not meet the requirements for CNV identification using existing assembly and other techniques.

[0146] Implementation of the present invention can filter low-depth and repetitive sequence sites, and utilize a large site baseline to eliminate differences in site depth caused by batch effects, experimental methods, quality control methods, and regional population characteristics. Linkage calculations can improve the stability of CNV detection and mapping from the genomic position dimension. Therefore, it is well suited for CNV identification in low-depth sequencing data structures such as metagenomic sequencing, and can be used for the exploration and early screening of tumors and genetic diseases in infected patients, with significant clinical significance.

[0147] Moreover, by implementing the embodiments of the present invention, it is possible to identify whether the sample to be tested is a female sample. If so, half of the sequencing depth value of the X chromosome site and the sequencing depth value of the autosomal site are compared with their respective confidence intervals, thereby automatically and accurately performing CNV detection on all nuclear genome chromosomes, thereby improving the accuracy of sex chromosome CNV detection.

[0148] like Figure 3 As shown, the embodiment of the present invention discloses a nuclear genome copy number variation detection device, including a depth determination unit 301, a gender judgment unit 302, a site identification unit 303, an abnormality detection unit 304, a calculation unit 305, a clustering unit 306, a variation determination unit 307 and a merging unit 308, wherein:

[0149] The depth determination unit 301 is used to determine the sequencing depth values ​​of multiple designated sites from the sequencing data of the sample to be tested; wherein each designated site corresponds to a confidence interval and an RD set mean;

[0150] A gender determination unit 302 is configured to determine whether the sample to be tested is a female sample based on the sequencing depth values ​​of multiple designated sites;

[0151] The site identifying unit 303 is configured to identify an X chromosome site and an autosome site in the plurality of specified sites when the gender judging unit 302 judges that the sample to be tested is a female sample.

[0152] The abnormality detecting unit 304 is configured to determine that the X chromosome site is a low-copy abnormality site when the half value of the sequencing depth value of any X chromosome site is less than the confidence interval corresponding to the X chromosome site, and determine that the X chromosome site is a high-copy abnormality site when the half value of the sequencing depth value of any X chromosome site is greater than the confidence interval corresponding to the X chromosome site, when the gender judging unit 302 judges that the sample to be tested is a female sample.

[0153] The abnormality detecting unit 304 is further configured to determine that the autosome site is a low-copy abnormality site when the sequencing depth value of any autosome site is less than the confidence interval corresponding to the autosome site, and determine that the autosome site is a high-copy abnormality site when the sequencing depth value of any autosome site is greater than the confidence interval corresponding to the autosome site, when the gender judging unit 302 judges that the sample to be tested is a female sample, and take both the low-copy abnormality site and the high-copy abnormality site as a copy abnormality site.

[0154] The calculating unit 305 is configured to calculate the copy number value of each copy abnormality site according to the mean value and the sequencing depth value corresponding to each copy abnormality site, wherein the copy number value and the sequencing depth value are in a positive correlation.

[0155] The clustering unit 306 is configured to perform spatial clustering classification on the copy number values of all copy abnormality sites to obtain a normal copy class and a copy variation class, and the copy variation class includes two copy variation sub-classes, i.e., a high-copy variation sub-class and a low-copy variation sub-class.

[0156] The variation determining unit 307 is configured to determine the high-copy abnormality site belonging to the high-copy variation sub-class and the low-copy abnormality site belonging to the low-copy variation sub-class as a copy variation site, respectively.

[0157] The merging unit 308 is configured to merge the copy variation sites that are adjacent in position and belong to the same copy variation sub-class to obtain a copy variation fragment.

[0158] Optionally, the abnormality detecting unit 304 is further configured to determine that the sample to be tested is a male sample when the gender judging unit 302 judges that the sample to be tested is not a female sample, determine that the specified site is a low-copy abnormality site when the sequencing depth value of any specified site is less than the confidence interval corresponding to the specified site, and determine that the specified site is a high-copy abnormality site when the sequencing depth value of any specified site is greater than the confidence interval corresponding to the specified site.

[0159] Optionally, the nuclear genome copy number variation detection device may further include the following units (not shown):

[0160] A data acquisition unit is used to acquire sequencing data of a plurality of training samples; the sequencing data includes sequencing depth values ​​of M1 candidate sites; the plurality of training samples includes a plurality of female training samples and a plurality of male training samples;

[0161] Classification unit, used to perform unsupervised clustering classification on the sequencing depth values ​​of all candidate sites to obtain multiple classification categories;

[0162] An averaging unit, used to calculate the mean of the first sequencing depth of each classification category;

[0163] a noise identification unit, configured to identify a noise category from a plurality of classification categories according to the first sequencing depth mean;

[0164] A noise elimination unit is used to eliminate candidate sites included in the noise category to obtain M2 set sites;

[0165] A site distinguishing unit, used for identifying an X chromosome set site, a Y chromosome set site, and an autosomal set site among the M2 set sites;

[0166] a baseline calculation unit, configured to calculate the RD set mean and variance of each X chromosome set site based on half the values ​​of the sequencing data of the plurality of female training samples and the sequencing data of the plurality of male training samples; and to calculate the RD set mean and variance of each Y chromosome set site based on the sequencing data of the plurality of training samples; and to calculate the RD set mean and variance of each autosomal set site based on the sequencing data of the plurality of training samples;

[0167] The site designation unit is used to determine all or part of the set sites from the M2 set sites as the designated sites; and calculate the confidence interval of each designated site according to the RD set mean and variance.

[0168] like Figure 4 As shown, an embodiment of the present invention discloses an electronic device, including a memory 401 storing executable program code and a processor 402 coupled to the memory 401;

[0169] The processor 402 calls the executable program code stored in the memory 401 to execute the nuclear genome copy number variation detection method described in the above embodiments.

[0170] An embodiment of the present invention further discloses a computer-readable storage medium storing a computer program, wherein the computer program enables a computer to execute the nuclear genome copy number variation detection method described in the above embodiments.

[0171] The above examples are intended to illustrate and deduce the technical solutions of the present application, and to completely describe the technical solutions, objects and effects of the present application. The purpose is to make the public more thoroughly and comprehensively understand the disclosure of the present application, and does not limit the protection scope of the present application.

[0172] The above examples are not based on an exhaustive enumeration of the present application, and there can be many other unlisted embodiments. Any substitution and improvement made without violating the concept of the present application is within the scope of protection of the present application.

Claims

1. A method for detecting nuclear genome copy number variation, characterized in that: include: Determine the sequencing depth values ​​of multiple designated sites from the sequencing data of the sample to be tested; wherein each designated site corresponds to a confidence interval and an RD set mean; Determining whether the sample to be tested is a female sample based on sequencing depth values ​​of multiple designated sites; If the sample to be tested is a female sample, identifying an X chromosome locus and an autosomal locus among the multiple designated loci; If half of the sequencing depth value of any X chromosome site is less than the confidence interval corresponding to the X chromosome site, the X chromosome site is determined to be a low-copy abnormal site; if half of the sequencing depth value of any X chromosome site is greater than the confidence interval corresponding to the X chromosome site, the X chromosome site is determined to be a high-copy abnormal site; If the sequencing depth value of any autosomal site is less than the confidence interval corresponding to the autosomal site, the autosomal site is determined to be a low-copy abnormal site; if the sequencing depth value of any autosomal site is greater than the confidence interval corresponding to the autosomal site, the autosomal site is determined to be a high-copy abnormal site; The low copy abnormal site and the high copy abnormal site are both regarded as copy abnormal sites; Calculate the copy number value of each copy abnormality site according to the RD setting mean and sequencing depth value corresponding to each copy abnormality site; wherein the copy number value is positively correlated with the sequencing depth value; Performing spatial clustering classification on the copy number values ​​of all the copy abnormality sites to obtain a normal copy class and a copy variation class, wherein the copy variation class includes two copy variation subclasses, namely a high copy variation subclass and a low copy variation subclass; Determine the high copy abnormal sites belonging to the high copy variation subclass and the low copy abnormal sites belonging to the low copy variation subclass as copy variation sites respectively; The copy variation sites that are adjacent in position and belong to the same copy variation subclass are merged to obtain copy variation fragments.

2. The method for detecting nuclear genome copy number variation according to claim 1, wherein The method further comprises: If the sample to be tested is not a female sample, the sample to be tested is determined to be a male sample. When the sequencing depth value of any designated site is less than the confidence interval corresponding to the designated site, the designated site is determined to be a low-copy abnormal site; when the sequencing depth value of any designated site is greater than the confidence interval corresponding to the designated site, the designated site is determined to be a high-copy abnormal site.

3. The method for detecting nuclear genome copy number variation according to claim 1 or 2, wherein: The confidence interval and RD set mean corresponding to the specified site are calculated by the following steps: Acquire sequencing data of multiple training samples; the sequencing data includes sequencing depth values ​​of M1 candidate sites; The multiple training samples include multiple female training samples and multiple male training samples; Perform unsupervised clustering classification on the sequencing depth values ​​of all candidate sites to obtain multiple classification categories; Calculating the mean first sequencing depth of each classification category; identifying a noise category from a plurality of classification categories according to the first sequencing depth mean; Eliminate the candidate sites included in the noise category to obtain M2 set sites; Identify the X chromosome setting site, the Y chromosome setting site, and the autosomal setting site among the M2 setting sites; Calculating the RD set mean and variance of each of the X chromosome set sites based on half of the sequencing data of multiple female training samples and the sequencing data of multiple male training samples; Calculating the RD setting mean and variance of each Y chromosome setting site based on the sequencing data of multiple training samples; Calculating the RD setting mean and variance of each of the autosomal setting sites based on sequencing data of multiple training samples; Determine all or part of the set sites from the M2 set sites as designated sites; The confidence interval for each of the designated sites was calculated based on the RD setting mean and variance.

4. The method for detecting nuclear genome copy number variation according to claim 3, wherein: Calculating the RD setting mean and variance of each X chromosome setting site based on half of the sequencing data of multiple female training samples and the sequencing data of multiple male training samples includes: The RD set mean and variance of each X chromosome set site were calculated using the following formula: in, Set the mean RD value for the wth X chromosome site, Y1 is the number of female training samples, Y2 is the number of male training samples, and RD wy Set the sequencing depth value of the site on the wth chromosome X for the yth training sample, S sw 2 Set the variance of the locus for the wth X chromosome.

5. The method for detecting nuclear genome copy number variation according to claim 4, wherein: The step of calculating the RD setting mean and variance of each Y chromosome setting site based on sequencing data of a plurality of training samples includes: The RD set mean and variance of each Y chromosome set site were calculated by the following formula: in, Set the mean RD value for the mth Y chromosome site, RD my Set the sequencing depth value of the site on the mth Y chromosome for the yth training sample; S ym 2 Set the variance of the locus for the mth Y chromosome.

6. The method for detecting nuclear genome copy number variation according to claim 5, wherein: The step of setting the mean and variance according to the RD and calculating the confidence interval of each designated site comprises: If the designated site is an X chromosome set site, its confidence interval is expressed by the following formula: If the designated site is a Y chromosome set site, its confidence interval is expressed by the following formula: If the designated locus is an autosomal set locus, its confidence interval is expressed by the following formula: Where, represents the RD setting mean corresponding to the i-th autosomal setting site, S i 2 represents the variance corresponding to the i-th autosomal set site.

7. The method for detecting nuclear genome copy number variation according to claim 5, wherein: Calculating the copy number value of each copy abnormality site according to the RD setting mean value and sequencing depth value corresponding to each copy abnormality site includes: If the copy abnormality site is an X chromosome set site, its copy number value is calculated using the following formula: Where, CP j Represents the copy number of the jth abnormal site on chromosome X, RD j Represents the sequencing depth value of the j-th copy abnormal site located on chromosome X, represents the RD set mean of the j-th copy abnormal site located on chromosome X; If the copy abnormality site is a Y chromosome set site or an autosomal set site, the copy number value is calculated using the following formula: Where, CP i Represents the copy number of the i-th abnormal site located on the Y chromosome or autosome, RD i Represents the sequencing depth value of the i-th copy abnormal site located on the Y chromosome or autosome, represents the RD set mean of the i-th copy abnormal site located on the Y chromosome or autosome.

8. A nuclear genome copy number variation detection device, characterized in that: include: A depth determination unit is used to determine the sequencing depth values ​​of multiple designated sites from the sequencing data of the sample to be tested; wherein each of the designated sites corresponds to a confidence interval and an RD set mean; A gender determination unit, configured to determine whether the sample to be tested is a female sample based on the sequencing depth values ​​of multiple designated sites; a site identification unit, configured to identify an X chromosome site and an autosomal site among a plurality of designated sites when the gender determination unit determines that the sample to be tested is a female sample; an abnormality detection unit, configured to determine that any X chromosome locus is a low-copy abnormality locus when half of the sequencing depth value of any X chromosome locus is less than the confidence interval corresponding to the X chromosome locus; and to determine that any X chromosome locus is a high-copy abnormality locus when half of the sequencing depth value of any X chromosome locus is greater than the confidence interval corresponding to the X chromosome locus; The abnormality detection unit is further configured to determine that, when the sequencing depth value of any autosomal site is less than the confidence interval corresponding to the autosomal site, the autosomal site is a low-copy abnormal site; and, when the sequencing depth value of any autosomal site is greater than the confidence interval corresponding to the autosomal site, the autosomal site is determined to be a high-copy abnormal site; and, both the low-copy abnormal site and the high-copy abnormal site are regarded as copy abnormal sites; A calculation unit, configured to calculate a copy number value of each copy abnormality site according to an RD setting mean value and a sequencing depth value corresponding to each copy abnormality site; wherein the copy number value is positively correlated with the sequencing depth value; a clustering unit, configured to perform spatial clustering classification on the copy number values ​​of all the copy abnormality sites to obtain a normal copy class and a copy variation class, wherein the copy variation class includes two copy variation subclasses, namely a high copy variation subclass and a low copy variation subclass; a variation determination unit, configured to determine the high-copy abnormal sites belonging to the high-copy variation subclass and the low-copy abnormal sites belonging to the low-copy variation subclass as copy variation sites, respectively; The merging unit is used to merge copy variation sites that are adjacent in position and belong to the same copy variation subclass to obtain copy variation fragments.

9. An electronic device, characterized in that It comprises a memory storing executable program code and a processor coupled to the memory; the processor calls the executable program code stored in the memory to execute the nuclear genome copy number variation detection method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein the computer program enables a computer to execute the method for detecting nuclear genome copy number variation according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Single-sample whole-genome allele specific copy number variation prediction method

    CN112802548A

  • Single sample allele copy number variation detection method, probe set and kit

    CN113889187A