Chromosome copy number variation detection method, device, equipment, medium and product

By obtaining sequencing depth data and using a set of normal distribution functions to identify regions of chromosome copy number variation, the problem of insufficient accuracy in fetal CNV detection in maternal plasma was solved, achieving higher detection accuracy.

CN120600103APending Publication Date: 2025-09-05GENEMIND BIOSCIENCES CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510669935.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

The accuracy of fetal chromosome copy number variation (CNV) detection in maternal plasma is insufficient, limited by factors such as low sequencing depth and high signal-to-noise ratio, resulting in poor detection accuracy.

Method used

By obtaining the sequencing depth data of the nucleic acid sample to be tested, the copy number type set is determined using the normal distribution function set, and multiple nucleic acid windows are spliced ​​to identify the copy number variation regions on the chromosome, and detection is performed using a combination of hardware and software.

Benefits of technology

The accuracy of CNV detection is improved, the problem of poor detection precision is solved, and higher detection accuracy is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120600103A_ABST
    Figure CN120600103A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of biomedicine, and discloses a chromosome copy number variation detection method, device and equipment, a medium and a product, and the method comprises the following steps: obtaining sequencing depth data of a to-be-detected chromosome in a to-be-detected nucleic acid sample; determining a copy number type set according to the sequencing depth data and a normal distribution function set corresponding to the to-be-detected chromosome; wherein the normal distribution function set comprises normal distribution functions corresponding to a copy number deletion type, a copy number repetition type and a copy number normal type respectively, and the copy number type set comprises labeled copy number types of a plurality of second nucleic acid windows; according to the copy number type set, splicing the plurality of second nucleic acid windows to obtain at least one copy number variation region; wherein the second nucleic acid window is a first nucleic acid window of the to-be-detected chromosome or a recombinant nucleic acid window obtained by combining a plurality of first nucleic acid windows based on a preset window number and a preset window step length, and the detection precision of chromosome copy number variation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of biomedical technology, and in particular to a method, device, equipment, medium and product for detecting chromosome copy number variation. Background Art

[0002] In recent years, due to the discovery of fetal-free nucleic acid in maternal plasma and the development of NGS sequencing technology (Next Generation Sequencing technology), the application of whole-genome low-depth sequencing for prenatal aneuploidy screening has been realized, mainly used for triploidy detection of chromosomes 21, chromosomes 18 and chromosomes 13.

[0003] Studies have shown that among pregnant women, the prevalence of fetal pathogenic CNVs (chromosomal copy number variations) can reach 1.6% to 1.7%, far exceeding the approximately 0.2% incidence of triploidy. However, non-invasive testing for pregnant women is limited by factors such as the shorter CNV region compared to chromosomes, low sequencing depth due to sequencing costs, and high sequencing signal-to-noise ratios, resulting in significantly insufficient CNV detection accuracy. Summary of the Invention

[0004] The embodiments of the present invention provide a method, apparatus, device, medium and product for detecting chromosome copy number variation to solve the problem of poor precision in CNV detection and to improve the accuracy of CNV detection.

[0005] According to one embodiment of the present invention, a method for detecting chromosome copy number variation is provided, the method comprising:

[0006] Acquiring sequencing depth data of a chromosome to be tested in a nucleic acid sample to be tested; wherein the sequencing depth data includes first sequencing depths under a plurality of first nucleic acid windows corresponding to the chromosome to be tested;

[0007] Determining a copy number type set according to the sequencing depth data and a normal distribution function set corresponding to the chromosome to be tested; wherein the normal distribution function set includes normal distribution functions corresponding to a copy number deletion type, a copy number duplication type, and a copy number normal type, respectively, and the copy number type set includes a plurality of marker copy number types of the second nucleic acid window;

[0008] splicing the plurality of second nucleic acid windows according to the copy number type set to obtain at least one copy number variation region on the chromosome to be tested;

[0009] In which, the second nucleic acid window is the first nucleic acid window, or the second nucleic acid window is a recombinant nucleic acid window obtained by merging the multiple first nucleic acid windows based on a preset window number and a preset window step size, and the normal distribution function represents the probability density of the second sequencing depth corresponding to the second nucleic acid window as the value of the second sequencing depth changes.

[0010] According to another embodiment of the present invention, a device for detecting chromosome copy number variation is provided, the device comprising:

[0011] A sequencing depth data acquisition module is used to obtain sequencing depth data of a chromosome to be tested in a nucleic acid sample to be tested; wherein the sequencing depth data includes a first sequencing depth under a plurality of first nucleic acid windows corresponding to the chromosome to be tested;

[0012] a copy number type set determination module, configured to determine a copy number type set based on the sequencing depth data and a normal distribution function set corresponding to the chromosome to be tested; wherein the normal distribution function set includes normal distribution functions corresponding to a copy number deletion type, a copy number duplication type, and a copy number normal type, respectively, and the copy number type set includes a plurality of marker copy number types of a second nucleic acid window;

[0013] a copy number variation region determination module, configured to splice the plurality of second nucleic acid windows according to the copy number type set to obtain at least one copy number variation region on the chromosome to be tested;

[0014] In which, the second nucleic acid window is the first nucleic acid window, or the second nucleic acid window is a recombinant nucleic acid window obtained by merging the multiple first nucleic acid windows based on a preset window number and a preset window step size, and the normal distribution function represents the probability density of the second sequencing depth corresponding to the second nucleic acid window as the value of the second sequencing depth changes.

[0015] According to another embodiment of the present invention, an electronic device is provided, the electronic device including:

[0016] at least one processor; and

[0017] a memory communicatively connected to the at least one processor; wherein,

[0018] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can perform the method for detecting chromosome copy number variation described in any embodiment of the present invention.

[0019] According to another embodiment of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the method for detecting chromosome copy number variation described in any embodiment of the present invention when executed.

[0020] According to another embodiment of the present invention, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the method for detecting chromosome copy number variation according to any embodiment of the present invention.

[0021] The technical solution of this embodiment is to determine the copy number type set corresponding to the second nucleic acid window based on the normal distribution function set corresponding to the chromosome to be tested and the sequencing depth data corresponding to the first nucleic acid window, and splice multiple second nucleic acid windows according to the copy number type set to obtain at least one copy number variation region on the chromosome to be tested, wherein the normal distribution function set includes normal distribution functions corresponding to the copy number deletion type, the copy number duplication type and the copy number normal type, respectively, and the second nucleic acid window is the first nucleic acid window or a recombinant nucleic acid window obtained by merging multiple first nucleic acid windows based on a preset window number and a preset window step size, thereby achieving the purpose of CNV detection based on the window segment characteristics of the chromosome, solving the problem of poor CNV detection precision, and improving the accuracy of CNV detection.

[0022] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0024] Figure 1 A flowchart of a method for detecting chromosome copy number variation provided by one embodiment of the present invention;

[0025] Figure 2 A schematic diagram of a normal distribution function set provided by one embodiment of the present invention;

[0026] Figure 3 A flowchart of another method for detecting chromosome copy number variation provided by one embodiment of the present invention;

[0027] Figure 4A scatter plot of initial depth data and sequencing depth data provided by one embodiment of the present invention;

[0028] Figure 5 A flowchart of a specific example of a method for detecting chromosome copy number variation provided by one embodiment of the present invention;

[0029] Figure 6 A schematic structural diagram of a device for detecting chromosome copy number variation provided by one embodiment of the present invention;

[0030] Figure 7 The present invention provides a schematic structural diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0031] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0032] It should be noted that the terms "first", "second", "third", "target", "reference", "initial", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0033] Figure 1 This is a flow chart of a method for detecting chromosome copy number variation provided by one embodiment of the present invention. This embodiment is applicable to the case of detecting chromosome copy number variation. The method can be performed by a chromosome copy number variation detection device. The chromosome copy number variation detection device can be implemented in the form of hardware and / or software. The chromosome copy number variation detection device can be configured in a terminal device. Figure 1 As shown, the method includes:

[0034] S110: Obtain sequencing depth data of the chromosome to be tested in the nucleic acid sample to be tested.

[0035] Specifically, the nucleic acid sample to be tested is a cell-free DNA (cfDNA) sample, and the chromosome to be tested is any pair of the 23 pairs of chromosomes in the human body used to detect chromosomal copy number variation. cfDNA samples refer to DNA fragments that exist freely in body fluids such as blood, cerebrospinal fluid, and urine and are not contained within cellular structures.

[0036] In an optional embodiment, the nucleic acid sample to be tested is extracted from a peripheral blood sample of a pregnant woman. Specifically, the peripheral blood sample contains maternal cells and cfDNA samples, wherein the cfDNA samples are mainly derived from apoptosis and necrosis of placental cells and active release of fetal cells.

[0037] In this embodiment, the sequencing depth data includes the first sequencing depths under multiple first nucleic acid windows corresponding to the chromosome to be tested.

[0038] The first nucleic acid window represents a nucleic acid fragment corresponding to the chromosome to be tested in the human reference genome. In an optional embodiment, the method further includes: obtaining chromosome nucleic acid data corresponding to the chromosome to be tested from the reference genome nucleic acid data corresponding to the human reference genome; and dividing the chromosome nucleic acid data according to a preset window length to obtain a plurality of first nucleic acid windows.

[0039] For example, the source of the human reference genome can be the GRCH36 version, GRCH37 version, GRCh38 version of the National Center for Biotechnology Information (NCBI) database, the hg18 version, hg19 version or hg38 version of the University of California, Santa Cruz (UCSC) database, etc. The source of the human reference genome is not limited here and can be customized according to actual needs.

[0040] Specifically, the reference genome nucleic acid data refers to a set of information consisting of base sequence information, gene expression information, functional annotations and other information obtained by sequencing the human reference genome. The chromosome nucleic acid data is the nucleic acid sequence corresponding to the chromosome to be tested in the human reference genome.

[0041] Exemplarily, the preset window length is 20 kbp, but is not limited to the example case.

[0042] In an optional embodiment, the chromosome nucleic acid data is divided according to a preset window length to obtain multiple first nucleic acid windows corresponding to the chromosome to be tested, including: dividing the chromosome nucleic acid data according to the preset window length to obtain a window division result; performing a deletion operation on the nucleic acid window that does not contain a known base in the window division result to obtain multiple first nucleic acid windows corresponding to the chromosome to be tested.

[0043] Specifically, the window division result includes multiple nucleic acid windows of preset window lengths. The nucleic acid windows in the window division result are traversed. If the nucleic acid window does not contain a known base, it means that the nucleic acid window contains all unknown bases, and the nucleic acid window is deleted from the window division result.

[0044] The advantage of this setting is that it avoids noise interference caused by nucleic acid windows that are all unknown bases, thereby further improving the accuracy of CNV detection.

[0045] Specifically, sequencing depth refers to the number of sequences in a nucleic acid sample to be tested that can be uniquely aligned to a certain region in the human reference genome.

[0046] In an optional embodiment, obtaining sequencing depth data of a chromosome to be tested in a nucleic acid sample to be tested includes: obtaining chromosome sequencing data corresponding to the chromosome to be tested in the genome sequencing data of the nucleic acid sample to be tested; performing sequence alignment on the chromosome sequencing data and multiple first nucleic acid windows corresponding to the chromosome to be tested to obtain sequence alignment results; and determining sequencing depth data based on the sequence alignment results.

[0047] Specifically, the genome sequencing data is the nucleic acid sequence data obtained by whole genome sequencing of the nucleic acid sample to be tested, which represents an information set consisting of multiple information such as base sequence information, gene expression information, and functional annotations obtained by sequencing the chromosome to be tested. The chromosome sequencing data is the nucleic acid sequence corresponding to the chromosome to be tested in the nucleic acid sample to be tested.

[0048] Based on the above embodiment, optionally, the detection method further includes: using a polymerase chain reaction (PCR) nucleic acid amplifier to perform PCR amplification on the nucleic acid sample to be tested, and then performing sample pretreatment to obtain a nucleic acid library; performing whole genome sequencing on the nucleic acid library to obtain genome sequencing data of the nucleic acid sample to be tested.

[0049] Exemplarily, sample preprocessing includes, but is not limited to, cfDNA fragmentation, end repair, ligation of adapters, quality control, and quantitative analysis. An Illumina library construction kit is used to construct a nucleic acid library based on the amplified nucleic acid sample to be tested. The sequencing technology used for whole genome sequencing includes, but is not limited to, second-generation sequencing technology, nanopore sequencing technology, or third-generation sequencing technology.

[0050] In an optional embodiment, the detection method further includes: performing data quality control on the genome sequencing data of the nucleic acid sample to be tested to obtain genome sequencing data that has passed quality control. The data quality control referred to in the embodiment of the present application refers to filtering the data to screen out sequencing data that does not meet the quality control conditions. For example, sequencing data that does not meet the quality control conditions include but are not limited to sequencing data such as an unknown base N that accounts for too high a proportion, a low-quality base that accounts for too high a proportion, and a length of the remaining sequence that is too short after the connector is removed. Exemplarily, the quality control tool used for data quality control can be a fastp tool, a Trimmomatic tool, or a FastQC tool. The quality control tool used is not limited here, and can be customized according to actual needs.

[0051] For example, the alignment tools used for sequence alignment include, but are not limited to, Tmap tools, BWA tools, SOAP tools, or samtools tools.

[0052] Specifically, the sequencing depth data is determined based on the sequence alignment results, including: performing a PCR duplicate removal operation on the sequence alignment results to remove duplicates introduced by PCR amplification, and determining the sequencing depth data based on the sequence alignment results after removing the PCR duplicates.

[0053] The duplication here means that after alignment with the human reference genome, only one sequence with exactly the same comparison parameters such as chromosomes, chromosome position, length, direction, and base sequence is retained, thereby reducing the deviation and error caused by PCR amplification.

[0054] In an optional embodiment, determining sequencing depth data according to the sequence alignment result includes: correcting the sequencing depth in the sequence alignment result, and determining the sequencing depth data according to the corrected sequence alignment result.

[0055] Among them, the correction processing includes at least one of effective base length correction, outlier correction, mappability correction and GC (guanine and cytosine) base content correction. Among them, the mappability value can be used to characterize the alignment tool's ability to correctly align chromosome sequencing data to the nucleic acid window of the human reference genome, and mappability correction refers to performing a local polynomial regression fitting correction on the sequencing depth based on the mappability value. Among them, since the sequencing depth corresponding to the sequence alignment results of high GC content or low GC content will be lower than the sequencing depth corresponding to the sequence alignment results with a small difference or consistency between GC content and AT content, GC base content correction refers to performing a standardized correction or a local polynomial regression fitting correction on the sequencing depth based on the GC content corresponding to the sequence alignment result.

[0056] The advantage of this setting is that it can eliminate the error interference caused by different effective base lengths, different outliers, different mappability values ​​and different GC contents on the sequencing depth of the target chromosome, thereby further improving the accuracy of CNV detection.

[0057] S120. Determine a copy number type set based on the sequencing depth data and a normal distribution function set corresponding to the chromosome to be tested.

[0058] In this embodiment, the normal distribution function set includes normal distribution functions corresponding to the copy number deletion type, the copy number duplication type, and the copy number normal type, respectively.

[0059] Among them, the copy number deletion type indicates that the copy number of a certain region on the chromosome or a certain gene is lower than the normal level, the copy number duplication type indicates that the copy number of a certain region on the chromosome or a certain gene is higher than the normal level, and the copy number normal type indicates that the copy number of a certain region on the chromosome or a certain gene is equal to the normal level.

[0060] In this embodiment, the normal distribution function represents how the probability density of the second sequencing depth corresponding to the second nucleic acid window changes with the value of the second sequencing depth, and the copy number type set includes marker copy number types of multiple second nucleic acid windows.

[0061] In a specific embodiment, the second nucleic acid window is the first nucleic acid window, or the second nucleic acid window is a recombinant nucleic acid window obtained by merging multiple first nucleic acid windows based on a preset window number and a preset window step size.

[0062] Specifically, the preset window number indicates the number of first nucleic acid windows included in the second nucleic acid window, and the preset window step indicates the number of overlapping first nucleic acid windows in two adjacent second nucleic acid windows.

[0063] In an optional embodiment, the number of preset windows is twice the preset window step size. For example, the number of preset windows is 10 and the preset window step size is 5.

[0064] In an optional embodiment, a copy number type set is determined based on sequencing depth data and a normal distribution function set corresponding to the chromosome to be tested, including: obtaining the normal distribution function set in the first iteration process; for each second nucleic acid window, determining the second sequencing depth corresponding to the second nucleic acid window based on the sequencing depth data, and determining the probability density values ​​corresponding to the second nucleic acid window and the copy number deletion type, copy number duplication type and copy number normal type respectively based on the normal distribution function set and the second sequencing depth, and taking the copy number type corresponding to the maximum probability density value as the marker copy number type of the second nucleic acid window; updating the normal distribution function set based on the second sequencing depth of at least one second nucleic acid window corresponding to the copy number normal type to obtain the normal distribution function set for the next iteration process; returning to execute the step of determining the second sequencing depth corresponding to the second nucleic acid window based on the sequencing depth data for each second nucleic acid window; until the number of second nucleic acid windows corresponding to the copy number normal type in the chromosome to be tested converges, determining the copy number type set based on the marker copy number type of each second nucleic acid window in the current iteration process.

[0065] Specifically, the standard deviations corresponding to the three normal distribution functions in the normal distribution function set are the same. In the first iteration, the standard deviations corresponding to the three normal distribution functions in the normal distribution function set are preset values. Exemplarily, the preset value can be 0.1, but is not limited to the example case.

[0066] In an optional embodiment, the second sequencing depth corresponding to the second nucleic acid window is determined based on the sequencing depth data, including: when the second nucleic acid window is the first nucleic acid window, the first sequencing depth corresponding to the second nucleic acid window in the sequencing depth data is used as the second sequencing depth corresponding to the second nucleic acid window.

[0067] In another optional embodiment, the second sequencing depth corresponding to the second nucleic acid window is determined based on the sequencing depth data, including: when the second nucleic acid window is a recombinant nucleic acid window, obtaining the first sequencing depth of the first nucleic acid window of a preset number of windows corresponding to the second nucleic acid window in the sequencing depth data; and taking the average value of the first sequencing depth of the preset number of windows as the second sequencing depth corresponding to the second nucleic acid window.

[0068] Figure 2This is a schematic diagram of a normal distribution function set provided by an embodiment of the present invention. Specifically, the green curve, black curve, and purple curve respectively represent the normal distribution functions corresponding to the copy number deletion type, the copy number normal type, and the copy number duplication type in the normal distribution function set, the horizontal axis represents the sequencing depth, and the vertical axis represents the probability density value.

[0069] by Figure 2 For example, the second sequencing depth sd is substituted into the normal distribution function (green curve) corresponding to the copy number deletion type to obtain the probability density value pd1 corresponding to the copy number deletion type, the second sequencing depth sd is substituted into the normal distribution function (black curve) corresponding to the copy number normal type to obtain the probability density value pd2 corresponding to the copy number normal type, the second sequencing depth sd is substituted into the normal distribution function (purple curve) corresponding to the copy number duplication type to obtain the probability density value pd3 corresponding to the copy number duplication type, and the probability density values ​​pd1, pd2 and pd3 are compared to obtain the maximum probability density value pd3. The copy number duplication type is used as the marker copy number type for the second nucleic acid window.

[0070] In an optional embodiment, the normal distribution function set is updated according to the second sequencing depth of at least one second nucleic acid window corresponding to the normal copy number type to obtain the normal distribution function set for the next iterative process, including: determining the standard deviation corresponding to the second sequencing depth of at least one second nucleic acid window corresponding to the normal copy number type; and updating the normal distribution function set according to the standard deviation to obtain the normal distribution function set for the next iterative process.

[0071] Specifically, according to the standard deviation, the standard deviations corresponding to the three normal distribution functions in the normal distribution function set are replaced respectively to obtain the normal distribution function set for the next iterative process.

[0072] In an optional embodiment, the mean of the normal distribution function corresponding to the normal copy number type in the normal distribution function set is 1, and the mean of the normal distribution function corresponding to the copy number deletion type is The mean of the normal distribution function corresponding to the copy number repeat type is Among them, FF represents the placenta-derived nucleic acid fraction corresponding to the nucleic acid sample to be tested.

[0073] In this embodiment, the first sequencing depth is the sequencing depth after normalization. Exemplarily, the normalization algorithm used in the normalization process is a median normalization algorithm.

[0074] Specifically, the placental-derived nucleic acid fraction (Placental-derived Nucleic Acid Fraction) represents the ratio of the nucleic acid content from the placenta to the total nucleic acid content in the nucleic acid sample to be tested. Exemplarily, the placental-derived nucleic acid fraction may also be referred to as placental DNA or placental DNA fraction.

[0075] In some instances, when the placenta-derived nucleic acid fraction FF remains unchanged, repeated testing experiments on nucleic acid samples revealed that the sequencing depth of normal diploid samples satisfies the normal distribution function with a mean of 1, and the sequencing depth of copy number deletion samples satisfies the normal distribution function with a mean of The normal distribution function of the copy number duplication sample is normalized to the mean sequencing depth of The normal distribution function of .

[0076] In another optional embodiment, the mean of the normal distribution function corresponding to the normal copy number type in the normal distribution function set is mean, and the mean of the normal distribution function corresponding to the copy number deletion type is The mean of the normal distribution function corresponding to the copy number repeat type is Here, mean represents the median of the second sequencing depths corresponding to the multiple second nucleic acid windows.

[0077] The advantage of setting the median instead of 1 is that due to the influence of factors such as differences in experimental technology and the characteristics of the sample itself, the distribution center of the second sequencing depth may fluctuate greatly. The median can more accurately reflect the central trend of the second sequencing depth, thereby making the normal distribution function better adapt to fit the actual data distribution of different nucleic acid samples to be tested, thereby further improving the accuracy and stability of CNV detection.

[0078] Based on the above embodiments, optionally, the detection method further includes: determining comparison feature data based on the genome sequencing data of the nucleic acid sample to be tested and the reference genome nucleic acid data of the human reference genome; inputting the comparison feature data into a pre-trained nucleic acid concentration prediction model to obtain the output placenta-derived nucleic acid score of the nucleic acid sample to be tested.

[0079] Illustratively, the comparison feature data includes but is not limited to the compared chromosomes, the compared genes on the chromosomes, the compared positions on the chromosomes, and gene annotations, etc., but is not limited to the example cases.

[0080] S130 , splicing multiple second nucleic acid windows according to the copy number type set to obtain at least one copy number variation region on the chromosome to be tested.

[0081] In an optional embodiment, the copy number variation region satisfies the following splicing conditions:

[0082] 1) The marker copy number types of the first second nucleic acid window and the last second nucleic acid window in the copy number variation region are both copy number abnormality types;

[0083] 2) the number of second nucleic acid windows in the copy number variation region that are labeled as copy number abnormality types is not less than a preset threshold;

[0084] 3) the number of second nucleic acid windows in the copy number variation region that are labeled as a normal copy number type is not greater than a preset threshold;

[0085] 4) The multiple copy number abnormality types corresponding to the copy number variation region are all copy number deletion types or copy number duplication types.

[0086] Exemplarily, the preset quantity threshold may be 6, but is not limited to the example case.

[0087] The technical solution of this embodiment is to determine the copy number type set corresponding to the second nucleic acid window based on the normal distribution function set corresponding to the chromosome to be tested and the sequencing depth data corresponding to the first nucleic acid window, and splice multiple second nucleic acid windows according to the copy number type set to obtain at least one copy number variation region on the chromosome to be tested, wherein the normal distribution function set includes normal distribution functions corresponding to the copy number deletion type, the copy number duplication type and the copy number normal type, respectively, and the second nucleic acid window is the first nucleic acid window or a recombinant nucleic acid window obtained by merging multiple first nucleic acid windows based on a preset window number and a preset window step size, thereby achieving the purpose of CNV detection based on the window segment characteristics of the chromosome, solving the problem of poor CNV detection precision, and improving the accuracy of CNV detection.

[0088] Figure 3 This is a flow chart of another method for detecting chromosome copy number variation provided by one embodiment of the present invention. This embodiment further refines the "obtaining sequencing depth data of the chromosome to be tested in the nucleic acid sample to be tested" in the above embodiment. Figure 3 As shown, the method includes:

[0089] S210: Obtain initial depth data of the chromosome to be tested in the nucleic acid sample to be tested.

[0090] In this embodiment, the initial depth data includes initial sequencing depths under multiple first nucleic acid windows corresponding to the chromosome to be tested.

[0091] Specifically, the method for obtaining the initial depth data is the same as or similar to the method for obtaining the sequencing depth data provided in the above embodiment, and will not be repeated in this embodiment.

[0092] S220. For each first nucleic acid window, obtain an initial sequencing depth and a reference sequencing depth corresponding to the first nucleic acid window, and use the ratio of the initial sequencing depth to the reference sequencing depth as the first sequencing depth corresponding to the first nucleic acid window.

[0093] In this embodiment, the reference sequencing depth represents the average of the third sequencing depths corresponding to the test nucleic acid sample and at least one third nucleic acid window in the reference nucleic acid window set, where the third nucleic acid window is a nucleic acid window corresponding to a chromosome other than the test chromosome. Specifically, the reference nucleic acid window set represents a collection of third nucleic acid windows that have similar sequencing depths to the first nucleic acid window.

[0094] In an optional embodiment, the detection method also includes: obtaining target depth data corresponding to the negative nucleic acid sample set and the chromosome to be tested, as well as reference depth data corresponding to other chromosomes; for each first nucleic acid window, determining depth distance data based on the reference depth data and the target depth sequence corresponding to the first nucleic acid window, and screening multiple third nucleic acid windows based on the depth distance data to obtain a reference nucleic acid window set corresponding to the first nucleic acid window.

[0095] The target depth data includes target depth sequences under multiple first nucleic acid windows corresponding to the chromosome to be tested, and the target depth sequence includes the negative sequencing depth under the first nucleic acid window of at least one negative nucleic acid sample in the negative nucleic acid sample set. A negative nucleic acid sample is a nucleic acid sample that does not have a CNV region.

[0096] For example, the target depth data is represented as SD g ={sd1, sd2, ..., sd n}, where sd1 represents the target depth sequence corresponding to the first nucleic acid window, sd2 represents the target depth sequence corresponding to the second first nucleic acid window, and sd n Represents the target depth sequence corresponding to the nth first nucleic acid window.

[0097] Take the target depth sequence sd n For example, when the chromosome to be tested is not the X chromosome, When the chromosome to be tested is the X chromosome, in, Indicates the negative sequencing depth of the first negative nucleic acid sample under the nth first nucleic acid window, Indicates the negative sequencing depth of the second negative nucleic acid sample under the nth first nucleic acid window, It represents the negative sequencing depth of the sth negative nucleic acid sample under the nth first nucleic acid window, represents the negative sequencing depth of the qth negative nucleic acid sample under the nth first nucleic acid window, s represents the total number of negative nucleic acid samples in the negative nucleic acid sample set, q represents the number of negative nucleic acid samples of female fetuses in the negative nucleic acid sample set, and q <s。

[0098] In particular, since only the negative nucleic acid samples of male fetuses have the sequencing depth of the Y chromosome, and the sequencing depth varies with the change of the placenta-derived nucleic acid fraction, the chromosome to be tested in this embodiment is not the Y chromosome.

[0099] In this embodiment, the reference depth data includes reference depth sequences corresponding to multiple third nucleic acid windows for other chromosomes, and the reference depth sequences include the negative sequencing depth of at least one negative nucleic acid sample in the third nucleic acid window. Specifically, the number of other chromosomes corresponding to the reference depth data is one or more.

[0100] In this embodiment, the depth distance data includes depth distances corresponding to the target depth sequence and at least one reference depth sequence. Specifically, the depth distance represents the degree of similarity between the target depth sequence and the reference depth sequence.

[0101] For example, the depth distance D(i, j) between the i-th target depth sequence and the j-th reference depth sequence is 2 Satisfies the following formula:

[0102]

[0103] When the chromosome to be tested is a non-X chromosome, n=s; when the chromosome to be tested is an X chromosome, n=q. represents the negative sequencing depth of the mth negative nucleic acid sample under the i-th first nucleic acid window, It represents the negative sequencing depth of the mth negative nucleic acid sample under the jth third nucleic acid window.

[0104] In an optional embodiment, multiple third nucleic acid windows are screened according to the depth distance data to obtain a reference nucleic acid window set corresponding to the first nucleic acid window, including: retaining third nucleic acid windows that meet the specified number of windows in order from small to large according to the depth distance data to obtain an initial nucleic acid window set corresponding to the first nucleic acid window; and performing data cleaning on the initial nucleic acid window set to obtain a reference nucleic acid window set corresponding to the first nucleic acid window.

[0105] Exemplarily, the number of designated windows may be 300 or 400, but is not limited to the example case.

[0106] In an optional embodiment, data cleaning is performed on the initial nucleic acid window set to obtain a reference nucleic acid window set corresponding to the first nucleic acid window, including: when the initial nucleic acid window set contains two adjacent third nucleic acid windows, obtaining the depth distances corresponding to the two third nucleic acid windows in the depth distance data; deleting the third nucleic acid window corresponding to the larger depth distance from the initial nucleic acid window set to obtain a reference nucleic acid window set corresponding to the first nucleic acid window.

[0107] Specifically, two adjacent third nucleic acid windows indicate that the two third nucleic acid windows are located on the same chromosome and are adjacent to each other on the chromosome.

[0108] For example, assuming that in the initial nucleic acid window set, if the depth distance D(i, j) corresponding to the i-th target depth sequence and the j-th reference depth sequence is 2 Greater than the depth distance D(i, j+1) between the i-th target depth sequence and the j+1-th reference depth sequence 2 , then the jth third nucleic acid window is deleted from the initial nucleic acid window set, if the depth distance D(i, j) between the i-th target depth sequence and the j-th reference depth sequence is 2 Less than the depth distance D(i, j+1) between the i-th target depth sequence and the j+1-th reference depth sequence 2 , then the j+1th third nucleic acid window is deleted from the initial nucleic acid window set.

[0109] In the nucleic acid sample to be tested, two adjacent third nucleic acid windows may be located in the same copy number variation region, which will cause significant noise interference to the accuracy of the first sequencing depth obtained by internal reference correction of the initial sequencing depth. This embodiment further improves the accuracy of CNV detection by screening the third nucleic acid windows adjacent to the initial nucleic acid window set.

[0110] In an optional embodiment, data cleaning is performed on the initial nucleic acid window set to obtain a reference nucleic acid window set corresponding to the first nucleic acid window, including: determining coefficient of variation data based on the reference depth sequence of each third nucleic acid window in the initial nucleic acid window set; wherein the coefficient of variation data includes the coefficient of variation of each third nucleic acid window; determining an outlier interval based on the coefficient of variation data; deleting the third nucleic acid window corresponding to the coefficient of variation that does not meet the outlier interval in the coefficient of variation data from the initial nucleic acid window set to obtain a reference nucleic acid window set corresponding to the first nucleic acid window.

[0111] Specifically, the coefficient of variation is the ratio of the standard deviation to the mean corresponding to the reference depth sequence, which is used to measure the degree of dispersion of the reference depth sequence. For example, the coefficient of variation CV of the i-th third nucleic acid window in the initial nucleic acid window set is i Satisfies the following formula:

[0112]

[0113] Among them, σ i represents the standard deviation of the reference depth sequence corresponding to the i-th third nucleic acid window, μ i represents the mean corresponding to the reference depth sequence of the i-th third nucleic acid window.

[0114] Specifically, the coefficient of variation in the coefficient of variation data is sorted from smallest to largest. The coefficient of variation in the top 25% of the sorted results is used as the lower quartile, denoted as Q1, and the coefficient of variation in the top 75% of the sorted results is used as the upper quartile, denoted as Q3. Accordingly, the interquartile range IRQ = |Q1-Q3|.

[0115] Exemplarily, the maximum value Up_threhold corresponding to the abnormal value interval is Q3+1.5×IRQ, and the minimum value Down_threhold corresponding to the abnormal value interval is Q1−1.5×IRQ, but this is not limited to the exemplary case.

[0116] The advantage of this setting is that it ensures the stability of the sequencing depth corresponding to the third nucleic acid window in the reference nucleic acid window set, thereby reducing the impact of the volatility of the sequencing depth of the third nucleic acid window on the first sequencing depth, and further ensuring the stability of CNV detection.

[0117] Based on the above embodiment, optionally, the detection method also includes: after performing data cleaning on the initial nucleic acid window set, if the difference window number between the specified nucleic acid number and the window number corresponding to the initial nucleic acid window set is greater than 0, then according to the depth distance data, the remaining third nucleic acid windows are retained in order from small to large by the difference window number of third nucleic acid windows, and the difference window number of third nucleic acid windows are added to the initial nucleic acid window set; return to execute the step of performing data cleaning on the initial nucleic acid window set until the difference window number is equal to 0, and use the initial nucleic acid window set as the reference nucleic acid window set corresponding to the first nucleic acid window.

[0118] The advantage of this setting is that it avoids the situation where the number of third nucleic acid windows in the reference nucleic acid window set is too small due to data cleaning operations, ensures the accuracy of the first sequencing depth, and thus improves the accuracy of CNV detection.

[0119] S230 , determining sequencing depth data of the chromosome to be tested in the nucleic acid sample to be tested according to the first sequencing depth corresponding to each first nucleic acid window.

[0120] Figure 4 A scatter plot of initial depth data and sequencing depth data provided by an embodiment of the present invention. Specifically, Figure 4 Taking chromosome 1 as an example, the horizontal axis represents the region on chromosome 1, and the vertical axis represents the sequencing depth. Figure 4 The upper figure in the figure shows the scatter plot of the initial depth data, and the lower figure shows the scatter plot of the sequencing depth data.

[0121] from Figure 4 It can be seen that the sequencing depth data has a higher degree of concentration than the initial depth data, indicating that the discreteness of the sequencing depth data is reduced and the sequencing depth is more stable.

[0122] S240. Determine a copy number type set based on the sequencing depth data and a normal distribution function set corresponding to the chromosome to be tested.

[0123] S250 , splicing multiple second nucleic acid windows according to the copy number type set to obtain at least one copy number variation region on the chromosome to be tested.

[0124] S240-S250 in this embodiment are the same as those in the above embodiment. Figure 1 The steps S120 - S130 shown are identical or similar to each other and are not described in detail in this embodiment.

[0125] Based on the above embodiment, optionally, after splicing multiple second nucleic acid windows according to the copy number type set to obtain at least one copy number variation region on the chromosome to be tested, the detection method further includes: obtaining copy number parameter data corresponding to each copy number variation region; in response to the copy number parameter data satisfying the parameter screening condition, using the copy number variation region as the final copy number variation region.

[0126] In this embodiment, the copy number parameter data includes at least one of a copy number abnormality rate, a window density, and a regional nucleic acid length.

[0127] Among them, the copy number abnormality rate represents the ratio of the number of second nucleic acid windows of the copy number abnormality type in the copy number variation region to the total number of second nucleic acid windows in the copy number variation region, the window density represents the number of first nucleic acid windows contained in a unit nucleic acid length, and the copy number abnormality type is a copy number deletion type or a copy number duplication type.

[0128] Specifically, the unit nucleic acid length is greater than the window length of the first nucleic acid window. Exemplarily, the unit nucleic acid length can be 1M, but is not limited to the exemplary case.

[0129] In an optional embodiment, the parameter screening conditions include at least one of a copy number abnormality rate greater than a preset abnormality rate threshold, a window density greater than a preset density threshold, and a regional nucleic acid length greater than a preset length threshold.

[0130] Among them, the larger the window density, the greater the possibility that the copy number variation region is a confident copy number variation region. Conversely, the smaller the window density, the smaller the possibility that the copy number variation region is a confident copy number variation region.

[0131] The benefit of setting parameter screening conditions is that it increases the confidence level of copy number variation regions and further improves the accuracy of CNV detection.

[0132] Figure 5 The present invention provides a flowchart of a specific example of a method for detecting chromosome copy number variation provided by one embodiment of the present invention. Specifically, PCR amplification is performed on a cfDNA sample to obtain a nucleic acid library, whole genome sequencing is performed on the nucleic acid library to obtain genome sequencing data, and data quality control is performed on the genome sequencing data. Exemplarily, the quality control tool used for data quality control can be the fastp tool, the Trimmomatic tool, or the FastQC tool. The genome sequencing data that passes the quality control is aligned to the reference genome nucleic acid data of the human reference genome hg19, and the obtained sequence alignment results are filtered and PCR duplications are removed.

[0133] The sequencing depth corresponding to a 20 kbp window length in the reference genome nucleic acid data of the human reference genome hg19 was calculated. The sequencing depth was corrected and median-normalized to obtain initial depth data for each chromosome. Corrections can include correction for effective base length, outlier correction, mappability correction, and GC base content correction. For each chromosome, the initial depth data was corrected for internal references to obtain sequencing depth data. Based on the sequencing depth data and the normal distribution function set, a copy number type set was determined. Based on the copy number type set, 200 kbp windows were spliced ​​to obtain copy number variation regions.

[0134] Because CNV regions are typically small parts of chromosomes, their sequencing depth varies less and fluctuates more than that of entire chromosomes. The difference between the sequencing depth of CNV regions and that of negative regions is also small. Therefore, the degree of fluctuation in sequencing depth is a key factor affecting the accuracy of CNV detection. Furthermore, in different nucleic acid samples, the possible locations of CNV regions on chromosomes vary significantly, making it difficult to construct an external parameter set as a reference for correcting sequencing depth, and the effectiveness of the correction cannot be guaranteed.

[0135] The technical solution of this embodiment corrects the initial sequencing depth corresponding to the first nucleic acid window for each first nucleic acid window according to the reference nucleic acid window set corresponding to the first nucleic acid window to obtain a first sequencing depth, wherein the reference nucleic acid window set includes the third sequencing depth of the nucleic acid sample to be tested on the third nucleic acid window corresponding to other chromosomes other than the chromosome to be tested, thereby achieving the purpose of internal reference correction of the initial sequencing depth of the nucleic acid sample to be tested, reducing the degree of fluctuation of the deep sequencing data, and further improving the accuracy of CNV detection.

[0136] Example 1

[0137] The genome sequencing data of 356 negative nucleic acid samples were provided by the NIPT plus test of the second-generation sequencing with medical device license. Among them, there were 196 negative nucleic acid samples from male fetuses and 160 negative nucleic acid samples from female fetuses.

[0138] The following Table 1 shows the CNV detection results of 356 negative nucleic acid samples provided in Example 1 of the present invention.

[0139] Table 1

[0140]

[0141]

[0142]

[0143]

[0144]

[0145]

[0146]

[0147]

[0148]

[0149] In the test results, del / dup indicates copy number deletion / copy number duplication, chrN indicates chromosome N, p / q indicates the short arm / long arm region, Mb indicates the length of the copy number in million bases, and % indicates the copy number abnormality rate.

[0150] As can be seen from Table 1, the false positive rate of the chromosome copy number variation detection method provided in this embodiment is 10 / 356=2.81%, which meets the requirements of CNV detection.

[0151] Example 2

[0152] A total of 178 nucleic acid samples were used for testing, including 9 positive nucleic acid samples and 169 negative nucleic acid samples, which were provided by NIPT product testing that has obtained medical device license.

[0153] The following Table 2 shows the CNV detection results of 178 nucleic acid samples to be tested provided in Example 2 of the present invention.

[0154] Table 2

[0155]

[0156]

[0157]

[0158]

[0159]

[0160]

[0161] As can be seen from Table 2, the positive detection rate of the chromosome copy number variation detection method provided in this embodiment is 100%, and the false positive rate is 12 / 178=6.74%, which meets the requirements of CNV detection.

[0162] Table 3 below is a contingency table obtained by statistically analyzing Tables 1 and 2 above.

[0163] Table 3

[0164]

[0165] Among them, 503 represents the number of true negative samples, 0 represents the number of false negative samples, 22 represents the number of false positive samples, and 9 represents the number of true positive samples.

[0166] Table 4 below is a summary table of performance indicators obtained by performing statistical analysis on the performance indicators in Table 3 above.

[0167] Table 4

[0168] Positive predictive value Negative predictive value Sensitivity Specificity 29.03% 100% 100% 95.81%

[0169] Among them, the positive predictive value represents the proportion of truly positive samples among samples with positive test results, the negative predictive value represents the proportion of truly negative samples among samples with negative test results, the sensitivity is also called the true positive rate, which represents the proportion of actual positive samples that are correctly detected as positive, and the specificity is also called the true negative rate, which represents the proportion of actual negative samples that are correctly detected as negative.

[0170] It can be seen from Tables 3 and 4 that the detection method for chromosome copy number variation provided in this example has high accuracy.

[0171] The following is an embodiment of a device for detecting chromosome copy number variation provided by an embodiment of the present invention. This device and the method for detecting chromosome copy number variation of the above-mentioned embodiment belong to the same inventive concept. For details not fully described in the embodiment of the device for detecting chromosome copy number variation, reference can be made to the content regarding the method for detecting chromosome copy number variation in the above-mentioned embodiment.

[0172] Figure 6 This is a schematic diagram of a device for detecting chromosome copy number variation provided by one embodiment of the present invention. Figure 6 As shown, the apparatus includes: a sequencing depth data acquisition module 310 , a copy number type set determination module 320 and a copy number variation region determination module 330 .

[0173] The sequencing depth data acquisition module 310 is used to obtain sequencing depth data of the chromosome to be tested in the nucleic acid sample to be tested; wherein the sequencing depth data includes the first sequencing depth under multiple first nucleic acid windows corresponding to the chromosome to be tested;

[0174] A copy number type set determination module 320 is configured to determine a copy number type set based on the sequencing depth data and a normal distribution function set corresponding to the chromosome to be tested; wherein the normal distribution function set includes normal distribution functions corresponding to a copy number deletion type, a copy number duplication type, and a copy number normal type, respectively, and the copy number type set includes a plurality of marker copy number types of the second nucleic acid window;

[0175] a copy number variation region determination module 330 for splicing a plurality of second nucleic acid windows according to the copy number type set to obtain at least one copy number variation region on the chromosome to be tested;

[0176] In which, the second nucleic acid window is the first nucleic acid window, or the second nucleic acid window is a recombinant nucleic acid window obtained by merging multiple first nucleic acid windows based on a preset window number and a preset window step size, and the normal distribution function represents the probability density of the second sequencing depth corresponding to the second nucleic acid window as the value of the second sequencing depth changes.

[0177] The technical solution of this embodiment is to determine the copy number type set corresponding to the second nucleic acid window based on the normal distribution function set corresponding to the chromosome to be tested and the sequencing depth data corresponding to the first nucleic acid window, and splice multiple second nucleic acid windows according to the copy number type set to obtain at least one copy number variation region on the chromosome to be tested, wherein the normal distribution function set includes normal distribution functions corresponding to the copy number deletion type, the copy number duplication type and the copy number normal type, respectively, and the second nucleic acid window is the first nucleic acid window or a recombinant nucleic acid window obtained by merging multiple first nucleic acid windows based on a preset window number and a preset window step size, thereby achieving the purpose of CNV detection based on the window segment characteristics of the chromosome, solving the problem of poor CNV detection precision, and improving the accuracy of CNV detection.

[0178] In an optional embodiment, the copy number type set determination module 320 includes:

[0179] A normal distribution function set acquisition unit, used for acquiring a normal distribution function set in a first iteration process;

[0180] a marker copy number type determination unit, configured to determine, for each second nucleic acid window, a second sequencing depth corresponding to the second nucleic acid window based on the sequencing depth data, and determine, based on the normal distribution function set and the second sequencing depth, probability density values ​​corresponding to the second nucleic acid window and the copy number deletion type, the copy number duplication type, and the copy number normal type, respectively, and use the copy number type corresponding to the maximum probability density value as the marker copy number type of the second nucleic acid window;

[0181] a normal distribution function set updating unit, configured to update the normal distribution function set according to a second sequencing depth of at least one second nucleic acid window corresponding to a normal copy number type, to obtain a normal distribution function set for a next iteration process;

[0182] The returning execution unit is configured to return to executing the step of determining, for each second nucleic acid window, a second sequencing depth corresponding to the second nucleic acid window according to the sequencing depth data;

[0183] The copy number type set determining unit is configured to determine the copy number type set according to the labeled copy number type of each second nucleic acid window in the current iteration until the number of second nucleic acid windows corresponding to the normal copy number type in the chromosome to be tested converges.

[0184] In an optional embodiment, the marker copy number type determination unit is specifically configured to:

[0185] When the second nucleic acid window is a recombinant nucleic acid window, obtaining a first sequencing depth of a first nucleic acid window of a preset number of windows corresponding to the second nucleic acid window in the sequencing depth data;

[0186] The average of the first sequencing depths of the preset number of windows is used as the second sequencing depth corresponding to the second nucleic acid window.

[0187] In an optional embodiment, the normal distribution function set updating unit is specifically configured to:

[0188] Determine a standard deviation corresponding to a second sequencing depth of at least one second nucleic acid window corresponding to a normal copy number type;

[0189] According to the standard deviation, the normal distribution function set is updated to obtain the normal distribution function set for the next iterative process.

[0190] In an optional embodiment, the mean of the normal distribution function corresponding to the normal copy number type in the normal distribution function set is 1, and the mean of the normal distribution function corresponding to the copy number deletion type is The mean of the normal distribution function corresponding to the copy number repeat type is or,

[0191] The mean of the normal distribution function corresponding to the normal copy number type in the normal distribution function set is mean, and the mean of the normal distribution function corresponding to the copy number deletion type is The mean of the normal distribution function corresponding to the copy number repeat type is

[0192] Here, mean represents the median of the second sequencing depths corresponding to the multiple second nucleic acid windows, and FF represents the placenta-derived nucleic acid fraction corresponding to the nucleic acid sample to be tested.

[0193] In an optional embodiment, the copy number variation region satisfies the following splicing conditions:

[0194] 1) The marker copy number types of the first second nucleic acid window and the last second nucleic acid window in the copy number variation region are both copy number abnormality types;

[0195] 2) the number of second nucleic acid windows in the copy number variation region that are labeled as copy number abnormality types is not less than a preset threshold;

[0196] 3) the number of second nucleic acid windows in the copy number variation region that are labeled as a normal copy number type is not greater than a preset threshold;

[0197] 4) The multiple copy number abnormality types corresponding to the copy number variation region are all copy number deletion types or copy number duplication types.

[0198] In an optional embodiment, the device further comprises:

[0199] a copy number variation region screening module, configured to obtain, for each copy number variation region, copy number parameter data corresponding to the copy number variation region after splicing the plurality of second nucleic acid windows according to the copy number type set to obtain at least one copy number variation region on the chromosome to be tested;

[0200] In response to the copy number parameter data satisfying the parameter screening condition, taking the copy number variation region as the final copy number variation region;

[0201] Among them, the copy number parameter data includes at least one of the copy number abnormality rate, window density and regional nucleic acid length. The copy number abnormality rate represents the ratio of the number of second nucleic acid windows of the copy number abnormality type in the copy number variation region to the total number of second nucleic acid windows in the copy number variation region. The window density represents the number of first nucleic acid windows contained in the unit nucleic acid length. The copy number abnormality type is a copy number deletion type or a copy number duplication type.

[0202] In an optional embodiment, the parameter screening conditions include at least one of a copy number abnormality rate greater than a preset abnormality rate threshold, a window density greater than a preset density threshold, and a regional nucleic acid length greater than a preset length threshold.

[0203] In an optional embodiment, the sequencing depth data acquisition module 310 is specifically configured to:

[0204] Obtaining initial depth data of a chromosome to be tested in the nucleic acid sample to be tested; wherein the initial depth data includes initial sequencing depths under a plurality of first nucleic acid windows corresponding to the chromosome to be tested;

[0205] For each first nucleic acid window, obtaining an initial sequencing depth and a reference sequencing depth corresponding to the first nucleic acid window, and taking the ratio of the initial sequencing depth to the reference sequencing depth as the first sequencing depth corresponding to the first nucleic acid window;

[0206] Determining sequencing depth data of a chromosome to be tested in the nucleic acid sample to be tested according to a first sequencing depth corresponding to each first nucleic acid window;

[0207] The reference sequencing depth represents the average of the third sequencing depths corresponding to the nucleic acid sample to be tested and at least one third nucleic acid window in the reference nucleic acid window set, respectively. The third nucleic acid window is a nucleic acid window corresponding to chromosomes other than the chromosome to be tested.

[0208] In an optional embodiment, the device further comprises:

[0209] A reference depth data acquisition module is used to obtain target depth data corresponding to the negative nucleic acid sample set and the chromosome to be tested, as well as reference depth data corresponding to other chromosomes; wherein the target depth data includes target depth sequences under multiple first nucleic acid windows corresponding to the chromosome to be tested, and the target depth sequence includes the negative sequencing depth of at least one negative nucleic acid sample in the negative nucleic acid sample set under the first nucleic acid window; the reference depth data includes reference depth sequences under multiple third nucleic acid windows corresponding to other chromosomes, and the reference depth sequence includes the negative sequencing depth of at least one negative nucleic acid sample under the third nucleic acid window;

[0210] a reference nucleic acid window set determination module, configured to determine, for each first nucleic acid window, depth distance data based on the reference depth data and the target depth sequence corresponding to the first nucleic acid window, and to screen a plurality of third nucleic acid windows based on the depth distance data to obtain a reference nucleic acid window set corresponding to the first nucleic acid window;

[0211] The depth distance data includes depth distances corresponding to the target depth sequence and at least one reference depth sequence.

[0212] In an optional embodiment, the reference nucleic acid window set determination module includes:

[0213] an initial nucleic acid window set determining unit, configured to retain, according to the depth distance data, third nucleic acid windows that meet a specified number of windows in ascending order, to obtain an initial nucleic acid window set corresponding to the first nucleic acid window;

[0214] The reference nucleic acid window set determination unit is used to perform data cleaning on the initial nucleic acid window set to obtain a reference nucleic acid window set corresponding to the first nucleic acid window.

[0215] In an optional embodiment, the reference nucleic acid window set determination unit includes:

[0216] A first reference nucleic acid window set determination subunit is configured to obtain depth distances corresponding to the two third nucleic acid windows in the depth distance data when the initial nucleic acid window set includes two adjacent third nucleic acid windows;

[0217] The third nucleic acid window corresponding to the larger depth distance is deleted from the initial nucleic acid window set to obtain a reference nucleic acid window set corresponding to the first nucleic acid window.

[0218] In an optional embodiment, the reference nucleic acid window set determination unit includes:

[0219] A second reference nucleic acid window set determination subunit is configured to determine coefficient of variation data based on the reference depth sequence of each third nucleic acid window in the initial nucleic acid window set; wherein the coefficient of variation data includes the coefficient of variation of each third nucleic acid window;

[0220] Based on the coefficient of variation data, determine the outlier interval;

[0221] The third nucleic acid window corresponding to the coefficient of variation that does not meet the abnormal value interval in the coefficient of variation data is deleted from the initial nucleic acid window set to obtain a reference nucleic acid window set corresponding to the first nucleic acid window.

[0222] The apparatus for detecting chromosome copy number variation provided in the embodiments of the present invention can execute the method for detecting chromosome copy number variation provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0223] Figure 7 A schematic diagram of the structure of an electronic device provided for one embodiment of the present invention. The electronic device 10 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0224] like Figure 7 As shown, the electronic device 10 includes at least one processor 11 and a memory connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., wherein the memory stores a computer program that can be executed by the at least one processor 11, and the processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 to the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12 and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0225] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information or data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0226] The processor 11 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors that run machine learning model algorithms, digital signal processors (DSP), and any appropriate processors, controllers, microcontrollers, etc. The processor 11 executes the various methods and processes described above, such as the detection method of chromosome copy number variation provided in the above embodiment.

[0227] In some embodiments, the detection method of chromosome copy number variation provided in the above embodiments may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps in the detection method of chromosome copy number variation described above may be performed. Alternatively, in other embodiments, the processor 11 may be configured to perform the detection method of chromosome copy number variation by any other appropriate means (e.g., by means of firmware).

[0228] In particular, according to an embodiment of the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product that includes a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication unit 19, or installed from the storage unit 18, or installed from the ROM 12. When the computer program is executed by the processor 11, the above-mentioned functions defined in the method of the embodiment of the present invention are performed.

[0229] Various embodiments of the systems and techniques described herein can be implemented in the following systems or combinations thereof: digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard parts (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0230] The computer program for implementing the detection method of chromosome copy number variation of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the computer program is executed by the processor, the functions / operations specified in the flow chart and / or block diagram are implemented. The computer program can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0231] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. A computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable storage medium. Examples of machine-readable storage media can include an electrical connection based on at least one line, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0232] To provide interaction with a user, the systems and techniques described herein can be implemented on a terminal device having: a display device (e.g., a cathode ray tube (CRT) or a liquid crystal display (LCD) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball), through which the user can provide input to the terminal device. Other types of devices can also provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0233] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0234] A computing system may include a client and a server. The client and server are generally remote from each other and typically interact via a communication network. The client-server relationship arises through computer programs running on the respective computers and establishing a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, a host product within a cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosts and virtual private server (VPS) services.

[0235] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.

[0236] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.

Claims

1. A method for detecting chromosome copy number variation, characterized in that: include: Acquiring sequencing depth data of a chromosome to be tested in a nucleic acid sample to be tested; wherein the sequencing depth data includes first sequencing depths under a plurality of first nucleic acid windows corresponding to the chromosome to be tested; Determining a copy number type set according to the sequencing depth data and a normal distribution function set corresponding to the chromosome to be tested; wherein the normal distribution function set includes normal distribution functions corresponding to a copy number deletion type, a copy number duplication type, and a copy number normal type, respectively, and the copy number type set includes a plurality of marker copy number types of the second nucleic acid window; splicing the plurality of second nucleic acid windows according to the copy number type set to obtain at least one copy number variation region on the chromosome to be tested; In which, the second nucleic acid window is the first nucleic acid window, or the second nucleic acid window is a recombinant nucleic acid window obtained by merging the multiple first nucleic acid windows based on a preset window number and a preset window step size, and the normal distribution function represents the probability density of the second sequencing depth corresponding to the second nucleic acid window as the value of the second sequencing depth changes.

2. The method according to claim 1, characterized in that Determining the copy number type set according to the sequencing depth data and the normal distribution function set corresponding to the chromosome to be tested includes: Get the normal distribution function set in the first iteration process; For each second nucleic acid window, determining a second sequencing depth corresponding to the second nucleic acid window according to the sequencing depth data, and determining probability density values ​​corresponding to the second nucleic acid window and the copy number deletion type, the copy number duplication type, and the copy number normal type, respectively, according to the normal distribution function set and the second sequencing depth, and using the copy number type corresponding to the maximum probability density value as the marker copy number type of the second nucleic acid window; updating the normal distribution function set according to the second sequencing depth of at least one second nucleic acid window corresponding to the normal copy number type to obtain a normal distribution function set for the next iterative process; Returning to the step of determining, for each second nucleic acid window, a second sequencing depth corresponding to the second nucleic acid window according to the sequencing depth data; until the number of second nucleic acid windows corresponding to normal copy number types in the chromosome to be tested converges, determining a copy number type set according to the labeled copy number type of each second nucleic acid window in the current iteration process; Optionally, determining a second sequencing depth corresponding to the second nucleic acid window according to the sequencing depth data includes: When the second nucleic acid window is a recombinant nucleic acid window, obtaining a first sequencing depth of a first nucleic acid window of a preset number of windows corresponding to the second nucleic acid window in the sequencing depth data; Taking the average of the first sequencing depths of the preset number of windows as the second sequencing depth corresponding to the second nucleic acid window; Optionally, updating the normal distribution function set according to the second sequencing depth of at least one second nucleic acid window corresponding to the normal copy number type to obtain the normal distribution function set for the next iterative process includes: Determine a standard deviation corresponding to a second sequencing depth of at least one second nucleic acid window corresponding to a normal copy number type; According to the standard deviation, the normal distribution function set is updated to obtain a normal distribution function set for the next iterative process; Optionally, the mean of the normal distribution function corresponding to the normal copy number type in the normal distribution function set is 1, and the mean of the normal distribution function corresponding to the copy number deletion type is The mean of the normal distribution function corresponding to the copy number repeat type is or, The mean of the normal distribution function corresponding to the normal copy number type in the normal distribution function set is mean, and the mean of the normal distribution function corresponding to the copy number deletion type is The mean of the normal distribution function corresponding to the copy number repeat type is Here, mean represents the median of the second sequencing depths corresponding to the multiple second nucleic acid windows, and FF represents the placenta-derived nucleic acid fraction corresponding to the nucleic acid sample to be tested.

3. The method according to claim 1 or 2, characterized in that The copy number variation region meets the following splicing conditions: 1) The marker copy number types of the first second nucleic acid window and the last second nucleic acid window in the copy number variation region are both copy number abnormality types; 2) the number of second nucleic acid windows in the copy number variation region that are labeled as copy number abnormality types is not less than a preset threshold; 3) the number of second nucleic acid windows in the copy number variation region that are labeled as a normal copy number type is not greater than a preset threshold; 4) The multiple copy number abnormality types corresponding to the copy number variation region are all copy number deletion types or copy number duplication types.

4. The method according to any one of claims 1 to 3, characterized in that After splicing the plurality of second nucleic acid windows according to the copy number type set to obtain at least one copy number variation region on the chromosome to be tested, the method further includes: For each copy number variation region, obtaining copy number parameter data corresponding to the copy number variation region; In response to the copy number parameter data satisfying a parameter screening condition, taking the copy number variation region as a final copy number variation region; The copy number parameter data includes at least one of a copy number abnormality rate, a window density, and a regional nucleic acid length; the copy number abnormality rate represents the ratio of the number of second nucleic acid windows of the copy number abnormality type in the copy number variation region to the total number of second nucleic acid windows in the copy number variation region; the window density represents the number of first nucleic acid windows contained in a unit nucleic acid length; and the copy number abnormality type is a copy number deletion type or a copy number duplication type. Optionally, the parameter screening condition includes at least one of the copy number abnormality rate being greater than a preset abnormality rate threshold, the window density being greater than a preset density threshold, and the regional nucleic acid length being greater than a preset length threshold.

5. The method according to any one of claims 1 to 4, characterized in that The step of obtaining sequencing depth data of a chromosome to be tested in a nucleic acid sample to be tested comprises: Obtaining initial depth data of a chromosome to be tested in a nucleic acid sample to be tested; wherein the initial depth data includes initial sequencing depths under a plurality of first nucleic acid windows corresponding to the chromosome to be tested; For each first nucleic acid window, obtaining an initial sequencing depth and a reference sequencing depth corresponding to the first nucleic acid window, and taking the ratio of the initial sequencing depth to the reference sequencing depth as the first sequencing depth corresponding to the first nucleic acid window; Determining sequencing depth data of a chromosome to be tested in the nucleic acid sample to be tested according to a first sequencing depth corresponding to each first nucleic acid window; The reference sequencing depth represents the average of the third sequencing depths corresponding to the nucleic acid sample to be tested and at least one third nucleic acid window in the reference nucleic acid window set, respectively, and the third nucleic acid window is a nucleic acid window corresponding to a chromosome other than the chromosome to be tested; Optionally, the method further comprises: Obtain target depth data corresponding to the negative nucleic acid sample set and the chromosome to be tested, and reference depth data corresponding to the other chromosomes; wherein the target depth data includes target depth sequences under multiple first nucleic acid windows corresponding to the chromosome to be tested, and the target depth sequence includes the negative sequencing depth of at least one negative nucleic acid sample in the negative nucleic acid sample set under the first nucleic acid window, and the reference depth data includes reference depth sequences under multiple third nucleic acid windows corresponding to the other chromosomes, and the reference depth sequence includes the negative sequencing depth of at least one negative nucleic acid sample under the third nucleic acid window; For each first nucleic acid window, determining depth distance data according to the reference depth data and a target depth sequence corresponding to the first nucleic acid window, and screening the plurality of third nucleic acid windows according to the depth distance data to obtain a reference nucleic acid window set corresponding to the first nucleic acid window; The depth distance data includes depth distances corresponding to the target depth sequence and at least one reference depth sequence.

6. The method according to claim 5, characterized in that The step of screening the plurality of third nucleic acid windows according to the depth distance data to obtain a reference nucleic acid window set corresponding to the first nucleic acid window includes: retaining third nucleic acid windows that meet a specified number of windows in ascending order according to the depth distance data to obtain an initial nucleic acid window set corresponding to the first nucleic acid window; performing data cleaning on the initial nucleic acid window set to obtain a reference nucleic acid window set corresponding to the first nucleic acid window; Optionally, performing data cleaning on the initial nucleic acid window set to obtain a reference nucleic acid window set corresponding to the first nucleic acid window includes: In a case where the initial nucleic acid window set includes two adjacent third nucleic acid windows, obtaining depth distances respectively corresponding to the two third nucleic acid windows in the depth distance data; Deleting the third nucleic acid window corresponding to the larger depth distance from the initial nucleic acid window set to obtain a reference nucleic acid window set corresponding to the first nucleic acid window; Optionally, performing data cleaning on the initial nucleic acid window set to obtain a reference nucleic acid window set corresponding to the first nucleic acid window includes: Determining coefficient of variation data based on a reference depth sequence of each third nucleic acid window in the initial nucleic acid window set; wherein the coefficient of variation data includes a coefficient of variation of each third nucleic acid window; Determining an outlier interval based on the coefficient of variation data; The third nucleic acid window corresponding to the coefficient of variation that does not satisfy the abnormal value interval in the coefficient of variation data is deleted from the initial nucleic acid window set to obtain a reference nucleic acid window set corresponding to the first nucleic acid window.

7. A device for detecting chromosome copy number variation, characterized in that: include: A sequencing depth data acquisition module is used to obtain sequencing depth data of a chromosome to be tested in a nucleic acid sample to be tested; wherein the sequencing depth data includes a first sequencing depth under a plurality of first nucleic acid windows corresponding to the chromosome to be tested; a copy number type set determination module, configured to determine a copy number type set based on the sequencing depth data and a normal distribution function set corresponding to the chromosome to be tested; wherein the normal distribution function set includes normal distribution functions corresponding to a copy number deletion type, a copy number duplication type, and a copy number normal type, respectively, and the copy number type set includes a plurality of marker copy number types of a second nucleic acid window; a copy number variation region determination module, configured to splice the plurality of second nucleic acid windows according to the copy number type set to obtain at least one copy number variation region on the chromosome to be tested; In which, the second nucleic acid window is the first nucleic acid window, or the second nucleic acid window is a recombinant nucleic acid window obtained by merging the multiple first nucleic acid windows based on a preset window number and a preset window step size, and the normal distribution function represents the probability density of the second sequencing depth corresponding to the second nucleic acid window as the value of the second sequencing depth changes.

8. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can perform the method for detecting chromosome copy number variation according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the method for detecting chromosome copy number variation according to any one of claims 1 to 6 when executed.

10. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the method for detecting chromosome copy number variation according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Copy number variation detection method based on high-throughput sequencing and Gaussian hybrid model

    CN108875311A

  • Method, device and medium for detecting copy number variation

    CN114758720A

  • Model construction for detecting copy number variation of genetic sequence and application thereof

    CN118016150A