Methods for detecting copy number variations and their applications

By dividing genomic regions into windows and analyzing sequencing data from control windows, the method addresses the limitations of conventional copy number variation detection, enhancing accuracy and stability in genomic analysis.

JP7745090B2Active Publication Date: 2025-09-26GUANGZHOU BURNING ROCK DX CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2024514011
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-09-17
Filing Date
2022-08-29
Publication Date
2025-09-26
Estimated Expiration
2042-08-29

AI Technical Summary

Technical Problem

Conventional methods for detecting copy number variations in the human genome suffer from low throughput, high cost, batch effects, and technical errors, particularly in analyzing the entire genome, which complicates the detection of copy number status, especially in tumor samples.

Method used

A method involving dividing genomic regions into windows, obtaining sequencing data from control window regions, and determining copy number status based on this data, using a sliding window approach to reduce batch effects and improve detection stability.

Benefits of technology

The method enhances the accuracy and stability of copy number variation detection, reducing batch effects and errors, thereby improving the reliability of genomic analysis, especially in tumor samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007745090000024
    Figure 0007745090000024
  • Figure 0007745090000025
    Figure 0007745090000025
  • Figure 0007745090000026
    Figure 0007745090000026
Patent Text Reader

Abstract

A method for detecting copy number variation is provided. A method for detecting copy number variation and its application is provided. Also provided is a method for analyzing copy number status, comprising the steps of: dividing a target section of a test sample into several window regions; obtaining sequencing data of control window regions in a group of test samples; and determining the copy number status of a target gene of the test sample based on the sequencing data of the control window regions.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present application relates to the field of bioinformatics, and in particular to methods for detecting copy number variations and their applications. [Background technology]

[0002] Copy number variation (CNV) is a common type of variation in the human genome. It encompasses both gene copy number amplifications and deletions. Detection of copy number variations can be used to monitor a subject's genomic status and to identify associations between specific genomic variations and specific diseases. For example, gene copy number variations can cause many common genetic disorders. For example, deletions of the BRCA1 / 2 gene can increase the risk of hereditary breast cancer. Gene copy number variations can also affect tumor development and progression. For example, HER2 gene amplification is not only related to tumor development and progression, but also serves as an important clinical treatment monitoring and prognostic indicator and an important target for targeted tumor therapy. Therefore, methods for detecting copy number variations play an important role in monitoring a subject's genomic status, genome-wide association studies, the prevention of genetic diseases, and precision tumor therapy. For example, subjects with specific copy number variations may have a higher lifetime risk of developing diseases (e.g., tumors) compared with normal individuals. Therefore, copy number variation detection methods can be used to screen high-risk subjects and individually monitor those subjects for disease, enabling earlier diagnosis and treatment.

[0003] Conventional copy number variation detection methods, such as droplet digital PCR (ddPCR), multiplexed ligation probe amplification (MLPA), and fluorescence in situ hybridization (FISH), can only detect the copy number status of one or a few genes at a time, or only detect the copy number status of specific genes, without being able to analyze the entire genome. These methods are characterized by low throughput and high cost. While numerous high-throughput techniques for copy number variation detection exist, results vary significantly depending on the method, and detection sensitivity and specificity are limited. High-throughput sequencing techniques are prone to batch effects and technical errors during library construction and sequencing. Furthermore, the complexity of tumor samples poses significant challenges to the stability of copy number detection results, making high-throughput sequencing-based copy number variation detection difficult in the field of precision medicine. Analytical methods that can reduce batch effects and errors and / or improve the stability of copy number detection results are urgently needed in this field. Summary of the Invention [Problem to be solved by the invention]

[0004] The purpose of this application is to provide a method for detecting gene copy number abnormalities that addresses the shortcomings of existing technologies. This method can at least reduce batch effects, errors, and / or improve the stability of copy number detection results, which is important for detecting driving events related to copy number abnormalities and interpreting tumor genome evolution information. This application provides a method for detecting copy number variations and its applications. [Means for solving the problem]

[0005] In one aspect, the present application provides a method for analyzing copy number status, the method comprising: dividing a target interval of test samples into several window regions; obtaining sequencing data of control window regions in a group of test samples; and determining the copy number status of a target gene of the test samples based on the sequencing data of the control window regions, and optionally, the control window regions include window regions with a low coverage variation level.

[0006] In one aspect, the present application provides an apparatus for analyzing copy number status, including a receiving module that acquires sequencing data of a group of test samples, a determination module that determines target genes in the test samples, and a judgment module that determines the copy number status of the target genes in the test samples based on the sequencing data of the group of test samples.

[0007] In one aspect, the present application provides a storage medium having stored thereon a program capable of carrying out the methods described herein.

[0008] Those skilled in the art will be able to readily discern other aspects and advantages of the present application from the following detailed description. The following detailed description shows and describes only exemplary embodiments of the present application. As will be recognized by those skilled in the art, the contents of the present application will enable those skilled in the art to make changes to the particular embodiments disclosed without departing from the spirit and scope of the invention to which the present application pertains. Accordingly, the accompanying drawings and description of the present application are illustrative only and are not intended to be limiting. [Brief explanation of the drawings]

[0009] Particular features of the invention to which this application relates are set forth in the appended claims. The features and advantages of the invention to which this application relates can be better understood with reference to the exemplary embodiments described in detail below and the accompanying drawings, a brief description of which follows. [Figure 1]Figures 1A and 1B show diagrams of detection results by a method based on the construction of a reference baseline and the method of the present application. Each box plot shows the distribution of exon copy numbers of the BRCA1 gene for 30 samples. Groups A and B represent batches captured by different probes, respectively. Figure 1A shows the distribution of copy numbers of each exon of the BRCA1 gene calculated by a method based on the construction of a reference baseline. Figure 1B shows the distribution of copy numbers of each exon of the BRCA1 gene calculated by the method of the present application. [Figure 2] Figures 2A and 2B show the results of detecting samples using different NGS (next-generation sequencing) library construction methods, using a method based on the construction of a reference baseline and the method of the present application. The horizontal axis represents chromosomal coordinates, and the vertical axis represents the evaluated copy number (CN) value. Figure 2A shows the results of detecting copy number variations using a method based on the construction of a reference baseline, and Figure 2B shows the results of detecting copy number variations using the method of the present application. The boxes indicate the detected copy number variations. [Figure 3] Figures 3A and 3B show the detection results for samples based on different threshold settings for the stable window in screening. The horizontal axis shows chromosomal coordinates, and the vertical axis shows the evaluated copy number (CN) value. Figure 3A shows the detection results for copy number mutations for samples when the threshold was set at 0.05, and Figure 3B shows the detection results for copy number mutations for samples when the threshold was set at 0.15. The detected copy number mutations are indicated within the boxes. [Figure 4] Figures 4A to 4J show the detection results after establishing a batch baseline for 10 mock samples positive for copy number mutations. The horizontal axis shows the chromosomal coordinates, and the vertical axis shows the evaluated copy number (CN) value. The boxes in the figures indicate the detected copy number mutations. [Figure 5] 5A to 5F show examples of distribution diagrams of copy numbers of some of the data of the detection results of the method of the present application for simulated samples. [Figure 6] 6A to 6C show examples of copy number distribution graphs of some of the data of the detection results of the method of the present application for a standard sample. [Figure 7]7A to 7C show examples of copy number distribution graphs of some of the data on the detection results of the method of the present application for real samples. [Figure 8] 8A to 8F show examples of distribution graphs of copy numbers of detection results of the method of the present application using different baselines for standard sample 1. DETAILED DESCRIPTION OF THE INVENTION

[0010] The following embodiments of the present invention are described by certain specific embodiments, and other benefits and advantages of the present invention may be readily apparent to those skilled in the art from the disclosure herein.

[0011] <Definition> In this application, the terms "second-generation gene sequencing," "high-throughput sequencing," or "next-generation sequencing" generally refer to second-generation high-throughput sequencing technology and subsequent higher-throughput sequencing methods. Next-generation sequencing platforms include, but are not limited to, existing sequencing platforms such as Illumina. As sequencing technology continues to evolve, those skilled in the art will understand that other sequencing methods and devices may also be adopted for use in the methods of the present invention. For example, second-generation gene sequencing may have the advantages of high sensitivity, high throughput, high sequencing depth, or low cost. Depending on the development history, influence, sequencing principles, and technology, there are the following main types of sequencing methods: These include massively parallel signature sequencing (MPSS), polony sequencing, 454 pyrosequencing, Illumina (Solexa) sequencing, Ion semi-conductor sequencing, DNA nano-ball sequencing, and Complete Genomics' DNA nanoarray and probe-anchor ligation hybrid sequencing method. Second-generation sequencing allows for detailed and comprehensive analysis of a species' transcriptome and genome, and is therefore also known as deep sequencing. For example, the method of the present application can be applied to first-generation gene sequencing, second-generation gene sequencing, third-generation gene sequencing, or single-molecule sequencing (SMS).

[0012] In this application, the term "database" generally refers to an organized entity of related data, regardless of how the data or organized entity is represented. For example, the organized entity of related data may take the form of a table, map, grid, group, datagram, file, document, list, or other form. In this application, the database may include any data that is collected and stored in a computer-accessible manner.

[0013] In this application, the term "computing module" generally refers to a functional module for calculation. The computing module may calculate an output value based on an input value to obtain a conclusion or result. For example, the computing module may be primarily used to calculate an output value. The computing module may be a tangible entity such as a processor in an electronic computer, a computer or electronic device equipped with a processor, or a computer network, or may be a program stored on an electronic medium, a command line, or a software package.

[0014] In this application, the term "processing module" generally refers to a functional module for data processing. The processing module may process input values ​​into statistically significant data, for example, classifying the input values. The processing module may be tangible, for example, an electronic or magnetic medium for storing data, a processor in an electronic computer, a computer or electronic device equipped with a processor, or a computer network, or may be a program stored on an electronic medium, a command line, or a software package.

[0015] In this application, the term "determination module" generally refers to a functional module for obtaining a related determination result. In this application, the determination module may calculate an output value based on an input value and obtain a conclusion or result. For example, the determination module may be primarily used to obtain a conclusion or result. The determination module may be a tangible entity such as a processor of an electronic computer, a computer or electronic device equipped with a processor, or a computer network, or may be a program stored on an electronic medium, a command line, or a software package.

[0016] In the present application, the term "sample acquisition module" generally refers to a functional module for acquiring the sample from a subject. For example, the sample acquisition module may include reagents and / or equipment necessary to acquire the sample (e.g., tissue sample, blood sample, saliva, pleural fluid, peritoneal fluid, cerebrospinal fluid, etc.). For example, a blood collection needle, a blood collection tube, and / or a blood sample transport box may be included. For example, a device of the present application may include zero, one, or more of the sample acquisition modules, and may optionally have the functionality to output sample measurements described herein.

[0017] In the present application, the term "receiving module" generally refers to a functional module for obtaining the measurement values ​​in the sample. In the present application, the receiving module may input a sample described in the present application (e.g., a tissue sample, a blood sample, saliva, pleural fluid, peritoneal fluid, cerebrospinal fluid, etc.). In the present application, the receiving module may input measurements of a sample described in the present application (e.g., a tissue sample, a blood sample, saliva, pleural fluid, peritoneal fluid, cerebrospinal fluid, etc.). The receiving module may detect the state of the sample. For example, the data receiving module may optionally perform gene sequencing (e.g., second-generation gene sequencing) described in the present application on the sample. For example, the data receiving module may optionally include reagents and / or equipment necessary to perform the gene sequencing. The data receiving module may optionally detect sequencing depth, counting of sequencing read length, or copy number.

[0018] In the present application, the term "copy number variation" generally refers to the amplification or deletion of the copy number of a target section, a target gene, or a target section in a target gene.For example, the analysis method for copy number variation provided in the present application may be for therapeutic or diagnostic purposes.For example, the analysis method for copy number variation provided in the present application may be used for non-therapeutic or diagnostic purposes, such as determining whether copy number variation exists through sequencing results.

[0019] In this application, the term "sliding window method" generally refers to a method of dividing a window region, for example, a full-length region can be divided into several windows with the same or different window region lengths. For example, a full-length region can be divided into several windows with the same or different step lengths. For example, a full-length region can be divided into several windows with the same window region length and the same step length.

[0020] In this application, the term "quality-passing sample" generally refers to a sample that passes quality control standards. For example, quality-passing sample can refer to a sample that passes average sequencing depth, minimum sequencing depth, and / or coverage uniformity. For example, passing average sequencing depth can refer to a sample whose average sequencing depth is about 100 times or more. For example, minimum sequencing depth passing sample can refer to a sample whose sequencing depth is about 30 times or more. For example, coverage uniformity passing sample can refer to a sample whose base number is 20% or more of the average sequencing depth of the sample and the ratio of the total base number in the sample is about 90% or more.

[0021] In this application, the term "rejected target section" generally refers to the section with low sequencing quality.For example, the rejected section may not be suitable for use in analyzing copy number variation.For example, the rejected section may be inappropriate for reference or baseline construction.In some cases, the rejected section can be excluded to improve the accuracy of detection results.In other cases, the detection results can be obtained with a certain degree of accuracy without screening the rejected section.For example, the rejected section may refer to the section with low sequencing depth, for example, the rejected section may refer to the section that the section varies greatly from sample to sample.

[0022] In this application, the term " low capture efficiency section " generally refers to the section that is difficult to be captured by the probe used.For example, if there is a specific sequence combination in this section, the sequence in a certain section may be difficult to be captured by nucleic acid probe.For example, the section that is low capture efficiency can refer to the section that has low sequencing depth.For example, the section that is low capture efficiency can refer to the section that has sequencing read length count of about 5 or less.

[0023] In this application, the term "unstable section" generally refers to the section where sequencing results vary greatly from sample to sample.For example, it may be the section where results vary greatly across multiple sequencing results of the same sample.For example, it may be the section where sequencing results vary greatly within the same batch of different samples.For example, it may be the section where sequencing results vary greatly between different batches of different samples.For example, it may be the section where sequencing results vary greatly across different reference samples.For example, the method for determining unstable section is to calculate the ratio of the standard deviation of the sequencing depth of a certain section across different samples to the average, and determine whether this ratio is greater than a certain threshold value, for example, the threshold value may be 0.8, or may be adjusted by those skilled in the art based on actual sequencing.

[0024] In this application, the term "test sample" generally refers to a sample that is detected to determine whether copy number variations exist in one or more gene regions on the sample. For example, the test sample or its data can be stored in a storage device before detection.

[0025] In this application, the term "human reference genome" generally refers to the human genome that can serve as a reference in gene sequencing.The information about the human reference genome can be found in UCSC (University of California, Santa Cruz).The human reference genome is available in different versions, for example, hg19, GRCH37 or ensembl 75.

[0026] In this application, the term "GC content" generally refers to the ratio of guanine G and cytosine C in a gene sequence (base sequence) to the total nucleotides contained in the sequence.

[0027] In this application, the term "target sequencing panel" or "panel" generally refers to a group / set of detection targets.For example, in the process of sequencing, one or more target sections are captured and detected by designing one or more probes, and such one or more probes can form a target sequencing panel.For example, a target sequencing panel can be optionally designed for a target gene, target section, or region of interest, such as any number of exon regions.For example, a probe can refer to the oligonucleotide of the target section or the oligonucleotide complementary to the target nucleic acid in research.For example, a target section is the section designed as the target of a probe.

[0028] In this application, the term " sequencing depth " generally refers to the number of times a specific region (for example, a specific gene, a specific section, or a specific base) is detected.Sequencing depth can refer to the sequence of bases detected by sequencing.For example, by comparing sequencing depth with human reference genome and optionally removing overlaps, the number of sequencing read lengths in a specific gene, a specific section, or a specific base position can be determined and counted, and this can be used as sequencing depth.In some cases, sequencing depth can be related to sequencing depth.For example, sequencing depth can be affected by copy number status.

[0029] In this application, the term "sequencing data" generally refers to data of short sequences obtained by sequencing. For example, sequencing data includes the base sequence of the sequenced short sequence (sequencing read length), the number of sequencing read lengths, etc.

[0030] In this application, the term " sequencing bias " generally refers to the bias in the sequencing data generated by different sections.For example, the specific arrangement or base ratio of the sequences in a section can affect the read length that is counted in that section.For example, if a section contains high or low GC content, the counting of the sequencing read length of that section may be biased compared with the section that has GC content close to 50%.

[0031] In this application, the term " distribution similarity " can refer to the distribution similarity between two sets of data.For example, in this application, distribution similarity can refer to the similarity between the counts of the sequencing read lengths of reference sample groups and test samples over one or more intervals.

[0032] In this application, the term " statistical distance " can refer to the distance between the data values ​​of two data groups.For example, in this application, statistical distance can be the statistical quantity of the difference between the counts of the sequencing read length of the reference sample group and the test sample over one or more intervals.For example, statistical distance can be calculated by Euclidean distance, Chebyshev distance, Mahalanobis distance, etc.

[0033] In this application, the term "statistical value" can refer to an analytical value calculated from the data values ​​of a sample. For example, in this application, a statistical value can refer to a mean value, a variance value, a standard deviation, a median value, a plurality of values, etc. Those skilled in the art will appropriately select one or more statistical values ​​to analyze data.

[0034] In this application, the term "probability distribution" generally refers to the distribution of values ​​of a random variable. For example, a probability distribution can take different forms depending on the type of random variable it belongs to. For example, a normal distribution can be used as the probability distribution of a random variable.

[0035] In this application, the term "smoothing" generally refers to a data processing method that reduces the deviation between the differences described herein. For example, it can refer to a method of fitting scattered data to a smoothed line. For example, the smoothing process can be analyzed using a locally weighted regression method. For example, after smoothing, bias caused by a variable (e.g., GC content) in the sample sequencing data can be removed or attenuated by removing the inherent influence of that variable (e.g., GC content) on the sample sequencing data. For example, the smoothing process can include averaging a certain number of the differences described in the specification of this application. For example, the smoothing process can include selecting data values ​​corresponding to different lengths based on a certain interval length and calculating the differences between the different data values. For example, the smoothing process can include dividing the cumulative value of the difference values ​​within a certain length range by the interval length again to obtain a ratio value. For example, the ratio can be considered as the average difference of the difference values ​​in that length range.

[0036] In this application, the term "regression" generally refers to a method for statistically analyzing the relationship between variables. For example, this application can derive a linear or nonlinear relationship between a sample's sequencing data and a certain variable (e.g., GC content) through regression analysis. For example, the relationship between a sample's sequencing data and a certain variable (e.g., GC content) can be obtained through locally weighted regression, and the sample's sequencing data can be adjusted / corrected according to this relationship. For example, correction in this application can refer to processing the sample's sequencing data based on the relationship between the sample's sequencing data and a specific variable to remove or attenuate the bias caused by that variable in the sample's sequencing data.

[0037] In this application, the term "locally weighted regression" generally refers to a method of regression analysis in which weights are locally introduced in the regression analysis of input variables and target variables. For example, locally weighted regression is analyzed and processed by the algorithm (loess(X~Y)) by locally weighting the regression of X according to Y.

[0038] In this application, the term "denoising" generally refers to removing or reducing noise data from data. For example, based on the fact that noise data generally appears as high frequency signals, noisy data can be denoised by extracting useful signals through methods such as transform analysis, principal component analysis algorithms, singular value decomposition and / or Gaussian filtering.

[0039] In this application, the term "cluster analysis" generally refers to the classification of similar objects into groups such that members of the same group share some similar attributes.

[0040] In this application, the term "K-means clustering" generally refers to a method of cluster analysis. For example, K-means clustering is a method of cluster analysis that can classify a set of data into multiple (K) categories based on K clustering centers, where each piece of data has the smallest sum of distances from its nearest clustering center.

[0041] In this application, the term "transform analysis" generally refers to a method of analyzing data. For example, transform analysis can analyze data and use it for further processing by converting the original distribution of the data into a distribution in a transform domain that can be easily solved or manipulated. For example, transform analysis can include a discrete wavelet transform.

[0042] In this application, the term "discrete wavelet transform" generally refers to discretizing the scale and translation of a base wavelet. For example, the discrete wavelet transform can be used as a method for noise removal.

[0043] In this application, the term " standardization " or " normalization " generally refers to the method of converting data.For example, standardization can refer to the process of converting different sets of data into a certain range.For example, standardization can also refer to the process of converting different sets of data into the same median value.For example, standardization in this application can refer to the process of converting the sequencing data of different samples into data with similar median levels.

[0044] In this application, the term "significance test" generally refers to a method for determining whether the difference between a sample and a hypothetical distribution is significant. For example, significance testing can be used to determine whether the copy number variation of a test sample is significant.

[0045] In this application, the term " normal probability distribution " generally refers to the probability distribution of random variables.For example, the occurrence probability of random variables can be determined by normal probability distribution and normal probability distribution density function.For example, the probability of the existence of copy number variation in target section of test sample can be confirmed by normal probability distribution based on the sequencing data of reference sample group.

[0046] In this application, the term "Grubbs test" generally refers to a method for determining and / or filtering outliers. For example, whether a value is an outlier can be determined by determining whether the value fits within the overall distribution range.

[0047] In this application, the term "t-test" generally refers to a form of statistical hypothesis testing using Student's t-distribution. For example, a t-test confirms the significance of copy number variation of a target gene in a test sample.

[0048] In this application, the term "comprising" generally means the inclusion of explicitly specified features but not the exclusion of other elements.

[0049] In this application, the term "about" generally means a variation within ±0.5-10% of the specified value, for example, ±0.5%, ±1%, ±1.5%, ±2%, ±2.5%, ±3%, ±3.5%, ±4%, ±4.5%, ±5%, ±5.5%, ±6%, ±6.5%, ±7%, ±7.5%, ±8%, ±8.5%, ±9%, ±9.5%, or ±10% of the specified value.

[0050] MODE FOR CARRYING OUT THE INVENTION In one aspect, the present application provides a method for analyzing copy number status.

[0051] In one aspect, the present application provides a method for analyzing copy number status, which may include steps of obtaining sequencing data for a group of test samples and determining the copy number status of a target gene in the test samples based on the sequencing data for the group of test samples.

[0052] In one aspect, the present application provides a method for detecting a target gene comprising the steps of: (S1) dividing a region where the target gene is present into several window regions; and obtaining sequencing data of control window regions in the test sample group; (S2) determining the copy number status of the target gene in the test sample based on the sequencing data of the control window region.

[0053] In one aspect, the present application provides a method for manufacturing a method of a semiconductor device comprising: (S1) dividing the region where the target gene is present into several window regions and obtaining sequencing data of the control window regions in the test sample group; (S2) determining the copy number status of the target gene of the test sample based on the sequencing data of the control window region, and arranging the window regions of the quality-passed samples in ascending order of the coverage variation level, wherein the control window region may include the first four or more windows in terms of coverage variation level, and the coverage variation level may be determined based on the ratio of the median absolute deviation to the median of the sequencing data of the window region of the quality-passed samples, or the ratio of the median absolute deviation to the median of the sequencing data of all the quality-passed samples in the control window region may be approximately 0.15 or less.

[0054] In one aspect, the present application provides a method for analyzing copy number status, which may include the following steps: (S1) Step (S1-1): Obtain sequencing data for the window region of all samples in the test sample group. Step (S1-2): Obtain quality-qualified samples in the test sample group, and the quality-qualified samples may include samples that pass the average sequencing depth, minimum sequencing depth, and / or coverage uniformity. Step (S1-3): Normalize the sequencing data for the window region of all samples in the test sample group. (S2) Determine the copy number status of the target gene of the test sample based on the sequencing data of the control window region, and arrange the window regions of the quality-passed samples in ascending order of the coverage variation level, where the control window region may include the top four or more windows of the coverage variation level, and the coverage variation level may be determined based on the ratio of the median to the median of the sequencing data of the window region of the quality-passed samples, or the ratio of the median to the median of the sequencing data of all the quality-passed samples in the control window region may be approximately 0.15 or less. Step (S2-1): Determine a normalization coefficient based on the sequencing data of the control window region. Step (S2-2): Determine the copy number of each window region of the test sample based on the normalization coefficient. Step (S2-3): Determine the significance of copy number variation of the test sample based on the sequencing data of each window region of the test sample and the sequencing data of other samples in the test sample group for the corresponding window region.

[0055] For example, the sequencing data may include sequencing depth. For example, the copy number state may include copy number amplifications and / or deletions. For example, the copy number state may include exon copy number states.

[0056] For example, the test samples may include about 10 or more samples. For example, the test samples may include about 10 or more, about 12 or more, about 15 or more, about 20 or more, about 25 or more, about 50 or more, or about 100 or more samples. For example, the present application does not require a larger number of samples in the same batch. For example, the test sample group may include about 10 or less, about 12 or less, about 15 or less, about 20 or less, about 25 or less, or about 50 or less samples. For example, the copy number status analysis method of the present application can be highly tolerant to the copy number variation level of the test samples. For example, samples containing about 30% copy number variation can be evaluated using the analysis method of the present application. For example, samples containing 10% or less, 15% or less, 20% or less, 25% or less, or 30% or less copy number variation can be evaluated using the analysis method of the present application. For example, the source of the sample in this application may be any sample that contains nucleic acid, such as tissue, blood, saliva, pleural fluid, peritoneal fluid, cerebrospinal fluid, and the like.

[0057] For example, step (S1) of the method of the present application may further include step (S1-1) of obtaining sequencing data of the window region of all samples in the test sample group.For example, the gene sequencing of the present application may include any high-throughput sequencing method, module, or device.For example, the sequencing may be selected from Solexa sequencing technology, 454 sequencing technology, SOLiD sequencing technology, Complete Genomics sequencing method, and semiconductor (Ion Torrent) sequencing technology, and their corresponding devices.

[0058] For example, step (S1-1) of the method of the present application may include dividing the region where the target gene is present into the window regions by a sliding window method. For example, the step length of the sliding window method may be about 24 bases. For example, the length of the window region may be about 120 bases.

[0059] For example, step (S1-1) of the method of the present application may include determining the average sequencing depth of each of the window regions after removing overlapping sequenced fragments.

[0060] For example, step (S1) of the method of the present application may further include step (S1-2) of acquiring quality-passing samples from the test sample group, and the quality-passing samples may include samples that pass the average sequencing depth, minimum sequencing depth, and / or coverage uniformity. For example, the average sequencing depth-passing samples may include samples with an average sequencing depth of about 100 times or more. For example, the minimum sequencing depth-passing samples may include samples with a minimum sequencing depth of about 30 times or more. For example, the individual quality-passing thresholds may be adjusted based on sequencing.

[0061] For example, the coverage uniformity can be related to the sequencing depth of each base of the sample. For example, the coverage uniformity can be calculated by the ratio of the number of bases with an average sequencing depth of 20% or more of the sample to the total number of bases of the sample. For example, the coverage uniformity passing sample can include a sample with a coverage uniformity of about 90% or more. For example, the coverage uniformity passing sample can include a sample with a coverage uniformity of about 90% or more, about 92% or more, about 95% or more, about 97% or more, or about 99% or more.

[0062] For example, the number of quality-qualified samples in the test sample group may be 10 or more.

[0063] For example, step (S1) described in the method of the present application may further include a step (S1-3) of normalizing the sequencing data of the window region of all samples in the test sample group.

[0064] For example, the normalization may include normalizing the sequencing data for each window region of the sample based on the average sequencing depth of all window regions of the sample, and / or normalizing the sequencing data for each window region of the sample based on the GC content of each window region of the sample.

[0065] For example, the normalization may include dividing the sequencing data in each window region of the sample by the sum of the sequencing data in all window regions of the sample, and then multiplying by a factor. For example, the factor may be set based on the size of all intervals. For example, the factor may optionally be 1E+07. For example, the factor may optionally be 1E+100, 1E+20, 1E+10, 1E+09, 1E+08, 1E+07, 1E+06, 1E+05, 1E+04, 1E+03, or 1E+02.

[0066] For example, the normalization may include normalizing the sequencing data for each window region of the sample by regression based on GC content. For example, the regression may include locally weighted regression.

[0067] For example, the control window region may include a window region with a low level of coverage fluctuation.

[0068] For example, the coverage fluctuation level may be determined based on the statistical value of the sequencing data of the window region of the quality-passing samples. For example, the coverage fluctuation level may be determined based on the deviation of the sequencing data of the window region of the quality-passing samples. For example, the coverage fluctuation level may be determined based on the median absolute deviation and / or median of the sequencing data of the window region of the quality-passing samples. For example, the coverage fluctuation level may be determined based on the ratio between the median absolute deviation and the median of the sequencing data for the window region of the quality-passing samples.

[0069] For example, the window regions of the quality-accepting samples may be arranged in ascending order of the coverage variation level, and the control window region may include the first two or more windows of the coverage variation level.

[0070] For example, the window regions of the quality-accepting samples may be arranged in ascending order of the coverage variation level, and the control window region may include the top four or more windows of the coverage variation level.

[0071] For example, the ratio of the median absolute deviation to the median of the sequencing data of all the quality-accepting samples in the control window region may be about 0.15 or less. For example, the ratio of the median absolute deviation to the median of the sequencing data of all the quality-accepting samples in the control window region may be about 0.15 or less, about 0.14 or less, about 0.13 or less, about 0.12 or less, about 0.11 or less, about 0.10 or less, about 0.09 or less, about 0.08 or less, about 0.07 or less, about 0.06 or less, or about 0.05 or less. For example, the ratio of the median absolute deviation to the median of the sequencing data for all the quality-accepting samples within the control window region may be about 0.05 to about 0.15, about 0.07 to about 0.15, about 0.10 to about 0.15, about 0.12 to about 0.15, about 0.05 to about 0.12, about 0.07 to about 0.12, about 0.10 to about 0.12, about 0.05 to about 0.10, about 0.07 to about 0.10, or about 0.05 to about 0.07.

[0072] For example, the step (S2) described in the present application may further include a step (S2-1) of determining a normalization coefficient based on the sequencing data of the control window region.

[0073] For example, the normalization factor may be determined by calculating the average value of the sequencing data of all the quality-passing samples within the control window region.

[0074] For example, before determining the normalization coefficient, coverage level values ​​of abnormal samples in the control window region can be excluded. For example, the abnormal coverage level values ​​may be coverage level values ​​for each of the control window regions that are determined to be abnormal samples using an outlier analysis method. For example, the outlier analysis method may include a Grubbs test. For example, each window may include coverage level values ​​of quality-passing samples within the batch in that window. Then, a Grubbs test may be used to determine whether those coverage level values ​​include outliers, and if so, the outliers may be removed. The Grubbs test may then be repeated for the remaining coverage level values ​​to determine whether there are any abnormalities until no outliers are detected. For example, outlier removal may be stopped when the number of remaining coverage level values ​​is 60% or less, 50% or less, or 40% or less of the number of quality-passing samples, and the remaining values ​​may be used to determine the normalization coefficient.

[0075] For example, the number of remaining samples after excluding the abnormal samples may be 40% or more, 70% or more, 80% or more, 90% or more, 95% or more, or 99% or more of the number of samples before excluding the abnormal samples.

[0076] For example, step (S2) described in the present application may further include a step (S2-2) of determining the copy number of each window region in the test sample based on the normalization coefficient.

[0077] For example, step (S2-2) described in the present application may include determining the copy number of each window region of the test sample by normalizing the sequencing data of each window region of the test sample based on the normalization coefficient.

[0078] For example, the normalization method may include dividing the sequencing data of the test sample in the window region by the normalization coefficient of the window region and multiplying by ploidy. For example, for a male X chromosome, the ploidy may be 1. If the subject is polyploid, the ploidy may be adjusted on a case-by-case basis, for example, to diploid.

[0079] For example, step (S2) described in the present application may further include a step (S2-3) of determining the significance of copy number variations in the test sample based on the sequencing data of each window region of the test sample and the sequencing data of other samples in the test sample group of corresponding window regions.

[0080] For example, step (S2-3) described in the present application may include determining a candidate region for copy number variation based on the copy number of each window region in the test sample.

[0081] For example, the copy number variation candidate region may be determined by region segmentation, which may include determining leading and trailing endpoints of the copy number variation candidate region by a cyclic binary segmentation algorithm.

[0082] For example, step (S2-3) described in the present application may include determining the significance of copy number variation based on the sequencing data of a window region in the copy number variation candidate region of the test sample and the sequencing data of other samples in the test sample group in the corresponding window region. For example, the significance of the copy number variation may be determined by a significance test. For example, the significance test may include a t-test.

[0083] In one aspect, the present application also provides an apparatus for analyzing copy number status, which may include: a receiving module for acquiring sequencing data of a group of test samples; a determination module for determining a target gene in the test samples; and a judgment module for determining the copy number status of the target gene in the test samples based on the sequencing data of the group of test samples.

[0084] For example, in the copy number status analysis device of the present application, the module may be configured to execute the copy number status analysis method described in the present application based on a program stored on the storage medium.

[0085] Methods for analyzing copy number status In one aspect, the present application provides a method for analyzing copy number status, which may include: (S1): obtaining sequencing data of a test sample and / or sequencing data of a plurality of reference samples; (S2): dividing the reference samples into two or more reference sample groups; (S3): determining a reference sample group that is closest to the test sample; and (S4): determining the copy number status of a target gene of the test sample based on the sequencing data of the reference sample group that is closest to the test sample.

[0086] In one aspect, the present application provides an apparatus for analyzing copy number status, which may include: (M1) a receiving module that acquires sequencing data of a test sample and / or sequencing data of a plurality of reference samples; (M2) a processing module that divides the reference samples into two or more reference sample groups; (M3) a calculation module that determines a reference sample group that is closest to the test sample; and (M4) a determination module that determines the copy number status of a target gene of the test sample based on the sequencing data of the reference sample group that is closest to the test sample.

[0087] In one aspect, the present application provides a method for analyzing copy number status, which may include the steps of: (S1) Sequencing data of a test sample and / or sequencing data of multiple reference samples is obtained. Step (S1-1): The sequencing data of the test sample and / or the reference sample is obtained by gene sequencing. Step (S1-2): The sequencing data of the test sample and / or the reference sample is corrected. (S2) Dividing the reference samples into two or more reference sample groups. Step (S2-1): Dividing the reference samples into groups. Step (S2-2): Determining statistics of the sequencing data for the reference sample groups. (S3) A group of reference samples that are closest to the test sample is determined. (S4) Determine the copy number state of the target gene of the test sample based on the sequencing data of the reference sample group closest to the test sample. Step (S4-1): Determine the copy number CN of the target gene of the test sample in the target section i i Step (S4-3): Determine the copy number CN of the target gene in the test sample. g Step (S4-4): Determine the probability of the presence of copy number variation in the test sample in the target interval. Step (S4-5): Determine the sigRatio, which is the ratio of the presence of significant copy number amplification or deletion in the test sample in the target gene. Step (S4-6): Determine the statistical test parameter p for the presence of copy number variation in the target gene in the test sample. The copy number status of the target gene in the test sample is determined by the following: CN g ≧CN thA , sigRatio ≥ sigRatio th , and p ttest ≦p th When the amplification of the copy number of the target gene in the test sample is confirmed, CN g ≦CN thD , sigRatio ≥ sigRatio th , and p ttest ≦p th When the above-mentioned step (a) is completed, it is confirmed that a copy number loss of the target gene in the test sample has occurred, and CN thA <CN g <CN thD, or sigRatio <sigRatio th , or p ttest >p th When the number of copies of the target gene in the test sample is normal, CN thA , C.N. thD , sigRatio th , and ,p th are set as thresholds independently of each other.

[0088] In one aspect, the present application provides a copy number status analysis device that may include modules for performing the copy number status analysis method of the present application.

[0089] In one aspect, the present application provides a method for analyzing copy number status, which may include the steps of: (S1) Sequencing data of a test sample and / or sequencing data of multiple reference samples is obtained. Step (S1-1): The sequencing data of the test sample and / or the reference sample is obtained by gene sequencing. Step (S1-2): The sequencing data of the test sample and / or the reference sample is corrected. (S2) Divide the reference samples into two or more reference sample groups. Step (S2-1): Divide the reference samples into groups. Step (S2-2): Check the statistics of the sequencing data of the reference sample groups. (S3) A group of reference samples that are closest to the test sample is determined. (S4) Based on the sequencing data of the reference sample group closest to the test sample, the copy number status of the target gene of the test sample is confirmed. Step (S4-1): The copy number CN of the target gene of the test sample in the target section i is determined. i Step (S4-2): The copy number of the test sample in the target section is denoised. Step (S4-3): The copy number CN of the test sample in the target gene is determined. gStep (S4-4): Determine the probability of the presence of a copy number variation in the test sample in the target interval. Step (S4-5): Determine the sigRatio, the ratio of the presence of significant copy number amplification or deletion in the test sample in the target gene. Step (S4-6): Determine the statistical test parameter p for the presence of a copy number variation in the target gene in the test sample. ttest The copy number status of the target gene in the test sample is determined by: g ≧CN thA , sigRatio ≥ sigRatio th , and p ttest ≦p th When the amplification of the copy number of the target gene in the test sample is confirmed, CN g ≦CN thD , sigRatio ≥ sigRatio th , and p ttest ≦p th When the above-mentioned step (a) is completed, it is confirmed that a copy number loss has occurred in the target gene of the test sample, and CN thA <CN g <CN thD , or sigRatio <sigRatio th , or p ttest >p th When the test sample is tested, it is confirmed that the copy number of the target gene is normal, and CN thA , C.N. thD , sigRatio th , and ,p th are set as thresholds independently of each other.

[0090] In one aspect, the present application provides a copy number status analysis device that may include modules for performing the copy number status analysis method of the present application.

[0091] For example, the sequencing data of the present application may include counts of sequencing read lengths. For example, the sequencing data of the present application may include counts of sequencing read lengths (reads) in a target gene or target interval.

[0092] For example, the step (S1) or module (M1) of the present application may include a step (S1-1) or module (M1-1) of obtaining the sequencing data of the test sample and / or the reference sample by gene sequencing. For example, the gene sequencing may include second-generation gene sequencing (NGS). For example, the gene sequencing of the present application may include any high-throughput sequencing method, module, or device. For example, the sequencing may be selected from Solexa sequencing technology, 454 sequencing technology, SOLiD sequencing technology, Complete Genomics sequencing method, and semiconductor (Ion Torrent) sequencing technology, and their corresponding devices.

[0093] For example, the test sample and / or the reference sample may include a sample containing nucleic acid. For example, the source of the sample in the present application may be any sample containing nucleic acid, such as tissue, blood, saliva, pleural fluid, peritoneal fluid, cerebrospinal fluid, etc.

[0094] For example, the step (S1-1) or module (M1-1) may include obtaining the sequencing data for each base within a target interval of the test sample and / or reference sample.

[0095] For example, the target interval may include an interval corresponding to a sequence of a targeted sequencing panel. For example, the length of the target interval may be about 20 to about 500 bases. For example, the length of the target interval may be about 20 to about 500 bases, about 50 to about 500 bases, about 100 to about 500 bases, about 200 to about 500 bases, about 20 to about 200 bases, about 50 to about 200 bases, about 100 to about 200 bases, about 20 to about 100 bases, about 50 to about 100 bases, or about 20 to about 50 bases.

[0096] For example, the number of target intervals may be at least about 100. For example, the number of target intervals may be at least about 100, at least about 200, at least about 500, at least about 1,000, or at least about 10,000.

[0097] For example, step (S1) or module (M1) may include step (S1-2) or module (M1-2) of correcting the sequencing data of the test sample and / or reference sample. For example, the method of the present application may not include step (S1-2) or may include only a part of step (S1-2). For example, the apparatus of the present application may not include module (M1-2) or may include only a part of module (M1-2). For example, the following steps in step (S1-2) of the method of the present application may be performed in any order: normalizing the sequencing data of the test sample and / or reference sample, smoothing the sequencing data of the test sample and / or reference sample, and excluding the target section with abnormal GC content. For example, the following modules in module (M1-2) of the apparatus of the present application may be performed in any order. A module for normalizing the sequencing data of the test sample and / or reference sample, a module for smoothing the sequencing data of the test sample and / or reference sample, and a module for excluding the target intervals with abnormal GC content.

[0098] For example, step (S1-2) or module (M1-2) may include normalizing or equalizing the sequencing data of the test sample and / or reference sample. For example, the normalization or equalization may include dividing the sequencing data in the target interval by the sum of the sequencing data in all target intervals of the sample corresponding to the target interval, and multiplying by a factor. For example, the factor may be set based on the size of all intervals. For example, the factor may optionally be 1E+07. For example, the factor may optionally be 1E+100, 1E+20, 1E+10, 1E+09, 1E+08, 1E+07, 1E+06, 1E+05, 1E+04, 1E+03, or 1E+02.

[0099] For example, step (S1-2) or module (M1-2) may include smoothing the sequencing data of the test sample and / or the reference sample. For example, the smoothing may include smoothing the sequencing data of the test sample and / or the reference sample based on sequencing bias using a regression method or an apparatus that describes such a procedure. For example, the regression may include locally weighted regression.

[0100] For example, the sequencing bias may include the number of probes covered in the target interval.

[0101] For example, the sequencing bias may include the GC content of the target interval.

[0102] For example, said step (S1-2) or module (M1-2) may optionally comprise excluding said target intervals with abnormal GC content.

[0103] For example, the target interval having an abnormal GC content may include the target interval having a GC content of about 25% or less and / or the target interval having a GC content of about 75% or more.

[0104] For example, the step (S2) or module (M2) may include a step (S2-1) or module (M2-1) of dividing the reference samples into groups. For example, the reference samples may be derived from the test samples or from samples other than the test samples. For example, a portion of the test samples may be divided into reference samples. For example, the reference samples may be updated. For example, each time sequencing data of a new sample is analyzed, the data of the new sample may be added to the existing database, and a database re-establishment process may be performed.

[0105] For example, the grouping may include grouping the reference samples based on the sequencing data of the target interval.

[0106] For example, the grouping may include grouping the reference samples by a cluster analysis method or an apparatus that describes the procedure.

[0107] For example, the cluster analysis method may include K-means clustering, hierarchical clustering, density clustering, grid clustering, probability model clustering, or neural network model clustering. For example, the cluster analysis method or apparatus describing a procedure therefor may include any of clustering, classification, and grouping methods or apparatus describing a procedure therefor.

[0108] For example, the number of reference samples may be about 30 or more. For example, the number of reference samples may be about 50 or more. For example, the number of reference samples may be about 30 or more, about 40 or more, about 50 or more, about 60 or more, about 70 or more, about 80 or more, about 90 or more, about 100 or more, about 200 or more, about 300 or more, about 400 or more, about 500 or more, or about 1000 or more.

[0109] For example, the grouping may include dividing into about two or more groups.For example, if the sequencing data of all reference samples are more similar, they can be divided into only one group.For example, the grouping may include dividing into about 2 or more, about 3 or more, about 4 or more, about 5 or more, about 6 or more, about 7 or more, about 8 or more, about 9 or more, about 10 or more, about 20 or more, about 30 or more, about 40 or more, about 50 or more, about 60 or more, about 70 or more, about 80 or more, about 90 or more, or about 100 or more.

[0110] For example, the number of the reference samples in each group may be about 30 or more. For example, the number of reference samples in each group may be about 30 or more, about 40 or more, about 50 or more, about 60 or more, about 70 or more, about 80 or more, about 90 or more, about 100 or more, about 200 or more, about 300 or more, about 400 or more, about 500 or more, or about 1000 or more.

[0111] For example, the step (S2) or module (M2) may include a step (S2-2) or module (M2-2) of checking the statistical values ​​of the sequencing data of the reference sample group. For example, the statistical values ​​of the sequencing data of the reference sample group may be provided as respective candidate baselines. For example, checking the statistical values ​​may include calculating the mean value and / or standard deviation of the reference samples of each group in the target interval.

[0112] For example, the step (S2) or module (M2) may include a step (S2-3) or module (M2-3) of excluding failed target sections in the reference sample. For example, the failed target sections may include capture inefficiency sections and / or unstable sections.

[0113] For example, the failed target interval may include a target interval with a sequencing read length count of about 5 or less. For example, the failed target interval may include a target interval with a sequencing read length count of about 30 or less, about 20 or less, about 10 or less, about 5 or less, about 4 or less, about 3 or less, about 2 or less, about 1 or less, or about 0 or less.

[0114] For example, the unacceptable target section may include a target section having a coefficient of variation of about 0.8 or more, where the coefficient of variation is the ratio of the standard deviation to the mean value of the sequencing data of the reference samples of each group in the target section.For example, the unacceptable target section may include a coefficient of variation of about 0.8 or more, about 0.9 or more, or about 1.0 or more.For example, the thresholds of the capture inefficiency section and / or the unstable section may be adjusted according to the sequencing situation.

[0115] For example, the step (S3) or module (M3) may include checking the similarity between the test sample and the group of reference samples.

[0116] For example, confirming the similarity may include confirming the distribution similarity between the reference sample group and the test sample based on the sequencing data of the reference sample group and the test sample in the target interval.

[0117] For example, the similarity may include the proximity between the sequencing data of the reference sample group and the test sample in the target interval.

[0118] For example, the confirmation of the similarity may include confirming the similarity between the distribution of the reference sample group and the distribution of the test sample group using a statistical distance calculation method, a similarity algorithm, or an apparatus that describes such a procedure. For example, the statistical distance may include a statistical value of the difference between the sequencing data of the reference sample group and the test sample in the target interval. For example, the statistical distance may include a statistical value of the absolute value of the difference between the sequencing data of the reference sample group and the test sample in the target interval. For example, the statistical distance may include a statistical value of the pth power of the absolute value of the difference between the sequencing data of the reference sample group and the test sample in the target interval, where p is 1 or greater. For example, the statistical value may include a summation value.

[0119] For example, the high degree of similarity may include a short statistical distance between the reference sample group and the test sample in the target interval.

[0120] For example, the statistical distance may include Minkowski distance. For example, the statistical distance may include European distance, Manhattan distance, Chebyshev distance, Minkowski distance (Manhattan distance when p=1, European distance when p=2, Chebyshev distance when p approaches infinity), etc. For example, the similarity algorithm may include cosine similarity, Pearson's correlation coefficient, Spearman's correlation coefficient, log-likelihood similarity, cross-entropy, etc.

[0121] For example, the copy number status of the target gene in the test sample may include the presence and / or number of copy number variations of the target gene in the test sample.

[0122] For example, the copy number variation may include copy number amplifications and / or deletions.

[0123] For example, the step (S4) or module (M4) may include determining the copy number CN of the target gene of the test sample in the target section i. i The method may include a step (S4-1) or module (M4-1) of determining:

[0124] For example, the CN i Determining the CN is performed by dividing the average value of the sequencing data in the target interval of the target gene of the test sample by the average value of the sequencing data in the target interval of a reference sample group closest to the test sample, and multiplying the result by ploidy. i The method may include obtaining:

[0125] For example, the ploidy may be 2, or, for example, in the case of a male X chromosome, the ploidy may be 1. If the subject is polyploid, the ploidy may be adjusted depending on the particular situation.

[0126] For example, the step (S4) or module (M4) may include a step (S4-2) or module (M4-2) of denoising the copy number in the target interval of the test sample.

[0127] For example, the denoising may include denoising the copy number in the target interval of the test sample using methods such as transform analysis, principal component analysis algorithms, singular value decomposition and / or Gaussian filtering, or devices incorporating such procedures.

[0128] For example, the denoising may include denoising the copy numbers in the target section of the test sample using a discrete wavelet transform or an apparatus that describes such a procedure. For example, the denoising may include denoising the copy numbers in the target section of the test sample using a method such as transform analysis, a principal component analysis algorithm, singular value decomposition, and / or Gaussian filtering, or an apparatus that describes such a procedure.

[0129] For example, the step (S4) or module (M4) may include determining the copy number CN of the target gene in the test sample. g It may include a step (S4-3) or module (M4-3) of determining:

[0130] For example, the target gene may include a gene in which copy number variation is to be determined to occur.

[0131] These include ABL1, ABL2, ABRAXAS1, ACVR1, ACVR1B, AKT1, AKT2, AKT3, AL K, ALOX12B, AMER1, APC, AR, ARAF, ARFRP1, ARID1A, ARID1B, ARID2, ARID5B. ASXL1, ASXL2, ASXL3, ATG5, ATM, ATR, ATRX, AURKA, AURKB, AXIN1, AXIN2, AX L, B2M, BAP1, BARD1, BBC3, BCL10, BCL2, BCL2L1, BCL2L11, BCL2L2, BCL6, BCO R, BCORL1, BIRC3, BLM, BMPR1A, BRAF, BRCA1, BRCA2, BRD4, BRD7, BRINP3, BR IP1, BTG1, BTG2, BTK, CALR, CARD11, CASP8, CBFB, CBL, CCND1, CCND2, CCND3 CCNE1, CD274, CD28, CD58, CD74, CD79A, CD79B, CDC73, CDH1, CDH18, CDK12 CDK4, CDK6, CDK8, CDKN1A, CDKN1B, CDKN1C, CDKN2A, CDKN2B, CDKN2C, CEBPA. CENPA, CHD1, CHD2, CHD4, CHD8, CHEK1, CHEK2, CIC, CIITA, CREBBP, CRKL, CR LF2, CRYBG1, CSF1R, CSF3R, CSMD1, CSMD3, CTCF, CTLA4, CTNNA1, CTNNB1, CUL 3. CUL4A, CXCR4, CYLD, CYP17A1, CYP2D6, DAXX, DCUN1D1, DDR1, DDR2, DDX3X DICER1, DIS3, DNAJB1, DNMT1, DNMT3A, DNMT3B, DOT1L, DPYD, DTX1, DUSP22 EED, EGFR, EIF1AX, EIF4E, EMSY, EP300, EPCAM, EPHA2, EPHA3, EPHA5, EPHA7 EPHB1, EPHB4, ERBB2, ERBB3, ERBB4, ERCC1, ERCC2, ERCC3, ERCC4, ERCC5, ER G、ERRFI1、ESR1、ETV4、ETV5、ETV6、EWSR1、EZH2、EZR、FANCA、FANCC、FANCD2 FANCE, FANCF, FANCG, FANCI, FANCL, FANCM, FAS, FAT1, FAT3, FBXW7, FGF10<h2 style=";text-align:left;direction:ltr">FGF12、FGF14、FGF19、FGF23、FGF3、FGF4、FGF6、FGF7、FGFR1、FGFR2、FGFR3、 FGFR4、FH、FLCN、FLT1、FLT3、FLT4、FOXA1、FOXL2、FOXO1、FOXO3、FOXP1、FRS 2、FUBP1、FYN、GABRA6、GALNT12、GATA1、GATA2、GATA3、GATA4、GATA6、GEN1、GID4、GLI1、GNA11、GNA13、GNAQ、GNAS、GPS2、GREM1、GRIN2A、GRM3、GSK3B、H3 F3A、H3F3B、H3F3C、HDAC1、HDAC2、HGF、HIST1H1C、HIST1H2BD、HIST1H3A、HI ST1H3B、HIST1H3C、HIST1H3D、HIST1H3E、HIST1H3G、HIST1H3H、HIST1H3I、HI ST1H3J、HIST2H3D、HIST3H3、HLA-A、HLA-B、HLA-C、HNF1A、HOXB13、HRAS、HSD3B1、HSP90AA1、ICOSLG、ID3、IDH1、IDH2、IFNGR1、IGF1、IGF1R、IGF2、IGHD、 IGHJ、IGHV、IKBKE、IKZF1、IL10、IL7R、INHA、INHBA、INPP4A、INPP4B、INSR、 IRF2、IRF4、IRS1、IRS2、ITK、ITPKB、JAK1、JAK2、JAK3、JUN、KAT6A、KDM5A、KD M5C、KDM6A、KDR、KEAP1、KEL、KIR2DL4、KIR3DL2、KIT、KLF4、KLHL6、KLRC1、K LRC2、KLRK1、KMT2A、KMT2C、KMT2D、KRAS、LATS1、LATS2、LMO1、LRP1B、LTK、LY N、MAF、MAGI2、MALT1、MAP2K1、MAP2K2、MAP2K4、MAP3K1、MAP3K13、MAP3K14、 MAPK1、MAPK3、MAX、MCL1、MDC1、MDM2、MDM4、MED12、MEF2B、MEN1、MERTK、MET、 MFHAS1、MGA、MIR21、MITF、MKNK1、MLH1、MLH3、MPL、MRE11、MSH2、MSH3、MSH6 、MST1、MST1R、MTAP、MTOR、MUTYH、MYC、MYCL、MYCN、MYD88、MYOD1、NAV3、NBN、NCOA3、NCOR1、NCOR2、NEGR1、NF1、NF2、NFE2L2、NFKBIA、NKX2-1、NKX3-1、NOTCH1、NOTCH2、NOTCH3、NOTCH4、NPM1、NRAS、NRG1、NSD1、NSD2、NSD3、NT5C2、NTHL1、NTRK1、NTRK2、NTRK3、NUP93、NUTM1、P2RY8、PAK1、PAK3、PAK5、PALB2、PALLD、PARP1、PARP2、PARP3、PAX5、PBRM1、PCDH11X、PDCD1、PDCD1LG2、PDGFRA、PDGFRB、PDK1、PGR、PHOX2B、PIK3C2B、PIK3C2G、PIK3C3、PIK3CA、PIK3CB、PIK3CD、PIK3CG、PIK3R1、PIK3R2、PIK3R3、PIM1、PLCG2、PLK2、PMS1、PMS2、PNRC1、POLD1、POLE、POM121L12、PPARG、PPM1D、PPP2R1A、PPP2R2A、PPP6C、PRDM1、PREX2、PRKAR1A、PRKCI、PRKDC、PRKN、PTCH1、PTEN、PTPN11、PTPRD、PTPRO、PTPRS、PTPRT、QKI、RAB35、RAC1、RAD21、RAD50、RAD51、RAD51B、RAD51C、RAD51D、RAD52、RAD54L、RAF1、RARA、RASA1、RB1、RBM10、RECQL4、REL、RET、RHEB、RHOA、RICTOR、RIT1、RNF43、ROS1、RPA1、RPS6KA4、RPS6KB2、RPTOR、RSPO2、RUNX1、RUNX1T1、SDC4、SDHA、SDHAF2、SDHB、SDHC、SDHD、SETD2、SF3B1、SGK1、SH2B3、SH2D1A、SHQ1、SLC34A2、SLIT2、SLX4、SMAD2、SMAD3、SMAD4、SMARCA4、SMARCB1、SMARCD1、SMO、SNCAIP、SOCS1、SOX10、SOX17、SOX2、SOX9、SPEN、SPI1、SPOP、SPTA1、SRC、SRSF2、STAG2、STAT3、STAT4、STAT5A、STAT5B、STAT6、STK11、STK40、SUFU、SYK、TAF1、TBX21、TBX3、TCF3、TCF7L2、TEK、TENT5C、The gene may include a gene selected from the group consisting of TERC, TERT, TET1, TET2, TGFBR1, TGFBR2, TIPARP, TMEM127, TMPRSS2, TNFAIP3, TNFRSF14, TOP1, TOP2A, TP53, TP63, TP73, TRAF2, TRAF3, TRAF7, TRIM58, TRPC5, TSC1, TSC2, TSHR, TYRO3, U2AF1, UGT1A1, VEGFA, VEGFB, VEGFC, VHL, WISP3, WRN, WT1, XIAP, XPO1, XRCC2, XRCC3, YAP1, YES1, ZAP70, ZBTB16, ZBTB2, ZNF217, ZNF703, and ZNRF3. For example, the target gene may include a gene selected from the group consisting of ALK (transcript number may be NM_004304.4), ERBB2 (transcript number may be NM_004448.3), EGFR (transcript number may be NM_005228.3), FGFR1 (transcript number may be NM_023110.2), FGFR2 (transcript number may be NM_000141.4), CDK4 (transcript number may be NM_000075.3), and MET (transcript number may be NM_000245.3).

[0132] For example, the sample in this application is selected from the group consisting of a tissue sample, a blood sample, saliva, pleural fluid, peritoneal fluid, and cerebrospinal fluid.

[0133] For example, the step (S4-3) or module (M4-3) may include determining the exon length of the target gene in the test sample and the copy number CN in the target section i of the test sample. i Based on the above, g may include determining:

[0134] For example, the step (S4-3) or module (M4-3) may calculate the CN based on the following formula: g may include determining:

number

[0135] For example, the step (S4) or module (M4) may include a step (S4-4) or module (M4-4) of determining the probability of the presence of a copy number variation in the test sample in the target interval.

[0136] For example, the probability of the presence of the copy number variation can be calculated by multiplying the probability of copy number amplification occurring in the test sample in the target interval (p a ) and / or the probability of deletion (p d ) may also be included.

[0137] For example, the step (S4-4) or module (M4-4) may include confirming the probability of the existence of the copy number variation based on the mean value and standard deviation of the sequencing data of the test sample in the target interval i and the sequencing data of a group of reference samples closest to the test sample in the corresponding target interval using a probability distribution method or an apparatus that describes the procedure.

[0138] For example, the probability distribution may include a normal probability distribution. For example, the probability distribution may include any general probability distribution. For example, the probability distribution may include any discrete probability distribution. For example, the probability distribution may include any continuous probability distribution.

[0139] For example, the step (S4) or module (M4) may include a step (S4-5) or module (M4-5) of determining the ratio sigRatio of the presence of significant copy number amplification or deletion in the test sample in the target gene.

[0140] For example, the step (S4-5) or module (M4-5) may include dividing the number of target intervals in the target gene in which significant copy number variation occurs by the number of all target intervals in the target gene to obtain the sigRatio.

[0141] For example, the target interval where significant copy number variation has occurred may include the target interval where the copy number variation rate is about 30% or more. For example, the target interval where significant copy number variation has occurred may include the target interval where the copy number variation rate is about 30% or more, about 40% or more, about 50% or more, about 60% or more, about 70% or more, about 80% or more, about 90% or more, or about 95% or more.

[0142] For example, the step (S4) or module (M4) may include a step (S4-6) or module (M4-6) of determining statistical test parameters for the presence of copy number variation in the target gene in the test sample.

[0143] For example, the statistical test parameters may include a p-value determined by a significance test.

[0144] For example, the significance test may include a t-test. For example, the significance test may be any significance test or may be a significance test modified according to actual circumstances.

[0145] For example, the step (S4-6) or module (M4-6) may be configured to calculate a p-value p using a t-test or an apparatus that describes the procedure thereof based on the number of target intervals of the test sample in the target gene, the sequencing data of the test sample in each target interval in the target gene, the standard deviation of the sequencing data of the test sample in each target interval in the target gene, and the mean value and standard deviation of the sequencing data in the target interval of the reference sample group whose corresponding target gene is closest to that of the test sample. ttest This may include verifying that:

[0146] For example, the step (S4) or module (M4) may determine the copy number status of the target gene of the test sample by: CN g ≧CN thA , sigRatio ≥ sigRatio th , and p ttest ≦p th and confirming that the copy number of the target gene in the test sample has been amplified when CN g ≦CN thD , sigRatio ≥ sigRatio th , and p ttest ≦p th and confirming that a copy number loss of the target gene in the test sample has occurred. CN thA <CN g <CN thD , or sigRatio <sigRatio th , or p ttest >p th When the number of copies of the target gene in the test sample is normal, CN thA , C.N. thD , sigRatio th , and p th are each independently threshold values.

[0147] For example, CN thA may be about 2.25 to about 4. For example, CN thA may be about 2.25, about 2.50, about 2.75, about 3.00, about 3.25, about 3.50, about 3.75, or about 4.00.

[0148] For example, CN thD may be about 1.0 to about 1.75. For example, CN thD may be about 0.25, about 0.50, about 0.75, about 1.00, about 1.25, about 1.50, or about 1.75.

[0149] For example, sigRatio thmay be about 0.3 to about 1. For example, sigRatio th may be about 0.3, about 0.4, about 0.5, about 0.6, about 0.7, about 0.8, about 0.9, or about 1.0.

[0150] For example, p th may be about 0.05 to about 0.00001. For example, p th may be about 0.05, about 0.01, about 0.001, about 0.0001, about 0.00001, about 0.000001, or about 0.0000001.

[0151] Building a database In one aspect, the present application provides a database construction method that may include obtaining sequencing data of a plurality of reference samples and dividing the reference samples into two or more reference sample groups.

[0152] For example, the database construction method may include: (S1): obtaining sequencing data of a test sample and / or sequencing data of multiple reference samples; and (S2): dividing the reference samples into two or more reference sample groups.

[0153] In one aspect, the present application provides a database construction device that may include a receiving module that acquires sequencing data of a test sample and / or sequencing data of a plurality of reference samples, and a processing module that divides the reference samples into two or more reference sample groups.

[0154] For example, the database construction device may include (M1) a receiving module that acquires sequencing data of a test sample and / or sequencing data of multiple reference samples, and (M2) a processing module that divides the reference samples into two or more groups of reference samples.

[0155] In one aspect, the present application provides a database construction method that may include the following steps. (S1) Sequencing data of a test sample and / or sequencing data of multiple reference samples is obtained. Step (S1-1): The sequencing data of the test sample and / or the reference sample is obtained by gene sequencing. Step (S1-2): The sequencing data of the test sample and / or the reference sample is corrected. (S2) Divide the reference samples into two or more reference sample groups. Step (S2-1): Divide the reference samples into groups. Step (S2-2): Check the statistics of the sequencing data of the reference sample groups. In one aspect, the present application provides a database construction apparatus that may include modules that implement the database construction method of the present application. In one aspect, the present application provides a database construction method that may include the following steps. (S1) Sequencing data of a test sample and / or sequencing data of multiple reference samples is obtained. Step (S1-1): The sequencing data of the test sample and / or the reference sample is obtained by gene sequencing. Step (S1-2): The sequencing data of the test sample and / or the reference sample is corrected. (S2) Divide the reference samples into two or more reference sample groups. Step (S2-1): Divide the reference samples into groups. Step (S2-2): Check the statistics of the sequencing data of the reference sample groups.

[0156] In one aspect, the present application provides a database construction apparatus that may include modules that implement the database construction method of the present application.

[0157] In one aspect, the present application provides a database construction device, which may include: (M1) a receiving module that acquires sequencing data of a test sample and / or sequencing data of a plurality of reference samples; and (M2) a processing module that divides the reference samples into two or more reference sample groups.

[0158] Methods for analyzing copy number status In one aspect, the present application provides a method for analyzing copy number status based on information from an existing database, which may include determining a reference sample group that is closest to a test sample from two or more reference sample groups; and determining the copy number status of a target gene in the test sample based on sequencing data of the reference sample group that is closest to the test sample.

[0159] For example, the method for analyzing copy number status may include: (S3) determining a group of reference samples that are closest to the test sample; and (S4) determining the copy number status of the target gene in the test sample based on sequencing data of the group of reference samples that are closest to the test sample.

[0160] In one aspect, the present application provides a copy number status analysis device, which may include a calculation module that determines a reference sample group that is closest to the test sample from two or more reference sample groups, and a determination module that determines the copy number status of a target gene of the test sample based on sequencing data of the reference sample group that is closest to the test sample.

[0161] For example, the copy number status analysis device may include (M3) a calculation module that determines a group of reference samples that are closest to the test sample, and (M4) a determination module that determines the copy number status of a target gene of the test sample based on sequencing data of a group of reference samples that are closest to the test sample.

[0162] In one aspect, the present application provides a method for analyzing copy number status, which may include the following steps. (S3) A group of reference samples that are closest to the test sample is determined. (S4) Determine the copy number state of the target gene of the test sample based on the sequencing data of the reference sample group closest to the test sample. Step (S4-1): Determine the copy number CN of the target gene of the test sample in the target section i i Step (S4-3): Determine the copy number CN of the target gene in the test sample. gStep (S4-4): Determine the probability of the presence of copy number variation in the test sample in the target interval. Step (S4-5): Determine the ratio sigRatio of the presence of significant copy number amplification or deletion in the test sample in the target gene. Step (S4-6): Determine statistical test parameters for the presence of copy number variation in the target gene in the test sample. The copy number status of the target gene in the test sample is determined by: CN g ≧CN thA , sigRatio ≥ sigRatio th , and p ttest ≦p th When the amplification of the copy number of the target gene in the test sample is confirmed, CN g ≦CN thD , sigRatio ≥ sigRatio th , and p ttest ≦p th When the above-mentioned step (a) is completed, it is confirmed that a copy number loss of the target gene in the test sample has occurred, and CN thA <CN g <CN thD , or sigRatio <sigRatio th , or p ttest >p th When the number of copies of the target gene in the test sample is normal, CN thA , C.N. thD , sigRatio th , and p th are each independently threshold values.

[0163] In one aspect, the present application provides a copy number status analysis device that may include modules for performing the copy number status analysis method of the present application.

[0164] In one aspect, the present application provides a method for analyzing copy number status, which may include the following steps. (S3) A group of reference samples that are closest to the test sample is determined. (S4) Determine the copy number state of the target gene of the test sample based on the sequencing data of the reference sample group closest to the test sample. Step (S4-1): Determine the copy number CN of the target gene of the test sample in the target section i i Step (S4-2): The copy number of the test sample in the target interval is denoised. Step (S4-3): The copy number CN of the test sample in the target gene is determined. g Step (S4-4): Determine the probability of the presence of copy number variation in the test sample in the target interval. Step (S4-5): Determine the sigRatio, the ratio of the presence of significant copy number amplification or deletion in the test sample in the target gene. Step (S4-6): Determine the parameters of a statistical test for the presence of copy number variation in the target gene in the test sample. The copy number status of the target gene in the test sample is determined by: CN g ≧CN thA , sigRatio ≥ sigRatio th , and p ttest ≦p th When the amplification of the copy number of the target gene in the test sample is confirmed, CN g ≦CN thD , sigRatio ≥ sigRatio th , and p ttest ≦p th When the above-mentioned step (a), it is confirmed that a copy number loss of the target gene in the test sample has occurred, and CNthA <CN g <CN thD , or sigRatio <sigRatio th , or p ttest >p th When the number of copies of the target gene in the test sample is normal, CN thA , C.N. thD , sigRatio th , and p th are each independently threshold values.

[0165] In one aspect, the present application provides a copy number status analysis device that may include modules for performing the copy number status analysis method of the present application.

[0166] Databases, Instruments, and Application Methods In one aspect, the present application provides a database constructed by the method of analyzing copy number status or the method of constructing a database described herein.

[0167] In one aspect, the present application further provides a storage medium having stored thereon a program capable of performing the methods described herein.

[0168] In one aspect, the present application further provides an apparatus that may include the storage medium described herein. For example, the non-volatile computer-readable storage medium may include a floppy disk, a flex disk, a hard disk, a solid-state storage (SSS) (e.g., a solid-state drive (SSD)), a solid-state card (SSC), a solid-state module (SSM)), an enterprise flash drive, a magnetic tape, or any other non-transitory magnetic medium. Non-volatile computer-readable storage media also include punch cards, paper tape, photomarker sheets (or other physical media having hole patterns or other optically identifiable markings), compact disk read-only memory (CD-ROM), compact disk re-writable (CD-RW), digital universal disk (DVD), Blu-ray disk (BD), and / or other non-transitory optical media.

[0169] For example, the apparatus of the present application may further include a processor coupled to the storage medium, the processor being configured to implement the method described in the present application based on execution of a program stored in the storage medium.

[0170] In one aspect, the present application further provides applications of the methods of the present application in the diagnosis, prevention and / or treatment of diseases.

[0171] In one aspect, the present application further provides an application of the method of the present application in monitoring the copy number status of a target gene.

[0172] In one aspect, the present application further provides the application of the methods of the present application in genome-wide association studies.

[0173] In the present application, the method may be used to determine whether the subject has a copy number variation. For example, any one or more of the methods of the present application may be for non-diagnostic purposes. For example, any one or more of the methods of the present application may be for diagnostic purposes.

[0174] In the present application, the method can be used in clinical practice (e.g., to predict whether a particular tumor treatment regimen is suitable for a subject) by detecting the copy number variation. In some cases, the level of copy number variation detected by the method can be used in clinical practice in combination with biomarkers known in the art.

[0175] Without wishing to be limited by any theory, the examples described below are merely intended to illustrate the methods and uses of the present application and are not intended to limit the scope of the invention of the present application.

[0176] <Example> Example 1 1.1 Data Preparation Thirty negative peripheral blood samples were selected, and DNA was extracted from the peripheral blood using the same batch of experimental reagents. A whole-genome prelibrary was then constructed through the experimental steps of fragmentation, linker addition, and PCR amplification. The prepared prelibrary was then divided into two aliquots, designated batch A and batch B. The aliquots were hybridized with probes from different batches, designated batch A and batch B, to specifically capture the BRCA1 gene in the human genome, yielding final libraries A and B. These two final libraries were then subjected to high-throughput sequencing using a sequencer. Finally, the sequencing data was aligned with the human genome reference sequence hg19, yielding an aligned BAM file.

[0177] 1.2 Conventional methods for detecting copy number variations based on the construction of a reference baseline A reference baseline was established in advance using a reference set of sufficient negative samples (e.g., 50 or more cases) with normal copy numbers collected in the first trimester. Copy number values ​​for each exon of the BRCA1 gene were then calculated using the baseline established from this reference set for two groups of experimental samples, and copy number variations were detected. The exon copy number calculation results (shown in Figure 1A) indicated that the experimental data captured with the Batch A probe were highly uniform, approaching the theoretical copy number of 2. However, the results captured with the Batch B probe were relatively poor, particularly for exon 8 of the BRCA1 gene, where the preference for all samples was significantly lower. Meanwhile, the copy number variation detection results for the experimental group using the Batch B probe revealed that two BRCA1-derived false-positive copy number variations were detected in 30 samples. This suggests that the use of conventional baseline reference methods is prone to reduced accuracy in detecting copy number variations due to differences in probe batches.

[0178] 1.3 Detection of copy number variations by the methods of the present application Therefore, copy number variations were then detected using the method of the present application.

[0179] (1) Data preparation The copy number variation detection algorithm of this application requires the selection of a sufficient number of samples, for example, 15 samples of the same sample type and experimental method, with the same reagent batches and experimental equipment used as closely as possible. The sample data must be obtained from the BAM file after alignment of the NGS sequencing data.

[0180] (2) Statistics based on the deduplication of BAM files and the coverage depth of DNA sequence fragments For each sample, the BAM file was first depleted of duplicated DNA fragments introduced by PCR during NGS library construction to obtain uniquely aligned DNA fragments. The target DNA region was then divided into windows with a fixed probe length of 120 bp and sliding window size of 24 bp each, and the average coverage level of the uniquely aligned DNA fragments within each window was calculated.

[0181] (3) Quality control of sample sequencing coverage Optionally, quality control was performed on each sample to determine whether the average sequencing depth, minimum sequencing depth, and coverage uniformity met the requirements. The average sequencing depth was ≥ 100x, the minimum sequencing depth was ≥ 30x, and coverage uniformity was ≥ 90% (referring to the percentage of bases whose sequencing depth is ≥ 20% of the sample's average sequencing depth, as determined by the formula: coverage uniformity = (number of bases whose sequencing depth is ≥ 20% of the sample's average sequencing depth / total number of bases in the sample) × 100%). If the quality of the sample data does not meet the requirements, it will not be used to construct the modified baseline. The detection method of the present application can detect at least 10 samples that meet the quality requirements.

[0182] (4) Data correction and normalization To minimize the impact of noise and systematic bias on the detection of copy number variations, the coverage level of each window region can be corrected using various methods, including initial coverage level correction (based on the sample average coverage level), GC correction, and batch correction.

[0183] (5) Initial coverage level correction The coverage level initial correction is to correct for differences in coverage depth among different sequenced samples by correcting the coverage levels of all samples in a batch to the same specified coverage level. Specifically, for each window region of each sample in the batch, the average coverage level by sequencing is divided by the sum of the average coverage levels of all window regions in that sample, and then multiplied by a fixed coefficient (the coefficient is 1E+07).

[0184] (6) GC correction To correct for differences in sequencing coverage depth due to GC preference, GC correction was performed by calculating the GC content of each window and correcting the coverage level of each window region within a sample for GC preference using Loess regression.

[0185] (7) Batch correction i. Obtain GC calibration data for all quality control passing samples in the batch. ii. For samples involved in batch baseline construction, calculate the median and median absolute deviation (MAD) of coverage levels within each window. If MAD / median > set threshold (e.g., set threshold can be about 0.05 to about 0.15), this indicates that the coverage level of that window is unstable and should be excluded. iii. The windows with MAD / median < the set threshold or the top four windows with the smallest MAD / median are kept as the window regions with stable coverage levels. iv. Next, for each window region where the retained coverage level has stabilized, the Grubbs test is used to remove any abnormal coverage level values ​​within that window, and the average of the remaining coverage level values ​​is calculated as the batch-corrected reference coefficient. v. Finally, for each test sample, the coverage level of each window region is normalized based on the batch-corrected reference coefficient calculated above, and the copy number CN value is calculated. The calculation formula for the copy number CN value of each window is as follows:

number

[0186] (8) Identification of copy number variations Using the CBS algorithm, identify the breakpoint position of the target region of the sample, and obtain copy number variation region candidate.Then, perform significance test for each copy number variation candidate region, specifically, use t-test to determine whether the window coverage level of the test sample in the copy number variation candidate region is significantly different from the coverage level of other samples in the batch in that region, and determine the reliability of the copy number variation candidate.

[0187] The distribution of copy numbers of BRCA1 gene exons is shown in Figure 1B. Compared with conventional methods based on constructing a reference baseline, the method of the present application exhibited superior uniformity in copy number results, particularly for group B probes, which have large batch variations. No false-positive copy number variations were detected in either of the two sets of experimental data.

[0188] Example 2 Twenty cell line samples were selected, 19 of which were negative and one was a known exon copy number variation (LGR) sample (BRCA1: Exon 12 amp). In the experiment, high-throughput sequencing data was obtained using an automated library construction method using an instrument. Finally, the sequencing data was aligned with the human genome reference sequence hg19 to obtain aligned BAM files. Copy number variations were detected from the sample BAM files using a conventional method based on the construction of a reference baseline and the method of the present application. Here, the baseline used in the method based on the construction of a reference baseline may be established using sample data from a previous manual library construction method (e.g., the reference baseline used in Example 1).

[0189] The results for positive samples containing copy number variations are shown in Figures 2A-2B. The results for the conventional method based on constructing a reference baseline (shown in Figure 2A) show very high background noise and are unable to detect copy number variations, whereas the data from the method used in this application show significantly lower background noise and are able to detect copy number variations (shown in Figure 2B). This indicates that NGS data generated by different experimental methods can vary significantly, and baselines constructed from data from manual library construction methods cannot be applied to data from automated library construction methods. This indicates that when changing experimental methods, using conventional reference baseline methods requires collecting sufficient sample data using the experimental method in advance and manually constructing a new baseline, which increases experimental costs and wastes manpower.

[0190] Example 3 To detect BRCA1 and BRCA2 exon copy number variations (LGR), 669 peripheral blood samples were selected and subjected to experiments using RNA probes specifically capturing the BRCA1 and BRCA2 gene regions, followed by high-throughput sequencing. The sequencing data was aligned with the human genome reference sequence hg19 to obtain an aligned BAM file. Copy number variations were then detected using a method based on the construction of a reference baseline and the method proposed in this application. Copy number variations were then confirmed in all samples using the BRCA MASTR Plus Dx kit (based on multiplex PCR capture methodology), resulting in a total of 17 LGR-positive samples and 679 LGR-negative samples.

[0191] Using the detection results of the BRCA MASTR Plus Dx kit as the standard, the sensitivity and specificity of the detection results of the conventional method based on the construction of a reference baseline and the detection results of the method of the present application were determined for these 696 peripheral blood samples, and are shown in Tables 1 and 2, respectively. [Table 1] [Table 2]

[0192] Comparing Table 1 with Table 2, it can be seen that compared with the conventional method of establishing a baseline, the method of the present application can significantly reduce false positives in samples without compromising sensitivity, and can improve detection accuracy from 75.3% to 98.9%.

[0193] Example 4 To construct the batch baselines, we selected data from 14 cell line samples from the sequencing alignment. We set the thresholds representing the coverage variation levels of the window coverage at 0.05 and 0.15, respectively, and constructed two batch baselines. We then performed batch correction using each of the two batch baselines for 14 samples with known copy number mutations in LGR (BRCA1: exon 4-6 del) to detect copy number mutations.

[0194] Figures 3A-3B show the results of a positive sample containing copy number variations. The batch baselines constructed based on thresholds with different window coverage variation levels were able to clearly detect copy number variations, demonstrating that the threshold range of the screening stability interval of this application can achieve the detection of copy number variations.

[0195] Example 5 Ten negative cell line samples were selected as the mock sample background, and ten LGR copy number mutations in the BRCA1 and BRCA2 genes reported in the literature (shown in Table 3) were then selected as the mutations to be simulated (five copy number amplification mutations and five copy number deletion mutations). After simulation, the above copy number amplification mutations and copy number deletion mutations were artificially added to the mock sample background data, ultimately obtaining 10 positive LGR mock sample data.

[0196] A batch baseline was constructed for 10 mock-positive samples, and batch correction and identification of copy number variations were performed for the 10 mock samples using the constructed batch baseline. The results for the 10 mock samples are shown in Figures 4A to 4J. All 10 mock copy number variations were accurately detected, demonstrating that the present application achieved accurate detection of copy number variations in all regions. [Table 3]

[0197] Example 6 The method of the present application uses a clustering method to divide a large amount of real sample data into different sample sets according to the trend clustering of sequencing depth, construct baselines (average depth and depth variation range) for each, and dynamically screen background baselines according to the similarity between the sample and the baseline, eliminating batch effects while improving detection specificity and sensitivity. Meanwhile, an optional discrete wavelet transform method can be used to smooth copy numbers, reduce noise, and improve the S / N ratio of sequencing data.

[0198] The present application aims to provide a method and medium for detecting copy number variations with high sensitivity and accuracy in high-throughput sequencing. This method detects copy number variations based on differences in the specificity of sample coverage features. Specifically, it establishes multiple control baselines based on cluster analysis of large-scale samples, avoiding the problem of baseline discrepancies caused by inconsistencies in coverage depth features due to differences in sequencing experiments or samples, and integrates various coverage depth correction strategies to reduce data variability in sample specificity. Finally, quantitative analysis and statistical difference assessment ensure the accuracy and stability of results. The copy number variation detection method described in the present application can be applied not only to targeted capture sequencing data of specific gene panels, but also to whole-exome capture sequencing data. Based on the above, the database construction method described in the present application may include the following steps.

[0199] 1. A data preparation module, including: a) Sequence alignment: The raw fastq data from high-throughput sequencing is posted to the human reference genome to determine the sequences of the target sections of the test sample and reference sample that match the reference sequences of the human reference genome. b) Removal of repeat sequences: Removal of repeat sequences generated during PCR amplification. c) Coverage depth calculation: Sequencing depth RD of each base in the target intervalBase Calculate.

[0200] 2. A coverage depth correction module, which includes three independent, sequential optional corrections. a) Normalization of the sequencing data: sample total sequencing depth correction, specifically, normalizing the coverage depth of each locus in each target interval according to the sample total sequencing depth, eliminating the difference in the amount of sequencing data between different samples, and RD normD get.

number

[0201] 3. A baseline construction module, including the steps of: a) Sample clustering: In existing methods, all reference samples are generally treated as one classification to establish baseline.The method of the present application groups reference samples, and specifically performs cluster analysis based on the consistency of the coverage depth of each reference sample that varies in target interval, for example, the similarity of the sequencing of reference samples in target interval, and divides reference samples into different categories of reference sample groups.Clustering method can be, for example, K-means clustering, hierarchical clustering method, etc. b) Baseline Construction: For each reference sample group, a baseline is constructed. Specifically, the mean sequencing depth (MeanRD) of all reference samples in each reference sample group in each target interval is calculated. i Baseline ) and standard deviation of sequencing depth (SdRD) i Baseline ) to the baseline i ) where i = {1, 2, 3, 4, ...}. For example, to be statistically significant, each reference sample group must have a sufficient number of samples, and the number of samples in each reference sample group should not be less than 30. The number of reference sample groups should be set taking into account the characteristics and sequencing quality of the tumor sample. The number of reference sample groups is determined according to the number of captured features. For example, the number of reference sample groups can be 2 or more, e.g., 2 to 10. c) Optionally, interval screening can be performed: calculate the coefficient of variation cv of sequencing depth in each interval and remove unstable intervals that vary greatly among samples.

number

[0202] Finally, the database of the present application is obtained according to this embodiment and includes two or more reference sample groups with consistent changes in the target interval.Compared to the prior art, the database of the present application has the advantage that, through cluster analysis of large-scale samples, the reference samples are divided into different categories of reference sample groups, and sample-specific background baselines are individually constructed, thereby significantly reducing false positives caused by batch effect in the detection of copy number variations in high-throughput sequencing data and improving the stability of results.In addition, the method of eliminating batch effect in the present application does not require a sufficient number of samples from the same gene panel in the same batch, greatly reducing the difficulty of practical application.

[0203] Example 7 To solve the problem of baseline discrepancies in sequencing due to discrepancies in coverage feature depths between different experiments or samples, the present application also provides a method for analyzing copy number status. The method for analyzing copy number status of the present application may include the following steps: a) According to the similarity between the test sample and the reference sample group, the reference sample group closest to the test sample is determined, i.e., the baseline is dynamically screened, i.e., the sequencing depth of the test sample in each target interval is compared with the sequencing depth of each reference sample group in that target interval by a method of calculating statistical distance such as Minkowski distance, and the statistical distance is confirmed.

number

number

number

number

number

number

[0204] The copy number status analysis method of the present application dynamically screens a group of reference samples that are closest to the test sample as a background baseline based on the similarity between the test sample and a group of reference samples, thereby eliminating batch effects and improving the specificity and sensitivity of detection.

[0205] Example 8 Database construction: A baseline is constructed using 655 reference samples, and the reference samples are divided into five reference sample groups and five different baseline candidates are constructed as a database using the database construction method of the present application, such as using the k-means clustering algorithm.

[0206] Construction of simulated data: Using the varBen tumor mutation data simulation software (github.com / nccl-jmli / VarBen), amplification of target genes with different copy numbers was simulated using gradients by inserting read segments of the target genes into the sequencing data based on benign tissue samples. A list of simulated samples is shown in Table 4. [Table 4]

[0207] The mock samples were assayed according to the copy number status analysis method of the present application, and the results are shown in Table 5. [Table 5]

[0208] Figures 5A-5F show example copy number distribution maps of some of the test results from the present application. Each dot represents a gene interval, with gray dots representing genes with normal copy numbers and black dots representing genes with amplified or deleted copy numbers. The corresponding gene names are also shown. The horizontal axis represents the chromosomal location of the gene, and the vertical axis represents the copy number calculated using the method of the present application (the horizontal line in the middle represents the copy number of the normal gene). The gray background represents the range of variation in each target interval relative to the background baseline (the reference sample group closest to the target sample). Figures 5A-5C show simulations of the ERBB2 gene with different degrees of copy number amplification, and Figures 5D-5F show simulations of the FGFR1 gene with different levels of copy number amplification, with copy number gradients of 2.5, 2.75, and 3.0, respectively. These results demonstrate that when the copy number status analysis method of the present application is used on simulated samples, copy number amplification at different gradients for all simulated genes can be stably detected, and copy number prediction is accurate.

[0209] Example 9 Positive standard samples: The measurements in this application included 30 positive standard samples derived from the NCI-BL2009 cell line. To obtain CNV positive data, these samples were transfected with corresponding proportions of target genes into the cell line and quantified using microtiter digital PCR (ddPCR). The plasmid numbers were Life RPCI11.C-433C10 BAC-EGFR, Life RPCI11.C-936I7 BAC-CDK4, Life RPCI11.C-163C9 BAC-MET, Life RPCI11.C-909L6 BAC-ERBB2, and Life RPCI11.C-957P17 BAC-FGFR1. The list of positive standard samples is shown in Table 6. [Table 6]

[0210] Database construction: A baseline was constructed using 655 reference samples, and the reference samples were divided into five reference sample groups using the database construction method of the present application, such as using the k-means clustering algorithm, to construct five different baseline candidates as a database.

[0211] According to the method for analyzing copy number status of the present application, the copy number status of the copy number amplification positive standard samples was detected, and the results are shown in Table 7. [Table 7]

[0212] Figures 6A to 6C show examples of copy number distribution diagrams for some of the detection result data of the present application. Figures 6A to 6C show the detection results of a plasmid-transfected CNV-positive cell line standard sample, with ddPCR calibration copy numbers of 3, 5, and 8, respectively. These results demonstrate that the method of the present application for the plasmid-transfected cell line standard sample can stably detect all genes and different copy number states with accurate copy number prediction.

[0213] Example 10 Actual data: The actual samples measured in this application included 20 ERBB2 amplification-positive samples confirmed by third-party immunohistochemistry (IHC) detection, and a list of the actual samples is shown in Table 8. [Table 8]

[0214] Database construction: A baseline was constructed using 443 reference samples, and the database construction method of the present application, such as using the k-means clustering algorithm, was used to divide the reference samples into reference sample groups, and different baseline candidates were constructed as a database.

[0215] According to the method of analyzing copy number status of the present application, the copy number status of the real samples was detected, and the results are shown in Table 9. [Table 9]

[0216] Figures 7A to 7C show examples of copy number distribution diagrams of some of the detection result data of the present application. Figures 7A to 7C show the detection results of actual ERBB2-positive samples. These results show that when the method of the present application was used on actual samples, 20 samples that were HER2-positive by IHC were consistently detected.

[0217] Example 11 Positive standard samples: The measurements in this application include three positive standard samples of the same origin as in Example 9 to measure results from different baselines. The list of positive standard samples is shown in Table 10. [Table 10]

[0218] Database construction: 655 reference samples were used to construct a baseline, and the database construction method of the present application, such as using the k-means clustering algorithm, was used to divide the reference samples into five reference sample groups and construct five different baseline candidates as a database. In addition, all the reference samples were constructed as one baseline without using the clustering method.

[0219] Five baselines obtained from the clustering algorithm and one baseline without clustering were used as reference controls to detect the copy number status of the amplified standard samples. The baseline selection and sample variation are shown in Table 11, and the detection results are shown in Table 12. [Table 11] [Table 12]

[0220] 8A to 8F are diagrams showing examples of copy number distributions for standard sample 1 using different baselines.

[0221] As a result, the optimal baseline matched using the inventive method of the present application was most similar to the test sample (the distance value between the test sample and the baseline was the smallest), the copy number variation (SD) of the entire test sample was the smallest, and the copy number distribution map was the most stable and had the least noise, indicating that the detection results of the method of the present application were more stable.In the present application, all genes can be stably detected even in different copy number states, but when the copy number is 3, they cannot be stably detected with other baselines.

[0222] The foregoing detailed description is provided by way of illustration and example, and is not intended to limit the scope of the appended claims. Many variations of the embodiments presently recited in this application will be apparent to those skilled in the art and are intended to fall within the scope of the appended claims and their equivalents.

Claims

1. A method for analyzing copy number status, comprising: dividing a target section of a test sample into several window regions; obtaining sequencing data of control window regions in a group of test samples; and determining the copy number status of a target gene of the test sample based on the sequencing data of the control window regions; The step of acquiring sequencing data of a control window region in the test sample group includes acquiring sequencing data of a window region of all samples in the test sample group, acquiring quality-qualified samples in the test sample group, and determining the control window region based on a coverage variation level of the window region of the quality-qualified samples; The quality-passing samples include samples that pass for average sequencing depth, minimum sequencing depth, and / or coverage uniformity; The coverage variation level is determined based on statistical values ​​of sequencing data in a window region of the quality-accepting sample; The step of determining the copy number status of the target gene of the test sample based on the sequencing data of the control window region includes: determining a normalization coefficient based on the sequencing data of the control window region; determining the copy number of each window region of the test sample based on the normalization coefficient; and determining the significance of the copy number variation of the test sample based on the sequencing data of each window region of the test sample and the sequencing data of other samples in the test sample group of the corresponding window region; The method further comprises correcting the coverage level of each window region, including an initial coverage level correction based on a sample average coverage level, a GC correction, and a batch correction. method.

2. (S1) obtaining sequencing data of the test sample and / or sequencing data of a plurality of reference samples; (S2) dividing the reference samples into two or more reference sample groups; (S3) determining a group of reference samples that are closest to the test samples; The method of claim 1, further comprising: (S4) determining the copy number status of the target gene of the test sample based on sequencing data of the reference sample group that is closest to the test sample.

3. sorting the window regions of the quality-accepting samples in ascending order of coverage variation level; 3. The method of claim 1 or 2, wherein the control window region includes the top two or more windows, or the top four or more windows of the coverage variation level, or the ratio of the median absolute deviation to the median of the sequencing data of all the quality-accepting samples in the control window region is 0.17 or less.

4. The method according to any one of claims 1 to 3, wherein the coverage variation level is determined based on the ratio of the median absolute deviation to the median of the sequencing data of the window region of the quality-accepting samples.

5. 2. The method of claim 1, wherein the normalization factor is determined by calculating the average value of the sequencing data of all the quality-passing samples in the control window region.

6. 2. The method of claim 1, wherein the normalization comprises dividing the sequencing data of the test sample in the window region by a normalization factor for the window region and multiplying by a ploidy.

7. The method according to any one of claims 1 to 6, comprising a step of determining the significance of copy number variations in the test sample based on sequencing data of each window region of the test sample and sequencing data of other samples in the test sample group for the corresponding window region.

8. The method of claim 7, wherein the significance of the copy number variation is determined by a t-test significance testing method.

9. The method according to any one of claims 1 to 8, wherein the test sample is selected from the group consisting of a tissue sample, a blood sample, saliva, pleural fluid, peritoneal fluid, and cerebrospinal fluid.

10. The step (S2) includes a step (S2-1) of grouping the reference samples, and the grouping includes grouping the reference samples by a cluster analysis method based on the sequencing data of the target interval; The method according to any one of claims 2 to 9, wherein the step (S2) includes a step (S2-2) of checking the statistics of the sequencing data of the reference sample group.

11. The method of claim 10, wherein the cluster analysis method includes K-means clustering and / or hierarchical clustering.

12. The method described in claim 10, wherein verifying the statistical value includes calculating the mean value and / or standard deviation of the reference samples for each group in the target interval.

13. The method according to any one of claims 2 to 12, wherein the step (S3) includes confirming the distribution similarity between the reference sample group and the test sample by calculating a statistical distance based on the sequencing data of the reference sample group and the test sample in the target interval.

14. The method described in claim 13, wherein the high distribution similarity includes the statistical distance between the reference sample group and the test sample in the target interval being short.

15. The method of claim 13 or 14, wherein the statistical distance comprises a statistical value of the pth power of the absolute value of the difference in the sequencing data of the reference sample group and the test sample in the target interval, where p is 1 or greater.

16. The method of claim 15, wherein the statistical value includes a sum value.

17. The method of claim 15, wherein the statistical distance includes a Minkowski distance.

18. The method according to any one of claims 2 to 17, wherein step (S4) comprises determining the copy number CNg of the test sample in the target gene based on the exon length of the target gene in the test sample and the copy number CNi of the test sample in the target section i.

19. The step (S4) includes determining the CNg based on the following formula: [Equation 1] wherein i represents the target interval, j represents the target exon, n represents the number of target intervals on the target exon j, m represents the number of target exons, CNi represents the copy number in the target interval i, and Lenj represents the length of the target exon j. The method of claim 18.

20. The method according to any one of claims 2 to 19, wherein step (S4) comprises determining the probability of the presence of a copy number variation of the test sample in the target interval, and the probability of the presence of the copy number variation comprises the probability of amplification (pa) and / or the probability of deletion (pd) of the copy number of the test sample in the target interval.

21. The method described in claim 20, wherein step (S4) includes confirming the probability of the existence of the copy number variation by a probability distribution based on the mean value and standard deviation of the sequencing data of the test sample in the target section i and the sequencing data of a group of reference samples closest to the test sample in the corresponding target section.

22. The method of claim 21, wherein the probability distribution comprises a normal probability distribution.

23. The step (S4) includes determining a sigRatio, which is a ratio of the presence of significant copy number amplification or deletion in the target gene in the test sample; The method according to any one of claims 2 to 22, wherein the step (S4) further comprises determining a statistical test parameter for the presence of a copy number variation in the target gene in the test sample.

24. The method described in claim 23, wherein the sigRatio is obtained by dividing the number of target sections in the target gene in which significant copy number variation has occurred by the number of all target sections in the target gene, and the target sections in which significant copy number variation has occurred include target sections in which the copy number variation ratio is 27% to 33% or more.

25. The method of claim 23, wherein a p-value p ttest is determined by a t-test method based on the number of target intervals of the test sample in the target gene, sequencing data of each target interval of the test sample in the target gene, the standard deviation of sequencing data of each target interval of the test sample in the target gene, and the mean value and standard deviation of sequencing data of the target interval of a reference sample group that is closest to the test sample in the corresponding target gene.

26. The step (S4) includes determining the copy number status of the target gene in the test sample by: CN g ≧CN thA , sigRatio≧sigRatio th , and p ttest ≦p th and confirming that the copy number of the target gene in the test sample has been amplified when CN g ≦CN thD , sigRatio≧sigRatio th , and p ttest ≦p th and confirming that a copy number loss of the target gene in the test sample has occurred. CN thA <CN g <CN thD , or sigRatio<sigRatio th , or p ttest >p th When the number of copies of the target gene in the test sample is normal, thA , C.N. thD , sigRatio th , and p th The method of any one of claims 23 to 25, wherein each of the is independently a threshold value.

27. The method according to claim 26, wherein CN thA is 2.05 to 4.

4.

28. The method according to claim 26, wherein CN thD is 0.9 to 1.

925.

29. The method of claim 26, wherein sigRatio th is 0.27 to 1.

1.

30. The method according to claim 26, wherein p th is 0.06 to 0.000009.

31. The target gene is <h2 style=";text-align:left;direction:ltr">ABL1、ABL2、ABRAXAS1、ACVR1、ACVR1B、AKT1、AKT2、AKT3、ALK、ALOX12B、AME R1、APC、AR、ARAF、ARFRP1、ARID1A、ARID1B、ARID2、ARID5B、ASXL1、ASXL2、A SXL3、ATG5、ATM、ATR、ATRX、AURKA、AU RKB、AXIN1、AXIN2、AXL、B2M、BAP1、BA RD1、BBC3、BCL10、BCL2、BCL2L1、BCL2 L11、BCL2L2、BCL6、BCOR、BCORL1、BIRC 3、BLM、BMPR1A、BRAF、BRCA1、BRCA2、B RD4、BRD7、BRINP3、BRIP1、BTG1、BTG2 、BTK、CALR、CARD11、CASP8、CBFB、CBL 、CCND1、CCND2、CCND3、CCNE1、CD274、C D28、CD58、CD74、CD79A、CD79B、CDC73 、CDH1、CDH18、CDK12、CDK4、CDK6、CDK 8、CDKN1A、CDKN1B、CDKN1C、CDKN2A、C DKN2B、CDKN2C、CEBPA、CENPA、CHD1、CH D2、CHD4、CHD8、CHEK1、CHEK2、CIC、CI ITA、CREBBP、CRKL、CRLF2、CRYBG1、CS F1R、CSF3R、CSMD1、CSMD3、CTCF、CTLA 4、CTNNA1、CTNNB1、CUL3、CUL4A、CXCR 4、CYLD、CYP17A1、CYP2D6、DAXX、DCUN 1D1、DDR1、DDR2、DDX3X、DICER1、DIS3 、DNAJB1、DNMT1、DNMT3A、DNMT3B、DOT 1L、DPYD、DTX1、DUSP22、EED、EGFR、EIF 1AX、EIF4E、EMSY、EP300、EPCAM、EPHA 2、EPHA3、EPHA5、EPHA7、EPHB1、EPHB4 、ERBB2、ERBB3、ERBB4、ERCC1、ERCC2、 ERCC3、ERCC4、ERCC5、ERG、ERRFI1、ESR 1、ETV4、ETV5、ETV6、EWSR1、EZH2、EZR 、FANCA、FANCC、FANCD2、FANCE、FANCF 、FANCG、FANCI、FANCL、FANCM、FAS、FA T1、FAT3、FBXW7、FGF10、FGF12、FGF14、<h2 style=";text-align:left;direction:ltr">FGF19、FGF23、FGF3、FGF4、FGF6、FGF7 、FGFR1、FGFR2、FGFR3、FGFR4、FH、FLC N、FLT1、FLT3、FLT4、FOXA1、FOXL2、FO XO1、FOXO3、FOXP1、FRS2、FUBP1、FYN、 GABRA6、GALNT12、GATA1、GATA2、GATA 3、GATA4、GATA6、GEN1、GID4、GLI1、GN A11、GNA13、GNAQ、GNAS、GPS2、GREM1、GRIN2A、GRM3、GSK3B、H3F3A、H3F3B、H3 F3C、HDAC1、HDAC2、HGF、HIST1H1C、HI ST1H2BD、HIST1H3A、HIST1H3B、HIST1 H3C、HIST1H3D、HIST1H3E、HIST1H3G、 HIST1H3H、HIST1H3I、HIST1H3J、HIST2 H3D、HIST3H3、HLA-A、HLA-B、HLA-C、H NF1A、HOXB13、HRAS、HSD3B1、HSP90AA 1、ICOSLG、ID3、IDH1、IDH2、IFNGR1、I GF1、IGF1R、IGF2、IGHD、IGHJ、IGHV、IK BKE、IKZF1、IL10、IL7R、INHA、INHBA、INPP4A、INPP4B、INSR、IRF2、IRF4、IR S1、IRS2、ITK、ITPKB、JAK1、JAK2、JAK3、JUN、KAT6A、KDM5A、KDM5C、KDM6A、KD R、KEAP1、KEL、KIR2DL4、KIR3DL2、KIT 、KLF4、KLHL6、KLRC1、KLRC2、KLRK1、K MT2A、KMT2C、KMT2D、KRAS、LATS1、LATS2、LMO1、LRP1B、LTK、LYN、MAF、MAGI2、 MALT1、MAP2K1、MAP2K2、MAP2K4、MAP3 K1、MAP3K13、MAP3K14、MAPK1、MAPK3、 MAX、MCL1、MDC1、MDM2、MDM4、MED12、MEF2B、MEN1、MERTK、MET、MFHAS1、MGA、M IR21、MITF、MKNK1、MLH1、MLH3、MPL、MRE11、MSH2、MSH3、MSH6、MST1、MST1R、 MTAP、MTOR、MUTYH、MYC、MYCL、MYCN、M YD88、MYOD1、NAV3、NBN、NCOA3、NCOR1、NCOR2、NEGR1、NF1、NF2、NFE2L2、NFKBIA、NKX2-1、NKX3-1、NOTCH1、NOTCH2、NOTCH3、NOTCH4、NPM1、NRAS、NRG1、NSD1、NSD2、NSD3、NT5C2、NTHL1、NTRK1、NTRK2、NTRK3、NUP93、NUTM1、P2RY8、PAK1、PAK3、PAK5、PALB2、PALLD、PARP1、PARP2、PARP3、PAX5、PBRM1、PCDH11X、PDCD1、PDCD1LG2、PDGFRA、PDGFRB、PDK1、PGR、PHOX2B、PIK3C2B、PIK3C2G、PIK3C3、PIK3CA、PIK3CB、PIK3CD、PIK3CG、PIK3R1、PIK3R2、PIK3R3、PIM1、PLCG2、PLK2、PMS1、PMS2、PNRC1、POLD1、POLE、POM121L12、PPARG、PPM1D、PPP2R1A、PPP2R2A、PPP6C、PRDM1、PREX2、PRKAR1A、PRKCI、PRKDC、PRKN、PTCH1、PTEN、PTPN11、PTPRD、PTPRO、PTPRS、PTPRT、QKI、RAB35、RAC1、RAD21、RAD50、RAD51、RAD51B、RAD51C、RAD51D、RAD52、RAD54L、RAF1、RARA、RASA1、RB1、RBM10、RECQL4、REL、RET、RHEB、RHOA、RICTOR、RIT1、RNF43、ROS1、RPA1、RPS6KA4、RPS6KB2、RPTOR、RSPO2、RUNX1、RUNX1T1、SDC4、SDHA、SDHAF2、SDHB、SDHC、SDHD、SETD2、SF3B1、SGK1、SH2B3、SH2D1A、SHQ1、SLC34A2、SLIT2、SLX4、SMAD2、SMAD3、SMAD4、SMARCA4、SMARCB1、SMARCD1、SMO、SNCAIP、SOCS1、SOX10、SOX17、SOX2、SOX9、SPEN、SPI1、SPOP、SPTA1、SRC、SRSF2、STAG2、STAT3、STAT4、STAT5A、STAT5B、STAT6、STK11、STK40、SUFU、SYK、TAF1、TBX21、TBX3、TCF3、TCF7L2、TEK、TENT5C、TERC、TERT、TET1、TET2, TGFBR1, TGFBR2, TIPARP, TMEM127, TMPRSS2, TNFAIP3, TNFRSF14, TOP1, TOP2A, TP53 , TP63, TP73, TRAF2, TRAF3, TRAF7, TRIM58, TRPC5, TSC1, TSC2, TSHR, TYRO3, U2AF1, UGT1A1 , VEGFA, VEGFB, VEGFC, VHL, WISP3, WRN, WT1, XIAP, XPO1, XRCC2, XRCC3, YAP1, YES1, ZAP70, ZBTB16, ZBTB2, ZNF217, ZNF703, and ZNRF3.

32. An apparatus for analyzing copy number status, said apparatus comprising a module for carrying out the method according to any one of claims 1 to 31, a receiving module for acquiring sequencing data of a group of test samples; a determination module for determining a target gene in the test sample; and a determination module that determines the copy number status of the target gene in the test samples based on sequencing data of the test sample group.

33. (M1) a receiving module that acquires sequencing data of a test sample and / or sequencing data of a plurality of reference samples; (M2) a processing module for dividing the reference samples into two or more reference sample groups; (M3) a calculation module for determining a set of reference samples that are closest to the test sample; (M4) a determination module that determines the copy number status of the target gene of the test sample based on sequencing data of the reference sample group that is closest to the test sample. The copy number status analysis device of claim 32, further comprising:

34. A storage medium storing a program capable of executing the method according to any one of claims 1 to 31.

Citation Information

Patent Citations

  • Device for detecting DNA copy number variation of circulating tumor

    CN106650312A

  • Copy number variation detection apparatus

    CN108256292A

  • Methods for detecting gene mutations

    JP2015502749A

  • Method, system, and computer-readable storage medium for determining the presence or absence of copy number variations in a genome

    JP2015506684A

  • Method and system for detecting copy number polymorphisms

    JP2018523198A