Chromosome outer circular DNA identification method and device, computer equipment and storage medium
By segmenting captured fragments and calculating fragment coverage using single-cell sequencing technology, the problems of low detection efficiency and neglect of differences in extrachromosomal circular DNA in existing technologies are solved, enabling high-precision identification of extrachromosomal circular DNA and analysis of tumor cells.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN HUADA GENE INST
- Filing Date
- 2024-10-18
- Publication Date
- 2026-04-21
AI Technical Summary
Existing technologies for detecting extrachromosomal circular DNA include fluorescence microscopy imaging, which is inefficient and susceptible to interference from experimental procedures, while high-throughput sequencing ignores the differences between different tumor cells, resulting in a lack of biological interpretation of the detection results and long computation time.
By employing single-cell sequencing technology, capturing fragments are segmented into different fragment intervals, fragment coverage is calculated, and thresholds are set to achieve high-throughput, high-precision identification of extrachromosomal circular DNA. This avoids the shortcomings of microscopic imaging technology and takes into account intercellular differences.
It achieves high-throughput, high-precision detection of extrachromosomal circular DNA at the whole genome level, and can accurately identify DNA fragments with highly open and uniformly distributed chromatin, supporting subsequent tumor cell analysis.
Smart Images

Figure CN121905271A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of bioengineering, specifically to a method, apparatus, computer device, and computer-readable storage medium for identifying extrachromosomal circular DNA. Background Technology
[0002] This section is intended to provide background or context for implementing the embodiments of the invention as set forth in the claims and detailed description. The description herein is not intended to imply that it is prior art simply because it is included in this section.
[0003] Extrachromosomal circular DNA (ecDNA or eccDNA) is a type of circular DNA that is widely distributed in nature and independent of chromosomes. It is more open than chromosomal DNA, leading to increased transcriptional activity. Lacking centromeres, extrachromosomal circular DNA is inherited in a non-Mendelian manner during mitosis, resulting in some tumor cells carrying high copy numbers of oncogenes. These oncogenes significantly affect tumor heterogeneity and drug resistance.
[0004] Currently, extrachromosomal circular DNA has been found in cancer cells of many cancers, but rarely in normal tissues. Therefore, developing effective methods for identifying extrachromosomal circular DNA is crucial for its application in clinical oncology. Summary of the Invention
[0005] Therefore, it is necessary to provide a method, apparatus, computer equipment, and computer-readable storage medium for identifying extrachromosomal circular DNA in response to the above-mentioned technical problems.
[0006] In a first aspect, this application provides a method for identifying extrachromosomal circular DNA, comprising:
[0007] Acquire a sample to be tested and its sequencing data, wherein the sample to be tested includes multiple cells and the sequencing data contains a capture fragment of each cell, wherein the capture fragment is a DNA fragment captured by single-cell sequencing of the sample to be tested.
[0008] The captured fragments are segmented to obtain a first fragment matrix located in a first fragment interval and a second fragment matrix located in a second fragment interval; the first fragment matrix contains the first fragment of each cell and the sequence number of the first fragment, the second fragment matrix contains the second fragment of each cell and the sequence number of the second fragment, the sequence number being the number of captured fragments contained therein, the second fragment interval being located within the first fragment interval, and the first fragment matrix including multiple consecutive second fragments;
[0009] The fragment coverage of the first fragment in the cell is determined based on the sequence number of the first fragment and the sequence number of multiple consecutive second fragments in the first fragment;
[0010] If the fragment coverage of the first fragment in the cell reaches a first preset threshold and satisfies uniform distribution, then the first fragment is identified as an extrachromosomal circular DNA fragment.
[0011] Secondly, this application also provides a device for identifying extrachromosomal circular DNA, comprising:
[0012] The data acquisition module is used to acquire the sample to be tested and the sequencing data of the sample to be tested. The sample to be tested includes multiple cells, and the sequencing data includes the number of captured fragments for each cell. The captured fragments are DNA fragments captured by single-cell sequencing of the sample to be tested.
[0013] The fragment segmentation module is used to segment the captured fragments to obtain a first fragment matrix located in a first fragment interval and a second fragment matrix located in a second fragment interval; the first fragment matrix contains the first fragment of each cell and the sequence number of the first fragment, the second fragment matrix contains the second fragment of each cell and the sequence number of the second fragment, the sequence number being the number of captured fragments contained therein, the second fragment interval being located within the first fragment interval, and the first fragment matrix including multiple consecutive second fragments;
[0014] The coverage determination module is used to determine the fragment coverage of the first fragment in the cell based on the sequence number of the first fragment and the sequence number of multiple consecutive second fragments in the first fragment;
[0015] The fragment identification module is used to identify the first unit fragment as an extrachromosomal circular DNA fragment if the fragment coverage of the first fragment in the cell reaches a first preset threshold and satisfies uniform distribution.
[0016] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:
[0017] Acquire a sample to be tested and its sequencing data, wherein the sample to be tested includes multiple cells and the sequencing data contains a capture fragment of each cell, wherein the capture fragment is a DNA fragment captured by single-cell sequencing of the sample to be tested.
[0018] The captured fragments are segmented to obtain a first fragment matrix located in a first fragment interval and a second fragment matrix located in a second fragment interval; the first fragment matrix contains the first fragment of each cell and the sequence number of the first fragment, the second fragment matrix contains the second fragment of each cell and the sequence number of the second fragment, the sequence number being the number of captured fragments contained therein, the second fragment interval being located within the first fragment interval, and the first fragment matrix including multiple consecutive second fragments;
[0019] The fragment coverage of the first fragment in the cell is determined based on the sequence number of the first fragment and the sequence number of multiple consecutive second fragments in the first fragment;
[0020] If the fragment coverage of the first fragment in the cell reaches a first preset threshold and satisfies uniform distribution, then the first fragment is identified as an extrachromosomal circular DNA fragment.
[0021] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the following steps:
[0022] Acquire a sample to be tested and its sequencing data, wherein the sample to be tested includes multiple cells and the sequencing data contains a capture fragment of each cell, wherein the capture fragment is a DNA fragment captured by single-cell sequencing of the sample to be tested.
[0023] The captured fragments are segmented to obtain a first fragment matrix located in a first fragment interval and a second fragment matrix located in a second fragment interval; the first fragment matrix contains the first fragment of each cell and the sequence number of the first fragment, the second fragment matrix contains the second fragment of each cell and the sequence number of the second fragment, the sequence number being the number of captured fragments contained therein, the second fragment interval being located within the first fragment interval, and the first fragment matrix including multiple consecutive second fragments;
[0024] The fragment coverage of the first fragment in the cell is determined based on the sequence number of the first fragment and the sequence number of multiple consecutive second fragments in the first fragment;
[0025] If the fragment coverage of the first fragment in the cell reaches a first preset threshold and satisfies uniform distribution, then the first fragment is identified as an extrachromosomal circular DNA fragment.
[0026] The above-described method, apparatus, computer equipment, and computer-readable storage medium for identifying extrachromosomal circular DNA acquire a test sample and its sequencing data. The test sample includes multiple cells, and the sequencing data contains a capture fragment for each cell. The capture fragment is a DNA fragment captured by single-cell sequencing of the test sample. The captured fragment is segmented to obtain a first fragment matrix located in a first fragment interval and a second fragment matrix located in a second fragment interval. The first fragment matrix contains the first fragment of each cell and the sequence number of the first fragment. The second fragment matrix contains the second fragment of each cell and the sequence number of the second fragment. The sequence number is the number of captured fragments included. The second fragment interval is located within the first fragment interval. The first fragment matrix includes multiple consecutive second fragments. The fragment coverage of the first fragment in the cell is determined based on the sequence number of the first fragment and the sequence number of multiple consecutive second fragments in the first fragment. If the fragment coverage of the first fragment in the cell reaches a first preset threshold and satisfies uniform distribution, the first fragment is identified as an extrachromosomal circular DNA fragment. In this way, by dividing the captured fragments into different fragment intervals and calculating the fragment coverage, the degree of chromatin openness of DNA fragments can be quantified. Fragments with highly open chromatin can be distinguished by setting a first preset threshold for fragment coverage, while fragments with uniform coverage indicate that the chromatin within them is open on an average basis. Thus, the characteristics of highly open and evenly open chromatin can be used to accurately identify extrachromosomal circular DNA. Attached Figure Description
[0027] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0028] Figure 1 This is a diagram illustrating the application environment of a method for identifying extrachromosomal circular DNA in one embodiment.
[0029] Figure 2 This is a flowchart illustrating a method for identifying extrachromosomal circular DNA in one embodiment;
[0030] Figure 3 This is a visual illustration of one embodiment;
[0031] Figure 4 This is a first schematic diagram of a probability distribution density map in one embodiment;
[0032] Figure 5This is a second schematic diagram of the probability distribution density map in one embodiment;
[0033] Figure 6 This is a first schematic diagram of the density distribution curve in one embodiment;
[0034] Figure 7 This is a second schematic diagram of the density distribution curve in one embodiment;
[0035] Figure 8 This is an internal structural diagram of a chromosome extracircular DNA identification device in one embodiment;
[0036] Figure 9 This is an internal structural diagram of the extrachromosomal circular DNA identification device in another embodiment;
[0037] Figure 10 This is an internal structural diagram of a computer device in one embodiment;
[0038] Figure 11 This is a diagram of the internal structure of a computer device in another embodiment. Detailed Implementation
[0039] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to be used in the overall description of this application. The use of terms like "first" and "second" is merely for distinguishing technical features and should not be construed as indicating or implying relative importance, the number of indicated technical features, or the sequential relationship between indicated technical features.
[0040] In the description of this application, unless otherwise expressly defined, terms such as "setup," "installation," and "connection" should be interpreted broadly, and those skilled in the art can reasonably determine the specific meaning of the above terms in this application in conjunction with the specific content of the technical solution.
[0041] Extrachromosomal circular DNA (ecDNA or eccDNA), a type of DNA with highly open chromatin, has been identified as a major vector for amplifying oncogenes. Through whole-genome sequencing (WGS) and cytogenetic methods, extrachromosomal circular DNA has been found in cancer cells of many cancers, but rarely in normal tissues. Mature, stable, and low-cost methods for identifying extrachromosomal circular DNA are crucial for its application in clinical oncology.
[0042] Some methods utilize fluorescence microscopy imaging to detect ecDNA. For example, software like ECdetect and EcSeg detect ecDNA using DAPI staining and / or FISH staining images. However, the irregular shape and small size of ecDNA, along with the high background noise and noise levels in mid-stage cell images, all affect the detection accuracy of these methods. Furthermore, fluorescence microscopy imaging is relatively inefficient, resulting in low detection rates.
[0043] Some protocols rely on high-throughput sequencing technologies to detect ecDNA, such as Circle-seq, AmpliconArchitect, circle_finder, and circlehunter. These protocols depend on whole-genome sequencing data and tissue-level ATAC sequencing data, treating all tumor cells as a whole and ignoring the differences in the distribution and structure of ecDNA among different tumor cells.
[0044] In summary, fluorescence microscopy imaging for ecDNA detection is inefficient and susceptible to negative impacts such as experimental manipulation, incomplete staining image information, and high image noise. High-throughput sequencing for ecDNA detection treats all tumor cells as a single entity, ignoring the differences between different tumor cells, which may lead to biologically uninterpretable ecDNA detection results and is computationally time-consuming.
[0045] Based on the above considerations, this application proposes a method for identifying extrachromosomal circular DNA, which does not rely on microscopic imaging technology, can detect whole-genome level ecDNA information in one go and improve the detection resolution to the single-cell level, achieving high-throughput and high-precision ecDNA detection.
[0046] Please see Figure 1 , Figure 1 This diagram illustrates the application environment of the extrachromosomal circular DNA identification method in one embodiment. The extrachromosomal circular DNA identification method provided in this application embodiment can be applied to, for example... Figure 1 The application environment shown is illustrated. Terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104, or it can be located in the cloud or on another server.
[0047] Both the terminal and the server can be used independently to execute the extrachromosomal circular DNA identification method provided in the embodiments of this application.
[0048] For example, the terminal acquires a test sample comprising multiple cells and its sequencing data. The sequencing data includes a capture fragment for each cell, which is a DNA fragment captured during single-cell sequencing of the test sample. The terminal segments the capture fragments to obtain a first fragment matrix located in a first fragment interval and a second fragment matrix located in a second fragment interval. The first fragment matrix contains the first fragment of each cell and its sequence number, while the second fragment matrix contains the second fragment of each cell and its sequence number, where the sequence number represents the number of captured fragments. The second fragment interval is located within the first fragment interval, and the first fragment matrix includes multiple consecutive second fragments. The terminal determines the fragment coverage of the first fragment in the cell based on the sequence number of the first fragment and the sequence number of multiple consecutive second fragments within the first fragment. If the fragment coverage of the first fragment in the cell reaches a first preset threshold and satisfies uniform distribution, the terminal identifies the first fragment as an extrachromosomal circular DNA fragment.
[0049] In addition, the terminal and server can also be used in conjunction to execute the extrachromosomal circular DNA identification method provided in the embodiments of this application.
[0050] For example, a terminal acquires a test sample comprising multiple cells and its sequencing data. The sequencing data includes a capture fragment for each cell, which is a DNA fragment captured during single-cell sequencing of the test sample. The terminal sends the acquired test sample and sequencing data to a server. The server segments the capture fragments, obtaining a first fragment matrix located in a first fragment interval and a second fragment matrix located in a second fragment interval. The first fragment matrix contains the first fragment of each cell and its sequence number, while the second fragment matrix contains the second fragment of each cell and its sequence number, where the sequence number represents the number of captured fragments. The second fragment interval is located within the first fragment interval, and the first fragment matrix includes multiple consecutive second fragments. The server determines the fragment coverage of the first fragment within the cell based on the sequence number of the first fragment and the sequence number of multiple consecutive second fragments within the first fragment. The server sends the cell coverage of the first fragment to the terminal. If the cell coverage of the first fragment reaches a first preset threshold and satisfies uniform distribution, the terminal identifies the first fragment as an extrachromosomal circular DNA fragment.
[0051] Terminal 102 can be a sequencer, smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, IoT device, or portable wearable device. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart in-vehicle systems, and projection devices. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted displays. Head-mounted displays can be virtual reality (VR) devices, augmented reality (AR) devices, or smart glasses.
[0052] Server 104 can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides cloud computing services.
[0053] Terminal 102 and server 104 can be connected via Bluetooth, USB (Universal Serial Bus) or network, etc., and this application does not impose any restrictions.
[0054] In one exemplary embodiment, such as Figure 2 As shown, a method for identifying extrachromosomal circular DNA is provided, which can be applied to... Figure 1 Taking the terminal in the example, the explanation includes the following steps S202 to S208. Wherein:
[0055] Step S202: Obtain the sample to be tested and its sequencing data.
[0056] The sample to be tested includes multiple cells, and the sequencing data contains capture fragments of each cell. The capture fragments are DNA fragments captured by single-cell sequencing of the sample to be tested.
[0057] The sample to be tested is the sample for which extrachromosomal circular DNA needs to be identified. The presence of extrachromosomal circular DNA is a characteristic of some cancers. The circular structure and relatively open chromatin structure of extrachromosomal circular DNA make the genes on it easier to be transcribed and expressed. In tumor cell screening, the cells or tissues to be screened can be used as the test sample. By identifying whether extrachromosomal circular DNA is present in the test sample, it can be determined whether the test sample contains tumor cells.
[0058] In one embodiment, single-cell sequencing can be performed on the sample to be tested to obtain sequencing data of the sample. The sequencing data includes captured fragments from each cell in the sample. These captured fragments, also called fragments, are DNA fragments captured during the single-cell sequencing process.
[0059] Single-cell sequencing methods include, but are not limited to, single-cell ATAC sequencing (single-cell Assay for Transposase-Accessible Chromatin using sequencing, scATAC-seq), which involves performing single-cell ATAC sequencing on the sample to be tested to obtain sequencing data containing captured fragments from each cell in the sample.
[0060] In one embodiment, the sequencing data can be a count matrix of captured fragments at single-cell precision, i.e., a cell × captured fragment matrix. The rows and columns of the count matrix represent the cells in the sample to be tested and the captured fragments within those cells, respectively. The values in the count matrix represent the number of times the captured fragment appears in the cell, i.e., the abundance of the captured fragment in the cell.
[0061] Furthermore, the reference genome version and chromosome symbol format can be determined based on the reference genome version number used during single-cell sequencing.
[0062] Step S204: Segment the captured segments to obtain a first segment matrix located in the first segment interval and a second segment matrix located in the second segment interval.
[0063] The first fragment matrix contains the first fragment of each cell and the sequence number of the first fragment, and the second fragment matrix contains the second fragment of each cell and the sequence number of the second fragment. The sequence number is the number of captured fragments contained. The second fragment interval is located within the first fragment interval, and the first fragment matrix includes multiple consecutive second fragments.
[0064] After obtaining sequencing data containing the captured fragments, the captured fragments in each cell of the test sample are segmented according to different interval lengths. Specifically, the captured fragments in each cell of the test sample are segmented according to a first interval length to obtain a first fragment matrix located in the first fragment interval; the captured fragments in each cell of the test sample are segmented according to a second interval length to obtain a second fragment matrix located in the second fragment interval.
[0065] For example, the length of the first interval is 100,000 pb (100 kb). The captured fragments in each cell of the sample are divided into 100 kb segments. The first 100 kb, the second 100 kb, and so on, each 100 kb segment is considered a first fragment interval. Each captured fragment is divided into multiple first fragments (100 kb fragments), resulting in a first fragment matrix containing the first fragment of each cell and the sequence number of each first fragment. The first fragment matrix is a cell × first fragment matrix. Each row in the first fragment matrix represents a first fragment, each column represents a cell, and the value of the first fragment matrix represents the number of captured fragments in the cell that fall within the corresponding first fragment interval. Taking 100 kb as an example, each row in the first fragment matrix represents a 100 kb fragment, each column represents a cell, and the value of the first fragment matrix represents the number of captured fragments in the cell that fall within the corresponding 100 kb fragment interval.
[0066] For example, the length of the second interval is 1000 pb (1 kb). The captured fragments in each cell of the sample are divided into 1 kb segments. The first 1 kb segment, the second 1 kb segment, and so on, each 1 kb segment constitutes a second fragment interval. Each captured fragment is divided into multiple second fragments (1 kb segments), resulting in a second fragment matrix containing the second fragments of each cell and the sequence number of each second fragment. The second fragment matrix is a cell × second fragment matrix. Each row in the second fragment matrix represents a second fragment, each column represents a cell, and the value of the second fragment matrix represents the number of captured fragments in a cell that fall within the corresponding second fragment interval. Taking 1 kb as an example, each row in the second fragment matrix represents a 1 kb segment, each column represents a cell, and the value of the second fragment matrix represents the number of captured fragments in a cell that fall within the corresponding 1 kb fragment interval.
[0067] The length of the first interval of the first segment is greater than the length of the second interval of the second segment. The first segment matrix contains multiple consecutive second segments, such that each row (one first segment) in the first segment matrix contains multiple consecutive second segments. The second segment interval, defined by the length of the second interval, lies within the first segment interval, defined by the length of the first interval. The sequence number is the count of captured segments, i.e., the number of captured segments contained. The sequence number of the first segment is the number of captured segments contained in the first segment, and the sequence number of the second segment is the number of captured segments contained in the second segment.
[0068] Optionally, the length of the first interval can be an integer multiple of the length of the second interval. For example, the first segment is a 100kb segment, the second segment is a 1kb segment, the length of the first interval is 100 times the length of the second interval, and each first segment contains 100 consecutive second segments.
[0069] In some cases, quality control of sequencing data can be performed before segmenting the captured fragments. For example, filtering conditions can be set to filter out captured fragments that meet the filtering conditions, and then segmenting the filtered captured fragments; or screening conditions can be set to screen captured fragments according to the screening conditions, and then segmenting the screened captured fragments.
[0070] Step S206: Determine the fragment coverage of the first fragment in the cell based on the sequence number of the first fragment and the sequence number of multiple consecutive second fragments in the first fragment.
[0071] The first fragment is the smallest unit used in subsequent steps to identify whether it is extrachromosomal circular DNA. The second fragment is an auxiliary unit used to help determine whether the first fragment is uniformly distributed. Based on the sequence number of the first fragment and the sequence number of multiple consecutive second fragments within the first fragment, the fragment coverage of the first fragment in the cell can be determined, and then the fragment coverage of the first fragment in the cell can be used to identify whether the first fragment is an extrachromosomal circular DNA fragment.
[0072] The fragment coverage of the first fragment in the cell is a quantitative result of the chromatin openness of the first fragment, which can be calculated in various ways. For example, the sum of the sequence numbers of all second fragments in the first fragment can be calculated as the proportion of the sequence number in the cell containing the first fragment, and this proportion can be used as the fragment coverage of the first fragment in the cell; or, other quantitative methods that can reflect the chromatin openness can be used to determine the fragment coverage, and this application does not limit this.
[0073] Step S208: If the fragment coverage of the first fragment in the cell reaches a first preset threshold and satisfies uniform distribution, then the first fragment is identified as an extrachromosomal circular DNA fragment.
[0074] In some embodiments, based on the fragment coverage of each first fragment in the cell, it can be further identified whether each first fragment is an extrachromosomal circular DNA fragment.
[0075] Because the chromatin of extrachromosomal circular DNA fragments is highly open and evenly open, if the fragment coverage of the first fragment in the cell reaches a certain threshold, it indicates that the chromatin of the first fragment is highly open; if the fragment coverage of the first fragment in the cell is evenly distributed, it indicates that the chromatin of the first fragment is evenly open. Therefore, the first fragment whose fragment coverage reaches the first preset threshold and satisfies the condition of even distribution can be identified as an extrachromosomal circular DNA fragment, while the first fragment whose fragment coverage does not reach the first preset threshold and / or does not satisfy the condition of even distribution can be identified as a non-extrachromosomal circular DNA fragment.
[0076] Based on the identified extrachromosomal circular DNA, further analysis of the function and evolution mechanism of tumor cell-specific extrachromosomal circular DNA can be conducted, which is of great significance for detecting circulating tumor cells, analyzing the complex heterogeneity of tumor cells, and elucidating the drug resistance mechanism of targeted therapy.
[0077] In the above embodiments, a test sample and its sequencing data are acquired. The test sample includes multiple cells, and the sequencing data contains a capture fragment for each cell. The capture fragment is a DNA fragment captured by single-cell sequencing of the test sample. The capture fragment is segmented to obtain a first fragment matrix located in a first fragment interval and a second fragment matrix located in a second fragment interval. The first fragment matrix contains the first fragment of each cell and the sequence number of the first fragment. The second fragment matrix contains the second fragment of each cell and the sequence number of the second fragment. The sequence number is the number of captured fragments contained therein. The second fragment interval is located within the first fragment interval. The first fragment matrix includes multiple consecutive second fragments. The fragment coverage of the first fragment in the cell is determined based on the sequence number of the first fragment and the sequence number of multiple consecutive second fragments in the first fragment. If the fragment coverage of the first fragment in the cell reaches a first preset threshold and satisfies uniform distribution, the first fragment is identified as an extrachromosomal circular DNA fragment. In this way, by dividing the captured fragments into different fragment intervals and calculating the fragment coverage, the degree of chromatin openness of DNA fragments can be quantified. Fragments with highly open chromatin can be distinguished by setting a first preset threshold for fragment coverage, while fragments with uniform coverage indicate that the chromatin within them is open on an average basis. Thus, the characteristics of highly open and evenly open chromatin can be used to accurately identify extrachromosomal circular DNA.
[0078] In one embodiment, before segmenting the captured segments to obtain a first segment matrix located in a first segment interval and a second segment matrix located in a second segment interval, the method further includes:
[0079] The fragment length of the captured fragment is determined based on the start and end points of the captured fragment; captured fragments whose fragment length is outside the preset upper and lower limits are filtered out; and / or the fragment length of the captured fragments in the cell is statistically analyzed to determine the peak value of the fragment length distribution in the cell, and captured fragments whose peak value of the fragment length distribution in the cell is outside the preset peak value range are filtered out; and / or the number of captured fragments in the cell is statistically analyzed to filter out captured fragments whose number of captured fragments in the cell is less than a second preset threshold.
[0080] Before segmenting the captured segments to obtain the first segment matrix located in the first segment interval and the second segment matrix located in the second segment interval, some captured segments can be filtered out based on pre-set filtering conditions. During the segmentation of the captured segments, the filtered captured segments are then segmented again.
[0081] The filtering conditions can be set based on the captured fragments in the sequencing data, including but not limited to at least one of the following: fragment length upper and lower limits, fragment distribution peak conditions, and fragment number conditions. Specifically, the fragment length upper and lower limits are used to define the upper and lower limits of the captured fragment length, the fragment distribution peak conditions are used to define the distribution peak of the captured fragments, and the fragment number conditions are used to define the number of captured fragments in the cell.
[0082] For example, the fragment length limit condition can be outside a preset range. If the length of a captured fragment is outside the preset range, the captured fragment is filtered out. For example, the fragment distribution peak value condition can be outside a preset range. If the peak value of a captured fragment is outside the preset range, the captured fragment is filtered out. For example, the fragment quantity condition can be less than a second preset threshold. If the number of captured fragments in a cell is less than the second preset threshold, the captured fragments in that cell are filtered out.
[0083] Understandably, to achieve quality control of sequencing data, one or more of the above filtering conditions can be used, or they can be combined arbitrarily, or other filtering conditions not listed can be set. For example, fragment length upper and lower limits, fragment distribution peak conditions, and fragment quantity conditions can all be set as filtering conditions. Taking a preset upper and lower limit range of 1-500bp, a peak range of 0-100dp, and a second preset threshold of 5000 as an example, then in cells with a captured fragment quantity of 5000, fragments with a fragment length between 1-500dp and a peak range between 0-100dp are considered to be of acceptable quality, and other captured fragments are filtered out.
[0084] In the above embodiments, by setting filtering conditions to perform quality control on sequencing data, the captured fragments are filtered before segmentation, which can improve the effectiveness of the captured fragments involved in segmentation and reduce the number of captured fragments involved in subsequent calculations, thereby saving the computing resources of the device and improving the identification efficiency of extrachromosomal circular DNA.
[0085] In one embodiment, determining the fragment coverage of the first fragment in a cell based on the sequence number of the first fragment and the sequence number of multiple consecutive second fragments within the first fragment includes:
[0086] Determine the sequence number of the cell; for the first fragment in the cell, determine the sum of the sequence numbers of all second fragments in the first fragment based on the sequence number of each second fragment in the first fragment; determine the fragment coverage of the first fragment in the cell based on the ratio of the sum of the sequence numbers of all second fragments in the first fragment of the cell to the sequence number of the cell.
[0087] When determining the fragment coverage of the first fragment in a cell, the sequence number of each first fragment can be obtained, and the sequence number of the cell containing the first fragment can be standardized. The fragment coverage of the first fragment in the cell can be obtained based on the result of the standardization process.
[0088] Standardization based on the sequence number of the cell containing the first fragment can, for example, involve dividing the sequence number of the first fragment by the sequence number of the cell. The sequence number of the first fragment can be the sum of the sequence numbers of all second fragments within the first fragment. That is, for a first fragment in a cell, the sum of the sequence numbers of all second fragments in the first fragment is determined based on the sequence number of each second fragment within the first fragment. The fragment coverage of the first fragment in the cell is then determined based on the ratio of the sum of the sequence numbers of all second fragments in the first fragment to the sequence number of the cell.
[0089] Assuming the first fragment is a 100kb fragment and the second fragment is a 1kb fragment, taking the calculation of the fragment coverage of the j-th 100kb fragment in the i-th cell of the test sample as an example, the fragment coverage of the 100kb fragment in the cell can be calculated using the following formula (1):
[0090]
[0091] Where, N frags (cell(i)) represents the sequence number of the i-th cell, N frags (cell(i),100kb(j),1kb(n)) represents the sequence number of the nth 1kb segment of the jth 100kb segment in the i-th cell. Mathematical symbols for any number, The domain of the independent variable 1kb(n) in formula (1) is (cell(i), 100kb(j)), meaning that any nth 1kb fragment in formula (1) belongs to the jth 100kb fragment of the i-th cell, and the value of n ranges from [1, 100]. -4 For the constant term, according to the actual situation, the constant term 10 in formula (1) -4 It can be replaced with other constant terms.
[0092] In one embodiment, determining the sequence number of a cell includes taking the sum of the number of various captured fragments in the cell as the sequence number of the cell.
[0093] In the above formula (1), the sequence number Nfrags(cell(i)) of the i-th cell can be determined by the following formula (1.1):
[0094]
[0095] Where frags(k) represents fragment k, In formula (1.1), segment k refers to the captured segment in the i-th cell. k is the accumulation function. The domain of the independent variable k is k∈[1,p1], where p1=N type (frags(k)) indicates that there are p1 possible fragments k.
[0096] In one embodiment, for a first fragment in a cell, determining the sum of sequence numbers of all second fragments in the first fragment, based on the sequence number of each second fragment in the first fragment, includes:
[0097] For the second segment in the first segment, the sum of the number of multiple captured segments in the second segment is taken as the sequence number of the second segment; based on the sequence number of each second segment in the first segment, the sum of the sequence numbers of all second segments in the first segment is determined.
[0098] In formula (1) above, the sum of the sequence numbers of all second segments in the first segment. It can be determined using the following formula (1.2):
[0099]
[0100] Where frags(l) represents fragment l, In formula (1.2), segment l refers to the captured segment in the j-th 100kb segment of the i-th cell. l is the accumulation function. The domain of the independent variable l is l∈[1,q1], where q1=N type (frags(l)) indicates that there are q1 possible fragments of l.
[0101] In one embodiment, a cell × first fragment coverage matrix of the sample to be tested can be obtained based on the sequence number of each first fragment and the sequence number of multiple consecutive second fragments within each first fragment. The coverage matrix contains the fragment coverage of each first fragment in the cells. The rows and columns of the coverage matrix represent cells and first fragments in the cells, respectively, and the fragment coverage value represents the fragment coverage of the first fragment in the cells.
[0102] Based on the coverage matrix, a coverage distribution density map of the sample to be tested can be plotted. A first preset threshold is set to filter out the first fragments with low coverage. Thus, based on the comparison between the first preset threshold and the fragment coverage, the first fragments with highly open chromatin can be selected. For example, the first preset threshold can be set to 6. After obtaining the fragment coverage of each first fragment in the cell, the first fragments with a coverage less than 6 are filtered out, and the first fragments with a coverage of 6 or higher are retained for further identification.
[0103] In the above embodiments, by classifying the captured fragments and calculating the fragment coverage based on the counting results of each type of captured fragment, the problem of differences in the quantification methods of the number of repetitions of captured fragments on different sequencing platforms can be solved, and the sequencing data types of multiple sequencing platforms can be compatible.
[0104] In one embodiment, if the fragment coverage of the first fragment in the cell reaches a first preset threshold and satisfies uniform distribution, before identifying the first fragment as an extrachromosomal circular DNA fragment, the method further includes:
[0105] For each first fragment, the transcription start site region and gene body region of the first fragment are determined; based on the distribution of captured fragments in the transcription start site region and gene body region of the first fragment, the transcription start site enrichment score of the first fragment is determined, and the transcription start site enrichment score is used to assess the average openness of chromatin; based on the transcription start site enrichment score of the first fragment, it is determined whether the direction of the density distribution curve of the first fragment tends to the transcription start site region, wherein the density distribution curve is the density distribution curve from the enrichment position of the first fragment to the transcription start site region, and the enrichment position is the position where the captured fragments in the start site region intersect with the first fragment; if so, the first fragment is filtered out.
[0106] The gene structure of a DNA fragment includes the transcription start site (TSS) region, the gene body region, and the intergenetic region.
[0107] The transcription start site region refers to the site where transcription begins in a gene. This region typically includes one or more promoters, regulatory elements, and transcription factor binding sites. In gene expression regulation, the transcription start site region influences the transcription rate, initiation location, and expression level.
[0108] The genomic body region is a specific region in the genome that contains the main functional parts of a gene, namely the transcription and protein translation processes. The function of the genomic body region is to control gene expression, thereby influencing cellular function.
[0109] Intergenic regions refer to chromosomal regions on chromosomes that do not control the shape of the genes between them. Because genes on chromosomes are not arranged continuously, but rather in a "gene-intergenic-gene" arrangement, intergenic regions occupy a considerable proportion of the genome, influencing gene expression and repression, and aiding in gene identification.
[0110] For each first fragment, the gene structure distribution of the first fragment is calculated, identifying the transcription start site region, gene body region, and intergene region of the first fragment. The enrichment sites where the captured fragments in the start site region intersect with the first fragment are determined, and density distribution curves from the enrichment sites to the transcription start site region are plotted. Based on whether the density distribution curve of the first fragment tends towards the transcription start site region, it is determined whether the first fragment is a possible extrachromosomal circular DNA fragment. If the density distribution curve of the first fragment tends towards the transcription start site region, it is determined that the first fragment is not an extrachromosomal circular DNA fragment and is excluded. If the density distribution curve of the first fragment does not tend towards the transcription start site region, it is determined that the first fragment may be an extrachromosomal circular DNA fragment, and further identification is carried out.
[0111] Among them, using the transcription start site enrichment score to quantify the distribution of captured fragments in the transcription start site region and gene body region of the first fragment can reflect the average openness of chromatin.
[0112] In one embodiment, the capture fragment includes a first capture fragment and a second capture fragment, wherein the first capture fragment is a capture fragment intersecting with the transcription start site region, and the second capture fragment is a capture fragment intersecting with the gene body region.
[0113] The capture fragments that intersect with the transcription start site region are identified from the captured fragments, and the captured fragments that intersect with the transcription start site region are designated as the first captured fragments; the capture fragments that intersect with the gene body region are identified from the captured fragments, and the captured fragments that intersect with the gene body region are designated as the second captured fragments; the transcription start site enrichment fraction of the first fragment is determined based on the first captured fragments and the second captured fragments.
[0114] In one embodiment, determining the transcription start site enrichment score of the first fragment based on the distribution of captured fragments in the transcription start site region and the gene body region of the first fragment includes: obtaining the number of first captured fragments in the first fragment and obtaining the number of second captured fragments in the first fragment; and determining the transcription start site enrichment score of the first fragment based on the number of first captured fragments and the number of second captured fragments.
[0115] Let the first fragment be a 100kb fragment and the second fragment be a 1kb fragment. Taking the j-th 100kb fragment of the i-th cell as an example, the enrichment fraction ES of the transcription start site of the j-th 100kb fragment of the i-th cell can be calculated using the following formula (2). tss :
[0116]
[0117] Where, N tss(cell(i), 100kb(j)) represents the number of segments in the j-th 100kb fragment of the i-th cell that are related to the first captured segment, N. genebody (cell(i), 100kb(j)) represents the number of second capture fragments contained in the j-th 100kb fragment of the i-th cell. The j-th 100kb fragment of the i-th cell can be any 100kb fragment of any cell in the sample to be tested.
[0118] In one embodiment, the number of all first capture segments in the first segment is the sum of the number of multiple different first capture segments, and the number of second capture segments in the first segment is the sum of the number of multiple different second capture segments.
[0119] In formula (2) above, N is the number of all first capture segments in the first segment. tss (cell(i), 100kb(j)) can be determined by the following formula (2.1):
[0120]
[0121] Where frags(s) represents fragment s, In formula (2.1), fragment 's' refers to the first captured fragment in the j-th 100kb fragment of the i-th cell. 's' is the accumulation function. The domain of the independent variable s is l∈[1,p2], p2=N type (frags(s)) indicates that there are p2 possible fragments s.
[0122] In formula (2) above, N is the number of all second capture segments in the first segment. genebody (cell(i), 100kb(j)) can be determined by the following formula (2.2):
[0123]
[0124] Where frags(t) represents fragment t, In formula (2.2), fragment t refers to the second captured fragment in the j-th 100kb fragment of the i-th cell. t is the accumulation function. The domain of the independent variable t is l∈[1,q2], where q2=N. type (frags(t)) indicates that there are q2 possible fragments t.
[0125] In one embodiment, determining whether the density distribution curve of the first fragment tends towards the transcription start site region based on the enrichment fraction of the transcription start site of the first fragment includes:
[0126] The enrichment fractions of transcription start sites for all first fragments in the test sample are sorted by size; the first k% of the first fragments from largest to smallest are identified as the first fragments whose density distribution curves tend to be in the transcription start site region.
[0127] The transcription start site enrichment score is a quantitative result of the distribution of captured fragments in the transcription start site region and the gene body region of the first fragment. The larger the transcription start site score, the closer the density distribution curve of the first fragment is to the transcription start site region.
[0128] Assuming the first fragment is a 100kb fragment and the second fragment is a 1kb fragment, to calculate the j-th 100kb fragment of the i-th cell in the test sample to represent any 100kb fragment in the test sample, the following formula (3) can be used to filter out the first fragment based on the enrichment fraction of the transcription start site of the first fragment:
[0129]
[0130] Among them, cut_off(ES tss ) is a filtering function that filters out the first fragment based on the enrichment fraction of the transcription start site of the first fragment, and top represents the upper quantile function. (k%) (ES tss (cell(i),100kb(j))) represents the first fragment in the top k% of the transcription start site enrichment fraction, i.e., the first fragment approaching the transcription start site region. Sample(a) represents the sample a to be tested. This indicates the presence of transcription start site enrichment score (ES). tss The first segment refers to the first segment in the sample a to be tested. P_value is the filtering ratio, i.e., the top k%, which can be set according to the actual situation. For example, P_value can be 0.05.
[0131] In the above embodiments, the distribution of captured fragments in the transcription start site region and gene body region of the first fragment is quantified as transcription start site enrichment score. Based on the transcription start site enrichment score of each first fragment, the first fragment with an excessively high transcription start site enrichment score is filtered out. This can effectively filter out the first fragment whose density distribution curve tends to the transcription start site region. This part of the first fragment does not conform to the characteristics of extrachromosomal circular DNA. The filtered first fragment can be further identified.
[0132] In one embodiment, if the fragment coverage of the first fragment in the cell reaches a first preset threshold and satisfies uniform distribution, then identifying the first fragment as an extrachromosomal circular DNA fragment includes:
[0133] Based on the sequence number of the first fragment and the sequence number of multiple consecutive second fragments in the first fragment, the nonparametric test method based on the cumulative distribution function in the Kolmogorov-Smirnov test is used to determine whether the fragment coverage of the first fragment in the cell conforms to the preset uniform distribution model; if the fragment coverage of the first fragment in the cell reaches the first preset threshold and conforms to the preset uniform distribution model, then the first fragment is identified as an extrachromosomal circular DNA fragment.
[0134] The fact that the fragment coverage of the first fragment reaches the first preset threshold indicates that the chromatin of the first fragment is highly open, and the fact that the fragment coverage of the first fragment is uniformly distributed indicates that the chromatin of the first fragment is averagely open. Extrachromosomal circular DNA can be identified from the filtered first fragment by the characteristics of high and average chromatin coverage.
[0135] Specifically, to determine whether the fragment coverage of the first fragment in the cell satisfies a uniform distribution, a non-parametric test based on the cumulative distribution function in the Kolmogorov-Smirnov test can be used to verify whether the fragment coverage of the first fragment in the cell conforms to the preset uniform distribution model. If the fragment coverage of the first fragment in the cell after filtering based on the first preset threshold conforms to the preset uniform distribution model, then the first fragment is identified as an extrachromosomal circular DNA fragment. If the fragment coverage of the first fragment in the cell after filtering based on the first preset threshold does not conform to the preset uniform distribution model, then the first fragment is identified as a non-extrachromosomal circular DNA fragment.
[0136] The Kolmogorov-Smirnov test (KS test) is a nonparametric test that assesses whether data conforms to a theoretical distribution based on the difference between the cumulative distribution function and the theoretical distribution function. In the KS test, a nonparametric method based on the cumulative distribution function, determining whether the fragment coverage of the first fragment in the cell conforms to a pre-defined uniform distribution model includes:
[0137] S1. For each first fragment, it is assumed that the fragment coverage of the first fragment in the cell conforms to a preset uniform distribution model.
[0138] This hypothesis is also called the null hypothesis (H0). Assuming the first fragment is a 100kb fragment and the second fragment is a 1kb fragment, taking the verification of whether the fragment coverage of the j-th 100kb fragment of the i-th cell in the test sample is uniformly distributed as an example, the null hypothesis is the following formula (4):
[0139]
[0140] Where Location is a location function, Location1kb(n) This represents the position of the nth 1kb segment. In formula (4), the nth 1kb fragment refers to the 1kb fragment in the jth 100kb fragment of the i-th cell. The domain of n is [1, 100]. frags(cell(i),100kb(j),1kb(n)) The sequence number of the nth 1kb segment of the jth 100kb segment in the i-th cell is represented by KS(uniform), which represents the pre-defined uniform distribution model in the KS test.
[0141] The alternative hypothesis (H1) is the opposite of the null hypothesis (H0), which states that the fragment coverage of the first fragment in the cell does not conform to the pre-defined uniform distribution model.
[0142] S2. Determine the significance parameter value of the first segment based on the sequence number of the first segment and the sequence number of multiple consecutive second segments in the first segment.
[0143] The significance parameter represents the probability that the KS test statistic or a more extreme value will appear in the second segment of the first segment. The significance parameter value for the first segment is determined based on the sequence number of the first segment and the sequence number of multiple consecutive second segments within the first segment, including:
[0144] S2.1 Determine the cumulative distribution function value and theoretical distribution function value of the first segment based on the sequence number of the first segment and the sequence number of multiple consecutive second segments in the first segment.
[0145] The cumulative distribution function (CDF) is calculated using the following formula (5):
[0146]
[0147]
[0148] Where F(x) is the cumulative distribution function value of the first segment. In formula (5), the nth 1kb refers to a consecutive 1kb segment within the jth 100kb segment of the i-th cell, where N... frags(cell(i),100kb(j),1kb(n)) x represents the sequence number of the 1kb fragment in the j-th 100kb fragment of the i-th cell. The value of x ranges from [a+1,b], where a = 0 and b = 100.
[0149] In the above formula (5), the sequence number of the 1kb fragment in the j-th 100kb fragment of the i-th cell can be determined by the following formula (5.2):
[0150]
[0151] Where frags(l) represents fragment l, In formula (5.2), segment l refers to the second segment in the j-th 100kb segment of the i-th cell. l is the accumulation function. The domain of the independent variable l is l∈[1,q3], where q3=N type (frags(l)) indicates that there are q3 possible fragments of fragment l.
[0152] The theoretical distribution function (PDF) is calculated using the following formula (6):
[0153]
[0154] Where G(x) is the theoretical distribution function value of the first segment, and x represents the independent variable of the theoretical distribution function G(x), that is, the xth consecutive 1kb segment of the 100kb segment, and the domain of the independent variable x is [1,100].
[0155] S2.2 Calculate the KS test statistic and the significance parameter of the KS test statistic based on the cumulative distribution function value and the theoretical distribution function value.
[0156] The formula for calculating the KS test statistic Dn is as follows (7):
[0157] Dn(x)=max(abs(F(x)-G(x))) (7)
[0158] Where x is the independent variable of the KS test statistic Dn(x), i.e. the xth consecutive 1kb segment of the 100kb segment, abs is the absolute value function, max is the maximum value function, F(x) is the cumulative distribution function value of the first segment, and G(x) is the theoretical distribution function value of the first segment.
[0159] Calculating the significance parameter (also known as the p-value) of the KS test statistic is equivalent to calculating the probability that a value of Dn or more extreme than the p-value will appear in the observed statistic (the distribution of 1kb segments within a 100kb segment) under the null hypothesis. Based on the calculated KS test statistic Dn, the significance parameter value of Dn can be calculated using the `scipy.stats.kstest` function in the Python programming language.
[0160] S3. Verify or falsify the hypothesis based on the significance parameter value.
[0161] The process of verifying or falsifying a hypothesis based on the significance parameter value includes: verifying the hypothesis if the significance parameter value is greater than or equal to the fourth preset threshold; and falsifying the hypothesis if the significance parameter value is less than the fourth preset threshold.
[0162] For example, the fourth preset threshold can be set to 0.05, meaning that 0.05 is considered the significance level of the significance parameter value. If the significance parameter value is less than 0.05, the null hypothesis H0 is rejected, that is, the hypothesis that conforms to the preset uniform distribution model is falsified, and it is determined that the fragment coverage of the first fragment in the cell does not conform to the preset uniform distribution model. If the significance parameter value is greater than or equal to 0.05, the null hypothesis is not rejected, that is, the hypothesis that conforms to the preset uniform distribution model is verified, and it is determined that the fragment coverage of the first fragment in the cell conforms to the preset uniform distribution model. If the fragment coverage of the first fragment in the cell simultaneously meets the condition of reaching the first preset threshold, then the first fragment can be determined to be an extrachromosomal circular DNA fragment.
[0163] In the above embodiments, the KS test is used to verify whether the fragment coverage of the first fragment in the cell conforms to a preset uniform distribution model.
[0164] In one embodiment, if the fragment coverage of the first fragment in the cell reaches a first preset threshold and satisfies uniform distribution, then after identifying the first fragment as an extrachromosomal circular DNA fragment, the method further includes:
[0165] Among the identified extrachromosomal circular DNA fragments, homologous and adjacent extrachromosomal circular DNA fragments are merged to obtain the merged extrachromosomal circular DNA fragment results. Based on the identified extrachromosomal circular DNA fragments in each cell and the merged extrachromosomal circular DNA fragment results, a single-cell extrachromosomal circular DNA matrix result for the sample to be tested is generated. The single-cell extrachromosomal circular DNA matrix result includes the coverage matrix of the extrachromosomal circular DNA fragments in each cell and the merged extrachromosomal circular DNA fragment results.
[0166] In this context, homology refers to the presence of fragments within the same cell, while adjacency refers to their adjacent positions on the genome. Homologous and adjacent extrachromosomal circular DNA fragments are those that exist in the same cell and are adjacent in position on the genome. Among the identified extrachromosomal circular DNA fragments, homologous and adjacent fragments are merged to obtain a merged extrachromosomal circular DNA fragment. If two extrachromosomal circular DNA fragments are homologous and adjacent, they are merged into a single merged extrachromosomal circular DNA fragment; if multiple extrachromosomal circular DNA fragments are homologous and adjacent to each other, they are merged into a single merged extrachromosomal circular DNA fragment.
[0167] For the merged extrachromosomal circular DNA fragments, the merged extrachromosomal circular DNA fragments can be treated as a whole, and the fragment coverage of the merged extrachromosomal circular DNA fragments can be obtained according to the fragment coverage calculation in the aforementioned embodiments, thereby generating the single-cell extrachromosomal circular DNA matrix result of the sample to be tested.
[0168] Unmerged extrachromosomal circular DNA fragments are designated as the first ecDNA, and merged extrachromosomal circular DNA fragments are designated as the second ecDNA. The resulting matrix includes the fragment coverage of the first and second ecDNA in each cell. Each row of the matrix represents a cell, and each column represents either a first or second ecDNA. The numerical values in the matrix represent the fragment coverage of the first or second ecDNA.
[0169] In one embodiment, merging homologous and adjacent extrachromosomal circular DNA fragments among the identified extrachromosomal circular DNA fragments includes:
[0170] For two adjacent extrachromosomal circular DNA fragments on the genome, the fragment coverage of the two extrachromosomal circular DNA fragments in each cell is calculated, and the correlation coefficient of the two extrachromosomal circular DNA fragments is determined based on the fragment coverage of the two extrachromosomal circular DNA fragments in each cell. If the correlation coefficient is greater than a third preset threshold, the two extrachromosomal circular DNA fragments are identified as homologous and adjacent extrachromosomal circular DNA fragments, and the two extrachromosomal circular DNA fragments are merged.
[0171] When identifying homologous and adjacent extrachromosomal circular DNA fragments, pairwise comparisons can be performed. Whether two extrachromosomal circular DNA fragments are adjacent can be determined based on their positions on the genome. For two extrachromosomal circular DNA fragments that are adjacent on the genome, their homology can be determined by calculating their correlation coefficient.
[0172] In one embodiment, the correlation coefficient is Spearman's rank correlation coefficient. The correlation coefficient r between two extrachromosomal circular DNA fragments... S It can be calculated using the following formula (8):
[0173]
[0174] The formulas for calculating x1 and x2 are as follows:
[0175] x1=cpr1(Coverage(cell(i),cpr1),…),i∈[1,n],n=N(cell) (8.1)
[0176] x2=cpr2(Coverage(cell(i),cpr2),…),i∈[1,n],n=N(cell) (8.2)
[0177] Here, cpr (circle-produced region, ecDNA circular production region) refers to a 100kb fragment identified as an ecDNA fragment, and cpr1 and cpr2 represent the two ecDNA fragments for which correlation coefficients are to be calculated. Coverage represents the corresponding fragment coverage. R(x1) and R(x2) represent the average rank numbers of x1 and x2 respectively, sorted in ascending order. cov(R(x1),R(x2)) represents the covariance of the rank variables R(x1) and R(x2). σ R(x1) σ R(x2) These represent the standard deviations of the rank variables R(x1) and R(x2), respectively.
[0178] Furthermore, the significance score p4 of the correlation coefficient can be calculated using the following formula (9):
[0179]
[0180] Where n is the total number of cells in the sample to be tested, and r S The correlation coefficient is the correlation coefficient between two extrachromosomal circular DNA fragments.
[0181] Optionally, the significance score p4 of the correlation coefficient can also be implemented using the Python programming language, as scipy.stats.spearmanr(x1,x2,alternative=′two-sided′), where the formulas for calculating the constraint parameters x1 and x2 are the same as those in equations (8.1) and (8.2) above.
[0182] In one embodiment, the length of the second segment interval is an integer multiple of the length of the first segment interval. After generating the single-cell extrachromosomal circular DNA matrix result of the sample to be tested based on the identified extrachromosomal circular DNA fragments in each cell and the combined result of the extrachromosomal circular DNA fragments, the method further includes:
[0183] The results of the single-cell extrachromosomal circular DNA matrix of the test sample are visualized, in which homologous and adjacent extrachromosomal circular DNA fragments in a single cell and the second fragment contained in the homologous and adjacent extrachromosomal circular DNA fragments are displayed on the same image.
[0184] This application identifies extrachromosomal circular DNA fragments from a first fragment with single-cell precision. After obtaining the matrix results, the contents of the matrix results can be visualized. Since the matrix results are based on single-cell precision, each cell can be displayed as a single image. The visualized content can include the first fragment contained in the cell, and the consecutive second fragments contained in each first fragment within the cell. The first fragment can include first fragments identified as ecDNA, or it can include first fragments not identified as ecDNA, and it can include homologous and adjacent ecDNA fragments, or it can include ecDNA fragments that have not been merged.
[0185] When the length of the second segment interval is an integer multiple of the length of the first segment interval, these contents can be distinguished by lines, colors, saturation, etc., and displayed at the corresponding position of the first segment, or at the corresponding position of the continuous second segment contained in the first segment.
[0186] Please see Figure 3 , Figure 3 The two images, one above the other, represent two cell types. The first image shows cells with high ecDNA abundance, where multiple ecDNA fragments were identified. The second image shows cells with low ecDNA abundance. In both images, each row represents a 100kb fragment, and each 100kb fragment comprises 100 consecutive 1kb fragments, represented by the horizontal axis 1-100. In cells identified as having high ecDNA abundance, the 100kb fragments, after being divided into 100 equal parts, exhibit an average high openness and continuous adjacent regions. In cells identified as having low ecDNA abundance, the 100kb fragments, after being divided into 100 equal parts, do not exhibit a continuous high openness (high coverage) state.
[0187] In the above embodiments, the matrix results were visually displayed through visualization. From the visualization results, it is clear whether the abundance of ecDNA in the cells is high or low, whether the ecDNA fragments in the cells are adjacent, and the degree of openness and whether the openness is average after staining in each first fragment.
[0188] This application also provides a specific embodiment. A specific embodiment of the method for identifying extrachromosomal circular DNA is as follows:
[0189] This embodiment analyzes single-cell ATAC sequencing data of human colorectal adenocarcinoma COLO320DM cell line, human chronic leukemia K562 cell line, human acute leukemia HL60 cell line, human triple-negative breast cancer MDAMB231 cell line, and human breast cancer BRCA tumor samples. The abundance of ecDNA gradually increases from front to back in these samples.
[0190] For each of the above samples, calculate the fragment coverage of the 100kb fragment in the cell, and plot the probability distribution density map of the fragment coverage of the 100kb fragment in a single cell. Please refer to [link to relevant documentation]. Figure 4 and Figure 5 ,in Figure 4 Visualization results of fragment coverage in K562, HL60, and MDAMB231 cell lines. Figure 5 Visualization results of fragment coverage for BRCA tumor samples. Figure 4 and Figure 5 In the data, the vast majority of 100kb fragments have a coverage of less than 6, with peak values ranging from 0 to 6. This aligns with the basic data characteristics of chromatin accessibility sequencing (i.e., background signals / normal chromosomal regions, copy number variation regions). The openness of these regions is reflected in a coverage range of 0-6. Therefore, filtering out 100kb fragments with low coverage (less than 6), the remaining 100kb fragments (corresponding to coverage greater than 6) are considered highly open DNA regions, consistent with the known characteristics of ecDNA: high chromatin openness.
[0191] Based on the fragment coverage of the 100kb fragment in the cells of each sample, density distribution curves from the enrichment locations of the 100kb fragment to the TSS region were plotted. Please refer to [link to relevant documentation]. Figure 6 and Figure 7 ,in Figure 6 This is a schematic diagram of the density distribution curves of the K562, HL60, and MDAMB231 cell lines. Figure 7 This is a schematic diagram of the density distribution curve of a BRCA tumor sample. Based on... Figure 6 and Figure 7 It is easy to see that, unlike the background signal / normal chromosome region and CNV region, which have a high density distribution of TSS enrichment scores in the range greater than 0.5 (0.5-1), especially the peak at the TSS enrichment score of 1, the TSS enrichment scores of ecDNA are generally low: most are less than 0.5, and there is almost no enrichment signal (no peak) at position 1. This proves that as the screening proceeds, the 100kb after screening further conforms to the known characteristics of ecDNA: average chromatin openness (tendency not to enrich in TSS region).
[0192] After completing the identification of ecDNA, the matrix results and process data obtained from the identification are visualized. Please refer to [link / reference]. Figure 3 . Figure 3The two images, one above the other, represent two cell types. The first image shows cells with high ecDNA abundance, where multiple ecDNA fragments were identified. The second image shows cells with low ecDNA abundance. In both images, each row represents a 100kb fragment, and each 100kb fragment comprises 100 consecutive 1kb fragments, represented by the horizontal axis 1-100. In cells identified as having high ecDNA abundance, the 100kb fragments, after being divided into 100 equal parts, exhibit an average high openness and continuous adjacent regions. In cells identified as having low ecDNA abundance, the 100kb fragments, after being divided into 100 equal parts, do not exhibit a continuous high openness (high coverage) state.
[0193] These results demonstrate the practicality and accuracy of this method in ecDNA identification.
[0194] The extrachromosomal circular DNA (ecDNA) identification method proposed in this application does not rely on microscopic imaging technology. It utilizes single-cell resolution information and highly open chromatin sequence information introduced by single-cell sequencing data to indiscriminately identify characteristic patterns of ecDNA fragments in individual tumor cells, and based on this, obtains quantitative statistical values of ecDNA abundance in individual tumor cells. Based on the identification of ecDNA, subsequent analysis can be performed on the tumor cell-specific ecDNA functions and evolutionary mechanisms, which is of great significance for detecting circulating tumor cells, analyzing the complex heterogeneity of tumor cells, and elucidating drug resistance mechanisms of targeted therapies. Furthermore, this method has low computational complexity and fast running time, making it a high-precision, high-throughput ecDNA identification method.
[0195] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0196] Based on the same inventive concept, this application also provides an extrachromosomal circular DNA identification device for implementing the above-described extrachromosomal circular DNA identification method. The solution provided by this device is similar to the solution described in the above method; therefore, the specific limitations of one or more extrachromosomal circular DNA identification device embodiments provided below can be found in the limitations of the extrachromosomal circular DNA identification method described above, and will not be repeated here.
[0197] In one exemplary embodiment, such as Figure 8 As shown, a device for identifying extrachromosomal circular DNA is provided, comprising: a data acquisition module 701, a fragment segmentation module 702, a coverage determination module 703, and a fragment identification module 704, wherein:
[0198] The data acquisition module 701 is used to acquire the sample to be tested and its sequencing data. The sample to be tested includes multiple cells, and the sequencing data includes the number of captured fragments for each cell. The captured fragments are DNA fragments captured by single-cell sequencing of the sample to be tested.
[0199] The fragment segmentation module 702 is used to segment the captured fragments to obtain a first fragment matrix located in a first fragment interval and a second fragment matrix located in a second fragment interval. The first fragment matrix contains the first fragment of each cell and the sequence number of the first fragment. The second fragment matrix contains the second fragment of each cell and the sequence number of the second fragment. The sequence number is the number of captured fragments contained therein. The second fragment interval is located within the first fragment interval. The first fragment matrix includes multiple consecutive second fragments.
[0200] The coverage determination module 703 is used to determine the fragment coverage of the first fragment in the cell based on the sequence number of the first fragment and the sequence number of multiple consecutive second fragments in the first fragment;
[0201] The fragment identification module 704 is used to identify the first unit fragment as an extrachromosomal circular DNA fragment if the fragment coverage of the first fragment in the cell reaches a first preset threshold and satisfies uniform distribution.
[0202] Please see Figure 9 In one embodiment, the extrachromosomal circular DNA identification device further includes a first filtering module 705, which, before segmenting the captured fragments to obtain a first fragment matrix located in a first fragment interval and a second fragment matrix located in a second fragment interval, is used to:
[0203] The length of the captured fragment is determined based on the start and end points of the captured fragment;
[0204] Filter out captured segments whose length falls outside the preset upper and lower limits; and / or
[0205] Statistical analysis of the fragment lengths captured in cells is performed to determine the peak values of fragment length distribution within the cells, and captured fragments whose peak values fall outside a preset peak range are filtered out; and / or
[0206] The number of captured fragments in the cells is counted, and captured fragments whose number of captured fragments in the cells is less than a second preset threshold are filtered out.
[0207] In one embodiment, when determining the fragment coverage of the first fragment in a cell based on the sequence number of the first fragment and the sequence number of multiple consecutive second fragments within the first fragment, the coverage determination module 703 is further configured to:
[0208] The sum of the number of various captured fragments in the cell is taken as the cell sequence number;
[0209] For the first segment in the cell, the sum of the sequence numbers of all second segments in the first segment is determined based on the sequence number of each second segment in the first segment;
[0210] The fragment coverage of the first fragment in the cell is determined by the ratio of the sum of the sequence numbers of all second fragments in the first fragment of the cell to the total number of sequences in the cell.
[0211] In one embodiment, when determining the sum of sequence numbers of all second segments in a first segment based on the sequence number of each second segment in the first segment, the coverage determination module 703 is further configured to:
[0212] For the second segment in the first segment, the sum of the number of multiple captured segments in the second segment is taken as the sequence number of the second segment;
[0213] Determine the sum of the sequence numbers of all second segments in the first segment based on the sequence number of each second segment in the first segment.
[0214] Please continue reading. Figure 9 In one embodiment, the extrachromosomal circular DNA identification device further includes a second filtering module 706, which, before identifying the first fragment as an extrachromosomal circular DNA fragment if the fragment coverage of the first fragment in the cell reaches a first preset threshold and satisfies uniform distribution, is used to:
[0215] For each first segment, determine the transcription start site region and gene body region of the first segment;
[0216] Based on the distribution of captured fragments in the transcription start site region and gene body region of the first fragment, the transcription start site enrichment score of the first fragment is determined. The transcription start site enrichment score is used to assess the average openness of chromatin.
[0217] Based on the enrichment fraction of the transcription start site of the first fragment, determine whether the direction of the density distribution curve of the first fragment tends to the transcription start site region. The density distribution curve is the density distribution curve from the enrichment position of the first fragment to the transcription start site region. The enrichment position is the position where the captured fragment in the start site region intersects with the first fragment.
[0218] If so, then filter out the first segment.
[0219] In one embodiment, the captured fragment includes a first captured fragment and a second captured fragment, wherein the first captured fragment is a captured fragment intersecting with the transcription start site region, and the second captured fragment is a captured fragment intersecting with the gene body region; when determining the transcription start site enrichment fraction of the first fragment based on the distribution of captured fragments in the transcription start site region and the gene body region of the first fragment, the second filtering module 706 is further configured to:
[0220] Get the number of first captured segments in the first segment, and get the number of second captured segments in the first segment;
[0221] The transcription start site enrichment fraction of the first fragment is determined based on the number of the first captured fragment and the number of the second captured fragment.
[0222] In one embodiment, the number of all first capture segments in the first segment is the sum of the number of multiple different first capture segments, and the number of second capture segments in the first segment is the sum of the number of multiple different second capture segments.
[0223] In one embodiment, when determining whether the density distribution curve of the first fragment tends towards the transcription start site region based on the enrichment fraction of the transcription start site of the first fragment, the second filtering module 706 is further configured to:
[0224] The enrichment scores of transcription start sites for all first fragments in the test sample were sorted by size.
[0225] The first fragment, which is the first k% from largest to smallest, is identified as the first fragment whose density distribution curve tends to be in the transcription start site region.
[0226] In one embodiment, when the fragment coverage of the first fragment in the cell reaches a first preset threshold and satisfies uniform distribution, and the first fragment is identified as an extrachromosomal circular DNA fragment, the fragment identification module 704 is further configured to:
[0227] Based on the sequence number of the first fragment and the sequence number of multiple consecutive second fragments in the first fragment, the nonparametric test method based on the cumulative distribution function in the Kolmogorov-Smirnov test is used to determine whether the fragment coverage of the first fragment in the cell conforms to the preset uniform distribution model.
[0228] If the fragment coverage of the first fragment in the cell reaches a first preset threshold and conforms to a preset uniform distribution model, then the first fragment is identified as an extrachromosomal circular DNA fragment.
[0229] Please continue reading. Figure 9 In one embodiment, the extrachromosomal circular DNA identification device further includes a fragment merging module 707. After identifying the first fragment as an extrachromosomal circular DNA fragment if the fragment coverage of the first fragment in the cell reaches a first preset threshold and satisfies uniform distribution, the fragment merging module 707 is used to:
[0230] Among the identified extrachromosomal circular DNA fragments, homologous and adjacent extrachromosomal circular DNA fragments are merged to obtain the merged results of extrachromosomal circular DNA fragments;
[0231] Based on the identified extrachromosomal circular DNA fragments in each cell and the combined results of the extrachromosomal circular DNA fragments, a single-cell extrachromosomal circular DNA matrix result for the sample to be tested is generated. The single-cell extrachromosomal circular DNA matrix result includes a coverage matrix of the combined results of extrachromosomal circular DNA fragments in each cell.
[0232] In one embodiment, when merging homologous and adjacent extrachromosomal circular DNA fragments among the identified extrachromosomal circular DNA fragments, the fragment merging module 707 is further configured to:
[0233] For two adjacent extrachromosomal circular DNA fragments on the genome, calculate the fragment coverage of the two extrachromosomal circular DNA fragments in each cell, and determine the correlation coefficient of the two extrachromosomal circular DNA fragments based on the fragment coverage of the two extrachromosomal circular DNA fragments in each cell.
[0234] If the correlation coefficient is greater than the third preset threshold, the two extrachromosomal circular DNA fragments are identified as homologous and adjacent extrachromosomal circular DNA fragments, and the two extrachromosomal circular DNA fragments are merged.
[0235] Please continue reading. Figure 9 In one embodiment, the length of the second fragment interval is an integer multiple of the length of the first fragment interval. The extrachromosomal circular DNA identification device further includes a visualization module 708. After generating a single-cell extrachromosomal circular DNA matrix result of the sample to be tested based on the identified extrachromosomal circular DNA fragments in each cell and the merging results of the extrachromosomal circular DNA fragments, the visualization module 708 is used for:
[0236] The results of the single-cell extrachromosomal circular DNA matrix of the test sample are visualized, in which homologous and adjacent extrachromosomal circular DNA fragments in a single cell and the second fragment contained in the homologous and adjacent extrachromosomal circular DNA fragments are displayed on the same image.
[0237] The aforementioned extrachromosomal circular DNA identification device includes a data acquisition module 701 that acquires the sample to be tested and its sequencing data. The sample to be tested includes multiple cells, and the sequencing data contains a captured fragment for each cell. The captured fragment is a DNA fragment captured by single-cell sequencing of the sample to be tested. A fragment segmentation module 702 segments the captured fragment to obtain a first fragment matrix located in a first fragment interval and a second fragment matrix located in a second fragment interval. The first fragment matrix contains the first fragment of each cell and the sequence number of the first fragment. The second fragment matrix contains the second fragment of each cell and the sequence number of the second fragment. The sequence number is the number of captured fragments included. The second fragment interval is located within the first fragment interval. The first fragment matrix includes multiple consecutive second fragments. A coverage determination module 703 determines the fragment coverage of the first fragment in the cell based on the sequence number of the first fragment and the sequence number of multiple consecutive second fragments in the first fragment. If the fragment coverage of the first fragment in the cell reaches a first preset threshold and satisfies uniform distribution, the fragment identification module 704 identifies the first fragment as an extrachromosomal circular DNA fragment. In this way, by dividing the captured fragments into different fragment intervals and calculating the fragment coverage, the degree of chromatin openness of DNA fragments can be quantified. Fragments with highly open chromatin can be distinguished by setting a first preset threshold for fragment coverage, while fragments with uniform coverage indicate that the chromatin within them is open on an average basis. Thus, the characteristics of highly open and evenly open chromatin can be used to accurately identify extrachromosomal circular DNA.
[0238] Each module in the aforementioned extrachromosomal circular DNA identification device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.
[0239] Terms such as “component,” “module,” and “system” are intended to refer to computer-related entities, which can be hardware, a combination of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, executable code, a thread of execution, a program, and / or a computer. For illustration, a running program on a server and the server itself can both be components. One or more components may reside within a process and / or a thread of execution, and components may be located within a single computer and / or distributed across two or more computers.
[0240] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 10 As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes a non-volatile computer-readable storage medium and internal memory. The non-volatile computer-readable storage medium stores the operating system, computer-readable instructions, and a database. The internal memory provides an environment for the operation of the operating system and computer-readable instructions in the non-volatile computer-readable storage medium. The database stores relevant data. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When the computer-readable instructions are executed by the processor, they implement a method for identifying extrachromosomal circular DNA.
[0241] In one exemplary embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 11As shown, the computer device includes a processor, memory, input / output interfaces, a communication interface, a display unit, and an input device. The processor, memory, and input / output interfaces are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interfaces. The processor provides computational and control capabilities. The memory includes a non-volatile computer-readable storage medium and internal memory. The non-volatile computer-readable storage medium stores the operating system and computer-readable instructions. The internal memory provides an environment for the operation of the operating system and computer-readable instructions in the non-volatile computer-readable storage medium. The input / output interfaces are used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When the computer-readable instructions are executed by the processor, they implement a method for identifying extrachromosomal circular DNA. The display unit is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.
[0242] Those skilled in the art will understand that Figure 10 , 11 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0243] In one embodiment, a computer device is also provided, including a memory and a processor, the memory storing computer-readable instructions, the processor executing the computer-readable instructions to implement the steps in the above method embodiments.
[0244] In one embodiment, a computer-readable storage medium is provided storing computer-readable instructions that, when executed by a processor, implement the steps in the above method embodiments.
[0245] In one embodiment, a computer program product is provided, the computer program product including computer-readable instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer-readable instructions from the computer-readable storage medium, and executes the computer-readable instructions, causing the computer device to perform the steps in the above-described method embodiments.
[0246] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware with computer-readable instructions. These computer-readable instructions can be stored in a non-volatile computer-readable storage medium. When executed, these computer-readable instructions can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0247] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0248] The above embodiments are merely illustrative of several implementation methods of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A method for identifying extrachromosomal circular DNA, characterized in that, include: Acquire a sample to be tested and its sequencing data, wherein the sample to be tested includes multiple cells and the sequencing data contains a capture fragment of each cell, wherein the capture fragment is a DNA fragment captured by single-cell sequencing of the sample to be tested. The captured fragments are segmented to obtain a first fragment matrix located in a first fragment interval and a second fragment matrix located in a second fragment interval; the first fragment matrix contains the first fragment of each cell and the sequence number of the first fragment, the second fragment matrix contains the second fragment of each cell and the sequence number of the second fragment, the sequence number being the number of captured fragments contained therein, the second fragment interval being located within the first fragment interval, and the first fragment matrix including multiple consecutive second fragments; The fragment coverage of the first fragment in the cell is determined based on the sequence number of the first fragment and the sequence number of multiple consecutive second fragments in the first fragment; If the fragment coverage of the first fragment in the cell reaches a first preset threshold and satisfies uniform distribution, then the first fragment is identified as an extrachromosomal circular DNA fragment.
2. The method according to claim 1, characterized in that, Before segmenting the captured segments to obtain a first segment matrix located in the first segment interval and a second segment matrix located in the second segment interval, the method further includes: The length of the captured fragment is determined based on the start and end points of the captured fragment; Filter out captured segments whose length falls outside the preset upper and lower limits; and / or Statistical analysis of the fragment lengths captured in cells is performed to determine the peak values of fragment length distribution within the cells, and captured fragments whose peak values fall outside a preset peak range are filtered out; and / or The number of captured fragments in the cells is counted, and captured fragments whose number of captured fragments in the cells is less than a second preset threshold are filtered out.
3. The method according to claim 1, characterized in that, Determining the fragment coverage of the first fragment in the cell based on the sequence number of the first fragment and the sequence number of the plurality of consecutive second fragments within the first fragment includes: The sum of the number of various captured fragments in the cell is taken as the sequence number of the cell; For the first segment in the cell, the sum of the sequence numbers of all second segments in the first segment is determined based on the sequence number of each second segment in the first segment; The fragment coverage of the first fragment in the cell is determined based on the ratio of the sum of the sequence numbers of all second fragments in the first fragment of the cell to the total sequence number of the cell.
4. The method according to claim 1, characterized in that, Before identifying the first fragment as an extrachromosomal circular DNA fragment if the fragment coverage of the first fragment in the cell reaches a first preset threshold and satisfies uniform distribution, the method further includes: For each first segment, determine the transcription start site region and gene body region of the first segment; Based on the distribution of captured fragments in the transcription start site region and gene body region of the first fragment, the transcription start site enrichment score of the first fragment is determined, and the transcription start site enrichment score is used to assess the average openness of chromatin; Based on the enrichment fraction of the transcription start site of the first fragment, determine whether the direction of the density distribution curve of the first fragment tends to the transcription start site region, wherein the density distribution curve is the density distribution curve from the enrichment position of the first fragment to the transcription start site region, and the enrichment position is the position where the captured fragment of the start site region intersects with the first fragment. If so, then the first segment is filtered out.
5. The method according to claim 4, characterized in that, The capture fragment includes a first capture fragment and a second capture fragment, wherein the first capture fragment is a capture fragment that intersects with the transcription start site region, and the second capture fragment is a capture fragment that intersects with the gene body region; The step of determining the transcription start site enrichment fraction of the first fragment based on the distribution of captured fragments in the transcription start site region and gene body region of the first fragment includes: Obtain the number of first captured segments in the first segment, and obtain the number of second captured segments in the first segment; The transcription start site enrichment fraction of the first fragment is determined based on the number of the first captured fragment and the number of the second captured fragment.
6. The method according to claim 5, characterized in that, The number of all first capture segments in the first segment is the sum of the number of various first capture segments, and the number of second capture segments in the first segment is the sum of the number of various second capture segments.
7. The method according to claim 4, characterized in that, The step of determining whether the density distribution curve of the first fragment tends to the transcription start site region based on the enrichment fraction of the transcription start site of the first fragment includes: The enrichment scores of transcription start sites for all first fragments in the test sample were sorted by size. The first fragment, which is the first k% from largest to smallest, is identified as the first fragment whose density distribution curve tends to be in the transcription start site region.
8. The method according to claim 1, characterized in that, The step of identifying the first fragment as an extrachromosomal circular DNA fragment if the fragment coverage of the first fragment in the cell reaches a first preset threshold and satisfies uniform distribution includes: Based on the sequence number of the first fragment and the sequence number of multiple consecutive second fragments in the first fragment, the nonparametric test method based on the cumulative distribution function in the Kolmogorov-Smirnov test is used to determine whether the fragment coverage of the first fragment in the cell conforms to the preset uniform distribution model. If the fragment coverage of the first fragment in the cell reaches a first preset threshold and conforms to a preset uniform distribution model, then the first fragment is identified as an extrachromosomal circular DNA fragment.
9. The method according to claim 1, characterized in that, If the fragment coverage of the first fragment in the cell reaches a first preset threshold and satisfies uniform distribution, then after identifying the first fragment as an extrachromosomal circular DNA fragment, the method further includes: Among the identified extrachromosomal circular DNA fragments, homologous and adjacent extrachromosomal circular DNA fragments are merged to obtain the merged results of extrachromosomal circular DNA fragments; Based on the identified extrachromosomal circular DNA fragments in each cell and the combined results of the extrachromosomal circular DNA fragments, a single-cell extrachromosomal circular DNA matrix result for the sample to be tested is generated. The single-cell extrachromosomal circular DNA matrix result includes a coverage matrix of the combined results of extrachromosomal circular DNA fragments in each cell.
10. The method according to claim 9, characterized in that, The process of merging homologous and adjacent extrachromosomal circular DNA fragments among the identified extrachromosomal circular DNA fragments includes: For two adjacent extrachromosomal circular DNA fragments on the genome, calculate the fragment coverage of the two extrachromosomal circular DNA fragments in each cell, and determine the correlation coefficient of the two extrachromosomal circular DNA fragments based on the fragment coverage of the two extrachromosomal circular DNA fragments in each cell; If the correlation coefficient is greater than the third preset threshold, the two extrachromosomal circular DNA fragments are identified as homologous and adjacent extrachromosomal circular DNA fragments, and the two extrachromosomal circular DNA fragments are merged.
11. A device for identifying extrachromosomal circular DNA, characterized in that, include: The data acquisition module is used to acquire the sample to be tested and the sequencing data of the sample to be tested. The sample to be tested includes multiple cells, and the sequencing data includes the number of captured fragments for each cell. The captured fragments are DNA fragments captured by single-cell sequencing of the sample to be tested. The fragment segmentation module is used to segment the captured fragments to obtain a first fragment matrix located in a first fragment interval and a second fragment matrix located in a second fragment interval; the first fragment matrix contains the first fragment of each cell and the sequence number of the first fragment, the second fragment matrix contains the second fragment of each cell and the sequence number of the second fragment, the sequence number being the number of captured fragments contained therein, the second fragment interval being located within the first fragment interval, and the first fragment matrix including multiple consecutive second fragments; The coverage determination module is used to determine the fragment coverage of the first fragment in the cell based on the sequence number of the first fragment and the sequence number of multiple consecutive second fragments in the first fragment; The fragment identification module is used to identify the first unit fragment as an extrachromosomal circular DNA fragment if the fragment coverage of the first fragment in the cell reaches a first preset threshold and satisfies uniform distribution.
12. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 10.
13. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 10.