Method, device, terminal, medium and product for analyzing outer circular DNA (deoxyribonucleic acid) of chromosome based on long-read-long sequencing data
By conducting tandem repeat sequence detection and three-way comparative analysis on long-read sequencing data, the problem that existing tools cannot make full use of HiFi sequencing data is solved, and efficient and accurate eccDNA detection and classification are achieved, improving analysis efficiency and result reliability.
Patent Information
- Application Number
- CN202510358872.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-03-25
AI Technical Summary
Existing extrachromosomal circular DNA analysis tools cannot fully utilize the high accuracy of HiFi sequencing data, lack systematic classification standards, low analysis efficiency and poor reliability of results, making it difficult to make standardized comparisons.
By obtaining long-read and long-sequencing data, establishing a reference genome database, performing tandem repeat sequence detection and filtering, and performing three comparative analysis, including the first, the second and the third comparative, the extrachromosomal circular DNA analysis results were obtained respectively, and combined and classified processing was performed.
It improves the sensitivity and accuracy of eccDNA detection, and can detect and classify multiple types of eccDNA simultaneously, providing comprehensive eccDNA analysis, improving the analysis efficiency and reliability of results.
Smart Images

Figure CN120340604A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of bioinformatics technology, and particularly to an extrachromosomal circular DNA analysis method, device, terminal, medium and product based on long-read sequencing data. Background Art
[0002] Extrachromosomal circular DNA (eccDNA) is a class of circular DNA molecules widely present in eukaryotic cells, and plays an important role in biological processes such as genomic plasticity, gene expression regulation, and cellular stress response. Research shows that eccDNA is closely related to the occurrence and development of various diseases. Therefore, the analysis of eccDNA is of great significance in disease research, especially in the occurrence and progression of tumors.
[0003] With the rapid development of high-throughput sequencing technology, the third-generation sequencing technology has become a popular choice in scientific research and clinical applications and has been widely adopted and applied. The third-generation sequencing platforms (such as nanopore sequencing and single-molecule real-time sequencing) have brought long-read data to genomics. However, the third-generation sequencing often has a relatively high sequencing error rate, which makes existing EccDNA analysis tools focus on error correction strategies in design, tolerate larger error matches, and to a certain extent sacrifice the accurate measurement ability for complex chimeric structures or low-abundance circular fragments. In recent years, PacBio HiFi sequencing technology can provide long-read data with an average accuracy of 99% or even higher (the length can reach thousands to tens of thousands of bases), bringing higher precision to the identification of complex genomic structures (including eccDNA). However, existing eccDNA analysis tools (such as CReSIL, SMOOTH-seq, etc.) have the following main problems:
[0004] (1) Most existing analysis tools have inherited the analysis process for dealing with high-error-rate data and cannot fully utilize the advantages of HiFi data;
[0005] (2) Existing analysis tools often treat all circular DNAs as the same type, lacking a systematic classification standard and unable to meet the needs of accurately studying the biological characteristics and functions of eccDNA;
[0006] (3) Existing analysis tools have low analysis efficiency, low reliability of analysis results and high repeatability;
[0007] (4) Researchers often need to use multiple tools in combination or make a large number of parameter adjustments, which not only increases the analysis complexity but may also introduce additional errors. At the same time, using different analysis processes makes it difficult to standardize and compare research results. Summary of the Invention
[0008] In view of the disadvantages of the above-mentioned prior art, the purpose of the present application is to provide an extrachromosomal circular DNA analysis method, device, terminal, medium and product based on long-read sequencing data, so as to solve at least one of the above-mentioned problems in the prior art.
[0009] To achieve the above object and other related objects, the first aspect of the present application provides an extrachromosomal circular DNA analysis method based on long-read sequencing data, including: obtaining long-read sequencing data to be analyzed and establishing a reference genome database; performing tandem repeat sequence detection on the long-read sequencing data to obtain candidate extrachromosomal circular DNA sequence data; filtering and processing the repetitive regions of the candidate extrachromosomal circular DNA sequence data to obtain potential extrachromosomal circular DNA sequence data; based on the reference genome database, performing a first alignment analysis on the potential extrachromosomal circular DNA sequence data to obtain a first extrachromosomal circular DNA analysis result; extracting first unrecognized data from the long-read sequencing data; based on the reference genome database, performing a second alignment analysis on the first unrecognized data to obtain a second extrachromosomal circular DNA analysis result; extracting second unrecognized data from the long-read sequencing data; based on the reference genome database, performing a third alignment analysis on the second unrecognized data to obtain a third extrachromosomal circular DNA analysis result; performing a combined classification process on the first extrachromosomal circular DNA analysis result, the second extrachromosomal circular DNA analysis result and the third extrachromosomal circular DNA analysis result to obtain a final extrachromosomal circular DNA analysis result.
[0010] In some embodiments of the first aspect of the present application, the specific process of filtering and processing the repetitive regions includes: sequentially performing integerization processing and average matching rate filtering processing on the candidate extrachromosomal circular DNA sequence data to obtain filtered extrachromosomal circular DNA sequence data; sequentially performing repetitive region merging processing and circularization processing on the filtered extrachromosomal circular DNA sequence data to obtain potential extrachromosomal circular DNA sequence data.
[0011] In some embodiments of the first aspect of the present application, based on the reference genome database, a first alignment analysis is performed on potential extrachromosomal circular DNA sequence data to obtain a first extrachromosomal circular DNA analysis result, including: aligning the potential extrachromosomal circular DNA sequence data with the reference gene database to obtain a first alignment result; performing circular DNA analysis on the first alignment result according to the first circular DNA analysis rule to obtain a first extrachromosomal circular DNA analysis result; wherein, the specific process includes: determining the first alignment parameter and the alignment strand parameter of each read in the first alignment result; and wherein, the first alignment parameter includes: alignment region length, identity length, gap from the target sequence, gap from the target sequence, and the number of alignment sites; the alignment strand parameter includes: the number of alignment strands, the length of the alignment strand, and the total length of the alignment strands; according to the first circular DNA analysis rule, based on the first alignment parameter and the alignment strand parameter of each read, circular DNA classification is performed on each read to obtain the circular DNA classification result of each read, and further obtain the first extrachromosomal circular DNA analysis result.
[0012] In some embodiments of the first aspect of the present application, the first circular DNA analysis rule includes: determining the category of a read that meets any one of the following UeccDNA classification conditions as UeccDNA: The first UeccDNA classification condition: the number of alignment sites of the read is equal to 1; the gap of the read from the target sequence is less than the set gap threshold, or the percentage of the gap of the read from the target sequence is less than the set percentage gap threshold; The second UeccDNA classification condition: the number of alignment sites of the read is greater than 1; the distance between the start positions or the end positions of every two adjacent alignment sites of the read is less than the set distance threshold; determining the category of a read that meets the MeccDNA classification condition as MeccDNA; wherein, the MeccDNA classification condition includes: the number of alignment sites of the read is multiple; the distance between the alignment sites of the read is greater than the set distance threshold; the gap of the read from the target sequence is less than the set gap threshold, or the percentage of the gap of the read from the target sequence is less than the set percentage gap threshold; determining the category of a read that meets the CeccDNA classification condition as CeccDNA; wherein, the CeccDNA classification condition includes: the number of alignment strands of the read is not zero; the length of the alignment strand of the read is greater than the identity length of the read; the number of alignment sites of the read is at least two; the ratio between the total length of the alignment strands of the read and the identity length of the read is within the set ratio range.
[0013] In some embodiments of the first aspect of the present application, based on the reference genome database, a second alignment and classification analysis is performed on the first unrecognized data, and the obtained extrachromosomal circular DNA analysis results include: aligning the first unrecognized data with the reference genome database to obtain a second alignment result; determining the second alignment parameters of each read in the second alignment result; based on the second alignment parameters of each read, performing a linear read recognition analysis and a circular DNA recognition analysis on the first unrecognized data to obtain a linear read recognition result and a circular DNA recognition result; calculating the recognition result parameters of each read in the circular DNA recognition result, and based on the recognition result parameters of each read, classifying the circular DNA recognition result according to the second circular DNA analysis rule to obtain the extrachromosomal circular DNA analysis results.
[0014] In some embodiments of the first aspect of the present application, second unrecognized data is extracted from the long-read sequencing data; based on the reference genome database, a third alignment and classification analysis is performed on the second unrecognized data to obtain third extrachromosomal circular DNA analysis results, including: using Minimap2 to align the second unrecognized data with the reference genome database to obtain a third alignment result, and sorting the third alignment result; identifying abnormal coverage regions according to the coverage map of the generated sorted result, and merging the identified abnormal coverage regions; processing the merged third alignment result by adopting an inference method based on split reads to obtain the third extrachromosomal circular DNA analysis results.
[0015] In some embodiments of the first aspect of the present application, processing the merged third alignment result by adopting an inference method based on split reads to obtain the third extrachromosomal circular DNA analysis results, including: outputting the reads that meet the split read recognition conditions in the merged third alignment result as the third extrachromosomal circular DNA analysis results; wherein, the split read recognition conditions include: the read contains a soft clipping operation and the read spans both ends of the target region.
[0016] To achieve the above object and other related objects, a second aspect of the present application provides an extrachromosomal circular DNA analysis device based on long-read sequencing data, including an acquisition module for acquiring long-read sequencing data to be analyzed and establishing a reference genome database; a first alignment and analysis module for identifying candidate extrachromosomal circular DNA sequence data from the long-read sequencing data through tandem repeat sequence detection, filtering and processing the repeated regions thereof to obtain potential extrachromosomal circular DNA sequence data; performing a first type of alignment and analysis on the potential extrachromosomal circular DNA sequence data based on the reference genome database to obtain a first extrachromosomal circular DNA analysis result; a second alignment and analysis module for extracting first unrecognized data from the long-read sequencing data; performing a second type of alignment and analysis on the first unrecognized data based on the reference genome database to obtain a second extrachromosomal circular DNA analysis result; a third alignment and analysis module for extracting second unrecognized data from the long-read sequencing data; performing a third type of alignment and analysis on the second unrecognized data based on the reference genome database to obtain a third extrachromosomal circular DNA analysis result; a merging module for performing a merging and classification process on the first extrachromosomal circular DNA analysis result, the second extrachromosomal circular DNA analysis result, and the third extrachromosomal circular DNA analysis result to obtain a final extrachromosomal circular DNA analysis result.
[0017] To achieve the above object and other related objects, a third aspect of the present application provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the extrachromosomal circular DNA analysis method based on long-read sequencing data is implemented.
[0018] To achieve the above object and other related objects, a fourth aspect of the present application provides a computer program product, and the computer program product includes computer program code, and when the computer program code runs on a computer, the computer is enabled to implement the extrachromosomal circular DNA analysis method based on long-read sequencing data.
[0019] To achieve the above object and other related objects, a fifth aspect of the present application provides an electronic terminal, including a memory, a processor, and a computer program stored on the memory; the processor executes the computer program to implement the extrachromosomal circular DNA analysis method based on long-read sequencing data.
[0020] As described above, the extrachromosomal circular DNA analysis method, device, terminal, medium, and product based on long-read sequencing data of the present application have the following beneficial effects:
[0021] In view of the high accuracy and long read length characteristics of HiFi sequencing data, the present application improves the sensitivity and accuracy of eccDNA detection; the method of the present application can simultaneously detect and classify multiple types of eccDNA, providing comprehensive eccDNA analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 Shown is a schematic flowchart of an extrachromosomal circular DNA analysis method based on long read length sequencing data in an embodiment of the present application.
[0023] Figure 2 Shown is a schematic diagram of the specific process of extrachromosomal circular DNA analysis based on long read length sequencing data in an embodiment of the present application.
[0024] Figure 3 Shown is a schematic block diagram of an extrachromosomal circular DNA analysis device based on long read length sequencing data in an embodiment of the present application.
[0025] Figure 4 Shown is a schematic diagram of the structure of an electronic terminal in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] The following specific examples illustrate the implementation manners of the present application. Those skilled in the art can easily understand other advantages and effects of the present application from the content disclosed in this specification. The present application can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present application. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0027] In the embodiments of the present application, terms such as "first" and "second" are used to distinguish identical or similar items with basically the same functions and effects. For example, the first alignment parameter and the second alignment parameter are only used to distinguish different alignment parameters, and do not limit their order. Those skilled in the art can understand that terms such as "first" and "second" do not limit the quantity and execution order, and "first", "second", etc. do not necessarily mean different.
[0028] It should be noted that in the embodiments of the present application, words such as "exemplary" or "for example" represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.
[0029] In the embodiments of the present application, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B may be singular or plural. The character " / " generally represents an "or" relationship between the front and back associated objects. "At least one (item)" or similar expressions refer to any combination of these items, including any combination of single item (item) or plural items (items). For example, at least one (item) of a, b, or c may represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c may be single or multiple.
[0030] Before further elaborating on the present invention in detail, the nouns and terms involved in the embodiments of the present invention are described. The nouns and terms involved in the embodiments of the present invention are applicable to the following explanations:
[0031] eccDNA can be classified into the following categories:
[0032] (1) U-eccDNA (Unique-locus eccDNA): Only a unique alignment site is found in the genome;
[0033] (2) M-eccDNA (Multi-locus eccDNA): The same eccDNA sequence can be found to match at multiple sites in the genome, but it is not "chimeric";
[0034] (3) C-eccDNA (Chimeric eccDNA): Composed of sequences spliced from multiple different genomic fragments;
[0035] (4) MC-eccDNA (Multiple-locus Chimeric eccDNA): Having both characteristics of "multiple-site matching" and "chimeric";
[0036] (5) X-eccDNA (Unknown eccDNA): An eccDNA type that cannot be determined currently.
[0037] To facilitate the understanding of the embodiments of the present application, first, in combination with Figure 1 Detailed description is as follows. Figure 1 The flowchart of an extrachromosomal circular DNA analysis method based on long-read sequencing data in the embodiments of the present invention is shown. The extrachromosomal circular DNA analysis method based on long-read sequencing data in this embodiment mainly includes the following steps:
[0038] Step S11: Obtain the long-read sequencing data to be analyzed and establish a reference genome database.
[0039] In one embodiment, the long-read sequencing data is the long-read (HiFi reads) data of the PacBio HiFi sequencing platform.
[0040] In one embodiment, the reference genome database is a BLAST nucleic acid database constructed using a reference genome (such as the human reference genome T2T-CHM13v2.0).
[0041] Step S12: Perform tandem repeat detection on the long-read sequencing data to obtain candidate extrachromosomal circular DNA sequence data; filter and process the repeat regions of the candidate extrachromosomal circular DNA sequence data to obtain potential extrachromosomal circular DNA sequence data; based on the reference genome database, perform a first alignment analysis on the potential extrachromosomal circular DNA sequence data to obtain a first extrachromosomal circular DNA analysis result.
[0042] In one embodiment, the TideHunter tool is used to perform tandem repeat detection on the long-read sequencing data to obtain candidate extrachromosomal circular DNA sequence data.
[0043] It should be understood that tandem repeats refer to DNA sequence fragments that are continuously repeated and identical or almost identical in the genome.
[0044] It should be noted that TideHunter is a tool specifically designed to efficiently and sensitively detect tandem repeats from noisy long-read sequence data and perform consensus sequence calling.
[0045] In one embodiment, the candidate extrachromosomal circular DNA sequence data output by the TideHunter tool includes multiple extrachromosomal circular DNA reads; each extrachromosomal circular DNA read includes: read name (readName), tandem repeat ID number (repN), tandem repeat copy number (copyNum), read length (readLen), tandem repeat start coordinate (start), tandem repeat end coordinate (end), consensus sequence length (consLen), average matching rate between each unit sequence and the consensus sequence (aveMatch), specific position of the tandem repeat in the original read (subPos), and consensus sequence (consSeq).
[0046] For example, the output of the TideHunter tool can refer to Table 1 below.
[0047] Table 1. Candidate extrachromosomal circular DNA sequence data
[0048] readName repN copyNum readLen start end conslen aveMatch subPos consSeq read1 rep1 3.2 1000 50 400 100 99.5 50,150,250 ATCG… read2 rep1 2.8 800 30 350 110 99.8 30,140 GCTA…
[0049] In one embodiment, the specific process of filtering and repetitive region processing includes: successively performing integerization processing and average matching rate filtering processing on the candidate extrachromosomal circular DNA sequence data to obtain the filtered extrachromosomal circular DNA sequence data; successively performing repetitive region merging processing and circularization processing on the filtered extrachromosomal circular DNA sequence data to obtain potential extrachromosomal circular DNA sequence data. It should be noted that filtering and repetitive region processing are used to reduce noise and highlight potential circular structures.
[0050] The specific process of filtering and repetitive region processing will be described below:
[0051] Successively perform integerization processing on the candidate extrachromosomal circular DNA sequence data. The integerization processing is to round the copyNum of each extrachromosomal circular DNA read in the candidate extrachromosomal circular DNA sequence data to an integer. It should be noted that the integerization processing facilitates subsequent judgment of MeccDNA or UeccDNA, and is conducive to the combined calculation of repetitive regions, avoiding errors caused by fractional copy numbers.
[0052] Furthermore, perform average matching rate filtering processing on the extrachromosomal circular DNA sequence data after integerization processing. The specific process of the average matching rate filtering processing includes: outputting the reads with an average matching rate greater than 99 (aveMatch>99) in the extrachromosomal circular DNA sequence data after integerization processing as the filtered extrachromosomal circular DNA sequence data.
[0053] Furthermore, successively perform repetitive region merging processing on the filtered extrachromosomal circular DNA sequence data, including: calculating the effective length Effective_Length of each extrachromosomal circular DNA read in the filtered extrachromosomal circular DNA sequence data according to the following formula 1;
[0054] Effective_Length = (end – start + 1) / readLen × 100%; (Formula 1)
[0055] where readLen is the length of the original read, start is the starting coordinate of the tandem repeat, and end is the ending coordinate of the tandem repeat;
[0056] According to the effective length of each extrachromosomal circular DNA read in the filtered extrachromosomal circular DNA sequence data, and according to the merging conditions, the repetitive regions of each extrachromosomal circular DNA read in the filtered extrachromosomal circular DNA sequence data are merged to obtain the extrachromosomal circular DNA sequence data after merging processing.
[0057] The merging conditions include: (1) the difference in consLen between two repetitive regions ≤ 15 bp (bp stands for base pair); the effective length (Effective_Length) of the merged region ≤ 102%. It should be noted that the effective length of the merged region can be calculated with reference to Formula 1, which will not be elaborated here. It should be noted that 15 bp is an exemplary threshold, and those skilled in the art can also set it to other values, and the present invention does not limit this.
[0058] Each read in the extrachromosomal circular DNA sequence data after merging processing includes: read name (readName), ID number of tandem repeat (repN), copy number of tandem repeat (copyNum), start coordinate of tandem repeat (start), end coordinate of tandem repeat (end), length of consensus sequence (consLen).
[0059] For example, read1 has two repetitive regions:
[0060] Region 1: start = 50, end = 400, consLen = 100
[0061] Region 2: start = 450, end = 750, consLen = 98
[0062] These two repetitive regions are merged according to the above merging conditions, and the merging result is shown in Table 2:
[0063] Table 2. Merged reads
[0064] readName repN copyNum start end consLen Effective_Length read1 repM0 6 50 750 100 70%
[0065] Furthermore, according to the first CtcR classification rule, the extrachromosomal circular DNA sequence data after merging processing is classified into four categories to obtain the first CtcR classification result; among them, the first CtcR classification rule includes:
[0066] Determine the category of reads with Effective_Length ≥ 99% as CtcR-perfect; specifically, reads with Effective_Length ≥ 99% indicate that the tandem repeat region almost covers the entire read length and is very close to the circular structure; such reads refer to reads with a very complete repetitive segment, a high average matching rate, and likely to fit well with the circular structure.
[0067] Reads with 70% ≤ Effective_Length < 99% are determined as CtcR-hybrid; specifically, such reads indicate a relatively large tandem repeat portion, but have not reached complete coverage or there are some mismatched regions; such reads have some characteristics of circularization and also contain a certain linear component, similar to the circularization possibility of a "hybrid" state.
[0068] Reads containing multiple repeat regions are determined as CtcR-multiple; specifically, such reads contain multiple relatively independent repeat regions (i.e., there are multiple segments of tandem repeats that can be combined), indicating multiple copies / multiple fragments stacking of the repeat regions; this read may show a multi-segment splicing form due to a large number of repeats and significant amplification, and is more inclined to have a high-copy (multi-ring) potential.
[0069] Reads with a single repN marker are determined as CtcR-inversion; specifically, such reads have a single repN marker but detect reverse (inversion) information, or show a certain inverted phenomenon after merging, and do not belong to the aforementioned perfect, hybrid, multiple categories; it may imply a special situation of local circularization or inversion of repeat units in this read; its circular structure has the characteristic of "inverted repeat / inversion".
[0070] It should be noted that CtcR (CtcReads, Concatemeric tandem copies reads) refers to reads containing tandem repeat sequences, which may be a characteristic of circular DNA.
[0071] Furthermore, circularization processing is performed on the extrachromosomal circular DNA sequence data after merging. The circularization processing of the extrachromosomal circular DNA sequence data after merging includes: performing circularization processing on each read in the extrachromosomal circular DNA sequence data after merging to obtain the circularized sequence data of each read; outputting the circularized sequence data of all reads as potential extrachromosomal circular DNA sequence data. It should be understood that DNA circularization processing refers to the process of processing DNA fragments through specific biochemical methods to make them form a circular structure.
[0072] For example, the sequence of a read in the extrachromosomal circular DNA sequence data after merging is: ATCGATCG (length 8bp); the sequence obtained after circularization processing is ATCGATCGATCGATCG (connected end to end, length 16bp).
[0073] Further, the output format of each read in the potential extrachromosomal circular DNA sequence data is query ID + specific sequence; wherein, the format of the query ID is read name / tandem repeat ID number / consensus sequence length / tandem repeat copy number. For example, one read in the potential extrachromosomal circular DNA sequence data is read1|repM0|100|6ATCGATCGATCGATCG.
[0074] It should be noted that after filtering and processing of the repetitive regions, a main file (in.csv format) containing all the filtered and merged sequence information, a file (in.csv format) containing the first CtcR classification results, and a file (in.fasta format) containing the potential extrachromosomal circular DNA sequence data will be obtained.
[0075] In one embodiment, based on the reference genome database, a first type of alignment analysis is performed on the potential extrachromosomal circular DNA sequence data to obtain a first extrachromosomal circular DNA analysis result, including: aligning the potential extrachromosomal circular DNA sequence data with the reference gene database to obtain a first alignment result; according to the first circular DNA analysis rule, performing circular DNA analysis on the first alignment result to obtain a first extrachromosomal circular DNA analysis result; wherein, the specific process includes: determining the first alignment parameter and the alignment strand parameter of each read in the first alignment result; and wherein, the first alignment parameter includes: alignment region length, identity length, gap from the target sequence, gap from the target sequence, and number of alignment sites; the alignment strand parameter includes: number of alignment strands, alignment strand length, and total alignment strand length; according to the circular DNA analysis rule, based on the first alignment parameter and the alignment strand parameter of each read, performing circular DNA classification on each read to obtain the circular DNA classification result of each read, and further obtaining the first extrachromosomal circular DNA analysis result.
[0076] In one embodiment, the first circular DNA analysis rule includes: The first circular DNA analysis rule includes: determining the category of reads that meet any one of the following UeccDNA classification conditions as UeccDNA: The first UeccDNA classification condition: the number of alignment sites of the reads is equal to 1; the difference between the reads and the target sequence is less than the set difference threshold, or the percentage difference between the reads and the target sequence is less than the set percentage difference threshold; The second UeccDNA classification condition: the number of alignment sites of the reads is greater than 1; the distance between the start positions or the end positions of every two adjacent alignment sites of the reads is less than the set distance threshold; determining the category of reads that meet the MeccDNA classification condition as MeccDNA; wherein, the MeccDNA classification condition includes: the number of alignment sites of the reads is multiple; the distance between the alignment sites of the reads is greater than the set distance threshold; the difference between the reads and the target sequence is less than the set difference threshold, or the percentage difference between the reads and the target sequence is less than the set percentage difference threshold; determining the category of reads that meet the CeccDNA classification condition as CeccDNA; wherein, the CeccDNA classification condition includes: the number of alignment strands of the reads is not zero; the length of the alignment strand of the reads is greater than the length of the identity of the reads; the number of alignment sites of the reads is at least two; the ratio between the total length of the alignment strands of the reads and the length of the identity of the reads is within the set ratio range.
[0077] The specific process of the first comparative analysis will be described below:
[0078] Set parameters for the BLAST tool; wherein, set word_size (minimum match length) to 100, set evalue (expectation threshold) to 1e-50, set perc_identity (sequence similarity threshold) to 99, and set the output format to 6-column table format.
[0079] Furthermore, use the BLAST tool to align the potential extrachromosomal circular DNA sequence data with the reference gene database to obtain the first alignment result. Each read in the first comparison result includes: query ID (query_id), target sequence ID (subject_id), percentage identity of the sequence alignment (identity), length of the alignment region (alignment_length), start site of the alignment region on the query sequence (q_start), end site of the alignment region on the query sequence (q_end), start site of the alignment region on the target sequence (s_start), and end site of the alignment region on the target sequence (Subject id).
[0080] For example, one read in the first comparison result is shown in Table 3:
[0081] Table 3. Reads in the first comparison result
[0082] query_id subject_id identity alignment_length q_start q_end s_start s_end read1|rep1|100|3 chr1 99.8 95 1 95 1000 1095
[0083] Furthermore, the query ID of each read in the first comparison result is parsed to obtain the ID parsing result of each read in the first comparison result.
[0084] For example, a query ID (query_id) is read1|rep1|100|3, and the parsing result of this query ID (query_id) is read name (readName): read1; tandem repeat ID number (repN): rep1; length of consensus sequence (consLen): 100; copy number of tandem repeats (copyNum): 3.
[0085] Furthermore, based on the ID parsing results of each read in the first comparison result, the alignment region length Rlength, the gap length gap_Length, the gap percentage Gap_Percentage with the target sequence, and the number of alignment sites of each read in the first comparison result are calculated with reference to the following Formulas 2, 3, and 4:
[0086] Rlength = s_end - s_start + 1; (Formula 2)
[0087] gap_Length = consLen – Rlength; (Formula 3)
[0088] Gap Percentage = |gap_Length| / consLen × 100%; (Formula 4)
[0089] Furthermore, reads with more than 100 alignment sites in the first comparison result are removed. It should be understood that more than 100 alignment sites usually indicate that the reads are extremely repetitive or noisy with respect to the reference sequence, which has a greater impact on the analysis efficiency and accuracy. The present invention preferably uses this threshold as a limit, but does not exclude the possibility of modifying the threshold in extreme cases.
[0090] Furthermore, the first comparison result obtained after removal is sorted according to q_start to obtain one or more alignment chains; the lengths of the respective alignment chains and the total length of the alignment chains are calculated.
[0091] Further, classify the reads in the first comparison result that meet any of the following UeccDNA classification conditions as UeccDNA:
[0092] Case 1: (1) The read has a single alignment site; (2) Gap_Percentage ≤ 10% or |gap_Length| ≤ 50bp.
[0093] Case 2: (1) The read has multiple alignment sites; (2) The difference in the starting positions of every two adjacent alignment sites is not less than 5bp, or the difference in the starting positions of every two adjacent alignment sites is not less than 5bp.
[0094] Classify the reads in the first comparison result that meet the following MeccDNA classification conditions as MeccDNA:
[0095] The read has multiple alignment sites; the distance between alignment sites is greater than 5bp; Gap_Percentage ≤ 10% or |gap_Length| ≤ 50bp.
[0096] Classify the reads in the first comparison result that meet the following CeccDNA classification conditions as CeccDNA:
[0097] The read has at least two alignment sites; the length of the alignment strand is greater than the length of the consensus sequence; the ratio of the total length of the alignment strands to the length of the consensus sequence (consLen) is between 0.8 and 1.2.
[0098] Further, CeccDNA can be further classified: classify the reads with the number of alignment strands equal to 1 as Cecc; classify the reads with the number of alignment strands greater than 1 as Ceccm. It should be understood that Cecc means that a single chimeric strand can form a loop; Ceccm means the complex of multiple chromosomal fragments, which may be seen in highly complex chimeric structures and is more in line with the amplification characteristics of some tumor cells in reality.
[0099] It should be noted that after the first alignment analysis, CSV files containing UeccDNA classification information, CSV files containing MeccDNA classification information, CSV files containing CeccDNA classification information, FASTA files containing UeccDNA sequences, FASTA files containing MeccDNA sequences, FASTA files containing CeccDNA sequences, and FASTA format record files are obtained.
[0100] It should be noted that UeccDNA (Unique eccDNA): the confirmed eccDNA obtained based on the analysis of concatenated tandem copies reads (CtcReads), and the eccDNA inferred based on split reads; MeccDNA (Multiple alignment eccDNA): the eccDNA with multiple alignment positions in the genome; CeccDNA (Chimeric eccDNA): the chimeric circular DNA composed of DNA fragments from different sources; MCeccDNA (Multiple alignment Chimeric eccDNA): the chimeric circular DNA with multiple alignment positions.
[0101] Step S13: Extract the first unrecognized data from the long-read sequencing data; based on the reference genome database, perform a second alignment analysis on the first unrecognized data to obtain a second extrachromosomal circular DNA analysis result.
[0102] In one embodiment, based on the candidate extrachromosomal circular DNA sequence data, use the seqkit tool to extract the first unrecognized data from the long-read sequencing data. Specifically, the candidate extrachromosomal circular DNA sequence data is the recognized data, and the data other than the recognized data is the first unrecognized data.
[0103] In one embodiment, based on the reference genome database, performing a second alignment analysis on the first unrecognized data to obtain a second extrachromosomal circular DNA analysis result includes: comparing the first unrecognized data with the reference genome database to obtain a second alignment result; determining the second alignment parameters of each read in the second alignment result; based on the second alignment parameters of each read, performing a linear read recognition analysis and a circular DNA recognition analysis on the first unrecognized data to obtain a linear read recognition result and a circular DNA recognition result; calculating the recognition result parameters of each read in the circular DNA recognition result, and based on the recognition result parameters of each read, classifying the circular DNA recognition result according to the second circular DNA analysis rule to obtain a second extrachromosomal circular DNA analysis result.
[0104] The specific process of the second alignment analysis is explained below:
[0105] Use the BLAST tool to compare the first unrecognized data with the reference genome database to obtain a second alignment result.
[0106] Furthermore, the samtools faidx tool is used to create a FASTA index for the first unrecognized data for quick access and processing.
[0107] Furthermore, the second alignment result is merged with the FASTA index to obtain a merged result; the consistency of the tname of the second alignment result is determined, and the repetitive regions are identified (tolerance = 30 Bp); the second alignment parameters of each read in the second alignment result are determined, and the second alignment parameters of each read in the second alignment result are added to the merged result. The second alignment parameters include: forward and reverse strand information (strand), efficiency ratio (eff_Ratio), and read occurrence frequency (frequency). Specifically, according to the position coordinates of the gene on the target sequence, if the start position tstart is less than the end position tend and the sequence direction is consistent with the forward strand, then this gene is on the forward strand; if the start position tstart is greater than the end position tend, or the sequence direction is consistent with the reverse strand, then this gene may be on the reverse strand. The tname usually indicates the name of the target sequence (Target). Checking its consistency helps ensure that all aligned fragments are from the same or the same batch of reference sequences, thereby guaranteeing the effectiveness of circular DNA splicing and analysis. The tolerance is usually used to define the range of variation or difference allowed when identifying repetitive sequences. The efficiency ratio (eff_Ratio) is used to measure the percentage of the aligned length to the read length, roughly reflecting the completeness of the comparison. The read occurrence frequency (frequency) is the number of times the read appears in the second alignment result and is used for subsequent judgment of linear reads and circular DNA reads.
[0108] Furthermore, the reads in the second alignment result with a read occurrence frequency equal to 1 and an efficiency ratio (eff_Ratio) > 98% are output as linear read identification results; the reads in the second alignment result with a read occurrence frequency greater than 2 and less than 100 are screened out, and the screened reads are sorted according to q_start to obtain an alignment chain; the repetitive regions of the alignment chain are merged, and the merged chain is marked (PreCtcRlist_number) to obtain the circular DNA identification result.
[0109] Further, determine the recognition result parameters of the circular DNA recognition result; the recognition result parameters include: circular DNA length eLength, number of repeat units eRepeatNum, match degree MatDegree, and whether there is an inversion. Specifically, the circular DNA length eLength is usually estimated based on the difference between the minimum start position and the maximum end position; the number of repeat units eRepeatNum is determined according to the occurrence of the repeat region; the match degree MatDegree is generally the average identity of each aligned fragment read; whether there is an inversion, 1 indicates the existence of an inversion, and 0 indicates the non-existence of an inversion.
[0110] Further, according to the second CtcR classification rule, classify the circular DNA recognition result into four categories to obtain the second CtcR classification result; among them, the second CtcR classification rule includes:
[0111] Determine the category of the read with a single PreCtcRlist and no inversion (inversion = 0) as CtcR-perfect;
[0112] Determine the category of the read with multiple PreCtcRlists and the difference between the maximum and minimum values of the circular lengths of these PreCtcRlists > 50bp as CtcR-hybrid;
[0113] Determine the category of the read with a merged PreCtcRlist (including the'm' mark) as CtcR-multiple;
[0114] Determine the category of the read with a single PreCtcRlist but with an inversion (inversion = 1) as CtcR-inversion.
[0115] Further, according to the second circular DNA analysis rule, classify the circular DNA recognition result to obtain a preliminary second extrachromosomal circular DNA analysis result. The second circular DNA analysis rule includes:
[0116] Classify the reads that meet the second MeccDNA classification conditions as MeccDNA; among them, the second MeccDNA classification conditions include: the reads have multiple PreCtcRlists; the difference between the maximum and minimum values of the circular DNA lengths eLength of the multiple PreCtcRlists is not greater than 50bp;
[0117] Classify the reads that meet the second UeccDNA classification condition as UeccDNA; wherein, the second UeccDNA classification condition includes: the read has only a single PreCtcRlist, or the PreCtcRlist of the read has an'm' mark.
[0118] Further, output the reads with an eRepeatNum difference of less than 2 in the preliminary extrachromosomal circular DNA analysis result of the second chromosome as the intermediate extrachromosomal circular DNA analysis result of the second chromosome; according to the intermediate extrachromosomal circular DNA analysis result of the second chromosome, perform extraction processing on the long-read sequencing data to be analyzed to obtain the final extrachromosomal circular DNA analysis result of the second chromosome.
[0119] Specifically, according to the qstart and qend of each read classified as UeccDNA in the preliminary extrachromosomal circular DNA analysis result of the second chromosome, extract the corresponding sequence data from the long-read sequencing data to be analyzed to obtain the UeccDNA classification result; according to the qstart and qend of each read classified as MeccDNA in the preliminary extrachromosomal circular DNA analysis result of the second chromosome, extract the corresponding sequence data from the long-read sequencing data to be analyzed to obtain the MeccDNA classification result; the UeccDNA classification result and the MeccDNA classification result constitute the final extrachromosomal circular DNA analysis result of the second chromosome.
[0120] Step S14: Extract the second unrecognized data from the long-read sequencing data; based on the reference genome database, perform a third alignment analysis on the second unrecognized data to obtain the extrachromosomal circular DNA analysis result of the third chromosome.
[0121] It should be noted that the third comparative analysis is used to detect circular DNA that was not recognized in the first two alignment analyses due to low abundance, high chimerism, or other factors.
[0122] In one embodiment, use the seqkit tool to extract the second unrecognized data from the long-read sequencing data.
[0123] In one embodiment, based on the reference genome database, perform a third alignment analysis on the second unrecognized data to obtain the extrachromosomal circular DNA analysis result of the third chromosome, including: using Minimap2, align the second unrecognized data with the reference genome database to obtain a third alignment result (file format is SAM or BAM), and use samtools or an equivalent tool to sort the third alignment result.
[0124] Further, based on the coverage map of the generated sorted results, identify abnormal coverage regions, and merge the identified abnormal coverage regions to obtain a merged third alignment result. Specifically,
[0125] Use statistical methods (such as MAD or Z-score) to evaluate the sequencing depth of each genomic position in the sorted results. If there are consecutive and significantly higher-than-background coverage peaks at the alignment coordinates, and the distance between peak segments is small (e.g., ≤10bp) or there is partial overlap, then these "adjacent regions" are merged into a candidate gain segment. This merging can be determined based on genomic coordinate thresholds (such as "gap threshold") or overlap ratios to decide whether two coverage peaks should be regarded as continuous. Similarly, if the overall coverage level is abnormally high in a large chromosomal segment, it can also be regarded as a whole as an amplified or potential circular fragment for subsequent analysis. Through this statistics of the coverage map and the merging of adjacent regions, low-frequency amplifications or multi-copy chimeric structures that are difficult to detect in the previous two alignment analyses can be identified.
[0126] Further, by adopting an inference method based on split reads, process the merged third alignment result to obtain a third extrachromosomal circular DNA analysis result.
[0127] In one embodiment, use the Minimap2 tool to align the second unrecognized data with the reference genome database to obtain a third alignment result. Then use the samtools tool to sort the third alignment result.
[0128] In one embodiment, by adopting an inference method based on split reads, process the merged third alignment result to obtain a third extrachromosomal circular DNA analysis result, including: outputting the reads that meet the split read recognition conditions in the merged third alignment result as the third extrachromosomal circular DNA analysis result; wherein, the split read recognition conditions include: the read contains a soft clipping operation and the read spans both ends of the target region.
[0129] It should be understood that the Minimap2 tool outputs the third alignment result in the form of a CIGAR string. The CIGAR string is a compact format used to describe the alignment results of DNA, RNA, or protein sequences in bioinformatics. In genomics, the CIGAR string details the alignment details between reads (query sequences) and reference sequences, including operations such as matches, mismatches, insertions, deletions, etc. The following are the meanings of the individual marker characters in the CIGAR string: M: Match or Mismatch; I: Insertion; D: Deletion; N: Skipped region from the reference; S: Soft clip; H: Hard clip; P: Padding; =: Perfect match (Match, where the bases must match the reference); X: Mismatch (Mismatch, where the bases may not match the reference).
[0130] Step S15: Combine and classify the first extrachromosomal circular DNA analysis result, the second extrachromosomal circular DNA analysis result, and the third extrachromosomal circular DNA analysis result to obtain the final extrachromosomal circular DNA analysis result.
[0131] Specifically, as Figure 2 shown, combine the classification results of UeccDNA in the first extrachromosomal circular DNA analysis result and the classification results of UeccDNA in the second extrachromosomal circular DNA analysis result, and remove the duplicate results to obtain the preliminary classification result of UeccDNA; then combine the preliminary classification result of UeccDNA with the third extrachromosomal circular DNA analysis result to obtain the final classification result of UeccDNA.
[0132] Combine the classification results of MeccDNA in the first extrachromosomal circular DNA analysis result and the classification results of MeccDNA in the second extrachromosomal circular DNA analysis result, and remove the duplicate results to obtain the final classification result of MeccDNA. Extract the classification result of CeccDNA from the first extrachromosomal circular DNA analysis result.
[0133] The final classification result of UeccDNA, the classification result of MeccDNA, and the classification result of CeccDNA constitute the final extrachromosomal circular DNA analysis result.
[0134] It should be noted that the method for analyzing extrachromosomal circular DNA based on long-read sequencing data of the present invention improves the sensitivity and accuracy of eccDNA detection in view of the high accuracy and long-read characteristics of HiFi sequencing data; the method of the present invention can simultaneously detect and classify multiple types of eccDNA, providing a comprehensive eccDNA analysis.
[0135] In one embodiment, the method further includes: summarizing all the above analysis results to generate an eccDNA analysis report to facilitate result interpretation and analysis.
[0136] Specifically, the eccDNA analysis report includes: read count statistics, CtcR classification statistics, and eccDNA statistics; among them, the read count statistics include the total number of reads, linear reads (obtained based on the linear recognition results), CtcR classification reads, and other reads (calculating the remaining reads); the CtcR classification statistics include: the quantity and percentage of four types of CtcR; the eccDNA statistics include: the quantity and length distribution of each type of eccDNA.
[0137] In a specific embodiment, the method of the present invention can run in a Python 3 computing environment installed with tools such as TideHunter, BLAST, Minimap2, and samtools. The present invention supports multi-threaded processing, makes full use of the computing power of multi-core CPUs, and speeds up the data analysis. In a 48-thread computing environment, the analysis of 10,000 reads can be completed in a short time. Compared with traditional methods, the processing speed is increased by about 2-3 times, significantly saving computing resources and time costs.
[0138] Figure 3 It is a schematic block diagram of an extrachromosomal circular DNA analysis device based on long-read sequencing data provided by an embodiment of the present application. As Figure 3 shown, the extrachromosomal circular DNA analysis device 3 based on long-read sequencing data includes:
[0139] An acquisition module 31, configured to acquire long-read sequencing data to be analyzed and establish a reference genome database;
[0140] A first alignment and analysis module 32, configured to identify candidate extrachromosomal circular DNA sequence data from the long-read sequencing data through tandem repeat sequence detection, filter and process the repeat regions thereof, and obtain potential extrachromosomal circular DNA sequence data; based on the reference genome database, perform a first type of alignment and analysis on the potential extrachromosomal circular DNA sequence data to obtain a first extrachromosomal circular DNA analysis result;
[0141] The second alignment and analysis module 33 is configured to extract first unrecognized data from the long-read sequencing data; perform a second type of alignment and analysis on the first unrecognized data based on the reference genome database to obtain a second extrachromosomal circular DNA analysis result;
[0142] The third alignment and analysis module 34 is configured to extract second unrecognized data from the long-read sequencing data; perform a third type of alignment and analysis on the second unrecognized data based on the reference genome database to obtain a third extrachromosomal circular DNA analysis result;
[0143] The merging module 35 is configured to perform a merging and classification process on the first extrachromosomal circular DNA analysis result, the second extrachromosomal circular DNA analysis result, and the third extrachromosomal circular DNA analysis result to obtain a final extrachromosomal circular DNA analysis result.
[0144] It should be understood that the specific processes for each module to execute the above corresponding steps have been described in detail in the above method embodiments. For the sake of brevity, they will not be repeated here.
[0145] It should also be understood that the division of modules in the embodiments of the present application is illustrative, merely a logical function division. In actual implementation, there may be other division methods. Additionally, in each embodiment of the present application, the various functional modules may be integrated in one processor, may exist separately physically, or two or more modules may be integrated in one module. The above integrated modules may be implemented in the form of hardware or in the form of software functional modules.
[0146] In one embodiment, the specific process of filtering and duplicate region processing includes: sequentially performing integerization processing and average matching rate filtering processing on the candidate extrachromosomal circular DNA sequence data to obtain filtered extrachromosomal circular DNA sequence data; sequentially performing duplicate region merging processing and circularization processing on the filtered extrachromosomal circular DNA sequence data to obtain potential extrachromosomal circular DNA sequence data.
[0147] In one embodiment, based on the reference genome database, a first alignment analysis is performed on potential extrachromosomal circular DNA sequence data to obtain a first extrachromosomal circular DNA analysis result, including: aligning the potential extrachromosomal circular DNA sequence data with the reference gene database to obtain a first alignment result; performing circular DNA analysis on the first alignment result according to the circular DNA analysis rules to obtain a first extrachromosomal circular DNA analysis result; wherein the specific process includes: determining the first alignment parameters and alignment strand parameters of each read in the first alignment result; and wherein the first alignment parameters include: alignment region length, identity length, gap from the target sequence, gap from the target sequence, and number of alignment sites; the alignment strand parameters include: number of alignment strands, alignment strand length, and total alignment strand length; according to the circular DNA analysis rules, based on the first alignment parameters and alignment strand parameters of each read, circular DNA classification is performed on each read to obtain the circular DNA classification result of each read, and further obtain the first extrachromosomal circular DNA analysis result.
[0148] In one embodiment, the circular DNA analysis rules include: determining the category of reads that meet any one of the following UeccDNA classification conditions as UeccDNA: The first UeccDNA classification condition: the number of alignment sites of the read is equal to 1; the gap from the target sequence of the read is less than the set gap threshold, or the gap percentage from the target sequence of the read is less than the set gap percentage threshold; The second UeccDNA classification condition: the number of alignment sites of the read is greater than 1; the distance between the start positions or the end positions of every two adjacent alignment sites of the read is less than the set distance threshold; determining the category of reads that meet the MeccDNA classification conditions as MeccDNA; wherein the MeccDNA classification conditions include: the number of alignment sites of the read is multiple; the distance between the alignment sites of the read is greater than the set distance threshold; the gap from the target sequence of the read is less than the set gap threshold, or the gap percentage from the target sequence of the read is less than the set gap percentage threshold; determining the category of reads that meet the CeccDNA classification conditions as CeccDNA; wherein the CeccDNA classification conditions include: the number of alignment strands of the read is not zero; the alignment strand length of the read is greater than the identity length of the read; the number of alignment sites of the read is at least two; the ratio between the total alignment strand length of the read and the identity length of the read is within the set ratio range.
[0149] In one embodiment, based on the reference genome database, the second alignment and classification analysis is performed on the first unrecognized data, and the second extrachromosomal circular DNA analysis result is obtained, including: aligning the first unrecognized data with the reference genome database to obtain a second alignment result; determining the second alignment parameters of each read in the second alignment result; based on the second alignment parameters of each read, performing linear read recognition analysis and circular DNA recognition analysis on the first unrecognized data to obtain a linear read recognition result and a circular DNA recognition result; calculating the recognition result parameters of each read in the circular DNA recognition result, and classifying the circular DNA recognition result based on the recognition result parameters of each read to obtain the second extrachromosomal circular DNA analysis result.
[0150] In one embodiment, second unrecognized data is extracted from the long-read sequencing data; based on the reference genome database, the third alignment and classification analysis is performed on the second unrecognized data to obtain a third extrachromosomal circular DNA analysis result, including: using Minimap2 to align the second unrecognized data with the reference genome database to obtain a third alignment result, and sorting the third alignment result; merging adjacent regions of the sorted result according to the coverage map of the generated sorted result; processing the merged third alignment result by adopting an inference method based on split reads to obtain the third extrachromosomal circular DNA analysis result.
[0151] In one embodiment, the third alignment result after sorting is processed by adopting an inference method based on split reads to obtain a third extrachromosomal circular DNA analysis result, including: outputting the reads that meet the split read recognition conditions in the third alignment result after sorting as the third extrachromosomal circular DNA analysis result; wherein, the split read recognition conditions include: the read contains a soft clipping operation and the read spans both ends of the target region.
[0152] Figure 4 It is a schematic block diagram of an electronic terminal provided by an embodiment of the present application. As Figure 4 shown, the electronic terminal includes: at least one processor 401, a memory 402, at least one network interface 403, and a user interface 405. Each component in the device is coupled together through a bus system 404. It can be understood that the bus system 404 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 404 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, in Figure 4 all kinds of buses are labeled as the bus system.
[0153] Among them, the user interface 405 may include a display, a keyboard, a mouse, a trackball, a pointing gun, a key, a button, a touchpad, or a touch screen, etc.
[0154] It can be understood that the memory 402 can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM, Read Only Memory), a programmable read-only memory (PROM, Programmable Read-Only Memory), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM, Static Random Access Memory), synchronous static random access memory (SSRAM, Synchronous Static Random Access Memory). The memory described in the embodiments of the present invention is intended to include but not be limited to these and any other suitable categories of memory.
[0155] The memory 402 in the embodiments of the present invention is used to store various categories of data to support the operation of the electronic terminal 400. Examples of these data include: any executable program for operating on the electronic terminal 400, such as the operating system 4021 and the application program 44022; the operating system 4021 contains various system programs, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks. The application program 4022 can contain various application programs, such as a media player (Media Player), a browser (Browser), etc., for implementing various application services. The method for analyzing extrachromosomal circular DNA 44 based on long-read sequencing data provided by the embodiments of the present invention can be included in the application program 4022.
[0156] The method disclosed in the embodiments of the present invention above can be applied to or implemented by the processor 401. The processor 401 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in hardware or instructions in software form in the processor 401. The above-mentioned processor 401 may be a general-purpose processor, a digital signal processor (DSP, Digital Signal Processor), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 401 can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor 401 may be a microprocessor or any conventional processor, etc. Combining the steps of the accessory optimization method provided by the embodiments of the present invention can be directly embodied as being executed and completed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module may be located in a storage medium, and this storage medium is located in the memory. The processor reads the information in the memory and combines its hardware to complete the steps of the foregoing method.
[0157] In an exemplary embodiment, the electronic terminal 400 may be an application-specific integrated circuit (ASIC, Application Specific Integrated Circuit), a DSP, a programmable logic device (PLD, ProgrammableLogic Device), or a complex programmable logic device (CPLD, Complex Programmable Logic Device) for executing the foregoing method.
[0158] According to the method provided by the embodiments of the present application, the present application also provides a computer program product, which includes: computer program code, when the computer program code runs on a computer, it causes the computer to execute Figure 1 the method for analyzing extrachromosomal circular DNA based on long-read sequencing data in the shown embodiments.
[0159] According to the method provided by the embodiments of the present application, the present application also provides a computer-readable storage medium, which stores program code, when the program code runs on a computer, it causes the computer to execute Figure 1 the method for analyzing extrachromosomal circular DNA based on long-read sequencing data in any one of the shown embodiments.
[0160] As used in this specification, the terms "component", "module", "system", etc. are used to denote computer-related entities, hardware, firmware, combinations of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program, and / or a computer. By way of illustration, both an application running on a computing device and the computing device can be components. One or more components can reside in a process and / or an execution thread, and a component can be located on one computer and / or distributed between two or more computers. In addition, these components can execute from various computer-readable media storing various data structures. A component can communicate, for example, through a signal according to one or more data packets (e.g., data from two components interacting with another component among a local system, a distributed system, and / or a network, e.g., the Internet interacting with other systems through a signal) through local and / or remote processes.
[0161] Those of ordinary skill in the art will appreciate that the various illustrative logical blocks and steps described in connection with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether such functions are implemented in hardware or software depends upon the particular application and design constraints of the technical solution. Skilled artisans may implement the described functions in different ways for each particular application, but such implementation should not be regarded as exceeding the scope of this application.
[0162] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0163] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.
[0164] The units described as separate components may or may not be physically separated, and the components presented as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0165] In addition, in each embodiment of this application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0166] In the above embodiments, the functions of the functional units can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions (programs). When the computer program instructions (programs) are loaded and executed on a computer, the processes or functions according to the embodiments of this application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a high-definition digital video disc (DVD)), or a semiconductor medium (for example, a solid state disk (SSD), etc.).
[0167] If a function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0168] As described above, the above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
[0169] In summary, this application provides a method, device, terminal, medium, and product for analyzing extrachromosomal circular DNA based on long-read sequencing data. After obtaining the long-read sequencing data to be analyzed and establishing a reference genome database, first, tandem repeat sequence detection and the first alignment analysis are sequentially performed on the long-read sequencing data to obtain the first extrachromosomal circular DNA analysis result. Then, the second alignment analysis and the third alignment analysis are performed on the data in the long-read sequencing data that have not been recognized, respectively obtaining the second extrachromosomal circular DNA analysis result and the third extrachromosomal circular DNA analysis result. Finally, the extrachromosomal circular DNA analysis results are merged to obtain the final extrachromosomal circular DNA analysis result. This application improves the accuracy of circular DNA analysis and the utilization rate of long-read data through three alignment analyses. Therefore, this application effectively overcomes various disadvantages in the prior art and has high industrial utilization value.
[0170] The above embodiments are only illustrative of the principles and effects of this application, rather than limiting this application. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of this application. Therefore, all equivalent modifications or changes completed by those with ordinary knowledge in the technical field without departing from the spirit and technical ideas disclosed in this application should still be covered by the claims of this application.
Claims
1. A method for analyzing extrachromosomal circular DNA based on long-read sequencing data, characterized in that, Including: Obtain long-read sequencing data to be analyzed and establish a reference genome database; Perform tandem repeat sequence detection on the long-read sequencing data to obtain candidate extrachromosomal circular DNA sequence data; perform filtering and repeat region processing on the candidate extrachromosomal circular DNA sequence data to obtain potential extrachromosomal circular DNA sequence data; based on the reference genome database, perform a first alignment analysis on the potential extrachromosomal circular DNA sequence data to obtain a first extrachromosomal circular DNA analysis result; Extract first unrecognized data from the long-read sequencing data; Based on the reference genome database, perform a second alignment analysis on the first unrecognized data to obtain a second extrachromosomal circular DNA analysis result; Extract second unrecognized data from the long-read sequencing data; based on the reference genome database, perform a third alignment analysis on the second unrecognized data to obtain a third extrachromosomal circular DNA analysis result; Perform a combined classification process on the first extrachromosomal circular DNA analysis result, the second extrachromosomal circular DNA analysis result, and the third extrachromosomal circular DNA analysis result to obtain a final extrachromosomal circular DNA analysis result.
2. The method for analyzing extrachromosomal circular DNA based on long-read sequencing data according to claim 1, wherein The specific process of filtering and repeat region processing includes: Perform integerization processing and average matching rate filtering processing on the candidate extrachromosomal circular DNA sequence data in sequence to obtain filtered extrachromosomal circular DNA sequence data; Perform repeat region merging processing and circularization processing on the filtered extrachromosomal circular DNA sequence data in sequence to obtain potential extrachromosomal circular DNA sequence data.
3. The method for analyzing extrachromosomal circular DNA based on long-read sequencing data according to claim 2, wherein Based on the reference genome database, perform a first alignment analysis on the potential extrachromosomal circular DNA sequence data to obtain a first extrachromosomal circular DNA analysis result, including: Align the potential extrachromosomal circular DNA sequence data with the reference gene database to obtain a first alignment result; According to the first circular DNA analysis rule, perform circular DNA analysis on the first alignment result to obtain a first extrachromosomal circular DNA analysis result; wherein, the specific process includes: Determine the first alignment parameters and alignment strand parameters of each read in the first alignment result; and wherein, the first alignment parameters include: alignment region length, identity length, gap from the target sequence, gap from the target sequence, and number of alignment sites; the alignment strand parameters include: number of alignment strands, alignment strand length, and total alignment strand length; According to the first circular DNA analysis rule, based on the first alignment parameters and alignment strand parameters of each read, perform circular DNA classification on each read to obtain the circular DNA classification result of each read, and further obtain a first extrachromosomal circular DNA analysis result; Wherein, the first circular DNA analysis rule includes: Determine the category of reads that meet any one of the following UeccDNA classification conditions as UeccDNA: The first UeccDNA classification condition: the number of alignment sites read is equal to 1; the difference between the read and the target sequence is less than the set difference threshold, or the percentage difference between the read and the target sequence is less than the set percentage difference threshold; The second UeccDNA classification condition: the number of alignment sites read is greater than 1; the distance between the start positions or the end positions of every two adjacent alignment sites read is less than the set distance threshold; Determine the category of the reads that meet the MeccDNA classification conditions as MeccDNA; wherein, the MeccDNA classification conditions include: the number of alignment sites read is multiple; the distance between the alignment sites read is greater than the set distance threshold; the difference between the read and the target sequence is less than the set difference threshold, or the percentage difference between the read and the target sequence is less than the set percentage difference threshold; Determine the category of the reads that meet the CeccDNA classification conditions as CeccDNA; wherein, the CeccDNA classification conditions include: the number of alignment strands read is not zero; the length of the alignment strand read is greater than the length of the consensus read; the number of alignment sites read is at least two; the ratio between the total length of the alignment strands read and the length of the consensus read is within the set ratio range.
4. The method for analyzing extrachromosomal circular DNA based on long-read sequencing data according to claim 1, wherein Based on the reference genome database, perform a second alignment classification analysis on the first unrecognized data, and the obtained extrachromosomal circular DNA analysis result includes: Align the first unrecognized data with the reference genome database to obtain a second alignment result; determine the second alignment parameters of each read in the second alignment result; Based on the second alignment parameters of each read, perform a linear read recognition analysis and a circular DNA recognition analysis on the first unrecognized data to obtain a linear read recognition result and a circular DNA recognition result; Calculate the recognition result parameters of each read in the circular DNA recognition result, and based on the recognition result parameters of each read, classify the circular DNA recognition result according to the second circular DNA analysis rule to obtain an extrachromosomal circular DNA analysis result.
5. The method for analyzing extrachromosomal circular DNA based on long-read sequencing data according to claim 1, wherein Extract the second unrecognized data from the long-read sequencing data; Based on the reference genome database, perform a third alignment classification analysis on the second unrecognized data to obtain a third extrachromosomal circular DNA analysis result, including: Use Minimap2 to align the second unrecognized data with the reference genome database to obtain a third alignment result, and sort the third alignment result; according to the coverage map of the generated sorted result, identify abnormal coverage regions, and merge the identified abnormal coverage regions to obtain a merged third alignment result; by adopting an inference method based on split reads, process the merged third alignment result to obtain a third extrachromosomal circular DNA analysis result.
6. The method for analyzing extrachromosomal circular DNA based on long-read sequencing data according to claim 5, wherein By adopting an inference method based on split reads, process the merged third alignment result to obtain a third extrachromosomal circular DNA analysis result, including: Output the reads that meet the split read recognition conditions in the merged third comparison result as the third extrachromosomal circular DNA analysis result; wherein, the split read recognition conditions include: the read contains a soft clipping operation and the read spans both ends of the target region.
7. An extrachromosomal circular DNA analysis device based on long-read sequencing data, characterized in that, including: an acquisition module, configured to acquire long-read sequencing data to be analyzed and establish a reference genome database; a first comparison and analysis module, configured to identify candidate extrachromosomal circular DNA sequence data from the long-read sequencing data through tandem repeat sequence detection, filter and process the repeat regions thereof, and obtain potential extrachromosomal circular DNA sequence data; based on the reference genome database, perform a first comparison and analysis on the potential extrachromosomal circular DNA sequence data to obtain a first extrachromosomal circular DNA analysis result; a second comparison and analysis module, configured to extract first unrecognized data from the long-read sequencing data; based on the reference genome database, perform a second comparison and analysis on the first unrecognized data to obtain a second extrachromosomal circular DNA analysis result; a third comparison and analysis module, configured to extract second unrecognized data from the long-read sequencing data; based on the reference genome database, perform a third comparison and analysis on the second unrecognized data to obtain a third extrachromosomal circular DNA analysis result; a merging module, configured to perform a merging and classification process on the first extrachromosomal circular DNA analysis result, the second extrachromosomal circular DNA analysis result, and the third extrachromosomal circular DNA analysis result to obtain a final extrachromosomal circular DNA analysis result.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the method described in any one of claims 1 to 6 is implemented.
9. A computer program product, characterized in that, The computer program product includes computer program code, and when the computer program code runs on a computer, the computer is caused to implement the method described in any one of claims 1 to 6.
10. An electronic terminal, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the method described in any one of claims 1 to 6.
Citation Information
Patent Citations
Novel third-generation sequencing method
CN114540472A
Single-cell multi-omics parallel sequencing method based on third-generation sequencing platform and application of single-cell multi-omics parallel sequencing method
CN117106873A
Multi-strain whole genome analysis method combining second-generation and third-generation sequencing
CN117126950A
Quality evaluation method, device and equipment for long-read-length gene sequencing data and medium
CN118692565A
METHODS AND COMPOSITIONS FOR DETECTING ecDNA
US20220364182A1