Genome type detection method for mixed library paired read segments and related product
By extracting features and detecting the data source of paired reads from mixed libraries, the problem of distinguishing between genomic and epigenomic data in low-volume tumor samples was solved, achieving efficient and accurate data analysis and detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-09
- Publication Date
- 2026-04-14
AI Technical Summary
Existing technologies cannot effectively distinguish between genomic and epigenomic data in low-volume tumor samples, leading to data splitting errors and low detection efficiency, which makes it difficult to meet the accuracy and efficiency requirements of clinical testing.
By extracting features from the paired reads to be tested, the number of cytosines at non-CG sites is counted and a set of feature parameters is generated. Based on these feature parameters, the data source is detected to distinguish between genomic reads and epigenomic reads.
It enables efficient and accurate differentiation between genomic and epigenomic reads, reduces experimental costs, and improves the parsing efficiency and application value of mixed sequencing data.
Smart Images

Figure CN121862200A_ABST
Abstract
Description
Technical Field
[0001] The embodiments disclosed herein relate to the field of data classification, specifically to a method for detecting genome type of paired reads from a mixed library and related products. Background Technology
[0002] With the rapid development of precision medicine and tumor molecular diagnostic technologies, the combined detection of genomics and epigenomics plays a crucial role in elucidating tumorigenesis mechanisms, formulating targeted therapy regimens, and assessing efficacy and prognosis. Currently, clinical testing typically requires two separate DNA samples: one for genomic testing to obtain genetic information such as gene mutations and fusions, and the other for epigenomic testing to analyze epigenetic characteristics. However, in real-world clinical scenarios, many tumor samples (such as biopsy samples and circulating tumor DNA samples) have extremely low total DNA content, which cannot meet the requirements for separate testing of two samples. Therefore, technologies that achieve "one-time library preparation, one-time hybridization capture, and one-time sequencing" based on a single DNA sample, while simultaneously acquiring genomic and epigenomic data, have become an important research direction in the field of tumor molecular testing.
[0003] However, in practical applications of this technology, genome sequencing data and epigenome sequencing data are often mixed into a single dataset due to experimental procedures. Existing data splitting methods still have many shortcomings: On the one hand, some methods rely on experimental means to achieve splitting, distinguishing data sources by the difference between cytosine and thymine after chemical or enzymatic transformation. However, this method is limited by the number of inserted cytosines, and the efficiency of chemical or enzymatic transformation cannot reach 100%, easily leading to data splitting errors. On the other hand, existing splitting methods either require additional experimental design, increasing procedural complexity, or rely on genome re-alignment, increasing analytical resource consumption, making it difficult to balance splitting efficiency and accuracy. For example, some methods, when performing sequence feature analysis, do not accurately classify cytosine sites, but only generally count the number of cytosines in the entire sequence. They cannot distinguish the state differences between CG sites and non-CG sites, easily confusing unconverted cytosine from the genome with residual unconverted cytosine from the epigenome, further exacerbating the risk of misjudgment in data splitting, and failing to meet the dual requirements of clinical testing for accuracy and efficiency.
[0004] Therefore, it is necessary to propose a method for splitting mixed data to solve at least one of the above-mentioned technical problems. Summary of the Invention
[0005] This disclosure provides a method and related products for detecting genome types of paired reads from mixed libraries. In a first aspect, this disclosure provides a method for detecting genome types of paired reads from mixed libraries, comprising:
[0006] For at least one read of the paired read to be tested, the following feature extraction operations are performed: For each candidate cytosine site in the read, based on the dinucleotide context in which the candidate cytosine site is located, it is determined whether the candidate cytosine site is a non-CG site; the number of all candidate cytosine sites in the read that are determined to be non-CG sites is counted and determined as the number of characteristic bases of the read; and a set of feature parameters of the read is generated based on the number of characteristic bases of the read. The paired read to be tested is a paired read in the mixed sequencing data corresponding to the mixed library.
[0007] Based on the feature parameter set of at least one of the paired reads to be tested, the data source of the paired reads to be tested is detected to obtain read detection results that indicate whether the paired reads to be tested are genomic reads or epigenomic reads.
[0008] In some optional implementations, generating the feature parameter set of the read segment based on the number of characteristic bases in the read segment includes:
[0009] The length of the reading segment is obtained by extracting its length feature.
[0010] Based on the number of characteristic bases and the length of the read segment, the distribution features of the read segment are extracted to obtain the target distribution features of the read segment.
[0011] Based on the number of characteristic bases and the target distribution characteristics of the read segment, a set of characteristic parameters for the read segment is generated.
[0012] In some optional implementations, the step of performing data source detection on the paired reads to be tested based on a set of feature parameters of at least one of the paired reads to obtain a read detection result indicating whether the paired reads to be tested are genomic reads or epigenomic reads includes:
[0013] In response to determining that the number of characteristic bases of at least one read in the paired reads to be tested meets the preset epigenome characteristic base number standard, a read detection result is generated to indicate that the paired reads to be tested are epigenome reads;
[0014] In response to determining that the number of characteristic bases in at least one of the paired reads to be tested meets a preset standard for the number of characteristic bases in a genome, a read detection result is generated to indicate that the paired reads to be tested are genome reads.
[0015] In some optional implementations, the step of performing data source detection on the paired reads to obtain read detection results indicating whether the paired reads are genomic or epigenomic reads based on a set of feature parameters of at least one of the paired reads to be tested further includes:
[0016] The sum of the characteristic bases in the two reads of the paired reads to be tested is determined as the number of double-stranded characteristic bases;
[0017] In response to determining that the number of double-stranded characteristic bases meets the preset epigenome double-stranded characteristic base number standard, a read detection result is generated to indicate that the paired read to be tested is an epigenome read;
[0018] In response to determining that the number of double-stranded characteristic bases meets the preset standard for the number of double-stranded characteristic bases in the genome, a read detection result is generated to indicate that the paired read to be tested is a genomic read.
[0019] In some optional embodiments, the step of extracting distribution features of the read segment based on the number of characteristic bases and the read segment length to obtain the target distribution features of the read segment includes:
[0020] The ratio of the number of characteristic bases in the read segment to the length of the read segment is determined as the percentage of characteristic bases in the read segment.
[0021] The proportion of characteristic bases in the read segment is determined as the target distribution characteristic of the read segment; and
[0022] The step of detecting the data source of the paired reads based on the feature parameter set of at least one read in the paired reads to be tested, and obtaining read detection results indicating whether the paired reads to be tested are genomic reads or epigenomic reads, includes:
[0023] In response to determining that the proportion of characteristic bases in at least one of the paired reads to be tested meets a preset standard for the proportion of characteristic bases in epigenomes, a read detection result is generated to indicate that the paired reads to be tested are epigenome reads.
[0024] In response to determining that the proportion of characteristic bases in at least one of the paired reads to be tested meets a preset standard for the proportion of characteristic bases in the genome, a read detection result is generated to indicate that the paired reads to be tested are genome reads.
[0025] In some optional embodiments, the step of extracting distribution features of the read segment based on the number of characteristic bases and the read segment length to obtain the target distribution features of the read segment includes:
[0026] The epigenomic origin matching probability and genomic origin matching probability of the read segment are determined based on the read length and the number of characteristic bases; and
[0027] The step of detecting the data source of the paired reads based on the feature parameter set of at least one read in the paired reads to be tested, and obtaining read detection results indicating whether the paired reads to be tested are genomic reads or epigenomic reads, includes:
[0028] In response to determining that the epigenomic origin matching probability of at least one of the paired reads to be tested meets a preset epigenomic origin matching probability standard or does not meet the preset genome origin matching probability standard, a read detection result is generated to indicate that the paired reads to be tested are epigenomic reads;
[0029] In response to determining that the genomic origin matching probability of at least one of the paired reads to be tested meets a preset genomic origin matching probability standard or does not meet a preset epigenome origin matching probability standard, a read detection result is generated to indicate that the paired reads to be tested are genomic reads.
[0030] In some optional implementations, the step of performing data source detection on the paired reads to be tested based on a set of feature parameters of at least one of the paired reads to obtain a read detection result indicating whether the paired reads to be tested are genomic reads or epigenomic reads includes:
[0031] The sum of the characteristic bases in the two reads of the paired reads to be tested is determined as the number of double-stranded characteristic bases;
[0032] The sum of the lengths of the two read segments in the paired read segments to be tested is determined as the length of the double-chain read segment;
[0033] Based on the length of the double-stranded read and the number of double-stranded characteristic bases, the overall epigenome origin matching probability and the overall genome origin matching probability of the paired read to be tested are determined;
[0034] In response to determining whether the overall epigenome origin matching probability of the paired read segment to be tested meets the preset overall epigenome origin matching probability standard or does not meet the preset overall genome origin matching probability standard, a read detection result is generated to indicate that the paired read segment to be tested is an epigenome read segment;
[0035] In response to determining whether the overall genome source matching probability of the paired read to be tested meets a preset overall genome source matching probability standard or does not meet the preset overall epigenome source matching probability standard, a read detection result is generated to indicate that the paired read to be tested is a genomic read.
[0036] Secondly, this disclosure provides a hybrid library paired read genome type detection device, comprising:
[0037] The feature extraction module is configured to perform the following feature extraction operations for at least one read of the paired read to be tested: for each candidate cytosine site in the read, determine whether the candidate cytosine site is a non-CG site based on the dinucleotide context in which the candidate cytosine site is located, count the number of all candidate cytosine sites in the read that are determined to be non-CG sites and determine them as the number of characteristic bases of the read, and generate a set of feature parameters of the read based on the number of characteristic bases of the read, wherein the paired read to be tested is a paired read in the mixed sequencing data corresponding to the mixed library;
[0038] The source detection module is configured to perform data source detection on the paired reads to be tested based on a set of feature parameters of at least one of the paired reads to be tested, and obtain a read detection result indicating whether the paired reads to be tested are genomic reads or epigenomic reads.
[0039] In some optional implementations, generating the feature parameter set of the read segment based on the number of characteristic bases in the read segment includes:
[0040] The length of the reading segment is obtained by extracting its length feature.
[0041] Based on the number of characteristic bases and the length of the read segment, the distribution features of the read segment are extracted to obtain the target distribution features of the read segment.
[0042] Based on the number of characteristic bases and the target distribution characteristics of the read segment, a set of characteristic parameters for the read segment is generated.
[0043] In some optional implementations, the source detection module is further configured to:
[0044] In response to determining that the number of characteristic bases of at least one read in the paired reads to be tested meets the preset epigenome characteristic base number standard, a read detection result is generated to indicate that the paired reads to be tested are epigenome reads;
[0045] In response to determining that the number of characteristic bases in at least one of the paired reads to be tested meets a preset standard for the number of characteristic bases in a genome, a read detection result is generated to indicate that the paired reads to be tested are genome reads.
[0046] In some optional implementations, the source detection module is further configured to:
[0047] The sum of the characteristic bases in the two reads of the paired reads to be tested is determined as the number of double-stranded characteristic bases;
[0048] In response to determining that the number of double-stranded characteristic bases meets the preset epigenome double-stranded characteristic base number standard, a read detection result is generated to indicate that the paired read to be tested is an epigenome read;
[0049] In response to determining that the number of double-stranded characteristic bases meets the preset standard for the number of double-stranded characteristic bases in the genome, a read detection result is generated to indicate that the paired read to be tested is a genomic read.
[0050] In some optional embodiments, the step of extracting distribution features of the read segment based on the number of characteristic bases and the read segment length to obtain the target distribution features of the read segment includes:
[0051] The ratio of the number of characteristic bases in the read segment to the length of the read segment is determined as the percentage of characteristic bases in the read segment.
[0052] The proportion of characteristic bases in the read segment is determined as the target distribution characteristic of the read segment; and
[0053] The source detection module is further configured to:
[0054] In response to determining that the proportion of characteristic bases in at least one of the paired reads to be tested meets a preset standard for the proportion of characteristic bases in epigenomes, a read detection result is generated to indicate that the paired reads to be tested are epigenome reads.
[0055] In response to determining that the proportion of characteristic bases in at least one of the paired reads to be tested meets a preset standard for the proportion of characteristic bases in the genome, a read detection result is generated to indicate that the paired reads to be tested are genome reads.
[0056] In some optional embodiments, the step of extracting distribution features of the read segment based on the number of characteristic bases and the read segment length to obtain the target distribution features of the read segment includes:
[0057] The epigenomic origin matching probability and genomic origin matching probability of the read segment are determined based on the read length and the number of characteristic bases; and
[0058] The source detection module is further configured to:
[0059] In response to determining that the epigenomic origin matching probability of at least one of the paired reads to be tested meets a preset epigenomic origin matching probability standard or does not meet the preset genome origin matching probability standard, a read detection result is generated to indicate that the paired reads to be tested are epigenomic reads;
[0060] In response to determining that the genomic origin matching probability of at least one of the paired reads to be tested meets a preset genomic origin matching probability standard or does not meet a preset epigenome origin matching probability standard, a read detection result is generated to indicate that the paired reads to be tested are genomic reads.
[0061] In some optional implementations, the source detection module is further configured to:
[0062] The sum of the characteristic bases in the two reads of the paired reads to be tested is determined as the number of double-stranded characteristic bases;
[0063] The sum of the lengths of the two read segments in the paired read segments to be tested is determined as the length of the double-chain read segment;
[0064] Based on the length of the double-stranded read and the number of double-stranded characteristic bases, the overall epigenome origin matching probability and the overall genome origin matching probability of the paired read to be tested are determined;
[0065] In response to determining whether the overall epigenome origin matching probability of the paired read segment to be tested meets the preset overall epigenome origin matching probability standard or does not meet the preset overall genome origin matching probability standard, a read detection result is generated to indicate that the paired read segment to be tested is an epigenome read segment;
[0066] In response to determining whether the overall genome source matching probability of the paired read to be tested meets a preset overall genome source matching probability standard or does not meet the preset overall epigenome source matching probability standard, a read detection result is generated to indicate that the paired read to be tested is a genomic read.
[0067] Thirdly, this disclosure provides an electronic device, including:
[0068] One or more processors;
[0069] Storage device, on which one or more programs are stored,
[0070] When the above-described one or more programs are executed by the above-described one or more processors, the above-described one or more processors implement the method as described in any embodiment of the first aspect of this disclosure.
[0071] Fourthly, this disclosure provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by one or more processors, implements the method described in any embodiment of the first aspect of this disclosure.
[0072] Fifthly, this disclosure provides a computer program product including a computer program / instructions that, when executed by a processor, implement the method described in any embodiment of the first aspect of this disclosure.
[0073] The hybrid library paired read genome type detection method, apparatus, electronic device, and storage medium provided in the embodiments of this disclosure first perform the following feature extraction operations on at least one read of the paired read to be tested: for each candidate cytosine site in the read, based on the dinucleotide context in which the candidate cytosine site is located, it is determined whether the candidate cytosine site is a non-CG site; the number of all candidate cytosine sites in the read determined to be non-CG sites is counted and determined as the number of characteristic bases of the read; and a set of feature parameters of the read is generated based on the number of characteristic bases of the read. Then, based on the set of feature parameters of at least one read of the paired read to be tested, the data source of the paired read to be tested is detected to obtain a read detection result used to indicate whether the paired read to be tested is a genomic read or an epigenomic read. In this way, by extracting feature parameters from the reads to be tested for source detection, efficient and accurate differentiation between genomic and epigenomic reads is achieved. This effectively avoids the limitations of single feature judgment and the risk of misjudgment caused by individual differences in reads. Thus, while reducing experimental costs and sample consumption, the parsing efficiency and application value of mixed sequencing data are significantly improved. Attached Figure Description
[0074] Other features, objects, and advantages of this disclosure will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings. The drawings are for illustrative purposes only and are not intended to limit the scope of this application. In the drawings:
[0075] Figure 1 This is an exemplary system architecture diagram to which one embodiment of this disclosure can be applied;
[0076] Figure 2A This is a flowchart of an embodiment of the hybrid library paired read genome type detection method according to the present disclosure;
[0077] Figure 2B This is a breakdown flowchart of one embodiment of the feature extraction operation according to the present disclosure;
[0078] Figure 2C This is an exploded flowchart of one embodiment of step 2013 according to the present disclosure;
[0079] Figure 2D This is an exploded flowchart of one embodiment of step 202 of the present disclosure;
[0080] Figure 3This is a correlation analysis graph showing the correlation between the abundance of detected variants and the abundance of variants at the gold standard in genomic variant detection according to one embodiment of this disclosure;
[0081] Figure 4 This is a graph showing the correlation between the abundance detected and the gold standard abundance in a qualitative detection of the presence of tumor DNA fragments in the epigenome according to one embodiment of this disclosure.
[0082] Figure 5 This is a schematic diagram of one embodiment of the hybrid library paired read genome type detection device according to the present disclosure;
[0083] Figure 6 A schematic diagram of the structure of a computer system suitable for implementing embodiments of the present disclosure. Detailed Implementation
[0084] The present disclosure will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings.
[0085] It should be noted that, unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other. This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.
[0086] Figure 1 An exemplary system architecture 100 is shown, in which embodiments of the hybrid library paired read genome type detection methods, apparatuses, electronic devices and storage media of this disclosure can be applied.
[0087] like Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, and 103, a network 104, and a server 105. Network 104 serves as the medium for providing communication links between terminal devices 101, 102, and 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.
[0088] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as mixed library paired read genome type detection applications, voice interaction applications, video conferencing applications, short video social applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.
[0089] Terminal devices 101, 102, and 103 can be either hardware or software. When terminal devices 101, 102, and 103 are hardware, they can be various electronic devices with microphones and speakers, including but not limited to smartphones, tablets, e-book readers, MP3 players (Moving Picture Experts Group Audio Layer III), MP4 players (Moving Picture Experts Group Audio Layer IV), portable computers, and desktop computers, etc. When terminal devices 101, 102, and 103 are software, they can be installed in the aforementioned electronic devices. They can be implemented as multiple software programs or software modules, or as a single software program or software module. No specific limitations are made here.
[0090] Server 105 can be a server that provides various services, such as a backend server that processes the paired read segments to be tested obtained from terminal devices 101, 102, and 103. The backend server can perform corresponding processing based on the set of feature parameters obtained from the terminal devices.
[0091] In some cases, the mixed library paired read genome type detection method provided in this disclosure can be jointly executed by terminal devices 101, 102, 103 and server 105. For example, the step of "performing feature extraction operation for at least one read of the paired read to be tested" can be executed by terminal devices 101, 102, 103, and the step of "based on the feature parameter set of at least one read in the paired read to be tested, performing data source detection on the paired read to be tested, and obtaining read detection results used to indicate whether the paired read to be tested is a genomic read or an epigenomic read" can be executed by server 105. This disclosure does not limit this. Correspondingly, the mixed library paired read genome type detection device can also be respectively set in terminal devices 101, 102, 103 and server 105.
[0092] In some cases, the mixed library paired read genome type detection method provided in this disclosure can be executed by server 105. Accordingly, the mixed library paired read genome type detection device can also be set in server 105. In this case, the system architecture 100 may not include terminal devices 101, 102, and 103.
[0093] In some cases, the hybrid library paired read genome type detection method provided in this disclosure can be executed by terminal devices 101, 102, and 103. Correspondingly, the hybrid library paired read genome type detection device can also be set in terminal devices 101, 102, and 103. In this case, the system architecture 100 may not include server 105.
[0094] It should be noted that server 105 can be either hardware or software. When server 105 is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. When server 105 is software, it can be implemented as multiple software programs or software modules (e.g., used to provide distributed services), or as a single software program or software module. No specific limitations are made here.
[0095] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0096] Continue to refer to Figure 2A , Figure 2A A flowchart 200 is shown as an embodiment of the hybrid library paired read genome type detection method according to this disclosure. Figure 2A The hybrid library paired read genome type detection method shown can be applied to Figure 1 The terminal device or server shown. The process 200 includes the following steps:
[0097] Step 201: Perform feature extraction on at least one read segment of the paired read segment to be tested.
[0098] In this embodiment, the execution entity of the mixed library paired read genome type detection method (e.g.) Figure 1 Server 105 can first acquire the paired reads to be tested. Here, the paired reads to be tested are the paired reads in the mixed sequencing data of the mixed library. This mixed sequencing data is generated by sequencing the mixed library, which is obtained by mixing the epigenome library and the genomic library. It is a library combination used to simultaneously acquire the genomic sequencing data and the epigenome sequencing data of the sample to be tested. The mixed sequencing data refers to the collection of all the original sequencing data from the genomic library and the epigenome library output by the sequencer after high-throughput sequencing of the obtained mixed library.
[0099] The epigenome library is obtained by performing targeted transformation on one copy of the sample library from the test sample, while the genomic library is obtained by not performing targeted transformation on the other copy. The paired reads to be tested can be any paired read from the mixed sequencing data. These paired reads serve as the basic units whose origins need to be distinguished. Through subsequent feature parameter extraction and data source detection, it is determined whether the entire read belongs to the genomic or epigenome sequence, thereby achieving precise separation of the mixed sequencing data and providing a reliable data foundation for subsequent independent analysis of the genomic and epigenome sequences.
[0100] Here, the mixed library and mixed sequencing data are predetermined through steps one through three:
[0101] The first step involves performing targeted transformation on one sample from the test sample to obtain an epigenome library, while the other sample is not subjected to targeted transformation to obtain a genome library.
[0102] Here, sample library construction refers to the process of transforming raw DNA from test samples (such as tumor tissue, blood, body fluids, etc.) into a library format suitable for high-throughput sequencing through a series of biochemical treatments. The sample library typically consists of a collection of DNA fragments containing specific adapter sequences. The core purpose of sample library construction is to fragment, repair ends, and ligate adapters from trace amounts of dispersed raw DNA, and then enrich the target fragments through steps such as PCR amplification. This ensures that the quality and quantity requirements of the template DNA for sequencing technology are met, providing a suitable template for subsequent mixed sequencing.
[0103] Targeted transformation refers to the process of selectively modifying or transforming specific bases in the DNA molecule using chemical or enzymatic methods (such as bisulfite treatment, enzymatic methylation conversion, etc.) for one sample library construction. For example, bisulfite treatment can ultimately yield an epigenomic library. The other sample library is constructed without the above targeted transformation process, directly retaining the base state of the original DNA, and is used as a genomic library to reflect the sequence information of the original genome.
[0104] The second step involves mixing the epigenome library and the genome library to obtain a mixed library.
[0105] As an example, mixing epigenomic and genomic libraries is typically performed under laboratory conditions. First, the concentrations and volumes of the epigenomic and genomic libraries are precisely measured to ensure they are mixed in a predetermined ratio (e.g., 1:1 or adjusted according to experimental needs). Then, the epigenomic and genomic libraries are transferred to the same sterile centrifuge tube and gently mixed using a pipette, avoiding vigorous shaking that could cause DNA fragmentation. Finally, a brief centrifugation can be performed to collect the library mixture at the bottom of the tube, resulting in a homogenized mixed library. The mixing process requires strict temperature control and avoidance of contamination to ensure complete fusion and preservation of biological activity of the two libraries.
[0106] The third step involves sequencing based on the mixed library to generate mixed sequencing data.
[0107] As an example, when generating mixed sequencing data based on a mixed library, the first step is to denature the constructed mixed library, causing the DNA double strands to unwind into single strands. Next, the denatured single-stranded DNA library is bound to primers on a sequencing flow cell, and in situ amplification is performed using bridge PCR or similar techniques, forming numerous DNA clusters, each containing multiple copies of the same original DNA fragment. Then, using the sequencer's optical detection system, fluorescently labeled dNTPs are added during DNA synthesis to record the base types incorporated into each DNA cluster in each cycle, thereby reading the base sequence of each DNA fragment. Finally, the sequencer converts the detected fluorescence signals into corresponding base sequence data and adds quality information to generate the original mixed sequencing data. The entire process requires strict control of experimental conditions, such as temperature, humidity, and reagent concentration, to ensure the accuracy and efficiency of sequencing. It should be noted that this method for feature extraction of paired reads can be applied to other paired reads without specific limitations.
[0108] Specifically, feature extraction operations can include, for example: Figure 2B The following steps 2011 to 2013 are shown:
[0109] Step 2011: For each candidate cytosine site in the read segment, determine whether the candidate cytosine site is a non-CG site based on the dinucleotide context in which the candidate cytosine site is located.
[0110] Here, the two reads in the paired reads to be tested refer to the first and second reads that constitute the paired reads. When sequencing a DNA fragment, the sequencer reads a sequence from one end of the fragment to obtain the first read; then, the sequencer reads the sequence from the other end of the DNA fragment (usually after special treatment, such as adapter ligation and fragment inversion) to obtain the second read. Therefore, the first and second reads are two short sequences from opposite ends of the same original DNA template molecule, with known orientation and spacing information. They are physically derived from the same DNA fragment, so in a mixed library, their origin is consistent, but in sequencing data, they are recorded as a pair of associated reads. Theoretically, their sequence contents are inversely complementary, or in other words, they jointly represent the sequence information of the original DNA fragment.
[0111] Here, if the read is the first read, the candidate cytosine site can be the specific location of the cytosine (C) base in the first read sequence, which is the core characteristic site for determining the read origin. Specifically, the CG site refers to the cytosine (C) site immediately downstream of guanine (G), i.e., the location of cytosine (C) with "CG" immediately adjacent in the base sequence; the non-CG site refers to the cytosine (C) site immediately downstream of adenine (A), thymine (T), or cytosine (C), i.e., the location of cytosine (C) with "CA," "CT," or "CC" immediately adjacent in the base sequence. By classifying the cytosine (C) sites in the first read into CG and non-CG sites, the dinucleotide context features of the cytosine (C) site can be clearly distinguished.
[0112] Accordingly, when the read segment is the first read segment, the determination of whether the candidate cytosine site is a non-CG site based on the dinucleotide context of the candidate cytosine site can be performed as follows: traverse the complete base sequence of the first read segment, identify all positions of the base type cytosine (C) to determine the cytosine (C) site, and then for each identified cytosine (C) site, further detect the type of its downstream adjacent base: if the downstream adjacent base of the cytosine (C) is guanine (G), that is, the base sequence forms a "CG" adjacent arrangement, then the cytosine site is determined to be a CG site; if the downstream adjacent base of the cytosine (C) is adenine (A), thymine (T) or cytosine (C), that is, the base sequence forms a "CA", "CT", or "CC" adjacent arrangement, then the cytosine (C) site is determined to be a non-CG site, thereby completing the classification of CG and non-CG sites among the candidate cytosine sites in the first read segment.
[0113] Here, if the read is the second read, the candidate cytosine site can be the specific location of the guanine (G) base in the second read sequence, which is also the core characteristic site for determining the read origin. Correspondingly, the CG site specifically refers to the guanine (G) site whose upstream base is cytosine (C), that is, the location of guanine (G) with "CG" arranged in adjacent positions in the base sequence; the non-CG site specifically refers to the guanine (G) site whose upstream base is adenine (A), thymine (T), or guanine (G), that is, the location of guanine (G) with "AG", "TG", or "GG" arranged in adjacent positions in the base sequence. By classifying the guanine (G) sites in the second read into CG sites and non-CG sites, the dinucleotide context features of the guanine (G) site can be clearly distinguished.
[0114] Accordingly, when the read is the second read, the determination of whether the candidate cytosine site is a non-CG site is based on the dinucleotide context in which the candidate cytosine site is located can be performed as follows: Traverse the complete base sequence of the second read, identify all positions of the base type guanine (G) one by one to determine the guanine (G) site, and then for each identified guanine (G) site, further detect the type of its upstream adjacent base: if the upstream adjacent base of the guanine (G) is cytosine (C), that is, the base sequence forms a "CG" adjacent arrangement, then the guanine (G) site is determined to be a CG site; if the upstream adjacent base of the guanine (G) is adenine (A), thymine (T) or guanine (G), that is, the base sequence forms an "AG", "TG" or "GG" adjacent arrangement, then the guanine (G) site is determined to be a non-CG site, thereby completing the classification of CG and non-CG sites among the candidate cytosine sites in the second read.
[0115] Step 2012: Count the number of candidate cytosine sites that are determined to be non-CG sites in the read segment and determine them as the number of characteristic bases of the read segment.
[0116] That is, firstly, the candidate cytosine sites in the non-CG sites of the read segment are counted, and then the number of candidate cytosine sites in the non-CG sites of the read segment is determined as the number of characteristic bases of the read segment.
[0117] Specifically, for the first read, the number of non-CG sites containing C can be counted as the characteristic base count of the first read. That is, for each C base in the first read, if its downstream (right) adjacent base is a G base, then the C base belongs to the "CG" context and is not counted. Conversely, if its downstream is an A, T, or C base, then the C base belongs to the "non-CG" context, and the characteristic base count of the first read is incremented by 1. Assuming r1C represents the characteristic base count of the first read, it can be understood that if the r1C value is high, it indicates that the first read contains many Cs that should have been converted but were not, and the first read is likely from a genomic library; if the r1C value of the first read is low (close to 0), it indicates that most of the non-CG Cs in the first read have been successfully converted to T, and the first read is likely from an epigenome library.
[0118] Specifically, for the second read segment, the number of G bases at non-CG sites can be counted as the characteristic base count of the second read segment. That is, for each G base in the second read segment, if its immediate upstream neighbor is a C base, then the G base belongs to the "CG" context and is not counted. Conversely, if the upstream neighbor of the G base is an A, T, or C base, then the G base belongs to the "non-CG" context, and the characteristic base count of the second read segment is incremented by 1. Assuming we use r2G to represent the number of characteristic bases in the second read, it's understandable that a high r2G value indicates that the complementary strand of the second read (i.e., the first read) contains many C atoms that should have been converted but were not. Consequently, both the first and second reads are likely from a genomic library. Conversely, a low r2G value (close to 0) indicates that most of the non-CG C atoms in the complementary strand of the second read (i.e., the first read) have been successfully converted to T atoms. Similarly, both the first and second reads are likely from an epigenome library.
[0119] By statistically analyzing the number of characteristic bases at non-CG sites, the differences in base state resulting from whether a read has undergone targeted transformation can be accurately captured, providing crucial quantitative evidence for distinguishing whether the paired read to be tested originates from a genomic library or an epigenomic library.
[0120] Step 2013: Generate a set of characteristic parameters for the read segment based on the number of characteristic bases in the read segment.
[0121] Here, various implementation methods can be used to generate a set of characteristic parameters for the read segment based on the number of characteristic bases in the read segment.
[0122] The feature parameter set of the paired read segments to be tested carries core information that can distinguish the source of the read segments. This provides comprehensive and systematic feature parameter support for data source detection of the paired read segments to be tested, ensuring the accuracy and reliability of source determination. In some optional implementations, step 2013 may include, for example... Figure 2C The following steps 20131 to 20133 are shown:
[0123] Step 20131: Extract the length feature of the reading segment to obtain the reading segment length.
[0124] Here, length feature extraction of the read segment refers to the process of accurately calculating and obtaining the length (i.e., the number of base pairs contained) of the first and second read segments that constitute the paired read segment to be tested, using bioinformatics methods and tools.
[0125] The main purpose of length feature extraction is:
[0126] (1) Obtain a basic, quantifiable feature parameter for subsequent data source detection;
[0127] (2) During the construction of libraries from different sources (such as genomic libraries and epigenomic libraries), factors such as enzyme digestion efficiency, fragment selection preferences, and PCR amplification bias may lead to subtle but distinguishable differences in the final read length distribution obtained from sequencing. For example, certain epigenetic modifications may affect the physical structure of DNA, making it easier to generate fragments of specific lengths during fragmentation. By extracting the read length feature, an auxiliary basis can be provided for subsequent determination of the read's origin.
[0128] As an example, when extracting length features from two reads in a pair of reads to be tested, sequencing data quality control and preprocessing tools (such as Trimmomatic, Cutadapt, FastQC, etc.) are first used to preprocess the paired reads, including removing sequencing adapter sequences, filtering out low-quality bases, and filtering out potentially contaminating sequences. After preprocessing, bioinformatics tools (such as those in Samtools, Picard Tools, or simple Linux commands like wc -l combined with awk to process FASTQ files) are used to traverse each read and count the total number of base identifiers it contains after quality trimming. This total number of base identifiers is the read length of that read.
[0129] Read length refers to the specific number of base markers contained in the first or second read after the aforementioned preprocessing and length measurement. For example, one read might be 150 bp long, and another might be 148 bp long. Under good read quality control and without significant length filtering, the read lengths of paired reads within the same batch are usually relatively consistent, but there may be minor length variations due to sequencing cycle termination or quality issues. In this application, the read lengths of the first and second reads are treated as independent feature parameters, collectively forming part of the set of feature parameters to be measured, thereby achieving effective splitting of mixed data.
[0130] By extracting length features from the two reads that make up the paired reads to be tested, the physical length of each read is accurately quantified. This captures the length feature differences that may arise between genomic and epigenomic reads due to differences in library fragmentation, thus supplementing the set of feature parameters to be tested with key information on the length dimension. This provides a foundation for subsequent data source detection based on multi-dimensional features and helps improve the comprehensiveness and accuracy of genome type detection for paired reads in mixed libraries.
[0131] Step 20132: Based on the number of characteristic bases and the length of the read segment, the distribution features of the read segment are extracted to obtain the target distribution features of the read segment.
[0132] Here, the target distribution features of a read segment are used to characterize the distribution of characteristic bases within that read segment. That is, by accurately counting the number of characteristic bases in the read segment and combining this with the distribution patterns of those characteristic bases, the distribution characteristics can be determined. This enables in-depth mining of the core differences between genomic and epigenomic read segments, clarifying the quantitative information of characteristic bases. This provides a richer and more discriminative basis for subsequent data source detection based on feature parameter sets, significantly improving the accuracy and reliability of genome type detection for paired reads from mixed libraries.
[0133] Specifically, when the read segment is the first read segment, the characteristic base in the read segment can refer to the C base in the non-CG site of the first read segment. When the read segment is the second read segment, the characteristic base in the read segment can refer to the G base in the non-CG site of the second read segment.
[0134] Step 20133: Based on the number of characteristic bases and the target distribution characteristics of the read segment, generate a set of characteristic parameters for the read segment.
[0135] Here, various implementation methods can be used to generate a set of feature parameters for the read segment based on the number of characteristic bases and the target distribution characteristics, according to actual needs. For example, the set of feature parameters for the read segment can be generated using the number of characteristic bases and the target distribution characteristics. Alternatively, new feature parameters can be generated based on the number of characteristic bases and the target distribution characteristics of the read segment, and then the set of feature parameters for the read segment can be generated using the number of characteristic bases, the target distribution characteristics, and the newly generated feature parameters.
[0136] Step 202: Based on the feature parameter set of at least one read in the paired reads to be tested, perform data source detection on the paired reads to be tested to obtain read detection results that indicate whether the paired reads to be tested are genomic reads or epigenomic reads.
[0137] In this embodiment, technical personnel can pre-define and store corresponding data source detection rules based on the parameter categories in the feature parameters of at least one read in the specific paired reads to be tested. Accordingly, during the execution of step 202, the corresponding data source detection rules can be obtained based on the parameter categories in the feature parameter set of at least one read in the paired reads to be tested. Then, based on the obtained data source detection rules, data source detection is performed on the paired reads to be tested to obtain read detection results that indicate whether the paired reads to be tested are genomic reads or epigenomic reads.
[0138] In some alternative implementations, step 202 may include, for example: Figure 2D The following steps 20201 to 20202 are shown:
[0139] Step 20201: In response to determining that the number of characteristic bases of at least one read in the paired reads to be tested meets the preset epigenomic characteristic base number standard, a read detection result is generated to indicate that the paired reads to be tested are epigenomic reads.
[0140] Here, the preset epigenomic characteristic base quantity standard is a pre-defined range of values representing the numerical range that the number of characteristic bases (e.g., the number of Cs in non-CG sites in the first read, or the number of Gs in non-CG sites in the second read) should meet when a read is determined to originate from an epigenomic library. The preset epigenomic characteristic base quantity standard is a range of characteristic values determined after analyzing a large number of reads from known epigenomic library sources, capable of effectively distinguishing the number of characteristic bases in epigenomic library reads from genomic library reads. The preset epigenomic characteristic base quantity standard serves as a criterion for quickly and effectively identifying whether the paired reads in mixed sequencing data originate from an epigenomic library during the data source detection phase.
[0141] In practice, the range of values corresponding to the preset epigenomic characteristic base quantity standard depends on the targeted transformation method, sequencing platform, and specific bioinformatics analysis process used.
[0142] Step 20202: In response to determining that the number of characteristic bases of at least one read in the paired reads to be tested meets the preset standard for the number of characteristic bases in the genome, a read detection result is generated to indicate that the paired reads to be tested are genome reads.
[0143] Here, the preset genomic characteristic base quantity standard is a pre-defined range of values representing the numerical range that the number of characteristic bases (e.g., the number of Cs in non-CG sites in the first read, or the number of Gs in non-CG sites in the second read) should meet when a read is determined to originate from a genomic library. The preset genomic characteristic base quantity standard is a range of characteristic values determined after analyzing a large number of reads from known genomic library sources, capable of effectively distinguishing the number of characteristic bases in genomic library reads from epigenomic library reads. The preset genomic characteristic base quantity standard serves as a criterion for quickly and effectively identifying whether the paired reads in mixed sequencing data originate from a genomic library during the data source detection phase.
[0144] In practice, the range of values corresponding to the preset standard for the number of characteristic bases in the genome depends on the targeted transformation method, sequencing platform, and specific bioinformatics analysis process used.
[0145] By setting and accurately determining the number of characteristic bases corresponding to the genome, efficient screening of genome reads was achieved, reducing the risk of misjudgment, ensuring the reliability of the test results, and providing high-quality data support for subsequent genome feature analysis.
[0146] As an example, assume that r1C is the number of characteristic bases in the first read segment and r2G is the number of characteristic bases in the second read segment.
[0147] The preset standard for the number of characteristic bases in the epigenome can be: r1C is less than or equal to r1C_L, and / or r2G is less than or equal to r2G_L.
[0148] The preset standard for the number of characteristic bases in the genome can be: r1C is greater than or equal to r1C_H, and / or r2G is greater than or equal to r2G_H.
[0149] r1C_L and r1C_H are preset thresholds, with a value range of [1, 15], and r1C_H is greater than r1C_L.
[0150] r2G_L and r2G_H are preset thresholds, with a value range of [1, 15], and r2G_H is greater than r2G_L.
[0151] In some alternative implementations, step 202 may also include, for example: Figure 2D The following steps 20203 to 20205 are shown:
[0152] Step 20203: The sum of the number of characteristic bases in the two reads in the paired reads to be tested is determined as the number of double-stranded characteristic bases.
[0153] Here, the number of double-stranded characteristic bases is used to comprehensively reflect the overall characteristics of the paired reads to be tested, avoiding the impact of randomness or characteristic deviations that may exist in a single read on the judgment results. This provides a more stable and comprehensive quantitative basis for distinguishing genomic reads from epigenomic reads based on the standard of the number of double-stranded characteristic bases, and helps to improve the accuracy and reliability of distinguishing genomic and epigenomic reads.
[0154] Step 20204: In response to determining that the number of double-stranded characteristic bases meets the preset epigenome double-stranded characteristic base number standard, a read detection result is generated to indicate that the paired read to be tested is an epigenome read.
[0155] Here, the preset epigenomic double-stranded characteristic base quantity standard is a pre-defined range of values. It represents the range of values that the sum of the characteristic base quantities of two reads should satisfy when a paired read is determined to originate from an epigenomic library. This preset epigenomic double-stranded characteristic base quantity standard is a range of characteristic values determined after analyzing a large number of paired reads from known epigenomic library sources. It effectively distinguishes the number of double-stranded characteristic bases in paired reads from epigenomic libraries from those from genomic libraries. The preset epigenomic double-stranded characteristic base quantity standard serves as a criterion for quickly and effectively identifying whether a paired read in mixed sequencing data originates from an epigenomic library during the data source detection phase.
[0156] In practice, the range of values corresponding to the preset standard for the number of double-stranded characteristic bases in the epigenome depends on the targeted transformation method, sequencing platform, and specific bioinformatics analysis process used.
[0157] By comprehensively considering the number of double-stranded characteristic bases in paired reads and comparing them with the preset standard for the number of double-stranded characteristic bases in epigenomes, the accuracy and stability of epigenome read identification are effectively improved, the risk of misjudgment caused by single read feature deviation is reduced, and a more reliable basis is provided for subsequent data analysis.
[0158] Step 20205: In response to determining that the number of double-stranded characteristic bases meets the preset standard for the number of double-stranded characteristic bases in the genome, a read detection result is generated to indicate that the paired read to be tested is a genomic read.
[0159] Here, the preset standard for the number of characteristic bases in double-stranded genome sequences is a pre-defined range of values. It represents the range of values that the sum of the characteristic bases in two reads must satisfy when a given paired read is determined to originate from a genomic library. This preset standard is based on the analysis of numerous paired reads from known genomic library sources and is a range of characteristic values that can effectively distinguish between paired reads from genomic libraries and paired reads from epigenomic libraries. The preset standard serves as a criterion for quickly and effectively identifying whether a paired read in mixed sequencing data originates from a genomic library during the data source detection phase.
[0160] In practice, the range of values corresponding to the preset standard for the number of characteristic bases in the genome depends on the targeted transformation method, sequencing platform, and specific bioinformatics analysis process used.
[0161] Here, the preset standard for the number of characteristic bases in the double-stranded epigenome can be: r1C+r2G is less than or equal to r1Cr2G_L.
[0162] The preset standard for the number of characteristic double-stranded bases in the genome can be: r1C+r2G > r1Cr2G_H.
[0163] r1Cr2G_L and r1Cr2G_H are preset thresholds, with a value range of [1, 30], and r1Cr2G_H is greater than r1Cr2G_L.
[0164] By using the number of double-stranded characteristic bases in paired reads as the criterion and comparing it with a preset standard for the number of double-stranded characteristic bases in the genome, the impact of fluctuations in single read characteristics is reduced, the stability and accuracy of genome read determination are improved, and reliable data support is provided for subsequent genome-related analyses.
[0165] In some optional implementations, step 20132, based on the number of characteristic bases and the length of the read segment, extracts the distribution features of the read segment to obtain the target distribution features of the read segment, which can be performed as follows:
[0166] First, the ratio of the number of characteristic bases in the read segment to the length of the read segment is determined as the percentage of characteristic bases in the read segment.
[0167] Then, the proportion of characteristic bases in the read segment is determined as the target distribution characteristic of the read segment.
[0168] Here, the percentage of characteristic bases in a read segment refers to the percentage of the number of bases identified as characteristic bases (e.g., the number of Cs in non-CG sites in the first read segment, or the number of Gs in non-CG sites in the second read segment) relative to the total number of bases in the read segment (i.e., the read segment length), which is used to reflect the relative enrichment of characteristic bases in the read segment.
[0169] In practice, step 20132 can be performed as follows:
[0170] When the reading segment is the first reading segment, the specific formula is as follows:
[0171]
[0172] Where r1ratio is the percentage of characteristic bases in the first read segment, C is the number of characteristic bases in the first read segment, and L1 is the length of the first read segment.
[0173] When the reading segment is the second reading segment, the specific formula is as follows:
[0174]
[0175] Where r2ratio is the percentage of characteristic bases in the second read segment, G is the number of characteristic bases in the second read segment, and L2 is the length of the second read segment.
[0176] Accordingly, in step 202, based on the feature parameter set of at least one read in the paired reads to be tested, the data source of the paired reads to be tested is detected to obtain a read detection result indicating whether the paired reads to be tested are genomic reads or epigenomic reads. This may also include, for example: Figure 2D The following steps 20206 and 20207 are shown:
[0177] Step 20206: In response to determining that the proportion of characteristic bases in at least one read in the paired reads to be tested meets the preset standard for the proportion of characteristic bases in epigenomic genomes, a read detection result is generated to indicate that the paired reads to be tested are epigenomic reads.
[0178] Here, the preset epigenomic characteristic base ratio standard is a pre-defined range of values. It represents the numerical range that the proportion of characteristic bases in a read, when determined to originate from an epigenomic library, must meet to be considered a valid source. This preset epigenomic characteristic base ratio standard is a range of values determined after analyzing a large number of reads from known epigenomic library sources. It effectively distinguishes the proportion of characteristic bases in epigenomic library reads from genomic library reads. The preset epigenomic characteristic base ratio standard serves as a criterion for quickly and effectively identifying whether the paired reads in mixed sequencing data originate from an epigenomic library during the data source detection phase.
[0179] In practice, the range of values corresponding to the pre-defined standard for the proportion of characteristic bases in the epigenomic genome depends on the targeted transformation method, sequencing platform, and specific bioinformatics analysis process used.
[0180] By responding to the determination that the proportion of characteristic bases in at least one read in the paired read to be tested meets the preset standard for the proportion of characteristic bases in the epigenome, a read detection result is generated to indicate that the paired read to be tested is an epigenome read. This can effectively distinguish the differences between reads of different lengths or reads with different characteristic base densities, thus improving the robustness and accuracy of the detection.
[0181] Step 20207: In response to determining that the proportion of characteristic bases in at least one read in the paired reads to be tested meets the preset standard for the proportion of characteristic bases in the genome, a read detection result is generated to indicate that the paired reads to be tested are genome reads.
[0182] Here, the preset standard for the proportion of characteristic bases in a genome is a pre-defined range of values. It represents the numerical range that the proportion of characteristic bases in a read, when determined to originate from a genomic library, must meet to be considered as such. This preset standard is a range of values determined after analyzing a large number of reads from known genomic library sources, effectively distinguishing the proportion of characteristic bases in genomic library reads from epigenomic library reads. The preset standard serves as a criterion for quickly and effectively identifying whether the paired reads in mixed sequencing data originate from a genomic library during the data source detection phase.
[0183] In practice, the range of the pre-defined standard for the proportion of characteristic bases in the genome depends on the targeted transformation method, sequencing platform, and specific bioinformatics analysis process used.
[0184] By responding to the determination that the proportion of characteristic bases in at least one read in the paired read to be tested meets the preset standard for the proportion of characteristic bases in the genome, a read detection result is generated to indicate that the paired read to be tested is a genome read. This can effectively distinguish the differences between reads of different lengths or reads with different characteristic base densities, thereby more accurately identifying reads with typical genome characteristics, improving the accuracy and reliability of genome read determination, and avoiding misclassification caused by factors such as read length.
[0185] As an example, assume that r1ratio and r2ratio are the proportions of characteristic bases in the first and second reads, respectively.
[0186] The preset standard for the proportion of characteristic bases in the epigenomic genome can be: r1ratio is less than or equal to r1ratio_L, and / or r2ratio is less than or equal to r2ratio_L.
[0187] The preset standard for the proportion of characteristic bases in the genome can be: r1ratio is greater than or equal to r1ratio_H, and / or r2ratio is greater than or equal to r2ratio_H.
[0188] r1ratio_L and r1ratio_H are preset thresholds, with a value range of [0, 0.15], and r1ratio_H is greater than r1ratio_L.
[0189] r2ratio_L and r2ratio_H are preset thresholds, with a value range of [0, 0.15], and r2ratio_H is greater than r2ratio_L.
[0190] In some alternative implementations, step 202 may also include, for example: Figure 2D The following steps are shown from 20208 to 20212:
[0191] Step 20208: The sum of the number of characteristic bases in the two reads in the paired reads to be tested is determined as the number of double-stranded characteristic bases.
[0192] Step 20209: Determine the sum of the lengths of the two read segments in the paired read segments to be tested as the double-chain read segment length.
[0193] Step 20210: Divide the number of double-stranded characteristic bases by the length of the double-stranded read to determine the overall proportion of characteristic bases in the paired read to be tested.
[0194] The specific formula is as follows:
[0195]
[0196] Where r1Cr2Gratio represents the percentage of characteristic bases in the paired read segment to be tested, C represents the number of characteristic bases in the first read segment, G represents the number of characteristic bases in the second read segment, L1 represents the read length of the first read segment, and L2 represents the read length of the second read segment.
[0197] Step 20211: In response to the determination that the overall characteristic base quantity ratio of the paired read to be tested meets the preset epigenome overall characteristic base quantity ratio standard, a read detection result is generated to indicate that the paired read to be tested is an epigenome read.
[0198] Here, the preset standard for the proportion of overall epigenomic characteristic bases is a pre-defined range of values. It represents the numerical range that the proportion of overall characteristic bases in a read, when determined to originate from an epigenomic library, must meet to be considered a valid source. This preset standard is based on analysis of numerous reads from known epigenomic library sources and is a range of characteristic values that effectively distinguishes the proportion of overall characteristic bases in epigenomic library reads from genomic library reads. This preset standard serves as a criterion for quickly and effectively identifying whether the paired reads in mixed sequencing data originate from an epigenomic library during the data source detection phase.
[0199] In practice, the range of values corresponding to the pre-defined standard for the proportion of bases in the overall epigenome characteristics depends on the targeted transformation method, sequencing platform, and specific bioinformatics analysis process used.
[0200] By responding to the determination that the overall characteristic base ratio of the paired read to be tested meets the preset standard for the overall characteristic base ratio of the epigenome, read detection results are generated to indicate that the paired read to be tested is an epigenome read. This avoids the impact of fluctuations in the characteristic base ratio of a single read, improves the stability and accuracy of epigenome read determination, and provides reliable data support for subsequent genome-related analyses.
[0201] Step 20212: In response to the determination that the overall characteristic base ratio of the paired read to be tested meets the preset standard for the overall characteristic base ratio of the genome, a read detection result is generated to indicate that the paired read to be tested is a genomic read.
[0202] Here, the preset standard for the proportion of overall genome characteristic bases is a pre-defined range of values. It represents the numerical range that the proportion of overall characteristic bases in a read, when determined to originate from a genomic library, must meet to be considered a valid source. This preset standard is based on analysis of numerous reads from known genomic library sources and is a range of characteristic values that effectively distinguishes the proportion of overall characteristic bases in epigenomic library reads from genomic library reads. The preset standard serves as a criterion for quickly and effectively identifying whether a paired read in mixed sequencing data originates from a genomic library during the data source detection phase.
[0203] In practice, the range of values corresponding to the pre-defined standard for the proportion of bases in the overall genome depends on the targeted transformation method, sequencing platform, and specific bioinformatics analysis process used.
[0204] By responding to the determination that the overall characteristic base ratio of the paired read to be tested meets the preset standard for the overall characteristic base ratio of the genome, a read detection result is generated to indicate that the paired read to be tested is a genomic read. This avoids the impact of fluctuations in the characteristic base ratio of a single read, improves the stability and accuracy of genomic read determination, and provides reliable data support for subsequent genome-related analyses.
[0205] As an example, the preset standard for the proportion of characteristic bases in the whole genome can be that r1Cr2Gratio is greater than or equal to r1Cr2G_ratio_H.
[0206] The preset standard for the proportion of bases in the overall epigenome characteristic can be r1Cr2Gratio less than or equal to r1Cr2G_ratio_L.
[0207] r1Cr2G_ratio_H and r1Cr2G_ratio_L are preset thresholds, with a value range of [0, 0.15], and r1Cr2G_H is greater than r1Cr2G_L.
[0208] In some optional implementations, step 20132, which involves extracting the distribution features of the read segment based on the number of characteristic bases and the read segment length, to obtain the target distribution features of the read segment, can also be performed as follows:
[0209] The epigenomic origin matching probability and genomic origin matching probability of the read segment are determined based on the read length and the number of characteristic bases.
[0210] Here, source matching probability refers to the likelihood that the features of a read (read length and number of characteristic bases) match the feature model of a known source (such as a genomic library or epigenomic library). It is a value between 0 and 1 (or a percentage between 0% and 100%) used to quantify the confidence that the read originates from a specific library. Source matching probability is primarily used in the data source detection phase to provide a refined and probabilistic basis for read classification. Unlike hard judgments based on a single threshold, source matching probability reflects the confidence level of the classification.
[0211] As an example, when determining the source-matching probability of a read based on its read length and the number of characteristic bases, firstly, a large amount of read data with known source labels (partly from genomic libraries and partly from epigenomic libraries) needs to be prepared as a training set. Then, the read length and the number of characteristic bases corresponding to each read data point are determined. Next, a machine learning classification model, such as logistic regression, random forest, or support vector machine (SVM), is trained using these labeled (i.e., known source) read lengths and characteristic base counts. This training process allows the machine learning classification model to learn how to distinguish reads from different sources based on the input read length and the number of characteristic bases, and output a probability value.
[0212] Subsequently, for each read in the current test read pair, the read length and the number of characteristic bases are input into the pre-trained and saved machine learning classification model. Finally, the machine learning classification model outputs the probability values of the read belonging to each possible source (such as genome or epigenome), and these values are the source matching probabilities of the read.
[0213] In practice, the above calculation of the source matching probability of the read segment can be performed as follows:
[0214] When the read is the first read, the epigenomic origin matching probability of the first read is determined based on the read length and the number of characteristic bases, using the following formula:
[0215] r1prob1=pbinom(k≥C|Size=L,prob=P a1 )
[0216] Where pbinom() is the binomial probability calculation function, r1prob1 is the epigenomic origin matching probability of the first read, and the smaller the value of r1prob1, the more likely the read is to originate from the genome. C is the number of characteristic bases in the read, L is the read length, and P... a1 is the prior probability value, representing the probability of observing a characteristic base (i.e., C in a non-CG site) in the first read of the epigenome, with a value range of [0, 0.05].
[0217] When the read is the first read, the genomic origin matching probability of the first read is determined based on the read length and the number of characteristic bases, using the following formula:
[0218] r1prob2=pbinom(k≤C|Size=L,prob=P b1 )
[0219] Where pbinom() is the binomial probability calculation function, r1prob2 is the genomic origin matching probability of the first read, a smaller value of r1prob2 indicates that the read is more likely to come from the epigenome, and a larger value of r1prob2 indicates that the read is more likely to come from the genome. C is the number of characteristic bases in the read, L is the read length, and P... b1 is the prior probability value, representing the probability of observing a matching characteristic base (i.e., C in non-CG sites) in the first read of the genome, with a value range of [0.05, 0.4].
[0220] When the read is the second read, the epigenomic origin matching probability of the second read is determined based on the read length and the number of characteristic bases, using the following formula:
[0221] r2prob1=pbinom(k≥G|Size=L,prob=P a2 )
[0222] Where pbinom() is the binomial probability calculation function, r2prob1 is the epigenomic origin matching probability of the second read, and the smaller the value of r2prob1, the more likely the read is to originate from the genome. G is the number of characteristic bases in the read, L is the read length, and P... a2 is the prior probability value, representing the probability of observing a characteristic base (i.e., G in a non-CG site) in the second read of the epigenome, with a value range of [0, 0.05].
[0223] When the read is the second read, the genomic origin matching probability of the second read is determined based on the read length and the number of characteristic bases, using the following formula:
[0224] r2prob2=pbinom(k≤G∣Size=L,prob=P b2 )
[0225] Where pbinom() is the binomial probability calculation function, r2prob2 is the genomic origin matching probability of the second read, and the smaller the value of r2prob2, the more likely the read is to originate from the epigenome. G is the number of characteristic bases in the read, L is the read length, and P... b2is the prior probability value, representing the probability of observing a matching characteristic base (i.e., G in a non-CG site) in the second read of the genome, with a value range of [0.05, 0.4].
[0226] Accordingly, step 202 may also include, for example: Figure 2D The following steps 20213 and 20214 are shown:
[0227] Step 20213: In response to determining that the epigenomic origin matching probability of at least one read in the paired reads to be tested meets the preset epigenomic origin matching probability standard or does not meet the preset genome origin matching probability standard, a read detection result is generated to indicate that the paired reads to be tested are epigenomic reads.
[0228] Here, the preset epigenomic origin matching probability standard is a pre-defined range of values. It represents the range of values that a read, when determined to originate from an epigenomic library, should satisfy in terms of similarity or matching probability between its features (such as read length and number of characteristic bases) and those of known epigenomic library reads. This preset epigenomic origin matching probability standard is determined by training a machine learning model (using algorithms such as Support Vector Machine (SVM), Random Forest, and Logistic Regression) on the read features (such as read length and number of characteristic bases) from a large number of known sources (epigenomic libraries and genomic libraries). It is a range of probabilities that can effectively distinguish between two types of reads. The preset epigenomic origin matching probability standard serves as a basis for judgment during the data source detection phase, enabling more accurate identification of whether the paired reads in mixed sequencing data originate from an epigenomic library based on the machine learning model.
[0229] In practice, the preset epigenomic origin matching probability standard is usually set between 0 and 1 (or 0% to 100%), with the specific value depending on the machine learning model used, the quality and quantity of training data, and the specific requirements for classification accuracy and recall. For example, to improve specificity (reduce the probability of misclassifying genomic reads as epigenomic reads), the standard may be set higher, such as 0.8 (80%) or higher.
[0230] By introducing a source matching probability determination method based on machine learning models, we can make more complex and intelligent judgments by comprehensively utilizing the features of the reading segments, providing a higher quality data foundation for subsequent accurate analysis.
[0231] Step 20214: In response to determining that the genomic origin matching probability of at least one read in the paired reads to be tested meets the preset genomic origin matching probability standard or does not meet the preset epigenome origin matching probability standard, a read detection result is generated to indicate that the paired reads to be tested are genomic reads.
[0232] Here, the preset genome origin matching probability standard is a pre-defined range of values. It represents the required similarity or matching probability (or a value within a range) between the features (such as read length and number of characteristic bases) of a read and those of known genome library reads when the read is determined to originate from a genome library. The preset genome origin matching probability standard is determined by training a machine learning model (using algorithms such as Support Vector Machine (SVM), Random Forest, and Logistic Regression) on the features of reads from a large number of known sources (genome libraries and epigenome libraries). This range of probabilities effectively distinguishes between two types of reads. The preset genome origin matching probability standard serves as a basis for judgment during the data source detection phase, enabling more accurate identification of whether the paired reads in mixed sequencing data originate from a genome library based on the machine learning model.
[0233] In practice, the range of values corresponding to the preset genome source matching probability standard is usually set between 0 and 1 (or 0% to 100%). The specific value depends on the machine learning model used, the quality and quantity of training data, and the specific requirements for classification accuracy and recall.
[0234] By employing a machine learning-based probabilistic model and a pre-defined genome source matching probability standard, the multidimensional features of reads can be integrated for accurate identification, effectively overcoming the limitations of the single feature threshold method and improving the accuracy and reliability of genome read identification. In particular, it performs better when dealing with complex or atypical reads, thereby ensuring the credibility and accuracy of downstream genome analysis results.
[0235] As an example, assume that r1prob1 and r1prob2 are the epigenomic origin matching probability and genomic origin matching probability of the first read, respectively, and r2prob1 and r2prob2 are the epigenomic origin matching probability and genomic origin matching probability of the second read, respectively.
[0236] The preset epigenome origin matching probability criteria can be: r1prob1 is greater than or equal to r1prob1_H, and / or r1prob2 is less than or equal to r1prob2_L.
[0237] The preset genome source matching probability criteria can be: r1prob1 is less than or equal to r1prob1_L, and / or r1prob2 is greater than r1prob2_H.
[0238] r1prob1_H and r1prob1_L are preset thresholds, with a value range of [0, 0.05], and r1prob1_H is greater than r1prob1_L.
[0239] r1prob2_H and r1prob2_L are preset thresholds, with a value range of [0, 0.05], and r1prob2_H is greater than r1prob2_L.
[0240] By adopting the above optional implementation method and introducing a machine learning classification model, probabilistic prediction of the source of the read segment is realized. Compared with simple threshold judgment, it can better handle noise and edge cases in the data, improve the accuracy and robustness of source determination, and provide valuable confidence information for subsequent decision-making.
[0241] In some alternative implementations, step 202 may also include, for example: Figure 2D The following steps 20215 to 20219 are shown:
[0242] Step 20215: The sum of the number of characteristic bases in the two reads of the paired reads to be tested is determined as the number of double-stranded characteristic bases.
[0243] Step 20216: Determine the sum of the lengths of the two read segments in the paired read segments to be tested as the length of the double-chain read segment.
[0244] Step 20217: Based on the double-stranded read length and the number of double-stranded characteristic bases, determine the overall epigenomic origin matching probability and the overall genome origin matching probability of the paired read to be tested.
[0245] The overall epigenomic origin matching probability of the paired reads to be tested is calculated using the following formula:
[0246] r1Cr2Gprob1=pbinom(k≤C+G∣Size=L1+L2,prob=P ab )
[0247] Where r1Cr2Gprob1 is the overall epigenomic origin matching probability of the paired reads to be tested, C is the number of characteristic bases in the first read, G is the number of characteristic bases in the second read, L1 is the read length of the first read, L2 is the read length of the second read, and P is the total number of characteristic bases in the second read. abis a priori probability value, representing the probability of observing a characteristic base marker (i.e., C in a non-CG site in the first read and G in a non-CG site in the second read) in the epigenome paired reads, with a value range of [0, 0.05].
[0248] The overall genomic origin matching probability of the paired reads to be tested is calculated using the following formula:
[0249] r1Cr2Gprob2=pbinom(k≤C+G∣Size=L1+L2,prob=P ab’ )
[0250] Where r1Cr2Gprob2 is the overall genomic origin matching probability of the paired reads to be tested, C is the number of characteristic bases in the first read, G is the number of characteristic bases in the second read, L1 is the read length of the first read, L2 is the read length of the second read, and P is the number of characteristic bases in the second read. ab ' is the prior probability value, representing the probability of observing a characteristic base (i.e., C in a non-CG site in the first read and G in a non-CG site in the second read) in a paired read of the genome, with a value range of [0.05, 4].
[0251] Step 20218: In response to determining whether the overall epigenomic origin matching probability of the paired read to be tested meets the preset overall epigenomic origin matching probability standard or does not meet the preset overall genome origin matching probability standard, a read detection result is generated to indicate that the paired read to be tested is an epigenomic read.
[0252] Here, the preset overall epigenomic origin matching probability standard is a pre-defined range of values. It represents the range of values that a paired read, when determined to originate from an epigenomic library as a whole, should satisfy in terms of similarity or matching probability between its features (such as double-stranded read length and number of double-stranded characteristic bases) and those of paired reads from known epigenomic libraries. This preset overall epigenomic origin matching probability standard is determined by training a machine learning model (using algorithms such as Support Vector Machine (SVM), Random Forest, and Logistic Regression) on the features (such as double-stranded read length and number of double-stranded characteristic bases) of paired reads from a large number of known sources (epigmogen and genomic libraries). This set of values effectively distinguishes between two types of paired reads. The preset overall epigenomic origin matching probability standard serves as a criterion for judgment, enabling more accurate identification of whether the paired reads in mixed sequencing data originate from an epigenomic library during the data source detection phase, based on the machine learning model.
[0253] In practice, the preset range of the overall epigenome origin matching probability standard is usually set between 0 and 1 (or 0% to 100%). The specific value depends on the machine learning model used, the quality and quantity of training data, and the specific requirements for classification accuracy and recall. For example, in order to improve specificity (reduce the probability of misclassifying genomic reads as epigenome reads), the standard may be set higher, such as 0.8 (80%) or higher.
[0254] By introducing a source matching probability determination method based on machine learning models, we can make more complex and intelligent judgments by comprehensively utilizing the features of the read segment (such as the overall read segment length and the overall number of feature bases), providing a higher quality data foundation for subsequent accurate analysis.
[0255] Step 20219: In response to determining whether the overall genome source matching probability of the paired read to be tested meets the preset overall genome source matching probability standard or does not meet the preset overall epigenome source matching probability standard, a read detection result is generated to indicate that the paired read to be tested is a genomic read.
[0256] Here, the preset overall genome origin matching probability standard is a pre-defined range of values. It represents the range of values that a paired read, when determined as a whole to originate from a genomic library, should satisfy in terms of similarity or matching probability between its features (such as double-stranded read length and number of double-stranded characteristic bases) and those of paired reads from known genomic libraries. This preset overall genome origin matching probability standard is determined by training a machine learning model (using algorithms such as Support Vector Machine (SVM), Random Forest, and Logistic Regression) on the features (such as double-stranded read length and number of double-stranded characteristic bases) of paired reads from a large number of known sources (epigeon libraries and genomic libraries). This set of values effectively distinguishes between two types of paired reads. The preset overall genome origin matching probability standard serves as a criterion for judgment, enabling more accurate identification of whether the paired reads in mixed sequencing data originate from an epigenomic library during the data source detection phase, based on the machine learning model.
[0257] In practice, the preset overall genome source matching probability standard is usually set between 0 and 1 (or 0% to 100%). The specific value depends on the machine learning model used, the quality and quantity of training data, and the specific requirements for classification accuracy and recall. For example, in order to improve specificity (reduce the probability of misclassifying genomic reads as epigenomic reads), the standard may be set higher, such as 0.8 (80%) or higher.
[0258] By introducing a source matching probability determination method based on machine learning models, we can make more complex and intelligent judgments by comprehensively utilizing the features of the read segment (such as the overall read segment length and the overall number of feature bases), providing a higher quality data foundation for subsequent accurate analysis.
[0259] As an example, assume that r1Cr2Gprob1 is the overall epigenomic origin matching probability of the paired reads, and r1Cr2Gprob2 is the overall genomic origin matching probability of the paired reads.
[0260] The overall epigenome origin matching probability criteria can be: r1Cr2Gprob1 is greater than or equal to r1Cr2Gprob1_H, and / or r1Cr2Gprob2 is less than or equal to r1Cr2Gprob2_L.
[0261] The overall genome origin matching probability criterion can be: r1Cr2Gprob1 is less than or equal to r1Cr2Gprob1_L, and / or r1Cr2Gprob2 is greater than or equal to r1Cr2Gprob2_H.
[0262] r1Cr2Gprob2_L and r1Cr2Gprob1_H are preset thresholds, with values ranging from [0, 0.05], and r1Cr2Gprob2_L is less than r1Cr2Gprob2_H.
[0263] r1Cr2Gprob1_L and r1Cr2Gprob1_L are preset thresholds, with a value range of [0.05, 4], and r1Cr2Gprob1_L is less than r1Cr2Gprob1_H.
[0264] The hybrid library paired read genome type detection method provided in the above embodiments of this disclosure performs the following feature extraction operations on at least one read of the paired read to be tested: For each candidate cytosine site in the read, based on the dinucleotide context of the candidate cytosine site, it is determined whether the candidate cytosine site is a non-CG site; the number of all candidate cytosine sites in the read that are determined to be non-CG sites is counted and determined as the number of characteristic bases of the read; and a set of feature parameters of the read is generated based on the number of characteristic bases of the read. Finally, based on the set of feature parameters of at least one read of the paired read to be tested, the source of the paired read to be tested is detected, and a read detection result indicating whether the paired read to be tested is a genomic read or an epigenomic read is obtained. In this way, by extracting feature parameters from the paired read to be tested for source detection, efficient and accurate differentiation between genomic reads and epigenomic reads is achieved, effectively avoiding the limitations of single feature determination and the risk of misjudgment caused by individual differences in reads. Thus, while reducing experimental costs and sample consumption, the parsing efficiency and application value of hybrid sequencing data are significantly improved.
[0265] Example 1
[0266] As an example, the experimental results of somatic cell mutation detection and qualitative detection of the presence of tumor DNA fragments according to the above scheme are listed below.
[0267] The clinical sample consisted of 20 clinical tumor-related samples, used to validate the accuracy of genomic variant detection and epigenomic tumor DNA residue detection.
[0268] The whole-genome data consisted of three pairs of genome and epigenome sequencing data from known sources, used to validate the performance of whole-genome-wide data splitting.
[0269] This experiment consists of four procedures:
[0270] 1. Sample processing and sequencing: After library construction, a single DNA sample is divided into two parts. One part is not enzymatically / chemically transformed (to obtain genomic fragments), and the other part is enzymatically / chemically transformed (to obtain epigenome fragments). After mixing, the samples are captured by probes, amplified by PCR, and sequenced to obtain mixed sequencing data.
[0271] 2. Data splitting: Extract the set of target feature parameters of paired reads from the sequencing data, calculate the proportion of feature bases and the probability of source matching using the corresponding formulas, and split the mixed data into genomic reads and epigenomic reads according to the preset standards.
[0272] 3. Subsequent detection: Somatic cell variation detection was performed on the split genomic reads using existing bioinformatics tools; qualitative detection of residual tumor DNA was performed on the split epigenomic reads using existing bioinformatics tools.
[0273] 4. Accuracy verification: Using the results of single-hybrid detection as the gold standard, head-to-head comparisons were performed on the detection results of clinical samples; the sensitivity and accuracy of genomic and epigenomic reads were statistically analyzed after splitting the results of mixing three pairs of whole genome data.
[0274] Please refer to Figure 3 This is a correlation analysis diagram between the abundance of detected variants and the gold standard variant abundance in genomic variant detection.
[0275] Here, the fitted equation (y = 6.78e-6 + 0.997x) and the adjusted correlation coefficient (R²) are obtained from the graph. 2 adj =1) It can be seen that the variation abundance of the two is highly consistent, which further proves the reliability of this application in the detection of genomic variation.
[0276] Please refer to Table 1 below. As shown in Table 1, in the head-to-head comparison of genomic variant detection in 20 clinical samples, using single heterozygotes as the gold standard, among the samples that were judged positive by the gold standard, 95 were positive and 2 were negative. Among the samples that were judged negative by the gold standard, the results of this application were all negative. Based on this, the accuracy of genomic variant detection in this application is calculated to be 98%, and the two missed variants were both ultra-low abundance variants, demonstrating excellent overall detection performance.
[0277] Table 1 Head-to-head comparison of genomic variation detection
[0278]
[0279] For epigenomic testing, please refer to [link / reference]. Figure 4 This is a correlation analysis diagram between the abundance of the detection method and the gold standard abundance in the qualitative detection of whether tumor DNA fragments exist in the epigenome.
[0280] Here, by Figure 4 The fitted equation is (y = -6.23 × 10). -7 +0.995x) and the adjusted correlation coefficient (R) 2 adj =1) It can be seen that the abundance detected in this application is highly consistent with the gold standard abundance, demonstrating excellent consistency in epigenome detection.
[0281] Please refer to Table 2. As shown in Table 2, in the qualitative detection of residual epigenomic tumor DNA in 20 clinical samples, using single heterozygotes as the gold standard, the 10 samples that were judged positive by the gold standard were all accurately detected as positive in this application, and the 10 samples that were judged negative by the gold standard were all accurately detected as negative in this application. There were no false positive or false negative results, and the detection accuracy reached 100%.
[0282] Table 2 Head-to-head comparison of epigenomic variation detection
[0283]
[0284] Regarding whole-genome sequencing, please refer to Table 3. As shown in Table 3, after splitting the first pair of whole-genome data, the total number of true genome reads was 12,183,824 (12,150,864 accurately split and 32,960 mistakenly split into epigenomes), with a calculated genome read sensitivity of 99.73%. The total number of true epigenome reads was 13,992,503 (13,925,050 accurately split and 67,353 mistakenly split into genomes), with an epigenome read sensitivity of 99.52%. At the same time, the accurate splitting rate of reads identified as genomes in this application reached 99.45%, and the accurate splitting rate of reads identified as epigenomes reached 99.76%, demonstrating high splitting sensitivity and accuracy.
[0285] Table 3. Head-to-head comparison of the first pair of whole-genome variation detections.
[0286]
[0287] Please refer to Table 4. As shown in Table 4, the splitting results of the second pair of whole genome data show that 12,403,025 real genome reads were accurately split and 33,434 were missplitted, with a genome sensitivity of 99.73%; 154,008,749 real epigenome reads were accurately split and 71,011 were missplitted, with an epigenome sensitivity of 99.5%; the accuracy of genome reads was 99.43%, and the accuracy of epigenome reads was 99.76%, which is highly consistent with the splitting performance of the first pair of data. It is clear that among the results identified as epigenome fragments in this application, 99.76% are true and valid, with a misclassification rate of only 0.24%.
[0288] Table 4. Second pair of head-to-head comparison table of whole-genome variation detection
[0289]
[0290] Please refer to Table 5. As shown in Table 5, the splitting results of the third pair of whole-genome data further validated the stability of the method: 12,128,452 genome reads were accurately split with 30,176 missplitting, achieving a sensitivity of 99.75%; 13,485,423 epigenome reads were accurately split with 67,137 missplitting, achieving a sensitivity of 99.5%; the accuracy of genome reads remained at 99.45%, and the accuracy of epigenome reads improved to 99.88%. The splitting results of the three sets of whole-genome data all met the requirements of genome sensitivity > 99.7% and accuracy > 99.4%, and epigenome sensitivity > 99.5% and accuracy > 99.7%, fully demonstrating the stable splitting capability of this application across the whole genome.
[0291] Table 5. Third pair of head-to-head comparison table of whole-genome variation detection
[0292]
[0293] Further reference Figure 5 As an implementation of the methods shown in the above figures, this disclosure provides an embodiment of a hybrid library paired read genome type detection device, which is similar to... Figure 2A Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0294] like Figure 5 As shown, the hybrid library paired read genome type detection device 500 of this embodiment includes: a feature extraction module 501 and a source detection module 502. The feature extraction module 501 is configured to perform the following feature extraction operations on at least one read of the paired read to be tested: for each candidate cytosine site in the read, based on the dinucleotide context of the candidate cytosine site, determine whether the candidate cytosine site is a non-CG site; count the number of all candidate cytosine sites in the read that are determined to be non-CG sites and determine them as the number of characteristic bases of the read; and generate a set of feature parameters for the read based on the number of characteristic bases of the read. The paired read to be tested is a paired read in the hybrid sequencing data corresponding to the hybrid library. The source detection module 502 is configured to perform data source detection on the paired read to be tested based on the set of feature parameters of at least one read in the paired read to be tested, and obtain a read detection result indicating whether the paired read to be tested is a genomic read or an epigenomic read.
[0295] In this embodiment, the specific processing of the feature extraction module 501 and the source detection module 502 of the mixed library paired read genome type detection device 500 and the resulting technical effects can be referred to respectively. Figure 2A The relevant descriptions of steps 201 and 202 in the corresponding embodiments will not be repeated here.
[0296] In some optional implementations, generating a set of characteristic parameters for the read segment based on the number of characteristic bases in the read segment may include:
[0297] The length of the reading segment is obtained by extracting its length feature.
[0298] Based on the number of characteristic bases and the length of the read segment, the distribution features of the read segment are extracted to obtain the target distribution features of the read segment.
[0299] Based on the number of characteristic bases and the target distribution characteristics of the read segment, a set of characteristic parameters for the read segment is generated. In some optional embodiments, the source detection module 502 can be further configured as follows:
[0300] In response to determining that the number of characteristic bases in at least one of the paired reads to be tested meets the preset epigenome characteristic base number standard, a read detection result is generated to indicate that the paired read to be tested is an epigenome read;
[0301] In response to determining that the number of characteristic bases in at least one of the paired reads to be tested meets the preset standard for the number of characteristic bases in the genome, a read detection result is generated to indicate that the paired read to be tested is a genome read.
[0302] In some optional implementations, the source detection module 502 may be further configured to determine the sum of the number of characteristic bases in the two reads in the paired reads to be tested as the number of double-stranded characteristic bases;
[0303] In response to the determination that the number of double-stranded characteristic bases meets the preset epigenome double-stranded characteristic base number standard, a read detection result is generated to indicate that the paired read to be tested is an epigenome read;
[0304] In response to the determination that the number of double-stranded characteristic bases meets the preset standard for the number of double-stranded characteristic bases in the genome, a read detection result is generated to indicate that the paired read to be tested is a genomic read.
[0305] In some optional implementations, based on the number of characteristic bases and the length of the read segment, distribution features of the read segment are extracted to obtain the target distribution features of the read segment, which may include:
[0306] The ratio of the number of characteristic bases in the read segment to the length of the read segment is determined as the percentage of characteristic bases in the read segment.
[0307] The proportion of characteristic bases in the read segment is determined as the target distribution characteristic of the read segment; and
[0308] The source detection module 502 can be further configured as follows:
[0309] In response to determining that the proportion of characteristic bases in at least one read in the paired reads to be tested meets the preset standard for the proportion of characteristic bases in epigenome, a read detection result is generated to indicate that the paired reads to be tested are epigenome reads;
[0310] In response to determining that the proportion of characteristic bases in at least one read in the paired reads to be tested meets the preset standard for the proportion of characteristic bases in the genome, a read detection result is generated to indicate that the paired reads to be tested are genomic reads.
[0311] In some optional implementations, based on the number of characteristic bases and the length of the read segment, distribution features of the read segment are extracted to obtain the target distribution features of the read segment, which may include:
[0312] The epigenomic origin matching probability and genomic origin matching probability of the read segment are determined based on the read length and the number of characteristic bases; and
[0313] The source detection module 502 can be further configured as follows:
[0314] In response to determining that the epigenomic origin matching probability of at least one read in the paired reads to be tested meets the preset epigenomic origin matching probability standard or does not meet the preset genome origin matching probability standard, a read detection result is generated to indicate that the paired reads to be tested are epigenomic reads;
[0315] In response to determining that the genomic origin matching probability of at least one read in the paired reads to be tested meets the preset genomic origin matching probability standard or does not meet the preset epigenome origin matching probability standard, a read detection result is generated to indicate that the paired reads to be tested are genomic reads.
[0316] In some alternative implementations, the source detection module 502 may be further configured as follows:
[0317] The sum of the characteristic bases in the two reads of the paired read to be tested is determined as the number of double-stranded characteristic bases;
[0318] The sum of the lengths of the two reads in the paired reads to be tested is determined as the length of the double-chain read;
[0319] Based on the length of the double-stranded read and the number of double-stranded characteristic bases, the overall epigenome origin matching probability and the overall genome origin matching probability of the paired read to be tested are determined.
[0320] In response to whether the overall epigenomic origin matching probability of the paired read to be tested meets the preset overall epigenomic origin matching probability standard or does not meet the preset overall epigenomic origin matching probability standard, a read detection result is generated to indicate that the paired read to be tested is an epigenomic read;
[0321] In response to whether the overall genome source matching probability of the paired read to be tested meets the preset overall genome source matching probability standard or does not meet the preset overall epigenome source matching probability standard, a read detection result is generated to indicate that the paired read to be tested is a genomic read.
[0322] It should be noted that the implementation details and technical effects of each module in the mixed library paired read genome type detection device provided in the embodiments of this disclosure can be referred to the descriptions of other embodiments in this disclosure, and will not be repeated here.
[0323] The following is for reference. Figure 6 It shows a schematic diagram of the structure of a computer system 600 suitable for implementing the terminal device of this disclosure. Figure 6 The computer system 600 shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.
[0324] like Figure 6 As shown, the computer system 600 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. The RAM 603 also stores various programs and data required for the operation of the computer system 600. The processing device 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0325] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 608 including, for example, magnetic tapes, hard disks, etc.; and communication devices 609. Communication device 609 allows computer system 600 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6 A computer system 600 with various electronic devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0326] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by a processing device 601, it performs the functions defined in the methods of embodiments of this disclosure.
[0327] It should be noted that the computer-readable medium described above in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0328] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0329] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the following functions: Figure 2AThe embodiments shown and their alternative implementations illustrate a method for detecting genome types of paired reads from mixed libraries.
[0330] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0331] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0332] The modules described in the embodiments of this disclosure can be implemented in software or hardware. The names of the modules do not necessarily limit the module itself; for example, the source detection module can also be described as "a module that performs data source detection on the paired reads to be tested based on a set of feature parameters of at least one read in the paired reads to be tested, and obtains a read detection result indicating whether the paired reads to be tested are genomic reads or epigenomic reads."
[0333] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
Claims
1. A method for detecting genome type of paired reads from a mixed library, characterized in that, The method includes: For at least one read of the paired read to be tested, the following feature extraction operations are performed: For each candidate cytosine site in the read, based on the dinucleotide context in which the candidate cytosine site is located, it is determined whether the candidate cytosine site is a non-CG site; the number of all candidate cytosine sites in the read that are determined to be non-CG sites is counted and determined as the number of characteristic bases of the read; and a set of feature parameters of the read is generated based on the number of characteristic bases of the read. The paired read to be tested is a paired read in the mixed sequencing data corresponding to the mixed library. Based on the feature parameter set of at least one of the paired reads to be tested, the data source of the paired reads to be tested is detected to obtain read detection results that indicate whether the paired reads to be tested are genomic reads or epigenomic reads.
2. The method according to claim 1, characterized in that, The generation of the feature parameter set for the read segment based on the number of characteristic bases in the read segment includes: The length of the reading segment is obtained by extracting its length feature. Based on the number of characteristic bases and the length of the read segment, the distribution features of the read segment are extracted to obtain the target distribution features of the read segment. Based on the number of characteristic bases and the target distribution characteristics of the read segment, a set of characteristic parameters for the read segment is generated.
3. The method according to claim 1, characterized in that, The step of detecting the data source of the paired reads based on the feature parameter set of at least one read in the paired reads to be tested, and obtaining read detection results indicating whether the paired reads to be tested are genomic reads or epigenomic reads, includes: In response to determining that the number of characteristic bases of at least one read in the paired reads to be tested meets the preset epigenome characteristic base number standard, a read detection result is generated to indicate that the paired reads to be tested are epigenome reads; In response to determining that the number of characteristic bases in at least one of the paired reads to be tested meets a preset standard for the number of characteristic bases in a genome, a read detection result is generated to indicate that the paired reads to be tested are genome reads.
4. The method according to claim 2, characterized in that, The step of detecting the data source of the paired reads based on a set of feature parameters of at least one read in the paired reads to be tested, and obtaining a read detection result indicating whether the paired reads to be tested are genomic reads or epigenomic reads, further includes: The sum of the characteristic bases in the two reads of the paired reads to be tested is determined as the number of double-stranded characteristic bases; In response to determining that the number of double-stranded characteristic bases meets the preset epigenome double-stranded characteristic base number standard, a read detection result is generated to indicate that the paired read to be tested is an epigenome read; In response to determining that the number of double-stranded characteristic bases meets the preset standard for the number of double-stranded characteristic bases in the genome, a read detection result is generated to indicate that the paired read to be tested is a genomic read.
5. The method according to claim 2, characterized in that, Based on the number of characteristic bases and the length of the read segment, the distribution features of the read segment are extracted to obtain the target distribution features of the read segment, including: The ratio of the number of characteristic bases in the read segment to the length of the read segment is determined as the percentage of characteristic bases in the read segment. The proportion of characteristic bases in the read segment is determined as the target distribution characteristic of the read segment; and The step of detecting the data source of the paired reads based on the feature parameter set of at least one read in the paired reads to be tested, and obtaining read detection results indicating whether the paired reads to be tested are genomic reads or epigenomic reads, includes: In response to determining that the proportion of characteristic bases in at least one of the paired reads to be tested meets a preset standard for the proportion of characteristic bases in epigenomes, a read detection result is generated to indicate that the paired reads to be tested are epigenome reads. In response to determining that the proportion of characteristic bases in at least one of the paired reads to be tested meets a preset standard for the proportion of characteristic bases in the genome, a read detection result is generated to indicate that the paired reads to be tested are genome reads.
6. The method according to claim 2, characterized in that, Based on the number of characteristic bases and the length of the read segment, the distribution features of the read segment are extracted to obtain the target distribution features of the read segment, including: The epigenomic origin matching probability and genomic origin matching probability of the read segment are determined based on the read length and the number of characteristic bases; and The step of detecting the data source of the paired reads based on the feature parameter set of at least one read in the paired reads to be tested, and obtaining read detection results indicating whether the paired reads to be tested are genomic reads or epigenomic reads, includes: In response to determining that the epigenomic origin matching probability of at least one of the paired reads to be tested meets a preset epigenomic origin matching probability standard or does not meet the preset genome origin matching probability standard, a read detection result is generated to indicate that the paired reads to be tested are epigenomic reads; In response to determining that the genomic origin matching probability of at least one of the paired reads to be tested meets a preset genomic origin matching probability standard or does not meet a preset epigenome origin matching probability standard, a read detection result is generated to indicate that the paired reads to be tested are genomic reads.
7. The method according to claim 2, characterized in that, The step of detecting the data source of the paired reads based on the feature parameter set of at least one read in the paired reads to be tested, and obtaining read detection results indicating whether the paired reads to be tested are genomic reads or epigenomic reads, includes: The sum of the characteristic bases in the two reads of the paired reads to be tested is determined as the number of double-stranded characteristic bases; The sum of the lengths of the two read segments in the paired read segments to be tested is determined as the length of the double-chain read segment; Based on the length of the double-stranded read and the number of double-stranded characteristic bases, the overall epigenome origin matching probability and the overall genome origin matching probability of the paired read to be tested are determined; In response to determining whether the overall epigenome origin matching probability of the paired read segment to be tested meets the preset overall epigenome origin matching probability standard or does not meet the preset overall genome origin matching probability standard, a read detection result is generated to indicate that the paired read segment to be tested is an epigenome read segment; In response to determining whether the overall genome source matching probability of the paired read to be tested meets a preset overall genome source matching probability standard or does not meet the preset overall epigenome source matching probability standard, a read detection result is generated to indicate that the paired read to be tested is a genomic read.
8. A device for detecting genome type of paired reads from a mixed library, comprising: The feature extraction module is configured to perform the following feature extraction operations for at least one read of the paired read to be tested: for each candidate cytosine site in the read, determine whether the candidate cytosine site is a non-CG site based on the dinucleotide context in which the candidate cytosine site is located, count the number of all candidate cytosine sites in the read that are determined to be non-CG sites and determine them as the number of characteristic bases of the read, and generate a set of feature parameters of the read based on the number of characteristic bases of the read, wherein the paired read to be tested is a paired read in the mixed sequencing data corresponding to the mixed library; The source detection module is configured to perform data source detection on the paired reads to be tested based on a set of feature parameters of at least one of the paired reads to be tested, and obtain a read detection result indicating whether the paired reads to be tested are genomic reads or epigenomic reads.
9. An electronic device, comprising: One or more processors; Storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the method as described in any one of claims 1-7.
10. A computer-readable medium having a computer program stored thereon, wherein, When the program is executed by the processor, it implements the method as described in any one of claims 1-7.
11. A computer program product comprising a computer program / instructions that, when executed by a processor, implement the method as described in any one of claims 1-7.