A method and system for detecting minimal residual disease based on ultra-deep sequencing
By employing ultra-high-depth sequencing and molecular tagging technology, the problem of noise suppression in MRD detection has been solved, enabling accurate assessment of NGS noise and precise quantification of tumor signals, thereby improving the specificity and sensitivity of MRD detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-27
- Publication Date
- 2026-04-10
AI Technical Summary
Existing minimal residual disease (MRD) detection technologies suffer from problems such as low ctDNA signal levels, difficulty in identification, high false negatives, and high false positives in clinical applications. In particular, the lack of effective suppression of background noise and technical noise leads to insufficient detection specificity and sensitivity.
By using ultra-high depth sequencing technology and molecular tagging technology to separate and evaluate noise in each step of NGS, and combining prior tumor knowledge, a noise and signal evaluation model is established to identify and distinguish different noise levels, construct a tumor variation atlas, and improve the specificity and sensitivity of MRD detection.
It enables precise differentiation and assessment of noise sources during NGS testing, improves the sensitivity and specificity of MRD testing, accurately reflects the proportion of tumor tissue in plasma, and reduces false positive results.
Smart Images

Figure CN121331226B_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application is based on Chinese invention application CN 202411782959.5, filed on December 6, 2024, entitled "A method and system for detecting tiny residual lesions based on ultra-high depth sequencing", and claims priority to the aforementioned application, the entire contents of which are incorporated herein by reference. Technical Field
[0003] This invention relates to the field of intelligent detection technology for minimal residual lesions, and more specifically, to a method and system for detecting minimal residual lesions based on ultra-high depth sequencing. Background Technology
[0004] The crucial role of circulating tumor DNA (ctDNA) as a biomarker in the detection of minimal residual disease (MRD) in solid tumors has been increasingly demonstrated in clinical cohorts. Detecting and identifying ctDNA signals using liquid biopsy techniques and bioinformatics methods can be used to predict the risk of recurrence after radical treatment of various solid tumors. However, ctDNA-based MRD detection still faces a number of technical and biological challenges in its clinical application.
[0005] In MRD testing, because cell-free DNA in blood has abundant sources, only trace amounts of DNA are often released from the tumor. Furthermore, variations from clonal hematopoiesis, subclonal variations due to tumor heterogeneity, and variations from other unknown sources all lead to low and difficult-to-identify true ctDNA signal levels. Low ctDNA levels are a significant factor contributing to false negatives in MRD testing. On the other hand, technical noise generated during sample processing and library preparation and sequencing, such as DNA strand damage during library preparation, damage during hybridization capture, amplification errors during PCR amplification, and sequencing errors during sequencing, are major factors causing false positives in MRD testing.
[0006] Current MRD detection strategies include two categories. One is the tumor-agonostic strategy, which uses the whole genome approach and does not rely on prior variant information. It enriches signals through methylation or transcriptomics. The other is the tumor-informed strategy, which simultaneously tracks multiple known tumor-derived genomic variants to determine the ctDNA level by identifying corresponding variant genotypes. The tumor-agonostic strategy, although with fewer operation steps, has a wide range and low-depth technical route, which requires complex bioinformatics algorithms and matching baseline models, making it difficult to generalize tumor-specific signals and lack specificity. The tumor-informed strategy, although relying on tumor prior, has very high sensitivity and specificity when combined with personalized variants after constructing the variant map.
[0007] Currently, MRD detection based on tumor prior has different focuses on suppressing background noise depending on the technical route. In population customization panel, the background baseline model constructed by healthy population is mainly used to distinguish background noise through statistical test. In individual customization panel, the consensus base is obtained by molecular marker technology, and the upstream and downstream context trimers on the genome of the individual are used as background, such as C[A>T]G, C[G>A]G, etc. for classification to evaluate the background error rate in the sample.
[0008] Using healthy human cohort data to establish a background noise database, for each detected variant, the significance between the frequency in the plasma sample and the frequency in the normal population background database is tested to determine whether the variant is likely to be tumor-derived, thereby excluding false positive results. This method ignores the differences between data samples and lacks hierarchical analysis of technical noise sources. Batch effects cannot be completely removed. Upgrading the sequencing link also requires a large cost to reestablish the background noise database as a baseline, which is time-consuming and labor-intensive.
[0009] Using genomic background context trimer, calculate the sample-specific background noise error rate, combine the depth and support number of the detection site to divide the threshold as the basis for site determination. This method is too general for background evaluation and does not fully utilize the data. It lacks decomposition of background noise sources and is difficult to balance specificity and sensitivity. In addition, the background noise evaluation and the model for enriching real signals are relatively separated, and there is a lack of effective model for ctDNA determination in plasma. There is also a lack of effective and interpretable model for ctDNA quantification. SUMMARY
[0010] In order to solve the above problems, the purpose of the present application is to systematically identify and evaluate the noise introduced in each link of NGS through ultra-saturation sequencing under the molecular tag technology, and to perform feature extraction and statistical modeling on background errors of different levels, while establishing a set of ctDNA evaluation model for inhibiting noise and enriching signal, thereby improving the specificity and sensitivity of MRD detection.
[0011] In order to achieve the above technical purpose, the present application provides a micro residual lesion detection method based on ultra-high depth sequencing, comprising the following steps:
[0012] Based on the sequencing library, the sample to be tested is obtained by splitting according to the sample index, and the consensus sequence is obtained according to the molecular identifier of the sample to be tested;
[0013] Based on the consensus sequence, the consensus sequence with a quality value less than 25 or a family size less than 3 is filtered out, and the subtype type and the distance from the fragment edge are combined as the noise level introduced on the original nucleic acid molecular chain; the context and the strand specificity are combined to evaluate the noise at the PCR level.
[0014] Based on the noise level, the ctDNA level is estimated by combining tumor prior knowledge, and the significance of the source of the molecular signal is verified to determine the MRD state.
[0015] Preferably, in the process of obtaining the sequencing library, sequencing adapters are added to the nucleic acid molecules by connecting adapters, and the adapter-nucleic acid combination is amplified by polymerase chain reaction to establish the sequencing library, wherein the adapter is a molecular identifier.
[0016] Preferably, in the process of obtaining the molecular identifier of the sample to be tested, the adapter of the sample to be tested and the low-quality sequencing reads are removed, and the molecular identifier is obtained after sequence alignment.
[0017] Preferably, in the process of obtaining the consensus sequence, the sequencing reads are grouped according to the molecular identifier and the position aligned to the reference genome, and the base sequences on the repeated sequencing reads contained in different groups are formed into consensus sequences by voting method.
[0018] Preferably, in the process of evaluating the noise level, the sequencing read pairs containing both positive and negative strand templates in the group are taken as the first consensus sequence duplex; and the consensus sequence is formed only by sequencing read pairs of single template strand, i.e. only containing positive or negative strand sequencing reads, as the second consensus sequence simplex.
[0019] According to the first consensus sequence duplex and the second consensus sequence simplex, the noise level is evaluated.
[0020] Preferably, in the process of evaluating the noise level, after excluding unknown variant sites with a variant allele frequency (VAF) greater than or equal to 1%, known tumor somatic variant sites, clonal hematopoietic variants, and germline variants, the noise level introduced into the original nucleic acid molecule chain is evaluated according to the duplex joint subtype type and the distance from the fragment edge of the first consensus sequence. The second consensus sequence simplex joint context and strand specificity evaluate the noise level of PCR.
[0021] Preferably, in the process of combining tumor prior knowledge, in the sample with a high tumor proportion (>=20%), the variant is identified, the somatic variant information from the tumor is determined by subtracting the germline variant and the clonal hematopoietic variant from the paired white blood cell control sample data; combine the information of the prior sample, classify and quantitatively sort the somatic variants, the classification categories are main clones and subclones, and the quantitative indicators are copy number and variant abundance. Filter and sort the variants according to the abundance to obtain a traceable variant set and construct a prior tumor variant map.
[0022] Preferably, in the process of obtaining the traceable variant set, single base point mutations, short fragment insertion or deletion mutations, and long fragment structure mutations are used as the traceable variant set, wherein the noise of point mutations, short fragment insertion and deletion is modeled and evaluated by the characteristics of genomic background context, edge distance, and strand variation direction; the noise of the structure variant fragment is modeled and evaluated by the sequence characteristics of the variant direction and split reads.
[0023] Preferably, in the process of determining the MRD state, based on the depth information of the to-be-tested sample, according to the evaluation model for evaluating the first noise and the second noise, the ctDNA level is estimated by calculating the maximum likelihood probability of the number of variant molecules supported by multiple sites, the expected number of molecules for evaluating different circulating tumor proportions is obtained and the probability is calculated, the significance is determined by likelihood ratio test, and thus the MRD state is determined.
[0024] The application discloses a micro residual lesion detection system based on ultra-high depth sequencing.
[0025] The data processing unit is used for splitting the sequencing library according to the sample index to obtain a to-be-tested sample, and obtaining a consensus sequence according to a molecular identifier of the to-be-tested sample.
[0026] The noise level evaluation unit is used for filtering out consensus sequences with a quality value less than 25 or a family size less than 3 based on the consensus sequence, and combining a subtype type and a distance from a fragment edge as a noise level introduced into an original nucleic acid molecule chain.
[0027] The detection unit is used for estimating the ctDNA level based on the noise level, combining tumor prior knowledge, determining the MRD state by testing the significance of the signal source.
[0028] The present application discloses the following technical effects:
[0029] The present application can distinguish the noise source in the NGS detection process, and evaluate the characteristics and level of noise introduced by different links, and can also be used for horizontal comparison of the noise level introduced by different experimental links, for example, the duplex horizontal noise can reflect the degree of jagged end and damage degree of the cfDNA sample, the simplex horizontal noise can reflect the noise level brought by the sample PCR amplification and capture process, and further can be used for the performance of reagents in the reaction library building, amplification and capture process.
[0030] The information of the copy number and clone situation other than the tumor tissue allele frequency can accurately reflect the tumor proportion of the tumor tissue, and the layered noise model is combined to analyze the significance of the signal source from the tumor tissue in the plasma sample. BRIEF DESCRIPTION OF DRAWINGS
[0031] In order to more clearly illustrate the technical solutions of the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed in the embodiments. Obviously, the drawings described below are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.
[0032] Figure 1 The main links of NGS sequencing and the noise signals introduced by each link are described
[0033] Figure 2 The present application discloses the method flowchart. DETAILED DESCRIPTION
[0034] In order to make the purpose, technical scheme and advantages of the embodiments of the present application more clear, the technical scheme of the embodiments of the present application will be described clearly and completely in the following with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all the embodiments. The components of the embodiments of the present application described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed present application, but only represents selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0035] As Figures 1-2 shown, the present application provides a method for detecting minimal residual lesions based on ultra-high depth sequencing, using molecular tags under ultra-high depth sequencing, distinguishing background noise introduced by different NGS links, extracting characteristics according to sources to characterize the background error rate of each link source, estimating the noise level of the sample itself, and developing an algorithm process to eliminate false positive signals without baseline dependence, maximize accurate determination of ctDNA state, and further determine the ctDNA level. The main process includes the following steps:
[0036] 1. Sequencing library preparation and sequencing:
[0037] Sequencing library preparation. In the library preparation process, sequencing adapters are added to nucleic acid molecules (e.g., DNA molecules) by connecting adapters, which include unique molecular identifiers (UMIs). After adapter ligation, adapter-nucleic acid conjugates are amplified by polymerase chain reaction (PCR). During PCR amplification, UMI is copied along with the additional DNA fragment, so that sequencing reads (reads) from the same original fragment can be identified, and can also be used to distinguish the noise level of different sources of NGS. After library preparation, enrichment, hybridization and amplification are performed to provide the sequencing instrument for sequencing.
[0038] 2. Preprocessing and alignment of data:
[0039] Sequencing data is split according to sample index using bcl2fastq software to obtain FASTQ files for testing samples, and then fastp software is used to remove adapters and low-quality reads to obtain filtered FASTQ files; sequence alignment tool BWA mem-M is used for sequence alignment to obtain bam files.
[0040] 3. Obtain consensus sequence by molecular tag
[0041] Read bam files using pysam module, group reads according to UMI tag and alignment to reference genome position, the number of repeated sequencing reads (reads) contained in the group is called the family size of the sequence, and the base sequence on the repeated reads is formed into a consensus sequence by voting. These reads are copied in the experimental and sequencing links, so the consensus sequence also represents the original nucleic acid fragment from which these reads are derived. If the group contains reads from both positive and negative strand templates at the same time, such a consensus sequence is called duplex, and if the consensus sequence is formed only by single template chain reads, i.e., only contains positive or negative chain reads, such a consensus sequence is called simplex.
[0042] 4. Estimate noise level of different NGS steps by hierarchical features:
[0043] Firstly, we excluded high frequency (VAF>=1%) unknown variants, known somatic variants, clonal hematopoiesis variants and germline variants in the process of noise evaluation. The base quality value reflects the error rate of the sequencing step. By base quality value 30, it represents 0.1% error rate. In order to accurately evaluate the background noise below 0.1%, we filtered out the consensus sequences with quality value less than 25 or family size less than 3. The consensus sequence may be consistent with the reference genome sequence, or it may not be consistent. The inconsistent consensus sequence will be used as the basis for subsequent background evaluation.
[0044] Since UMI undergoes library amplification and capture enrichment after being connected to nucleic acid molecules, the base changes at the duplex level are before PCR amplification, which may be derived from real mutations, or from linker connection and DNA damage in the early library stage. The base changes at the simplex level may contain real mutations and errors in PCR amplification, enrichment and capture process. The noise levels of different sources are different, and their noise levels and characteristics also differ, so it is necessary to construct a noise model in layers.
[0045] Studies have shown that most plasma DNA molecules need to be filled in when building a library because they have jagged ends. In this process, the synthesis bias of DNA polymerase and sample damage will cause base changes. The base changes produced in this process are on the original nucleic acid fragment molecules before PCR, and duplex can be used to reflect their characteristics and levels. We combined subtype type and distance from fragment edge as features to estimate the noise level introduced by sample extraction and library construction process on the original nucleic acid molecule chain.
[0046] In the process of PCR amplification, i.e. cluster generation, the signal of the original template is "amplified", and errors are also introduced. Studies have shown that the error types introduced by PCR mainly include mismatches, insertions and deletions (indels), etc., and the background bases upstream and downstream of the sequence are different, and the probability of error occurrence is also different. The trimer on the upstream and downstream of the base, the type of variation (subtype) and the direction of the chain (forward or reverse) will affect the probability of error occurrence, so these factors are included as features to measure the noise level produced in the process of PCR amplification and cluster generation.
[0047] 5. Estimate ctDNA level by combining tumor prior and noise model
[0048] In the sample with high tumor proportion (>=20%), usually surgical tissue, identify variants, subtract germline variants and clonal hematopoiesis variants by paired leukocyte control sample data, determine somatic variants information of tumor origin. Then combine the information of prior samples, classify and quantitatively sort somatic variants, classification categories are main clone and subclone, quantitative indicators are copy number and variant abundance. The proportion of tumor cells, tumor heterogeneity and the difference of genome copy number lead to the difference of variant abundance detection, these variants are filtered and sorted according to their abundance, and the top number of variants is selected as the trackable variant set to construct the tumor prior variant atlas.
[0049] Each variant information of tumor prior atlas is standardized to relative coefficient The calculation formula is as follows:
[0050]
[0051] Among them, refers to the proportion of tumor cells, that is, the proportion of tumor cells in the prior sample; refers to the proportion of circulating tumor, that is, the proportion of tumor molecules in cell-free DNA (cfDNA), is the tumor genome copy number, i refers to the i th variant site.
[0052] The trackable variant set (trackable mutations) includes single base point mutation (SNV), short fragment insertion or deletion mutation (INDEL), and long fragment structure mutation (SV). The technical noise of point mutation, short fragment insertion and deletion mainly comes from different links of NGS, and the change of base level on duplex and simplex molecules is modeled by context, edge distance, chain variation direction and other characteristics; The noise of structure variant fragment mainly comes from repeat sequence or random sequence, which causes the sequencing reads to be incorrectly mapped to different gene positions, and the sequence characteristics of variant direction and split reads (supporting structure variants) are modeled.
[0053] Combined with the depth information of sequencing data and hierarchical background noise model, the expected number of molecules of the i th variant is evaluated The formula is as follows:
[0054]
[0055] Among them, is the expected number of molecules of the i th variant based on the j th background noise model; is the sequencing depth of the i th variant corresponding to the j th background noise model; is the background error rate of the i th variant corresponding to the j th background noise model. Tumor Allele Frequency refers to the i th allele frequency
[0056] The maximum likelihood probability of the number of molecules supporting the variation is calculated by combining multiple variation sites, and the expected circulating tumor proportion is obtained :
[0057] .
[0058] 6. Test the significance of the source of the molecular signal:
[0059] By combining multiple tracking variations with hierarchical background noise models, the probability of the H0 hypothesis that the signal comes from background noise, i.e. , and the probability of the H1 hypothesis that the signal comes from the tumor (i.e. ) are calculated, and then the significance is determined by likelihood ratio test, thereby determining the MRD state.
[0060] The H0 hypothesis probability calculation formula is as follows:
[0061] ;
[0062] The H1 hypothesis probability calculation formula is as follows:
[0063] ;
[0064] In the formula, y represents the actual observed molecular signal of the site; Y represents the expected molecular signal of the site under the condition that cTF=0 or ; j represents three levels of signals, i.e. the positive strand of the simplex level, the negative strand of the simplex level, and the duplex level; X represents the expected molecular signal of each site at each level; n represents the tracked sites; P represents the probability of the H1 hypothesis that the signal comes from the tumor or the H0 hypothesis that the signal comes from background noise.
[0065] The sample level significance test formula is as follows:
[0066] .
[0067] The input files required by the present application include:
[0068] (1) sequencing data of high tumor proportion (>=20%) samples;
[0069] (2) sequencing data of control samples;
[0070] (3) sequencing data of postoperative or post-treatment plasma samples.
[0071] The output results of the present application include:
[0072] (1) Consensus sequence based on UMI tag and reference genome same alignment position;
[0073] (2) Inconsistent sequence based on UMI tag and reference genome same alignment position;
[0074] (3) Noise file based on position effect and different subtype characteristics (duplex level);
[0075] (4) Noise file based on chain direction, context base and variation direction (simplex level);
[0076] (5) Variation map information file, including information of these variations in white blood cells, tissues and plasma samples;
[0077] (6) ctDNA level estimated by combining prior variation map and hierarchical noise model;
[0078] (7) Significance level tested by combining prior variation map and hierarchical noise model.
[0079] The noise introduced at different stages of NGS data is analyzed hierarchically, the noise sources and their characteristic levels are compared and analyzed, which is used for subsequent noise reduction and MRD detection of tracking sites. Compared with the previous research on noise evaluation, the method is more detailed and accurate, and can improve the sensitivity and specificity of MRD detection.
[0080] In addition, the present application fully utilizes the information of clonal evolution in the development process of tumor, in addition to the allele frequency of tumor tissue, the tumor mutant allele frequency, copy number and clonality information are also included, which can more accurately estimate the proportion of ctDNA in plasma and improve the performance of MRD detection.
[0081] The present application first uses plasma leukocyte subtraction to eliminate the false positive sites introduced by clonal hematopoiesis, thereby improving the specificity of MRD detection. Then, the noise from different sources in the NGS sequencing process is analyzed hierarchically, which can more accurately reflect the noise source level and characteristics, and can eliminate the false positive rate caused by noise introduced in each link as much as possible without losing too much true positive signal. On this basis, combined with tumor tissue information, including mutant allele frequency, tumor proportion, copy number and clonal information, etc. Compared with the method of assuming that the related mutation in tumor tissue is single copy and the tumor proportion in tissue is 100%, the present application can more accurately reflect the tumor proportion in plasma. Combined with multiple tracking sites, the hierarchical background noise is used to test the significance level of signals in plasma derived from tumor tissue, which has very good effect on distinguishing the source of signals in plasma.
[0082] The computer program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other processing device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other processing device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions specified in the flowchart block or blocks or in conjunction with the flowcharts and / or block diagram. Figure 1 one or more of the blocks or steps in the flowchart or flow diagram. Figure 1 one or more of the blocks or steps in the flowchart or flow diagram.
[0083] In the description of the present application, it is to be understood that the terms "first", "second", "third" and the like, merely identify features belonging to distinct categories, and do not imply or imply a relative importance or a specific number thereof. Thus, a feature identified as "first" or "second" can implicitly or explicitly include one or more of the features. In the description of the present application, the meaning of "a plurality" is two or more, unless otherwise expressly specified.
[0084] Obviously, various modifications and changes can be made to the present application by those skilled in the art without departing from the spirit and scope of the present application. Thus, it is intended that the present application also include such modifications and changes insofar as they come within the scope of the claims of the present application and their equivalents.
Claims
1. A method for detecting minimal residual lesions based on ultra-high depth sequencing, characterized in that, Includes the following steps: Based on the sequencing library, the sample to be tested is obtained by splitting according to the sample index, and a consensus sequence is obtained according to the molecular identifier of the sample to be tested; Based on the consensus sequence, by filtering out consensus sequences with a quality value less than 25 or a family size less than 3, and combining the subtype and distance from the fragment edge, the noise level introduced into the original nucleic acid molecule chain is determined. Based on the noise level and combined with prior tumor knowledge, the ctDNA level is estimated, and the MRD status is determined by examining the significance of the molecular signal source. The process of estimating ctDNA levels by combining tumor priors and noise model tests includes: In the process of combining prior knowledge of tumors, variants are identified in samples with a tumor proportion of >=20%. Germline variants and clonal hematopoietic variants are subtracted from paired leukocyte control sample data to determine the somatic variant information of tumor origin. Then, combined with the information of prior samples, somatic variants are classified, clustered and quantitatively ranked. The classification categories are principal clones and subclones, and the quantitative indicators are copy number and variant abundance. The differences in the proportion of tumor cells, tumor heterogeneity and genome copy number lead to the differences in the abundance of detected variants. These variants are filtered and ranked according to their abundance, and the top number of variants are selected as the set of traceable variants to construct a tumor prior variant map. Each variant in the tumor prior map is standardized into a relative coefficient. The calculation formula is as follows: ; in, It refers to the proportion of tumor cells, that is, the proportion of tumor cells in the prior sample; This refers to the circulating tumor percentage, which is the proportion of tumor molecules in cell-free DNA. , where i represents the tumor genome copy number, and i refers to the i-th mutation site; The traceable variant set includes single-base point mutations, short-fragment insertion or deletion mutations, and long-fragment structural mutations. Among them, the noise on the original strand of point mutations and short-fragment insertions or deletions is estimated by subtype and edge distance level; the noise introduced by PCR amplification capture is modeled and evaluated based on the characteristics of context and strand mutation direction. Combining depth information from sequencing data with a hierarchical background noise model, through evaluation The expected number of molecules is given by the following formula: ; in, It is the expected number of molecules for the i-th mutation based on the j-th layer background noise model; It is the sequencing depth of the background noise model corresponding to the i-th mutation in layer j; It is the background error rate of the background noise model at layer j corresponding to the i-th mutation; This refers to the frequency of the i-th allele; The expected percentage of circulating tumor cells is obtained by calculating the maximum likelihood probability of the number of molecules supporting the variant by combining multiple variant sites. : 。 2. The method for detecting minimal residual lesions based on ultra-high depth sequencing according to claim 1, characterized in that: In the process of obtaining the sequencing library, sequencing adapters are added to nucleic acid molecules by ligating adapters, and the adapter-nucleic acid conjugate is amplified by polymerase chain reaction to establish the sequencing library, wherein the adapter is a molecular identifier.
3. The method for detecting minimal residual lesions based on ultra-high depth sequencing according to claim 2, characterized in that: In the process of obtaining the molecular identifier of the sample to be tested, the adapters and low-quality sequencing reads of the sample to be tested are removed, and the molecular identifier is obtained after sequence alignment.
4. The method for detecting minimal residual lesions based on ultra-high depth sequencing according to claim 3, characterized in that: In the process of obtaining consensus sequences, the sequencing reads are grouped according to their molecular identifiers and their positions aligned to the reference genome. The base sequences on the repetitive sequencing reads contained in different groups are then voted on to form the consensus sequences.
5. The method for detecting minimal residual lesions based on ultra-high depth sequencing according to claim 4, characterized in that: In assessing noise levels, sequencing read pairs that simultaneously contain both positive and negative templates are designated as the first consensus sequence duplex; and consensus sequences formed solely by sequencing read pairs from a single template strand, i.e., containing only positive or negative sequencing reads, are designated as the second consensus sequence simplex. The noise level is assessed based on the first consensus sequence duplex and the second consensus sequence simplex.
6. The method for detecting minimal residual lesions based on ultra-high depth sequencing according to claim 5, characterized in that: In assessing noise levels, after excluding unknown variants with allele frequencies >= 1%, tumor somatic variant sites of known origin, clonal hematopoietic variants, and germline variants, the noise level introduced onto the original nucleic acid molecular chain is assessed based on the first consensus sequence duplex combined with subtype and distance from the fragment edge; the second consensus sequence simplex, combined with context and chain specificity, is used to assess PCR-level noise.
7. The method for detecting minimal residual lesions based on ultra-high depth sequencing according to claim 6, characterized in that: In determining the MRD status, based on the depth information of the sample to be tested, and according to the evaluation models used to assess the first noise and the second noise, the maximum likelihood probability of the number of supporting molecules under multiple variants is calculated to estimate the ctDNA level, obtain the expected number of molecules to assess the proportion of different circulating tumors and calculate the probability, and determine the significance through the likelihood ratio test, thereby determining the MRD status.
8. A system for detecting minimal residual lesions based on ultra-high depth sequencing, characterized in that, This system is used to implement a method for detecting minimal residual lesions based on ultra-high depth sequencing as described in any one of claims 1-7. The system comprises: The data processing unit is used to split the sequencing library according to the sample index to obtain the test sample, and to obtain the consensus sequence according to the molecular identifier of the test sample; The noise level assessment unit is used to estimate the noise level on the original nucleic acid molecular chain based on the consensus sequence by filtering out consensus sequences with a quality value less than 25 or a family size less than 3, and by combining subtype and distance from the fragment edge; and to assess the noise at the PCR level by combining context and chain specificity. The detection unit is used to estimate the ctDNA level based on the noise level and in conjunction with prior tumor knowledge, and to determine the MRD status by examining the significance of the molecular signal source. The process of estimating ctDNA levels by combining tumor priors and noise model tests includes: In the process of combining prior knowledge of tumors, variants are identified in samples with a tumor proportion of >=20%. Germline variants and clonal hematopoietic variants are subtracted from paired leukocyte control sample data to determine the somatic variant information of tumor origin. Then, combined with the information of prior samples, somatic variants are classified, clustered and quantitatively ranked. The classification categories are principal clones and subclones, and the quantitative indicators are copy number and variant abundance. The differences in the proportion of tumor cells, tumor heterogeneity and genome copy number lead to the differences in the abundance of detected variants. These variants are filtered and ranked according to their abundance, and the top number of variants are selected as the set of traceable variants to construct a tumor prior variant map. Each variant in the tumor prior map is standardized into a relative coefficient. The calculation formula is as follows: ; in, It refers to the proportion of tumor cells, that is, the proportion of tumor cells in the prior sample; This refers to the circulating tumor percentage, which is the proportion of tumor molecules in cell-free DNA. , where i represents the tumor genome copy number, and i refers to the i-th mutation site; The traceable variant set includes single-base point mutations, short-fragment insertion or deletion mutations, and long-fragment structural mutations. Among them, the noise on the original strand of point mutations and short-fragment insertions or deletions is estimated by subtype and edge distance level; the noise introduced by PCR amplification capture is modeled and evaluated based on the characteristics of context and strand mutation direction. Combining depth information from sequencing data with a hierarchical background noise model, through evaluation The expected number of molecules is given by the following formula: ; in, It is the expected number of molecules for the i-th mutation based on the j-th layer background noise model; It is the sequencing depth of the background noise model corresponding to the i-th mutation in layer j; It is the background error rate of the background noise model at layer j corresponding to the i-th mutation; This refers to the frequency of the i-th allele; The expected percentage of circulating tumor cells is obtained by calculating the maximum likelihood probability of the number of molecules supporting the variant by combining multiple variant sites. : 。
Citation Information
Patent Citations
Ultra-high depth sequencing-based tiny residual focus detection method and system
CN119252324A