DNA-based adaptome profiling for minimal residual disease quantification in lymphoid malignancies

The method improves MRD detection in lymphoid malignancies by employing multiplex PCR and sequencing of T and B cell receptor gene rearrangements, addressing sensitivity limitations and enabling accurate quantification of residual disease.

WO2025207143A1PCT designated stage Publication Date: 2025-10-02MILABORATORIES INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/047671
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-25
Filing Date
2024-09-20
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

Current methods for detecting minimal residual disease (MRD) in lymphoid malignancies are limited by low sensitivity, especially in leukemias with unstable cell immunophenotypes, and existing DNA-based tests are restricted to detecting chromosomal translocations, lacking comprehensive detection of leukemic cells.

Method used

A method involving multiplex PCR and high-throughput sequencing of T and B cell receptor gene rearrangements, including incomplete rearrangements, using a designed oligonucleotide library to minimize primer dimer formation and correct for amplification bias, followed by semi-global alignment and clustering algorithms to accurately quantify MRD.

Benefits of technology

Enhances MRD detection sensitivity, enabling precise quantification of leukemic cells down to a low frequency, overcoming limitations of existing tests and providing reliable monitoring of disease status during and after therapy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024047671_02102025_PF_FP_ABST
    Figure US2024047671_02102025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to target sequencing of T and B cell receptor gene rearrangements at the DNA level and using this technology to detect and quantify lymphoid malignant T or B cells (minimal residual disease, MRD) during and after therapy.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] DNA-B ASED ADAPTOME PROFILING FOR MINIMAL RESIDUAL DISEASE QUANTIFICATION IN LYMPHOID MALIGNANCIES

[0002] CROSS-REFERENCE TO RELATED APPLICATIONS

[0003] This application claims priority to and the benefit of United States utility patent application no. 18 / 615,441, filed March 25, 2024, entitled DNA-B ASED ADAPTOME PROFILING FOR MINIMAL RESIDUAL DISEASE QUANTIFICATION IN LYMPHOID MALIGNANCIES, which is incorporated herein in its entirety.

[0004] FIELD

[0005] The present disclosure is directed to methods for determining the presence and quantification of minimal residual disease (MRD) level in patients. The methods incorporate multiplex amplification and sequencing of T and B cell receptor gene rearrangements, including incomplete rearrangements, to detect lymphoid malignant T or B cells related to MRD during and after therapy.

[0006] SEQUENCE LISTING

[0007] This application incorporates by reference a Sequence Listing submitted with the application as XML file entitled 15930.10009W001, created on September 17, 2024, and having a size of 61,964 bytes.

[0008] BACKGROUND

[0009] Minimal residual disease (MRD) refers to a small amount of cancer cells remaining in a patient following treatment that cannot be detected using standard cancer scans or laboratory tests. As MRD is likely to lead to relapse, its prompt detection and determining the MRD level (quantification, concentration of leukemic cells) is important for relapse prediction and therapy outcome evaluation.

[0010] Immunological based testing for surface proteins on white blood cells can be used for MRD testing in leukemias and lymphomas. However, these tests have a limit of detection of around 1 in 10,000 cells and can only be used in detecting leukemias with a stable cell immunophenotype.

[0011] Some current MRD tests are based on detecting a DNA sequence specific to the presence of the cancer. DNA markers tested for include chromosomal translocations, microsatellites, point mutations, immunoglobulin or T cell receptor sequences. Additional sequencing tests are RNA based, but these tests are essentially limited to detecting chromosomal translocations.

[0012] SUMMARY

[0013] In aspects, the disclosure provides a method of determining the presence and quantification of minimal residual disease (MRD) in patients with hematological malignancies, the method comprising: (a) detecting rearrangements of T cell receptor genes and B cell receptor genes characteristic for malignant clones in an initial sample from a patient using: a multiplex polymerase chain reaction (PCR) with isolated genomic DNA wherein the PCR reaction comprises a library of oligonucleotides amplifying one or more of the following: TRa / p / y / 5, IgH / K / X, DJ, DD, VD, and kappa-deleting element (KDE) rearrangements in a single multiplex mixture of oligonucleotides; high-throughput sequencing of the obtained PCR products; extracting complete and incomplete rearrangements by mapping potential rearrangements to a library comprising V, D, and J gene segments, kappa deletion element (KDE), and IGKC intron sequences, wherein the extraction is performed using a semi-global alignment algorithm to identify flanking sequences, followed by a clustering algorithm correcting for PCR and sequencing errors to assemble complete and / or incomplete rearrangement clonotypes; determining precise clonotype sizes from hematological malignancy -related repertoire; and malignancy -related clonotype detection; and (b) performing follow-up monitoring of MRD in at least 4 independent PCR reactions with isolated and quantified genomic DNA obtained at a follow-up time point; wherein the concentration of a malignant clone in the follow-up time point is determined based on the proportion of independent PCR reactions where the corresponding rearrangement characteristic for the malignant clone is observed.

[0014] In aspects of the method, the library of oligonucleotides for multiplex PCR is designed to decrease primer dimer formation and minimize nucleotide diversity.

[0015] In aspects of the method, the library of oligonucleotides is designed by a method comprising: a. dividing target genomic DNA regions into k-mers of a certain length; b. clustering the target genomic DNA regions using Hamming distances not exceeding 4, not exceeding 5, or not exceeding 6, or not exceeding 7 nucleotides; and c. identifying clusters of highly similar k-mers according to the selected Hamming distance. In aspects of the method, the k-mers have a length of at least 16 bases.

[0016] In aspects of the method, from the clusters identified, tables of 5-mers in the reverse orientation and 5-mers at the end of 3' end of the nucleotide sequences in the forward orientation are made. In aspects, the 5-mers tables are intersected pairwise, and the primer set with the lowest 5-mers overlap is selected for further analysis.

[0017] In aspects of the method, to equilibrate the annealing temperature of the primers, up to 14 additional target-site-related nucleotides are added to the 5'-end of 18-mers.

[0018] In aspects of the method, the intersection procedure using 5-mers is repeated.

[0019] In aspects of the method, at least 2 working sets of forward primers and at least 2 working sets of reverse primers are generated.

[0020] In aspects of the method, the reverse primer set comprises primers for J-genes, downstream introns for D-genes and KDE. In aspects, the forward primer set comprises primers for V-genes, upstream introns for D-genes and C-intron for IgK. In aspects, the multiplex primer sets generated can be used for separate amplification of complete VJ and / or VDJ rearrangements at TRa, TRP, TRy, TR5, IgH, IgK, IgL loci; partial DJ rearrangements at TRP, TR5, IgH loci; VD and DD rearrangements at TR5 and TRP loci; chimeric TRDV-TRAJ rearrangements; or Kappa deletion rearrangements.

[0021] In aspects of the method, dual indexing of each aliquot amplicon and introduction of adapters for sequencing are performed using "step-out" PCR.

[0022] In aspects of the method, the average lengths of target amplicons in the 4 multiplex primer sets are: set 1 - about 120 bp, set 2 - about 160 bp, set 3 - about 240 bp, and set 4 - about 280 bp.

[0023] In aspects of the method, amplicons obtained with a set of primers cannot be a matrix for PCR with the subsequent set of primers.

[0024] In aspects of the method, the concentrations of primers in the multiplex mixture are optimized by analyzing the frequency ratios of non-functional rearrangements of T- and B-cell receptor genes.

[0025] In aspects of the method, an overamplification rates (OAR) is determined for at least one rearrangement, and the concentration of the primer detecting the rearrangement is adjusted to correct for the amplification bias including overamplification and underamplifi cation). In aspects, an overamplification rates (OAR) is determined for at least one rearrangement to computationally correct for the amplification bias including overamplifi cation and underamplification. In aspects of the method, a ratio of weighted frequencies (proportion of reads) to unweighted frequencies (proportion of clonotypes) for V, D and J genes is determined. In aspects, if the ratio is a substantial deviation from 1, it is indicative of quantitative bias occurring during amplification.

[0026] In aspects of the method, the concentrations of primers in the multiplex mixture is selected to minimize the deviation from the expected value of one for the ratio of weighted frequencies to unweighted frequencies.

[0027] In aspects of the method, the MRD detection sensitivity depends exclusively on the available input DNA and sequencing coverage.

[0028] In aspects of the method, TRa / p / y / 5, IgH / K / X, DJ, DD, VD, and kappa-deleting element (KDE) rearrangements are detected in a single multiplex mixture of oligonucleotides.

[0029] BRIEF DESCRIPTION OF THE DRAWINGS

[0030] FIG. l is a schematic showing positions of designed multiplex primer sets relative to each other in rearranged TCR and BCR genes as described in Example 1.

[0031] FIG. 2 is a plot of sample intensity vs. size for a capillary electrophoresis separation of ready-for- sequencing amplicons (target amplicon + sequencing adapters) as described in Example 1. Ready- for-sequencing amplicons were obtained via PCR using four combinations of designed multiplex primer sets: set2+set3, set2+set4, setl+set3, setl+set4.

[0032] FIGs. 3A-M are plots of the frequency of rearrangements detected using the 7GENES multiplex mix as described in Example 2. The average usage of each gene segment (black columns) and proportion of samples in which the corresponding gene segment was identified (gray columns). Next-generation sequencing (NGS) data was obtained with the 7GENES multiplex mix from 190 human PBMC samples and are shown for IGHV, IGHD, IGKV, IGLV, TRB V and TRBD, TRGV,

[0033] TRDV, TRDD, IGHJ, IGKJ, IGLJ, TRAJ, TRBJ, and TRDJ gene segments as indicated in FIGs. 3 A-3M.

[0034] FIG. 4 is a plot showing the correlation between expected and observed concentration of leukemic cells in samples with added spike-in as described in Example 3. r - Pearson’s correlation coefficient, R2 - squared r. Black dots represent the mean value for MRD calculated using three different markers, whiskers show standard deviation. FIG. 5 shows a plot of the percentage of samples [log 10] having the MRD value indicated, in percent, as described in Example 6.

[0035] FIGs. 6A and 6B show plots of TRAV (6A) and TRAJ (6B) overamplification rates before (black) and after (grey) optimization of PCR primer concentrations as described in Example 7.

[0036] FIGs. 7A and 7B show plots of TRAV (7 A) and TRAJ (7B) overamplification rates before (black) and after (grey) residual amplification bias correction using a computational approach as described in Example 7.

[0037] DETAILED DESCRIPTION

[0038] The present disclosure provides improved methods for deterring minimal residual disease (MRD) using adaptome profiling. In aspects, the methods are used for detecting MRD in patients with hematological malignancies, e.g., blood cancers.

[0039] In this specification and the appended claims, the singular forms "a," "an" and "the" include plural reference unless context clearly dictates otherwise. Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood to a person of ordinary skill in the art in the field.

[0040] Where a range of values is provided, it is understood that each intervening value, to the tenth of the unit of the lower limit unless the context clearly dictates otherwise, between the upper and lower limit of that range, and any other stated or intervening value in that stated range, is encompassed within the disclosure. The upper and lower limits of these smaller ranges may independently be included in the smaller ranges, and are also encompassed within the disclosure, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the disclosure. Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood to one of ordinary skill in the art to which this disclosure belongs. Although any methods, devices and materials similar or equivalent to those described herein can be used in the practice or testing of the disclosure, the preferred methods, devices and materials are now described.

[0041] All publications and patents cited in this specification are herein incorporated by reference as if each individual publication or patent were specifically and individually indicated to be incorporated by reference. The citation of any publication is for its disclosure prior to the filing date and should not be construed as an admission that the present disclosure is not entitled to antedate such publication by virtue of prior disclosure.

[0042] It should be understood that the materials and methods proposed herein are not limited to particular compositions or process steps, as they may vary. It is pointed out that, as used in this specification and the appended claims, singular forms include the corresponding plural forms, unless the context clearly dictates otherwise.

[0043] The term "antigen-recognizing receptor" as used herein is equivalent to "immune receptor" and refers to T cell receptor (TCR) and B cell receptor (BCR).

[0044] The following abbreviations may be used when referring to genes for which rearrangements are detected: immunoglobulin heavy chain variable gene segment (IGHV); immunoglobulin heavy chain diversity gene segment (IGHD); immunoglobulin kappa variable gene segment (IGKV); immunoglobulin lambda variable gene segment (IGLV); T-cell receptor beta chain variable gene segment (TRBV); T-cell receptor beta chain diversity gene segment (TRBD); T-cell receptor alpha chain variable gene segment (TRAV); T-cell receptor gamma chain variable gene segment (TRGV); T-cell receptor delta chain variable gene segment (TRDV); T-cell receptor delta chain diversity gene segment (TRDD); immunoglobulin heavy chain joining gene segment (IGHJ); immunoglobulin kappa chain joining gene segment (IGKJ); immunoglobulin lambda chain joining gene segment (IGLJ); T-cell receptor alpha chain joining gene segment (TRAJ); T-cell receptor beta chain joining gene segment (TRBJ); T-cell receptor gamma chain joining gene segment (TRGJ); and T-cell receptor delta chain joining gene segment (TRDJ).

[0045] The term "immune receptor repertoires" as used herein refers to TCR and / or BCR repertoires. The term "BCR" as used herein refers to B cell receptors, antibodies, or immunoglobulins.

[0046] As used herein, the term "CDR" refers to complementarity determining regions - the regions of the antigen-recognizing receptors located in the variable domains of the polypeptide chains of TCRs or BCRs that largely determine the specificity of immune receptor. Variable segment of each chain of the antigen-recognizing receptor consists of three CDRs and four framework regions (FR) located from the amino- to carboxyl terminus in the following order: FR1, CDR1, FR2, CDR2, FR3, CDR3, FR4.

[0047] The terms "patient" and "individual" refer to a vertebrate, in particular to a representative of mammalian species, and include, but are not limited to, pets, sports animals, primates, including humans. In specific aspects, the patients are humans. As used herein, the term "cohort of individuals" refers to a group of patients of the same species suffering from the same disease, or individuals undergoing the same type of therapy or vaccination. In specific aspects, the term refers to a group of people.

[0048] As used herein, the term "experimental data" refers to the sequencing data of biological samples obtained from individuals belonging to the analyzed cohort.

[0049] As used herein, the term "biological sample" refers to a sample of peripheral blood, a sample of cells or a tissue sample, for example, a biopsy or puncture material taken from a patient, and to cell cultures derived from patient cells.

[0050] In some aspects, the cells of interest are cultured and differentiated in vitro before use.

[0051] The term "gene segments" is used to describe the segments of the genes that are involved in the generation of TCRs and BCRs.

[0052] The term "nucleotide sequences" refers to nucleic acid sequences, including DNA, such as genomic DNA or a cDNA molecule, or RNA. As used herein, the term "cDNA" refers to nucleic acids that have sequence elements complementary to the native mature mRNA species, where the sequence elements are exons and the 5' and 3' non-coding regions.

[0053] As used herein, the term "sample preparation" refers to all manipulations to which nucleic acids are subjected after purification and before sequencing.

[0054] The length of a nucleotide sequence is the number of nucleotides it consists of. The “Hamming distance” is a measure of similarity for two nucleotide sequences with the same length. It is equal to the number of nucleotide substitutions required to change a given nucleotide sequence to target one.

[0055] For the purposes of the present disclosure, the length of the compared sequences may be the same as the lengths of the CDR3 regions of the TCR or BCR chains being analyzed, or may be extended to the portion of or to the whole sequences of V and / or J gene segments, or may represent a portion of CDR3.

[0056] Sequence comparison is performed by methods known to those skilled in the art. For example, the algorithm described in [Altschul et al. J. Mol. Biol., 215, pp. 403-10 (1990)] may be employed for sequence comparison. As another example, to determine the level of identity and similarity between nucleotide or amino acid sequences, the Blast software package provided by National Center for Biotechnology Information (ncbi.nlm.nih.gov / blast). Reference to the nucleotide sequence coding amino acid sequence means that this amino acid sequence is produced from the nucleotide sequence during translation of mRNA. As is obvious to any person skilled in the art, the term also includes degenerate nucleotide sequences encoding the same amino acid sequence.

[0057] Immunoglobulins consist of heavy chain immunoglobulins (IgH) with constant regions (a, 5, a, y, or p) or light chain immunoglobulins (IgK or IgL) with constant regions or K. An antibody has two identical light chains and two identical heavy chains. Each chain is composed of a constant (C) and a variable region. For the heavy chain, the variable region is composed of a variable (V), diversity (D), and joining (J) segments. Several distinct sequences coding for each type of these segments are present in the genome. A specific VDJ recombination event occurs during the development of aB-cell, marking that cell to generate a specific heavy chain. Diversity in the light chain is generated in a similar fashion except that there is no D region so there is only VJ recombination. Between the joined segments, random addition or deletion of one or several nucleotides may occur, further increasing the diversity of heavy and light chains generated by naive B-cells. The possible diversity of the antibodies generated by B-cells is then the product of the different heavy and light chains. The variable regions of the heavy and light chains contribute to form the antigen recognition (or binding) region or site. Added to this diversity is a process of somatic hypermutation of VDJ regions, which can occur after a specific response is mounted against some epitope, in germinal centers of lymph nodes and other secondary lymphoid tissues.

[0058] A "T cell receptor", also referred to as "TCR", is a heterodimeric protein complex located on the T lymphocyte surface. This receptor is present only on T lymphocytes. The main function of TCR is the specific recognition of processed antigens represented by molecules of the main histocompatibility complex (MHC, or HLA). Some TCRs may recognize “nonclassical” MHC molecules, such as CD Id, CDlc, MR1, and others, that may present lipid or other low molecular weight molecules.

[0059] A human TCR consists of two subunits, TCRa (TCR alpha) and TCRP (TCR beta), or TCRy (TCR gamma) and TCRS (TCR delta) chains, interconnected by a disulfide bond and presented on the T cell membrane. Each of the TCR chains has an N- terminal variable (V) domain, a junction domain (J) and a constant (C) domain coupled to a transmembrane domain that fixes the receptor in the plasma membrane of the T lymphocyte. The TCR usually and mostly interacts with the MHC antigen complex with six complementarity determining regions (CDRs): three alpha chain and three beta chain regions. In some cases, TCRa chain interaction with antigen dominates. More usually, TCRP chain interaction with antigen dominates.

[0060] In addition to complete V-J and V-D-J rearrangements, the genomes of T and B cells can contain partial (incomplete) D-J, V-D, and D-D rearrangements and rearrangements with kappa deleting element (KDE). These rearrangements cannot generate functional TCR and BCR. Junctions of gene segments forming partial rearrangements similar to CDR3 have a high level of diversity due to random non-template insertions and deletions.

[0061] Adaptome (adaptive immunome, VDJ repertoire, immune repertoire) is defined as a set of all TCR and BCR loci rearrangements (complete and partial) identified in the sample. DNA adaptome or DNA-based adaptome can be referred to as an adaptome obtained using DNA as a template for sequencing library preparation.

[0062] Methods For Detection ofMRD

[0063] In aspects, disclosed herein is a method of determining the presence and quantification of minimal residual disease (MRD) in patients with hematological malignancies, the method comprising:

[0064] (a) detecting rearrangements of T cell receptor genes and B cell receptor genes characteristic for malignant clones in an initial sample from a patient using: a multiplex polymerase chain reaction (PCR) with isolated genomic DNA wherein the PCR reaction comprises a library of oligonucleotides amplifying one or more of the following: TRa / p / y / 5, IgH / K / k, DJ, DD, VD, and kappa-deleting element (KDE) rearrangements in a single multiplex mixture of oligonucleotides; high-throughput sequencing of the obtained PCR products; extracting complete and incomplete rearrangements by mapping potential rearrangements to a library comprising V, D, and J gene segments, kappa deletion element (KDE), and IGKC intron sequences, wherein the extraction is performed using a semi-global alignment algorithm to identify flanking sequences, followed by a clustering algorithm correcting for PCR and sequencing errors to assemble complete and / or incomplete rearrangement clonotypes; determining precise clonotype sizes from hematological malignancy-related repertoire; and malignancy-related clonotype detection; and (b) performing follow-up monitoring of MRD in at least 4 independent PCR reactions with isolated and quantified genomic DNA obtained at a followup time point; wherein the concentration of a malignant clone in the followup time point is determined based on the proportion of independent PCR reactions where the corresponding rearrangement characteristic for the malignant clone is observed.

[0065] In aspects of the method, the hematological malignancy is a leukemia, a lymphoma or a multiple myeloma. In aspects, the leukemia is Acute lymphocytic leukemia, Chronic lymphocytic leukemia, Chronic myeloid leukemia, Acute myeloid leukemia, Hairy cell leukemia, Acute promyelocytic leukemia, T-cell prolymphocytic leukemia, Lymphoid leukemia, Large granular lymphocytic leukemia, or B-cell prolymphocytic leukemia. In aspects, the lymphoma is Non-Hodgkin lymphoma, Hodgkin's lymphoma, Diffuse large B-cell lymphoma, Follicular lymphoma, Burkitt lymphoma, Mantle cell lymphoma, Waldenstrom macroglobulinemia, Peripheral T-cell lymphoma, B-cell lymphoma, Nodular lymphocyte predominant Hodgkin lymphoma, B-cell chronic lymphocytic leukemia, Cutaneous B-cell lymphoma, Angioimmunoblastic T-cell lymphoma or Central nervous system lymphoma. In aspects, the multiple myeloma is Light Chain Myeloma, Non-secretory Myeloma, Solitary Plasmacytoma, Extramedullary Plasmacytoma, Monoclonal Gammopathy of Undetermined Significance, Smoldering Multiple Myeloma, Immunoglobulin D (IgD) Myeloma or Immunoglobulin E (IgE) Myeloma.

[0066] In aspects of the method, the method is performed at the onset of the hematological malignancy in the patient. In aspects, the method is performed during a relapse of the hematological malignancy in the patient. In aspects, the method is performed following treatment of the hematological malignancy. In aspects, the method is performed during treatment of the hematological malignancy. In aspects, the method is performed when the patient is thought to be in remission as determined by a cancer scan or standard cancer laboratory test.

[0067] In aspects of the method, the sample is bone marrow, blood, plasma, tissue (including formalin- fixed paraffin-embedded (FFPE) tissue samples). DNA may be isolated from cells or may be cell- free DNA.

[0068] In aspects of the method, extracting the complete and incomplete (i.e., partial) rearrangements is performed using a semi-global alignment (or glocal) algorithm. In general, a semi-global alignment is a variant of global alignment that allows for gaps at the beginning and / or the end of one of the sequences. Semi-global alignment algorithms search for the best possible alignment between sequences. Semi-global alignment algorithms known in the art can be used in the methods described herein. In aspects, the semi-global alignment algorithm is MiXCR as described in Bolotin et al. (Nat Methods 12, 380-381 (2015) / / doi.org / 10.1038 / nmeth.3364).

[0069] In aspects of the method, precise clonotype sizes are determined using nonfunctional (e.g., non- productive or out-of-frame) clonotypes for quantitative bias correction. Non-functional clonotypes arise from non-functional rearrangements. As out- of-frame or stop codon containing TCR / BCR rearrangements do not form a functional receptor, they are not subjected to any specific clonal expansions and selection (Murugan et al., 2012, PNAS 109: 16161-16166). As a passenger genomic variation, these rearrangements change their initial (recombinational) clonal frequencies randomly following the frequency changes of the second functional (in-frame) TCR / BCR allele present in the same T / B cell clone. According to the TCR / BCR loci rearrangement mechanism, the formation of in-frame and out-of-frame allele combinations in the same cell is also a stochastic and independent process in terms of V- and J-genes frequency. Thus, V- and J-gene frequencies among out-of-frame rearrangements must be sufficiently stable and must be equal to the initial recombination frequencies. Reproducible deviation of out- of-frame V- and J-gene frequencies (for the same multiplex PCR primer set) from the initial recombinational frequencies observed in the sequenced repertoire dataset is determined to be a result of artificial aberration caused by PCR amplification rather than immune repertoire evolution. Thus out-of-frame clonotypes can be considered a natural calibrator that can be used to measure amplification bias and quantitatively correct immune repertoire data. Examples of algorithms that can perform this quantitative bias correction are known in the art. In aspects, the quantitative bias correction is performed using the multiplex PCR-specific bias evaluation and correction algorithm iROAR: immune Repertoire Over Amplification Removal (github.com / smiranast / iROAR; Smirnova et al., 2023, eLife 12:e69157). In other aspects, the quantitative bias correction is performed using a DNA spike in method such as that described by Carlson et al. (Nat Commun 4, 2680 (2013), doi.org / 10.1038 / ncomms3680).

[0070] In aspects of the method, follow-up monitoring is performed at a follow-up timepoint. In aspects, the follow-up timepoint is about 1, 2, 3, 4, 5, 6, 7, or 8 weeks following the initial detection. In aspects, the follow-up timepoint is about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 or 12 months following the initial detection. In aspects, the follow-up timepoint is about 1, 2 or 3 years following the initial detection. In aspects of the method, follow-up monitoring is performed in at least 4 independent PCR reactions. In aspects, follow-up monitoring is performed in at least 6 independent PCR reactions. In aspects, follow-up monitoring is performed in at least 8 independent PCR reactions. In aspects, follow-up monitoring is performed in at least 10 independent PCR reactions. In aspects, follow- up monitoring is performed in at least 12 independent PCR reactions. In aspects, follow-up monitoring is performed in at least 14 independent PCR reactions. In aspects, follow-up monitoring is performed in at least 16 independent PCR reactions.

[0071] In aspects of the method, the minimal frequency of a malignant clone in the followup time point is determined based on the proportion of independent PCR reactions where the corresponding rearrangement characteristic for the malignant clone is observed. In aspects, the minimal frequency is determined as follows: the number of positive signals is divided by the number of analyzed diploid human genomes, whereas the number of analyzed diploid human genomes is calculated as total amount of analyzed DNA in all analyzed samples divided by 6.5 pg - the weight of one diploid human genome.

[0072] Without wishing to be bound by theory, in aspects of the method, the number of aliquots containing malignancy-related clonotypes is roughly comparable to the number of detected malignant cells. Thus, in aspects of the method, it is possible to count positive and negative aliquots. The determined portion of positive aliquots is then estimated to be the portion of the total cells assayed that are detected as malignant cells.

[0073] In aspects, the method further comprises determining the presence and minimal number of detected malignant cells in the patient if the corresponding rearrangement characteristic for the malignant clone is observed in all independent PCR reactions.

[0074] In aspects, the methods disclosed herein use a designed library of oligonucleotides for multiplex PCR. In aspects, the library of oligonucleotides for multiplex PCR is designed to decrease primer dimer formation and minimize nucleotide diversity.

[0075] In aspects, the oligonucleotide library is designed by a method comprising: a. dividing target genomic DNA regions into k-mers of a certain length; b. clustering the target genomic DNA regions using Hamming distances not exceeding 4, not exceeding 5, or not exceeding 6, or not exceeding 7 nucleotides; and c. identifying clusters of highly similar k-mers according to the selected Hamming distance. In aspects, the k-mers have a length of at least 16 bases. In aspects, the k-mers have a length of at least 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24 or 25 bases.

[0076] In aspects, tables of 5-mers in the reverse orientation and 5-mers at the end of 3' end of the nucleotide sequences in the forward orientation are made from the clusters identified. In aspects, the 5-mers are intersected pairwise, and the primer set with the lowest overlap with the 5-mers is selected for further analysis. In aspects, this intersection procedure is repeated. In aspects, the intersection procedure is performed in an iterative manner.

[0077] In aspects of the method, after a set of primers is determined, to equilibrate the annealing temperature of the primers, up to 14 additional target- site-related nucleotides are added to the 5'- end of the k-mers. In aspects, up to 12, 13, 14, 15, 16, 17, 18, 19 or 20 additional target-site-related nucleotides are added to the 5'-end of the k-mers.

[0078] In aspects of the method, 2 working sets of forward primers and 2 working sets of reverse primers are generated. In aspects, 3 working sets of forward primers and 3 working sets of reverse primers are generated. In aspects, 4 working sets of forward primers and 4 working sets of reverse primers are generated.

[0079] In aspects of the method, the reverse primer set comprises primers for J-genes, downstream introns for D-genes and KDE. In aspects, the forward primer set comprises primers for V-genes, upstream introns for D-genes and C-intron for IgK.

[0080] In aspects of the method, the multiplex primer sets generated can be used for amplification of complete VJ and / or VDJ rearrangements at one or more of TCRa, TCRP, TCRy, TCR5, IgH, IgK, IgL loci; partial DJ rearrangements at TCRP, TCR5, IgH loci; VD and DD rearrangements at TCR5 and TCRP loci; chimeric TRDV-TRAJ rearrangements; or Kappa deletion rearrangements. In aspects, the multiplex primer sets generated can be used for amplification of complete VJ and / or VDJ rearrangements at two, three, four, five, six, seven eight, nine, ten or more of TCRa, TCRP, TCRy, TCR5, IgH, IgK, IgL loci; partial DJ rearrangements at TCRP, TCR5, IgH loci; VD and DD rearrangements at TCR5 and TCRP loci; chimeric TRDV-TRAJ rearrangements; or Kappa deletion rearrangements. In aspects of the method, the multiplex primer sets generated can be used for amplification of complete VJ and / or VDJ rearrangements at each of TCRa, TCRP, TCRy, TCR5, IgH, IgK, IgL loci; partial DJ rearrangements at TCRP, TCR5, IgH loci; VD and DD rearrangements at TCR5 and TCRP loci; chimeric TRDV-TRAJ rearrangements; or Kappa deletion rearrangements. In aspects of the method, primers for a particular V, D and J combination can be used to reduce unnecessary background for more specific amplification of target malignancy-related rearrangements and to reduce required sequencing depth (coverage).

[0081] In aspects of the method, indexing or dual indexing of each aliquot amplicon and introduction of adapters for sequencing are performed using "step-out" PCR. Step-out PCR is discussed in Wesley et al., (1997). Rapid Directional Walk Within DNA Clones by Step- Out PCR. In: White, B.A. (eds) PCR Cloning Protocols. Methods in Molecular Biology™, vol 67. Humana Press, Totowa, NJ.

[0082] In aspect of the method, indexing or dual indexing of each aliquot amplicon and introduction of adaptors for sequencing are performed using known in the art ligation reaction (Head at all., Library construction for next-generation sequencing: overviews and challenges. BioTechniques, 2014 Feb l;56(2):61-4. doi: 10.2144 / 000114133).

[0083] In aspects of the methods, the amplicons are sequenced using next-generation sequencing methods. Such methods are known in the art. In aspects, single side sequencing is used. In aspects, double side sequencing is used. In aspects, single side sequencing used can be at least 100 nt, at least 150 nt, at least 250 nt, or at least 300 nt length. In aspects, double side sequencing used can be at least 50 nt, at least 100 nt, at least 150 nt, at least 250 nt length. In aspects, first side sequencing length and second side sequencing length may differ. In aspects, sequencing coverage of 1, 2, 3 and more sequencing reads per cell (per each 6.5 pg of genomic DNA) in the analyzed sample is used to detect TCR and BCR rearrangements presented in the prepared adaptome library.

[0084] In aspects of the methods, follow-on monitoring is performed with less than the full set of oligonucleotides from the initial multiplex detection. In aspects, the follow-on monitoring PCR reactions comprises one or more, two or more, three or more, four or more, five or more or six or more oligonucleotides amplifying TRa / p / y / 5, IgH / K / k, DJ, DD, VD, and kappa-deleting element (KDE) rearrangements in a single multiplex mixture of oligonucleotides.

[0085] In aspects of the methods, the average length of target amplicons is from about 120 bp to about 280 bp. In aspects, the average length of target amplicons is about 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290 or 300 bp. In aspects of the methods, the average lengths of target amplicons in the 4 multiplex primer sets are: set 1 - about 120 bp, set 2 - about 160 bp, set 3 - about 240 bp, and set 4 - about 280 bp. In aspects of the methods, amplicons obtained with a set of primers cannot be a matrix for PCR with the subsequent set of primers. In aspects, the fact that a set of primers cannot be a matrix for PCR with the subsequent set of primers can prevent contaminating sequences in the sample from being further amplified. In aspects, the 1st set of primers cannot be a matrix for PCR with the 2nd set of primers. In aspects, the 1st set of primers cannot be a matrix for PCR with the 2nd set of primers or the 3rd set of primers. In aspects, the 1st set of primers cannot be a matrix for PCR with the 2nd set of primers, 3rd set of primers or 4th set of primers. In aspects, the 2nd set of primers cannot be a matrix for PCR with the 3rd set of primers. In aspects, the 2nd set of primers cannot be a matrix for PCR with the 3rd set of primers or 4th set of primers. In aspects, the 3rd set of primers cannot be a matrix for PCR with the 4th set of primers.

[0086] In aspects of the methods where there are more than four sets of primers, the set of primers for a PCR reaction cannot be a matrix for the directly following PCR reaction.

[0087] In aspects of the methods, the concentrations of primers in the multiplex mixture are optimized. In aspects, the concentrations of primers in the multiplex mixture are optimized by analyzing the frequency ratios of non-functional rearrangements of T- and B-cell receptor genes.

[0088] In aspects the primer concentration is optimized by determining the overamplifi cation rate (OAR) and adjusting the primer concentration to compensate for the OAR. In aspects, the OAR is determined using the method in Example 7 below.

[0089] In aspects of the methods, a ratio of weighted frequencies (proportion of reads) to unweighted frequencies (proportion of clonotypes) for V, D and J genes is determined. In aspects, if the ratio is a substantial deviation from 1, it is indicative of quantitative bias occurring during amplification. In aspects, the concentrations of primers in the multiplex mixture is selected to minimize the deviation from the expected value of one for the ratio of weighted frequencies to unweighted frequencies.

[0090] In aspects of the methods, the MRD detection sensitivity depends exclusively on the available input DNA and sequencing coverage.

[0091] In other aspects, the disclosure provides a method of determining the presence and quantification of Minimal residual disease (MRD) level in patients with hematological malignancies, the method comprising: a. detecting rearrangements of T cell receptor genes and / or B cell receptor genes characteristic for a malignant clones in an initial sample from a patient using a multiplex polymerase chain reaction (PCR) with isolated genomic DNA and using high-throughput sequencing to obtained a PCR library; wherein the PCR reaction comprises oligonucleotides amplifying one or more of the following: TCRa / p / y / 5, IgH / k / k, DJ, DD, VD, and kappa deleting element (KDE) rearrangements, in a single multiplex mixture of oligonucleotides; b. extracting complete and incomplete rearrangements from high-throughput sequencing data by mapping potential rearrangements to a library comprising V, D, and J gene segments, kappa deleting element (KDE), and IGKC intron RSS sequences, wherein the extracting is performed using a semi-global alignment algorithm to identify flanking sequences, followed by a clustering algorithm correcting for PCR and sequencing errors to assemble complete and / or incomplete rearrangement clonotypes; c. determining precise clonotype sizes from hematological malignancy- related VDJ repertoire data using non-functional clonotypes as a calibrator for quantitative bias correction; d. performing follow-up monitoring of MRD in at least 4 independent PCR reactions with isolated and quantified genomic DNA obtained at a followup time point; wherein the minimal frequency of a malignant clone in the follow-up time point is determined based on the proportion of independent PCR reactions where the corresponding rearrangement characteristic for the malignant clone is observed.

[0092] Methods of Designing of Oligonucleotide Libraries

[0093] In aspects, the disclosure also provides a method for generating a library of oligonucleotides for multiplex PCR. In aspects, the library of oligonucleotides for multiplex PCR is designed to decrease primer dimer formation and minimize nucleotide diversity.

[0094] In aspects, the oligonucleotide library is designed by a method comprising: a. dividing target genomic DNA regions into k-mers of a certain length; b. clustering the target genomic DNA regions using Hamming distances distances not exceeding 4, not exceeding 5, or not exceeding 6, or not exceeding 7 nucleotides; and c. identifying clusters of highly similar k-mers according to the selected Hamming distance. In aspects, the k-mers have a length of at least 16 bases. In aspects, the k-mers have a length of at least 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24 or 25 bases.

[0095] In aspects, tables of 5-mers in the reverse orientation and 5-mers at the end of 3' end of the nucleotide sequences in the forward orientation are made from the clusters identified. In aspects, the 5-mers tin the tables are intersected pairwise, and the primer set with the lowest overlap with the 5-mers is selected for further analysis. In aspects, this intersection procedure is repeated. In aspects, the intersection procedure is performed in an iterative manner.

[0096] In aspects of the method, after a set of primers is determined, to equilibrate the annealing temperature of the primers, up to 14 additional target- site-related nucleotides are added to the 5'- end of the k-mers. In aspects, up to 12, 13, 14, 15, 16, 17, 18, 19 or 20 additional target-site-related nucleotides are added to the 5'-end of the k-mers.

[0097] In aspects of the method, 2 working sets of forward primers and 2 working sets of reverse primers are generated. In aspects, 3 working sets of forward primers and 3 working sets of reverse primers are generated. In aspects, 4 working sets of forward primers and 4 working sets of reverse primers are generated.

[0098] In aspects, the forward primer set comprises primers for V-genes, upstream introns for D-genes and C-intron for IgK. In aspects of the method, the reverse primer set comprises primers for J- genes, downstream introns for D-genes and KDE.

[0099] In aspects of the method, the multiplex primer sets generated can be used for separate amplification of complete VJ and / or VDJ rearrangements at TRa, TRP, TRy, TR5, IgH, IgK, IgL loci; partial DJ rearrangements at TRP, TR5, IgH loci; VD and DD rearrangements at TR5 and TRP loci; VD rearrangements at IgH loci; chimeric TRDV- TRAJ rearrangements; or Kappa deletion rearrangements.

[0100] In aspects of the method, primers for a particular V, D and J combination can be used to reduce unnecessary background for more specific amplification of target malignancy-related rearrangements and to reduce required sequencing coverage.

[0101] It is to be understood that the disclosure is not limited to the particular aspects exemplified below, as variations of the particular aspects may be made and still fall within the scope of the appended claims. It is also to be understood that the terminology employed is for the purpose of describing particular aspects, and is not intended to be limiting. Examples

[0102] Example 1. Nested localization of oligos in sequential MRD detection and follow up time points. This example provides a scheme of annealing positions for four designed primer sets: forward primers setl, set2 and reverse primers set3, set4.

[0103] FIG. 1 is a schematic showing how nested localization of different primer sets is used to avoid undesired amplification of TCR / BCR amplicons obtained in the previous stage of MRD monitoring, for example during clonal marker detection in the onset sample. The amplicon is rich in molecules with target clonal markers so it might be one of the main sources of contamination, leading to potential false positive MRD results. The proposed primer positions are reliable protection against such cross-contamination.

[0104] FIG. 2 shows picks of four different tested target amplicons as separated using capillary electrophoresis (Tape Station, Agilent). All four picks correspond to ready-to- sequence TCR / BCR libraries differ one from another by amplicon length. As can be seen from the figure, all four combinations of forward and reverse primer sets can be used for adaptome library generation.

[0105] Example 2. Validation of multiplex primer mix

[0106] Primer mixes capable of amplifying multiple types of VDJ rearrangements were prepared. One example of such a primer mix is the 7GENES mixture sold by MiLaboratories, Inc. To validate the efficiency of the multiplex mix to amplify diverse VDJ rearrangements, a single tube 7GENES (setl+set3 primers) amplification was performed starting from 100-150 ng of human PBMC DNA derived from 190 healthy donors. Amplification was performed separately for each donor. Detections were performed for IGHV, IGHD, IGKV, IGLV, TRBV and TRBD, TRGV, TRDV + TRDD, IGHJ, IGKJ, IGLJ, TRAJ, TRBJ, and TRDJ gene segments.

[0107] The resulting libraries were sequenced on an IlluminaNextSeq550 sequencer using paired-end 150+150 nt sequencing, and analyzed using MiXCR software ( / pubmed. ncbi.nlm.nih.gov / 25924071 / ) with additional reference datasets required to extract incomplete rearrangements.

[0108] FIGs. 3A-M show the gene segment frequency in complete and incomplete rearrangements and the proportion of samples containing rearrangements with such gene segments across 190 healthy donors.

[0109] A total of 9,532,826 unique rearrangements with 5619 combinations of 296 V- genes, 29 D-genes, Kde, IGKC-intron and 97 J-genes were identified in DNA from the peripheral blood of 190 healthy individuals, confirming the successful performance of the designed multiplex primers. The results demonstrate that the designed multiplex primers are capable of amplifying the desired gene rearrangements and provide nearly comprehensive coverage of adaptome rearrangements for their initial detection and MRD follow up.

[0110] Example 3. Validation of method using a diluted spike-in leukemic clone carrying several VDJ rearrangements

[0111] This example demonstrates the precision of MRD detection based on the approach presented in this disclosure. To measure precision, observed and expected MRD levels (concentration of leukemic cells) were compared in six composite samples containing a counted number of cells of regenerating bone marrow (background cells) from one patient and a counted number of leukemic cells from another patient (spike-in cells). Cell counting and mixing were performed using FACS (Fluorescence Activated Cell Sorting). Spike-in leukemic cells (10, 100, 500, 1000, 5000 and 10 000 cells) were sorted into separate aliquots (1 mln cells each) of background cells to imitate MRD at six different levels: 1 leukemic cell per 1) 100 000 cells (0.001% MRD); 2) 10 000 cells (0.01%

[0112] MRD); 3) 5000 cells (0.02% MRD), 4) 1000 cells (0.1% MRD) 5) 500 cells (0.2% MRD); 6) 100 cells (1% MRD). DNA from the obtained mixtures was extracted and used for sequencing library preparation. The obtained DNA for each spike-in sample was distributed among 32 separate multiplex PCR reactions with a 320 ng DNA per reaction for 0.001% MRD, a 32 ng DNA per reaction for 0.01% MRD, 16 ng DNA per reaction for 0.02% MRD, a 3.2 ng DNA per reaction for 0.1% MRD, a 1.6 ng DNA per reaction for 0.2% MRD, and a 0.32 ng DNA per reaction for 1% MRD. Setl+Set4 primers were used for PCR to avoid target amplicon contamination by any previous PCR products. Three different rearrangements (IGH D-J, TRD D2-D3, TRD V2-D3) initially detected in pure leukemia sample were used as clonal markers to track leukemic cell presence in prepared test samples. A plot of the measured portion of leukemic cells vs. the expected portion of leukemic cells is shown in FIG. 4. As can be seen in the figure, the method showed high precision, with r=0.9906 and R2 = 0.9813. This experiment shows that the method can be used for high precision determination of malignant cells in a sample using several VDJ rearrangements.

[0113] Example 4. Comparison of NGS analysis with flow-cytometry using real clinical data This example demonstrates the level of consistency between the disclosed Nextgeneration sequencing (NGS) based approach and the conventional flow cytometry-based MRD monitoring method.

[0114] Multicolor flow cytometry for MRD measurement was performed using 8-12 surface antigens according to the consensus recommendations of the AIEOP-BFM group [M. Dworzak et. al., Standardization of flow cytometric minimal residual disease evaluation in acute lymphoblastic leukemia: Multicentric assessment is feasible. Cytometry B Clin Cytom. 2008 Nov;74(6):331-40. doi: 10.1002 / cyto.b.20430], Identification of CD 19-negative relapses during CD19-directed immunotherapy was performed using multicolor flow cytometry with a modified antibody panel and adjusted data analysis algorithm, according to the recommendations of Cherian et al. [Cherian et al. A novel flow cytometric assay for detection of residual disease in patients with B- lymphoblastic leukemia / lymphoma post anti -CD 19 therapy. Cytometry B Clin Cytom. 2018 Jan;94(l): 112-120. doi: 10.1002 / cyto.b.21482. Epub 2016 Sep 23.].

[0115] Libraries for NGS-based MRD analysis were prepared using forward and reverse primers from setl and set4, specific for V / D / J / Kde combinations present in pre-detected leukemic clone-related rearrangements. A total of 32 DNA aliquots (16 for 320 ng, 8 for 32 ng, 8 for 3.2 ng DNA) were distributed across separate multiplex PCR reactions for each analyzed time point. Amplicons from each of the 32 PCR reactions were indexed uniquely in the subsequent indexing PCR. Pooled libraries for each time point were sequenced on the MiSeq instrument in paired-ends mode 150+150 nt and ~1 mln sequencing reads per time-point.

[0116] Results are provided in Table 1 below. As can be seen from the results, the specificity and sensitivity of the NGS-based method are similar to flow-cytometry methods and, furthermore, in manycases, the NGS-based method is able to detect MRD that flow-cytometry is not. Table 1. Number of positive and negative MRD results obtained by disclosed NGS- based approach and conventional flow cytometry.

[0117] Example 5. Classic and non-classic examples of VDJ rearrangements in leukemic cells.

[0118] This example provides examples of classic and non-classic TCR and BCR rearrangements which are used as patient-specific clonal biomarkers for MRD monitoring.

[0119] The presented rearrangements were detected in DNA from actual clinical diagnostic samples of bone marrow from T-cell acute lymphoblastic leukemia and B-cell acute lymphoblastic leukemia patients. The detection was performed using the 7GENES kit (a mix of setl and set3 primers).

[0120] Classic VDJ rearrangements that are detected include: partial rearrangements between D and J genes (DJ rearrangements) in TRB, TRD and IGH loci; partial rearrangements between V and D, D and D genes in TRD locus; rearrangements between IGK C-intron or IGKV genes and Kappa deletion element (Kde); complete rearrangements between V, D and J genes in TRB, TRD and IGH loci and rearrangements between V and J genes in TRA, TRG, IGK and IGL loci.

[0121] Non-classic VDJ rearrangements that are detected include: partial rearrangements between DI and D2 genes in TRB locus; partial rearrangements between V and D genes in TRB and IGH loci; complete chimeric rearrangements between TRDV genes and TRAJ genes. The designed multiplex system contains primers for all TCR and BCR genes in the single common mix which allow detection of all chimeric VDJ combinations. Examples of specific rearrangements that can be detected are shown in Tables 2-18 below. Table 2. Examples of TRA V-J rearrangements

[0122] Table 3. Examples of TRB D-J rearrangements Table 4. Examples of TRB VDJ rearrangements

[0123] Table 5. Examples of TRG V-J rearrangements Table 6. Examples of TRD V-D rearrangements

[0124] Table 7. Examples of TRD D-D rearrangements Table 8. Examples of TRD D-J rearrangements

[0125] Table 9. Examples of TRDV-TRAJ rearrangements Table 10. Examples of IGH D-J rearrangements

[0126] Table 11. Examples of IGH V-D-J rearrangements Table 12. Examples of IGK V-J rearrangements

[0127] Table 13. Examples of IGK C-intron-Kde rearrangements Table 14. Examples of IGK V-Kde rearrangements

[0128] Table 15. Examples of IGL V-J rearrangements Table 16. Examples of human TRB V-D rearrangements

[0129] Table 17. Examples of human TRB D-D rearrangements. Table 18. Examples of human IGH V-D rearrangements.

[0130] The results indicate that the NGS-based method is able to determine multiple, leukemic clone- specific, classic and non-classic examples of VDJ rearrangements that can be analyzed in follow- up samples for MRD analysis.

[0131] Example 6. Analysis of real clinical data on a cohort of 58 B-ALL patients.

[0132] This example demonstrates detection and measurement (quantification) of MRD for actual clinical bone marrow samples obtained from acute leukemia patients during therapy. Observed patients were receiving anti -CD 19 CAR-T cell therapy with subsequent bone marrow transplantation. MRD analysis for each patient was performed for 1 -14 time points (6.5 average) for a total of 386 time-points. Timepoints included samples taken at 24 days, 2 months, 3 months, 6 months, 1 year, and 2 years after the onset of therapy. Quantification of MRD levels (concentration of leukemic cells) was performed by the technique described in this disclosure. A plot of the percentage of samples having the indicated MRD value is shown in FIG. 5.

[0133] The presence of target clonal rearrangements was analyzed by high throughput sequencing of VDJ libraries obtained from thirty -two independent DNA aliquots: 16 aliquots of 300 ng DNA, 8 aliquots of 30 ng DNA, and 8 aliquots of 3 ng DNA. Primary results of MRD marker extraction from VDJ repertoire data for one of the patients of the study are summarized in Table 19. The table contains the number of total sequencing reads obtained per aliquot and the number of reads covering target MRD markers in each analyzed aliquot. Subsequent analysis of obtained primary results includes under-sequenced aliquot exclusion, removal contamination between aliquots, and calculation of MRD level. Undersequencing is considered a source of false-negative results. Contamination is considered a source of false- positive results. Under-sequenced aliquots are those with sequencing coverage (number of reads) significantly lower than the average sequencing coverage of other aliquots. Contamination between aliquots can occur in sequencing instruments. Aliquots with a coverage of the target clonal marker that is significantly lower compared to the average coverage of the same marker in other aliquots are considered as contaminated aliquotas. Z-score is used as a statistical metric for the detection of both under-sequenced aliquots and possible contaminations. After the exclusion of under-sequenced aliquots from analysis and contamination removal, aliquots with at least one sequencing read corresponding to the target clonal marker are considered positive aliquots, other aliquots are considered negative aliquots. Each aliquot can be positive for one target marker and negative for another. In the present example, 4 out of 32 aliquots were considered to be under- sequenced. After filtering, fifteen 300 ng aliquots, seven 30 ng aliquots, and six 3 ng aliquots were taken to the downstream analysis.

[0134] MRD presence is determined when at least one positive aliquot at least for one marker is detected. MRD levels (concentration of leukemic cells) are calculated using each target clonal marker separately. Among equal aliquotes, MRD is the number of positive aliquots (proxy to the number of leukemic cells) divided by the number of analyzed cells (one cell contains ~ 6.5 pg DNA) as a first approximation. More precisely the number of leukemic cells can be calculated using the Poisson statistics model: N = -ln(l-n / m), where “N” is the number of leukemic cells in the sample (number of target molecules in the DNA sample), “n” is the number of positive aliquots, “m” - number of analyzed aliquots. When all aliquots are positive, a set of diluted aliquots is used to detect MRD levels. In the given example all fifteen 300 ng aliquots and all seven diluted 30 ng aliquots are positive for all target markers. Among the six remaining 3 ng aliquots TRBV20-1 / J2- 2 marker is present in 6 aliquots, one IGKVlD-43 / Kde marker is present in 4 aliquots, another IGKV1D- 43 / Kde marker is present in 3 aliquots, TRDV2 / D3 marker is present in 5 aliquots, IGKV1- 39 / J1 marker is present 4 aliquots, TRGV3 / J1 is present in 5 aliquots. The number of positive aliquots allows to conclude that MRD levels (concentration of leukemic cells) calculated using separate MRD markers are >0.5%, 0.14%, 0.1%, 0.18%, 0.14%, 0.18%. Integration of these results across all 6 MRD markers allows to conclude that consensus MRD (max value across the markers) is present at the minimal frequency 0.5%. Table 19. MRD level calculation for clinical samples.

[0135] Example 7. Quantitative bias correction via weighted versus unweighted nonfunctional rearrangement frequencies analysis. This example demonstrates how gene segment frequency correction in the VDJ rearrangement data studied in Example 6 can be performed based on the relative abundance of non-functional (e.g., non-productive or out-of-frame) clonotypes. Nonproductive (containing the stop codon) or out-of-frame TCR / BCR rearrangements cannot form a functional receptor, and therefore they are not subjected to antigen-specific selection and clonal expansion (Murugan et al., 2012, PNAS 109:16161-16166).

[0136] Therefore, such rearrangements can be used as passenger genomic variations, with clonal frequencies that change randomly, following the frequency changes of the second functional (in- frame, no stop codon) TCR / BCR allele present in the same T / B cell clone. The formation of in- frame and out-of-frame allele combinations in the same cell is also a stochastic and independent process in terms of V- and J-gene frequency.

[0137] Therefore, reproducible deviation of out-of-frame V- and J-gene frequencies from the initial recombinational frequencies observed in the sequenced repertoire dataset (for the same multiplex PCR primer set, same production lot, etc) can be determined as a result of artificial aberration caused by PCR amplification rather than immune repertoire evolution.

[0138] In this example, out-of-frame clonotypes were used as a natural calibrator to measure amplification bias. This allowed for changing the primer concentration in order to reduce such bias, and finally quantitatively correct the remaining amplification bias in immune repertoire data.

[0139] Changes in TRAV and TRAJ overamplification rate during primer concentration optimization are plotted in FIGs. 6A and 6B. Changes in TRAV and TRAJ overamplifi cation rate after final computational bias correction are plotted in FIGs. 7A and 7B. Overamplifi cation rates (OAR) are calculated according to the formula: where OAR - OverAmplification Rate, RC - read counts, UCN - unique out-of- frame clones, Vi - particular V gene, Ji - particular J gene.

[0140] The correction coefficient for computational bias removal is calculated by simply multiplying OAR(Vi) by OAR(Ji). To correct amplification bias, each clonotype should be divided by the corresponding correction coefficient.

[0141] As can be seen in both FIGs. 6 and 7, the methods were effective in overcoming amplification bias at the stage of wet-lab adaptome library preparation and computational analysis of the obtained adaptome after high-throughput sequencing.

Claims

CLAIMSWhat is claimed is:

1. A method of determining the presence and quantification of minimal residual disease (MRD) in patients with hematological malignancies, the method comprising: a. detecting rearrangements of T cell receptor genes and B cell receptor genes characteristic for malignant clones in an initial sample from a patient using: a multiplex polymerase chain reaction (PCR) with isolated genomic DNA wherein the PCR reaction comprises a library of oligonucleotides amplifying one or more of the following: TRa / p / y / 5, IgH / K / X, DJ, DD, VD, and kappa-deleting element (KDE) rearrangements in a single multiplex mixture of oligonucleotides; high-throughput sequencing of the obtained PCR products; extracting complete and incomplete rearrangements by mapping potential rearrangements to a library comprising V, D, and J gene segments, kappa deletion element (KDE), and IGKC intron sequences, wherein the extraction is performed using a semi-global alignment algorithm to identify flanking sequences, followed by a clustering algorithm correcting for PCR and sequencing errors to assemble complete and / or incomplete rearrangement clonotypes; determining precise clonotype sizes from hematological malignancy- related repertoire; and malignancy-related clonotype detection; b. performing follow-up monitoring of MRD in at least 4 independent PCR reactions with isolated and quantified genomic DNA obtained at a follow-up time point; wherein the concentration of a malignant clone in the follow-up time point is determined based on the proportion of independent PCR reactions where the corresponding rearrangement characteristic for the malignant clone is observed.

2. The method of claim 1, wherein the library of oligonucleotides for multiplex PCR is designed to decrease primer dimer formation and minimize nucleotide diversity.

3. The method of claim 1, wherein the library of oligonucleotides is designed by a method comprising: a. dividing target genomic DNA regions into k-mers of a certain length; b. clustering the target genomic DNA regions using Hamming distances not exceeding 4, not exceeding 5, or not exceeding 6, or not exceeding 7 nucleotides; and c. identifying clusters of highly similar k-mers according to the selected Hamming distance.

4. The method of claim 3, wherein the k-mers have a length of at least 16 bases.

5. The method of claim 3, wherein from the clusters identified, tables of 5- mers in the reverse orientation and 5-mers at the end of 3' end of the nucleotide sequences in the forward orientation are made.

6. The method of claim 5, wherein the 5-mers tables are intersected pairwise, and the primer set with the lowest 5-mers overlap is selected for further analysis.

7. The method of claim 4, wherein, to equilibrate the annealing temperature of the primers, up to 14 additional target- site-related nucleotides are added to the 5'-end of 18-mers.

8. The method of claim 5, wherein the intersection procedure using 5-mers is repeated.

9. The method of claim 3, wherein at least 2 working sets of forward primers and at least 2 working sets of reverse primers are generated.

10. The method of claim 9, wherein the reverse primer set comprises primers for J-genes, downstream introns for D-genes and KDE.

11. The method of claim 9, wherein the forward primer set comprises primers for V-genes, upstream introns for D-genes and C-intron for IgK.

12. The method of claim 9, wherein the multiplex primer sets generated can be used for separate amplification of complete VJ and / or VDJ rearrangements at TRa, TRp, TRy, TR5, IgH, IgK, IgL loci; partial DJ rearrangements at TRP, TR5, IgH loci; VD and DD rearrangements at TR5 and TRP loci; chimeric TRDV-TRAJ rearrangements; or Kappa deletion rearrangements.

13. The method of claim 12, wherein dual indexing of each aliquot amplicon and introduction of adapters for sequencing are performed using "step-out" PCR.

14. The method of claim 12, wherein the average lengths of target amplicons in the 4 multiplex primer sets are: set 1 - about 120 bp, set 2 - about 160 bp, set 3 - about 240 bp, and set 4 - about 280 bp.

15. The method of claim 9, wherein amplicons obtained with a set of primers cannot be a matrix for PCR with the subsequent set of primers.

16. The method of claim 1, wherein the concentrations of primers in the multiplex mixture are optimized by analyzing the frequency ratios of non-functional rearrangements of T- and B-cell receptor genes.

17. The method of claim 16, wherein an overamplification rates (OAR) is determined for at least one rearrangement, and the concentration of the primer detecting the rearrangement is adjusted to correct for the amplification bias including overamplifi cation and underamplification).

18. The method of claim 1, wherein an overamplification rates (OAR) is determined for at least one rearrangement to computationally correct for the amplification bias including overamplification and underamplification.

19. The method of claim 1, wherein a ratio of weighted frequencies (proportion of reads) to unweighted frequencies (proportion of clonotypes) for V, D and J genes is determined.

20. The method of claim 19, wherein, if the ratio is a substantial deviation from1, it is indicative of quantitative bias occurring during amplification.

21. The method of claim 20, wherein the concentrations of primers in the multiplex mixture is selected to minimize the deviation from the expected value of one for the ratio of weighted frequencies to unweighted frequencies.

22. The method of claim 1, wherein the MRD detection sensitivity depends exclusively on the available input DNA and sequencing coverage.

23. The method of claim 1, wherein TRa / p / y / 5, IgH / K / X, DJ, DD, VD, and kappa-deleting element (KDE) rearrangements are detected in a single multiplex mixture of oligonucleotides.

Citation Information

Patent Citations

  • Monitoring health and disease status using clonotype profiles

    US10865453B2

  • Amplification with primers of limited nucleotide composition

    US20180148775A1

  • Means and methods for accurately assessing clonal immunoglobulin (IG) / t cell receptor (TR) gene rearrangements.

    WO2020190138A1