A method of quantifying nucleic acids
The method improves disease diagnosis by quantifying and identifying nucleic acids using statistical modeling and scoring matrices, addressing low abundance and sensitivity issues in existing liquid biopsy techniques.
Patent Information
- Application Number
- PCT/SG2025/050383
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-05
- Filing Date
- 2025-06-05
- Publication Date
- 2025-12-11
AI Technical Summary
Existing methods for quantifying and identifying nucleic acids in liquid biopsy face challenges such as low abundance and the need for sensitive and specific detection, limiting their accuracy in disease diagnosis.
A method involving obtaining nucleic acid sequence reads, identifying boundaries using a statistical model, and generating scoring matrices to quantify and identify annotated and unannotated nucleic acids, allowing for improved detection and prediction of disease conditions by calculating sample scores based on read scores.
Enhances the accuracy of disease diagnosis by accurately quantifying and identifying nucleic acids, enabling precise detection and prediction of conditions like cancer through statistical modeling and scoring matrices.
Smart Images

Figure SG2025050383_11122025_PF_FP_ABST
Abstract
Description
[0001] A METHOD OF QUANTIFYING NUCLEIC ACIDS
[0002] TECHNICAL FIELD
[0003] The present disclosure relates broadly to a method of quantifying and / or identifying one or more nucleic acids associated with a condition.
[0004] BACKGROUND
[0005] In recent years, liquid biopsy has emerged as a revolutionary approach in disease diagnostics, offering a minimally invasive method to detect and monitor disease progression. Central to this innovation is the analysis of cell-free nucleic acids, which include both cell-free DNA (cfDNA) and cell-free RNA (cfRNA), circulating in the bloodstream of cancer patients. The application of Next Generation Sequencing (NGS) technologies to these cell-free nucleic acids has significantly advanced our understanding of their potential as cancer biomarkers. NGS has uncovered a variety of cell-free nucleic acids present in blood samples.
[0006] However, the challenges of the methods in the art that use liquid biopsy for quantifying I identifying nucleic acids in disease diagnosis include low abundance of target nucleic acids, need for sensitive and specific detection methods etc.
[0007] Given the increasing gravitation towards liquid biopsy for disease diagnosis, identification of nucleic acids / biomarkers with improved accuracy of disease diagnosis is needed. As such, it is essential to develop robust methods that are relatively simple to execute and have the potential to identify nucleic acids I uncover new biomarkers for disease diagnosis / detection.
[0008] Therefore, there is a need to provide an alternative method to quantify nucleic acids.
[0009] SUMMARY
[0010] In one aspect, there is provided method of quantifying and / or identifying one or more nucleic acids associated with a condition, the method comprising: i) obtaining one or more nucleic acids from a plurality of sample groups and generating one or more nucleic acid sequence reads; ii) identifying a plurality of boundaries of the one or more nucleic acid sequence reads using a statistical model; iii) obtaining a plurality of nucleic acid sequences bound by the selected boundaries; iv) quantifying and / or identifying the annotated nucleic acid and the unannotated nucleic acid that is associated with the condition.
[0011] In some examples, identifying one or more nucleic acids comprises comparing the expression level or copy number of the annotated nucleic acid and the unannotated nucleic acid in the plurality of sample groups; wherein an altered expression level or altered copy number as compared between the plurality of sample groups is indicative of the annotated nucleic acid and / or the unannotated nucleic acid to be associated with the condition.
[0012] In another aspect, there is provided a method of determining / predicting a condition in a subject comprising: i) obtaining one or more nucleic acids from a plurality of sample groups and generating one or more nucleic acid sequence reads from the plurality of sample groups; ii) aligning the one or more nucleic acid sequence reads to obtain one or more aligned portion, one or more empty portion, and / or one or more missing portion, iii) generating one or more scoring matrices from the alignment of the plurality of nucleic acid sequences; iv) calculating one or more read scores for the alignment using the generated one or more scoring matrices wherein the read score is determined by calculating the difference between the plurality of nucleic acid sequence read scores; and v) calculating the one or more sample scores based on the calculated one or more read scores; wherein the one or more sample scores provide an indication of the disease condition of the subject.
[0013] In some examples, each scoring matrix is represented by a modified positionspecific scoring matrix (PSSM).
[0014] In some examples, the modified PSSM comprises the likelihood of occurrence of each nucleotide sequence (A, T, G, C) and being empty / missing at different positions along the portion of an annotated genomic locus. In some examples, the method of the present disclosure comprises calculating the one or more sample scores based on the sum of the individual read scores of the one or more nucleic acid sequence reads, wherein the individual read score of the one or more nucleic acid sequence reads is calculated using one or more modified PSSM.
[0015] In some examples, the method of the present disclosure further comprises calculating a sample score for a plurality of sample groups according to the following equation: sample score wherein [31, [32, .... (3N represent the coefficients; r represents the sequence read;
[0016] R represents the total number of nucleic acid sequence reads in the sample; S1, S2, ... SN represent the read scores derived from the first, second, ... Nth modified PSSM, and wherein the sample score represents the condition likelihood.
[0017] In some examples, the method of the present disclosure further comprises calculating a sample score for two sample groups according to the following equation: wherein r represents the sequence read;
[0018] R represents the total number of sequence reads in the sample library;
[0019] 51 represents the first read score derived from the first modified PSSM; and
[0020] 52 represents the second read score derived from the second modified PSSM.
[0021] In some examples, the method of the present disclosure further comprises calculating a sample score for one sample group according to the following equation: s mpte score wherein r represents the sequence read;
[0022] R represents the total number of sequence reads in the sample library; and S represents the read score derived from the modified PSSM.
[0023] In yet another aspect, there is provided a method of determining whether a subject suffers from, or is at risk of developing a condition, the method comprising: i) determining the expression level or copy number of one or more nucleic acid sequence reads identified / quantified comprising: a) obtaining one or more nucleic acids from a plurality of samples and generating one or more nucleic acid sequence reads; b) identifying a plurality of boundaries of the one or more nucleic acid sequence reads using a statistical model; c) obtaining a plurality of nucleic acid sequences bound by the selected boundaries; d) quantifying and / or identifying the annotated nucleic acid and the unannotated nucleic acid that are associated with the condition; comparing the expression level or copy number of the one or more nucleic acid sequence reads in the subject with that in a control; wherein an altered expression level or altered copy number as compared between the plurality of samples is indicative of the subject suffering from or is at risk of developing the condition; and / or ii) calculating one or more sample scores for a sample obtained from a subject comprising a) obtaining one or more nucleic acids from a plurality of samples and generating one or more nucleic acid sequence reads from the plurality of samples; b) aligning the plurality of nucleic acid sequence reads to obtain one or more aligned portion, an empty portion, and / or one or more missing portion, c) generating one or more scoring matrices from the alignment of the plurality of nucleic acid sequence reads; d) calculating one or more read scores for the alignment using the generated one or more scoring matrices wherein the read score is determined by calculating the difference between the plurality of nucleic acid sequence read score; and e) calculating the one or more sample scores based on the calculated one or more read scores; wherein the one or more sample scores provide an indication of the condition of the subject, determining whether a subject suffers from, or is at risk of developing a condition based on the one or more calculated sample scores.
[0024] In some examples, the method comprises converting the alignment to a multiple sequence alignment format.
[0025] In some examples, the one or more nucleic acids comprise transfer ribonucleic acids (tRNAs), microRNAs (miRNAs), ribosomal ribonucleic acids (rRNAs), long noncoding RNAs (IncRNAs), small nucleolar RNAs (snoRNAs), small nuclear RNAs (snRNAs), yRNA, piwi-interacting RNA (piRNA) and / or fragments thereof.
[0026] In some examples, the one or more nucleic acid sequence reads comprise tRNA fragments.
[0027] In some examples, the one or more nucleic acid sequence reads is selected from the group consisting of tRNA-Lys or its fragments thereof, tRNA-Glu or its fragments thereof, and / or tRNA-Pro or its fragments thereof.
[0028] In some examples, the condition is a cancer selected from colorectal cancer, lung cancer, gastric cancer, breast cancer, liver cancer and / or prostate cancer.
[0029] In some examples, the condition is colorectal cancer.
[0030] In some examples, the sample comprises a tissue sample and / or a liquid biological sample.
[0031] DEFINITIONS
[0032] As used herein, the term “quantifying” or “quantification” of nucleic acids involves determining the number of copies of said nucleic acid molecules, or calculating, computing and / or assigning a score for the nucleic acids to infer the number of copies of said nucleic acid molecules.
[0033] As used herein, the term “score” refers to an integer or number, that can be determined mathematically, for example, by using computational models known in the art, which can include but are not limited to, linear regression, clustering techniques such as K-means, and the like, and that is calculated using any one of a multitude of mathematical equations and / or algorithms known in the art for the purpose of statistical classification. In some examples, a score can be the detected level of biomarkers unique to a certain sample type. In some examples, a score represents the likelihood of the sample type or origin. Such a score is used to enumerate one outcome on a spectrum of possible outcomes. For example, a read score may be determined for each nucleic acid sequence read obtained from a sample / biological sample using certain mathematical equations, followed by computing a sample score (based on the read scores) to determine the similarity of the sample / biological sample to a diseased sample. The relevance and statistical significance of such a score depends on the size and the quality of the underlying data set used to establish the results spectrum. For example, a blind sample may be inputted into an algorithm, which in turn calculates a score based on the information provided by the analysis of the blind sample. This results in the generation of a score for said blind sample. Based on this score, a decision can be made, for example, how likely the patient, from which the blind sample was obtained, has cancer or not. The ends of the spectrum may be defined logically based on the data provided, or arbitrarily according to the requirement of the experimenter. In both cases the spectrum needs to be defined before a blind sample is tested. As a result, the score generated by such a blind sample, for example the number “45” may indicate that the corresponding patient has cancer, based on a spectrum defined as a scale from 1 to 50, with “1” being defined as being cancer-free and “50” being defined as having cancer.
[0034] As used herein, the term “biological sample” or “sample” is intended to include any sampling of cells, cell extracts, tissues, organs, or bodily fluids isolated from a subject in which presence of a biomarker can be detected.
[0035] As used herein, the term “level” of a biomarker refers to the presence / absence, amount or concentration of the biomarker in a biological sample, and may be represented in any suitable forms or units as determined by a person skilled in the art. For example, the expression level of a nucleic acid (e.g. RNA) or the copy number of a nucleic acid (such as DNA) may be represented as, but not limited to, copy number (copy / mL), Ct (cycle threshold), Cq (quantification cycle), Ct / Cq, Iog2 scale expression levels or relative to other expression levels, and the like. In addition, the expression level or copy number of a nucleic acid may be expressed as a score constructed using any form of mathematical model or algorithm.
[0036] As used herein, the term “differential expression / differential copy number” or “altered expression / altered copy number” may be used interchangeably and refers to the measurement of a cellular component in comparison to a control or another sample, and thereby determining the difference in, for example, concentration, presence or intensity of said cellular component. The result of such a comparison can be given in the absolute, that is a component is present in the samples and not in the control, or in the relative, that is the expression or concentration or copy number of the component is increased or decreased compared to the control. The terms “increased” and “decreased” in this case can be interchanged with the terms “upregulated I over-expressed” and “downregulated / under-expressed” which are also used in the present disclosure.
[0037] As used herein, the term "nucleic acid / nucleic acid molecule" has its general meaning in the art and refers to a coding or non-coding nucleic sequence in a biological setting (i.e., the nucleic acid molecule is the biological molecule found in a sample). Nucleic acids include deoxyribonucleic acid (DNA) and ribonucleic acid (RNA). Examples of nucleic acid thus include but are not limited to DNA, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), microRNA (miRNA), PIWI- interacting RNA (piRNA), small nucleolar RNA (snoRNA), and small nuclear RNA (snRNA). Nucleic acids thus encompass coding and non-coding regions of a genome. In some examples, a nucleic acid I nucleic acid molecule may be used to infer a disease status of a subject.
[0038] As used herein, the term “nucleic acid sequence” refers to the nucleotide sequence of the nucleic acid molecule. In detail, a nucleic acid sequence is an arrangement of nucleotides within a nucleic acid molecule, such as a deoxyribose nucleic acid (DNA) or a ribonucleic acid (RNA).
[0039] As used herein, the term “nucleic acid sequence read / reads” refers to the string of nucleotides measured or obtained from a sequencing experiment that is meant to reflect the actual nucleic acid sequence. Nucleic acid sequence reads or sequencing data may be obtained from nucleic acid sequencing techniques, such as Sanger sequencing or Next-Generation Sequencing (NGS) of samples / biological samples and include information identifying nucleotide sequences and their order in each nucleic acid molecule. Sequencing data from a sequencer can include a series of nucleotides corresponding to these nucleic acid sequences and may be referred to as sequence reads. Each sequence read typically refers to the nucleic acid sequence from one fragment (refer herein as a small section / portion of DNA or RNA molecule) and may identify the order of nucleotides in a nucleic acid sequence. Further to any of the examples as provided herein, sequencing data from RNA may be obtained by isolating RNA from a sample, converting said RNA to complementary DNA (cDNA), and sequencing said cDNA. The nucleic acid sequence reads are allocated to be an annotated sequence and / or unannotated sequence after obtaining nucleic acid sequence bound by selected boundaries according to the methods of the present disclosure. The one or more nucleic acid sequence may be an annotated sequence and / or unannotated sequence. In some examples, the nucleic acid sequence reads may comprise i) the nucleic acid sequence reads of annotated sequences ii) the nucleic acid sequence reads of unannotated sequences iii) the nucleic acid sequence reads of irrelevant sequence / noise / background noise.
[0040] The term “annotated” as used herein refers to explanatory notes, labels or metadata that have been added to provide additional information, interpretation or context. In genomics, annotations identify genes, exons, regulatory elements, etc. An annotated genome includes gene names, coding regions, and known variants. An annotated nucleic acid sequence includes information about biologically significant features. This annotation provides context and functional insights beyond just the raw sequence of nucleotides. Annotations may include gene locations, coding regions (exons), regulatory elements (promoters, enhancers), transcription start / stop sites, known mutations / variants, functional elements (e.g., miRNA binding site, splice sites), protein translation information (e g., reading frames). For example, a GenBank entry showing a gene’s full DNA sequence along with labels for exons, introns, start codon, and known single nucleotide polymorphisms (SNPs). In some examples, annotated sequences are biologically relevant sequence reads from the nucleic acids present in the biological sample. In some examples, annotated sequences align or map to the reference genome i.e., it is possible to locate where in the genome the sequence is from. For the reference genome, some coordinates or locations carry “annotation”, which provides information such as the identity of a gene, functionality, disease implication etc. Annotation relies on previously reported or discovered knowledge of the sequence. Annotated sequences are therefore sequences whose biological function can be inferred based on available prior knowledge. Without wishing to be bound by theory, annotations may shift / change over time. However, the shifts I changes in annotations are not expected to impact its applicability to disease detection of the present disclosure.
[0041] The term “unannotated” as used herein refers to lacking added notes or descriptive information. It refers to raw or basic content that has not been interpreted, explained or labelled. In genomics, only the nucleotide sequence is present, with no information about function or structure. An unannotated nucleic acid sequence is just a string of nucleotides without any indication of where genes or regulatory elements are. In some examples, unannotated sequences are biologically relevant sequence reads from the nucleic acids present in the biological sample. In some examples, unannotated sequences include sequences that can be aligned or mapped to the reference genome (i.e., the location of the genetic sequence along the reference genome is known) but there is no known gene or genetic elements overlapping to the location. In other words, unannotated sequences are sequences which come from genomic coordinates without any known biological function or interpretation. In some examples, unannotated nucleic acid sequences (such as transfer RNA (tRF) sequences or their fragments thereof) are nucleotide sequences aligning to the predicted nucleic acid (such as transfer RNA (tRNA)) gene coordinates but have not been reported previously as a standalone nucleic acid (such as transfer RNA (tRNA)) or fragments thereof). In some examples, unannotated nucleic acid sequences align partially to the nucleic acid (such as tRNA) gene coordinates, but the nucleic acid gene coordinates have no prior nucleic acid (such as tRNA) annotation.
[0042] The term “irrelevant sequence / irrelevant sequence reads / noise / background noise” as used herein refers to technical artifacts introduced during sequencing. As used herein, the terms “tRNA fragments”, “transfer RNA fragments”, “tRNA-derived fragments”, “tRNA-derived RNA fragments”, “fragments of tRNA” or “tRFs” are used interchangeably, and refer to short or small non-coding RNAs generated from a tRNA gene locus. Typically, these fragments are often regarded as products of tRNA precursors or mature tRNAs cleaved by different enzymes at specific positions.
[0043] As used herein, the term “computer-readable storage medium” encompasses only a computer-readable medium that can be considered to be a manufacture (e.g., article of manufacture) or a machine. Alternatively, or additionally, methods or processes described herein may be embodied as a computer readable medium other than a computer-readable storage
[0044] The terms “program” or “software” are used herein in a generic sense to refer to any type of code or set of executable instructions that can be employed to program a computer or other processor to implement various aspects of the methods or processes described herein. Additionally, it should be appreciated that according to one aspect of this embodiment, one or more programs that when executed perform a method or process described herein need not reside on a single computer or processor but may be distributed in a modular fashion amongst a number of different computers or processors to implement various procedures or operations.
[0045] The term “subject” may include humans and animals. Animals may include, but are not limited to, mammals (for example non-human primates, canine, murine and the like), and the like. “Murine” refers to any mammal from the family Muridae, such as mouse, rat, rabbit, and the like. The subject may be a clinical patient, a clinical trial volunteer, an experimental animal, etc. The subject may be suspected of having or at risk of having one or more diseases, such as, but is not limited to cancer (e g. colorectal cancer, lung cancer, gastric cancer, breast cancer, liver cancer and prostate cancer) etc. The subject may be a healthy subject, a non-diseased subject and / or a cancer-free subject.
[0046] The term “relevant” as used herein refers to directly connected to or important for the matter at hand i.e., it implies significance or applicability to a particular subject, situation, or issue. The term “relevant” emphasizes importance or usefulness in a specific context e.g., a relevant nucleic acid sequence for disease diagnosis. A nucleic acid sequence relevant to a disease refers to a specific segment of nucleic acid molecules (such as DNA, RNA) that has direct or indirect association with the onset, progression, diagnosis, or treatment of a particular disease. This relevance may arise from mutations or variants that contribute to disease risk or pathology; gene expression changes observed in diseased versus healthy tissue; pathogen sequences present in infected individuals; or biomarkers used for diagnosis, prognosis or therapeutic targeting. For example, a BRCA1 gene mutation associated with hereditary breast and ovarian cancer, an mRNA sequence overexpressed in colorectal cancer, and the like. As used herein, the term “relevant” or “relevance” may be used interchangeably with the term “associated with”. In some examples, the “relevant” or “associated” nucleic acids as described herein may not be a nucleic acid that is known to have direct or indirect association with the onset, progression, diagnosis, or treatment of the disease.
[0047] The term “empty portion / empty / empty base / missing base / 5thbase / 5thnucleotide” as used herein refers to a portion of nucleic acid sequence reads with no alignment I absence of any alignment to a reference sequence. An “empty base” as part of the sequence profile is incorporated to accommodate for insertions and deletions (indels).
[0048] The term “boundary” corresponds to a nucleotide position along the reference genome. As will be appreciated, the term “probable” as used herein refers to having a high likelihood or probability of occurrence. As such, in the current context, “probable boundary” refers to identifying boundaries that are the most authentic or have a high likelihood of occurrence. In some examples, a boundary, or the fragmentation position, is determined based on the frequency of which the boundary read ends were observed among the observed or studied samples. This frequency threshold can be optimized for different datasets. For instance, in the study of the present disclosure, a boundary is considered as a “probable boundary” if more than 10% of the studied samples exhibited nucleic acid reads whose start or end occurred at the said position. As will be appreciated, bound nucleic acid sequences are obtained based on the selected boundaries. The selected boundary may correspond to either the start or end position of a nucleic acid sequence. A pair of boundaries are required to define the bounded nucleic acid sequence, in which each boundary would correspond to the start and end position of the sequence, respectively.
[0049] The term “mismatch” as used herein refers to when a base (A, T, C, or G) in a sequenced read does not match the corresponding base in the reference genome. Up to two mismatches are allowed in the alignment of the present disclosure i.e., no base mismatch, one base mismatch or two base mismatches in a sequenced read.
[0050] The term "associated with", used herein when referring to two elements refers to a broad relationship between the two elements. In genomics, “associated” typically refers to entities that are found in a correlative relationship (e.g., when A is upregulated I downregulated, B is also upregulated I downregulated). “Associated” is used when one cannot establish the causal or functional relationship based on correlation.
[0051] The term "and / or", e.g., "X and / or Y" is understood to mean either "X and Y" or "X or Y" and should be taken to provide explicit support for both meanings or for either meaning.
[0052] Further, in the description herein, the word “substantially” whenever used is understood to include, but not restricted to, "entirely" or “completely” and the like. In addition, terms such as "comprising", "comprise", and the like whenever used, are intended to be non-restricting descriptive language in that they broadly include elements / components recited after such terms, in addition to other components not explicitly recited. For example, when “comprising” is used, reference to a “one” feature is also intended to be a reference to “at least one” of that feature. Terms such as “consisting”, “consist”, and the like, may in the appropriate context, be considered as a subset of terms such as "comprising", "comprise", and the like. Therefore, in embodiments disclosed herein using the terms such as "comprising", "comprise", and the like, it will be appreciated that these embodiments provide teaching for corresponding embodiments using terms such as “consisting”, “consist”, and the like. Further, terms such as "about", "approximately" and the like whenever used, typically means a reasonable variation, for example a variation of + / - 5% of the disclosed value, or a variance of 4% of the disclosed value, or a variance of 3% of the disclosed value, a variance of 2% of the disclosed value or a variance of 1% of the disclosed value.
[0053] The phraseology and terminology used herein is for the purpose of description and should not be regarded as limiting. The use of “including,” “comprising,” “having,” “containing,” “involving,” and variations thereof, is meant to encompass the items listed thereafter and additional items. Use of ordinal terms such as “first,” “second,” “third,” etc., in the claims to modify a claim element does not by itself connote any priority, precedence, or order of one claim element over another or the temporal order in which acts of a method are performed. Ordinal terms are used merely as labels to distinguish one claim element having a certain name from another element having a same name (but for use of the ordinal term), to distinguish the claim elements. Additionally, the terms and expressions employed herein have been used as terms of description and not of limitation, and there is no intention in the use of such terms and expressions of excluding any equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the invention claimed. Thus, it should be understood that although the present invention has been specifically disclosed by preferred embodiments and optional features, modification and variation of the inventions embodied therein herein disclosed may be resorted to by those skilled in the art, and that such modifications and variations are considered to be within the scope of this invention.
[0054] The invention has been described broadly and generically herein. Each of the narrower species and subgeneric groupings falling within the generic disclosure also form part of the invention. This includes the generic description of the invention with a proviso or negative limitation removing any subject matter from the genus, regardless of whether or not the excised material is specifically recited herein.
[0055] Throughout this disclosure, certain embodiments may be disclosed in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the disclosed ranges. Accordingly, the description of a range should be considered to have specifically disclosed all the possible sub-ranges as well as individual numerical values within that range. For example, description of a range such as from 1 to 6 should be considered to have specifically disclosed sub-ranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6 etc., as well as individual numbers within that range, for example, 1 , 2, 3, 4, 5, and 6. This applies regardless of the breadth of the range.
[0056] Other embodiments are within the following claims and non-limiting examples. In addition, where features or aspects of the invention are described in terms of Markush groups, those skilled in the art will recognize that the invention is also thereby described in terms of any individual member or subgroup of members of the Markush group. Exemplary embodiments of the present invention are provided in the following examples. While the exemplary embodiments are described in sufficient detail to enable those skilled in the art to practice the invention, it should be understood that the present invention is not limited to these examples.
[0057] Furthermore, in the description herein, certain values may be disclosed in a range. The values showing the end points of a range are intended to illustrate a preferred range. Whenever a range has been described, it is intended that the range covers and teaches all possible sub-ranges as well as individual numerical values within that range. That is, the end points of a range should not be interpreted as inflexible limitations. For example, a description of a range of 1% to 5% is intended to have specifically disclosed sub-ranges 1% to 2%, 1% to 3%, 1 % to 4%, 2% to 3% etc., as well as individually, values within that range such as 1%, 2%, 3%, 4% and 5%. It is to be appreciated that the individual numerical values within the range also include integers, fractions and decimals. Furthermore, whenever a range has been described, it is also intended that the range covers and teaches values of up to 2 additional decimal places or significant figures (where appropriate) from the shown numerical end points. For example, a description of a range of 1% to 5% is intended to have specifically disclosed the ranges 1 .00% to 5.00% and also 1 .0% to 5.0% and all their intermediate values (such as 1 .01 %, 1.02% ... 4.98%, 4.99%, 5.00% and 1.1%, 1.2% ... 4.8%, 4.9%, 5.0% etc.,) spanning the ranges. The intention of the above specific disclosure is applicable to any depth / breadth of a range.
[0058] Additionally, when describing some embodiments, the disclosure may have disclosed a method and / or process as a particular sequence of steps. However, unless otherwise required, it will be appreciated that the method or process should not be limited to the particular sequence of steps disclosed. Other sequences of steps may be possible. The particular order of the steps disclosed herein should not be construed as undue limitations. Unless otherwise required, a method and / or process disclosed herein should not be limited to the steps being carried out in the order written. The sequence of steps may be varied and still remain within the scope of the disclosure.
[0059] Furthermore, it will be appreciated that while the present disclosure provides embodiments having one or more of the features / characteristics discussed herein, one or more of these features / characteristics may also be disclaimed in other alternative embodiments and the present disclosure provides support for such disclaimers and these associated alternative embodiments. DESCRIPTION OF EMBODIMENTS
[0060] The present disclosure discloses a method of quantifying nucleic acids. The present disclosure discloses methods for disease detection and / or diagnosis by quantifying and / or identifying fragments of nucleic acid sequences relevant to one or more diseases. The aim of the present disclosure is to provide methods to improve the accuracy (i.e. sensitivity and specificity) of disease diagnosis and detection.
[0061] Notably, two methods are disclosed herein: the first involves quantifying nucleic acid sequences from sequencing reads and identifying sequences that are relevant to a disease while the second method involves determining or predicting a condition / disease condition of a biological sample by scoring to determine the likelihood of a sample to be a diseased or non-diseased sample.
[0062] In some examples, quantifying the one or more nucleic acid sequences may include the use of any appropriate quantitative approaches to quantify the abundance of nucleic acid sequences or fragments thereof.
[0063] Method 1 : A method of quantifying nucleic acid sequences using a statistical model
[0064] In one aspect, there is provided a method of quantifying and / or identifying one or more nucleic acids associated with a condition, the method comprising: i) obtaining one or more nucleic acids from a plurality of samples and generating one or more nucleic acid sequence reads; ii) identifying a plurality of boundaries of the one or more nucleic acid sequence reads using a statistical model; iii) obtaining a plurality of nucleic acid sequences bound by the selected boundaries; iv) quantifying and / or identifying the annotated nucleic acid and the unannotated nucleic acid that are associated with the condition.
[0065] In some examples, identifying one or more nucleic acids comprises comparing the expression level or copy number of the annotated nucleic acid and the unannotated nucleic acid in the plurality of samples; wherein an altered expression level or altered copy number as compared between the plurality of samples is indicative of the annotated nucleic acid and / or the unannotated nucleic acid to be associated with the condition.
[0066] In detail, the first method includes quantifying nucleic acid sequences by unsupervised discovery of nucleic acid fragments. Without wishing to be bound by theory, the method of the present disclosure interprets / analyses the detected number and diversity of nucleic acid sequence reads, to filter out the nucleic acid sequence reads that are from technical errors (such as “background noise / irrelevant sequence / noise”) and to focus on the nucleic acid sequence reads that are from the biological target nucleic acid molecule of interest. With the statistical model that the inventors of the present disclosure designed, the nucleic acid sequence reads are processed to accurately, and with minimal error as possible, determine the amount / quantity of nucleic acid molecules, which levels indicate a disease status in a subject. The identification and allocation of nucleic acid sequence reads as annotated sequence and unannotated sequence allows improvement of the unsupervised model without reliance on genome annotation.
[0067] In another aspect, there is provided a method of quantifying nucleic acid sequences in one or more samples / biological samples, the method comprises: i) obtaining one or more nucleic acid sequence reads; ii) aligning the nucleic acid sequence reads to a reference genome to provide an alignment; iii) identifying a plurality of probable boundaries of the nucleic acid sequences using a statistical model, wherein each boundary corresponds to a nucleotide position of sequence ends along the reference genome; iv) selecting from the identified boundaries of (c) boundaries satisfying one or more pre-determined conditions; v) obtaining the nucleic acid sequences bound by the selected boundaries (d); and vi) quantifying the number of obtained nucleic acid sequences; wherein at least one of steps (i) to (vi) is performed by a suitably programmed computer.
[0068] In another aspect, the present disclosure comprises a method of identifying one or more nucleic acid sequences for use in determining and / or detecting a disease in a subject, the method comprises: i) quantifying the nucleic acid sequences from one or more biological samples according to the method of the present disclosure; and ii) identifying one or more nucleic acid sequences relevant to the disease; wherein an altered number of the one or more nucleic acid sequences, as compared to a control, is indicative of the one or more nucleic acid sequences relevant to the disease. Method 2: Quantifying the likelihood of disease status of a sample by scoring nucleic acid seguences using scoring matrices
[0069] In another aspect, there is provided a method of determining I predicting a condition in a subject comprising:
[0070] I) obtaining one or more nucleic acids from a plurality of samples and generating one or more nucleic acid sequence reads from the plurality of samples; ii) aligning the one or more nucleic acid sequence reads to obtain one or more aligned portion, one or more empty portion, and / or one or more missing portion, iii) generating one or more scoring matrices from the alignment of the plurality of nucleic acid sequences; iv) calculating one or more read scores for the alignment using the generated one or more scoring matrices wherein the read score is determined by calculating the difference between the plurality of nucleic acid sequence read scores; and v) calculating the one or more sample scores based on the calculated one or more read scores; wherein the one or more sample scores provide an indication of the disease condition of the subject.
[0071] The second method of quantifying the likelihood of a disease status of a sample by scoring nucleic acid sequences involves the quantitative or numeric scoring of nucleic acid sequence reads obtained from one or more samples / biological samples using scoring matrices. The quantitative scores derived from the scoring matrices are used for determining similarity (or dissimilarity) of the samples / biological samples to a diseased sample (e.g. tumour tissue).
[0072] Without wishing to be bound by theory, the present disclosure advantageously applies a scoring matrix to disease diagnosis by building two scoring matrices, wherein the scoring matrices are built from a plurality of samples. In some examples, the plurality of samples is from a plurality of biologically different groups. The disease status of an unknown sample can be determined based on the differences / discrepancy of the scores derived from the first and second samples. In another aspect, the present disclosure comprises a method of determining or predicting a condition of a biological sample, the method comprises:
[0073] I) obtaining one or more nucleic acid sequence reads of the biological sample; ii) aligning each nucleic acid sequence reads to a portion of a reference genome to provide an alignment having one or more aligned and empty portions; iii) calculating read scores for each sequence read using one or more scoring matrices; and iv) calculating a sample score for the biological sample based on the calculated read scores; wherein the sample score provides an indication of the condition of the biological sample; and wherein at least one of steps I) to iv) is performed by a suitably programmed computer.
[0074] In some examples, the method further comprises calculating one or more sample scores, wherein each sample score is calculated using different portions of a reference genome.
[0075] In some examples, the one or more sample scores provide an indication of the condition of the sample / biological sample.
[0076] Condition
[0077] In some examples, the condition comprises the likelihood or the possibility of the sample / biological sample to be a diseased or non-diseased sample.
[0078] In some examples, the condition is the likelihood or possibility of the sample / biological sample to be a cancer sample or non-cancer sample.
[0079] In some examples, the condition may include cancer such as but is not limited to colorectal cancer, lung cancer, gastric cancer, breast cancer, liver cancer and / or prostate cancer, and the like.
[0080] In some examples, the condition is colorectal cancer.
[0081] A method of determining whether a subject suffers from / is at risk of developing a disease
[0082] In another aspect, there is provided a method of determining whether a subject suffers from, or is at risk of developing a condition, the method comprising:
[0083] I) determining the expression level or copy number of one or more nucleic acid sequence reads identified / quantified comprising: a) obtaining one or more nucleic acids from a plurality of samples and generating one or more nucleic acid sequence reads; b) identifying a plurality of boundaries of the one or more nucleic acid sequence reads using a statistical model; c) obtaining a plurality of nucleic acid sequences bound by the selected boundaries; d) quantifying and / or identifying the annotated nucleic acid and the unannotated nucleic acid that are associated with the condition; comparing the expression level or copy number of the one or more nucleic acid sequence reads in the subject with that in a control; wherein an altered expression level or altered copy number as compared between the plurality of samples is indicative of the subject suffering from or is at risk of developing the condition; and / or ii) calculating one or more sample scores for a sample obtained from a subject comprising a) obtaining one or more nucleic acids from a plurality of samples and generating one or more nucleic acid sequence reads from the plurality of samples; b) aligning the plurality of nucleic acid sequence reads to obtain one or more aligned portion, an empty portion, and / or one or more missing portion, c) generating one or more scoring matrices from the alignment of the plurality of nucleic acid sequence reads; d) calculating one or more read scores for the alignment using the generated one or more scoring matrices wherein the read score is determined by calculating the difference between the plurality of nucleic acid sequence read score; and e) calculating the one or more sample scores based on the calculated one or more read scores; wherein the one or more sample scores provide an indication of the condition of the subject, determining whether a subject suffers from, or is at risk of developing a condition based on the one or more calculated sample scores.
[0084] In another aspect, there is provided a method of determining whether a subject suffers from, or is at risk of developing a condition, the method comprising determining the expression level or copy number of one or more nucleic acid sequence reads identified / quantified comprising: a) obtaining one or more nucleic acids from a plurality of samples and generating one or more nucleic acid sequence reads; b) identifying a plurality of boundaries of the one or more nucleic acid sequence reads using a statistical model; c) obtaining a plurality of nucleic acid sequences bound by the selected boundaries; d) quantifying and / or identifying the annotated nucleic acid and the unannotated nucleic acid that are associated with the condition; comparing the expression level or copy number of the one or more nucleic acid sequence reads in the subject with that in a control; wherein an altered expression level or altered copy number as compared between the plurality of samples is indicative of the subject suffering from or is at risk of developing the condition.
[0085] In another aspect, there is provided a method of determining whether a subject suffers from, or is at risk of developing a condition, the method comprising calculating one or more sample scores for a sample obtained from a subject comprising a) obtaining one or more nucleic acids from a plurality of samples and generating one or more nucleic acid sequence reads from the plurality of samples; b) aligning the plurality of nucleic acid sequence reads to obtain one or more aligned portion, an empty portion, and / or one or more missing portion, c) generating one or more scoring matrices from the alignment of the plurality of nucleic acid sequence reads; d) calculating one or more read scores for the alignment using the generated one or more scoring matrices wherein the read score is determined by calculating the difference between the plurality of nucleic acid sequence read score; and e) calculating the one or more sample scores based on the calculated one or more read scores; wherein the one or more sample scores provide an indication of the condition of the subject, determining whether a subject suffers from, or is at risk of developing a condition based on the one or more calculated sample scores. In another aspect, the present disclosure comprises a method of determining whether a subject suffers from, or is at risk of developing a disease, the method comprises: i) determining, in a sample / biological sample obtained from the subject, the expression levels of one or more nucleic acids identified or quantified according to the methods of the present disclosure; and ii) comparing the expression level of the one or more nucleic acids in the subject with that in a control; wherein an altered expression level of the one or more nucleic acids, as compared to a control, is indicative of the subject suffering from or is at risk of developing the disease.
[0086] In another aspect, the present disclosure comprises a method of determining whether a subject suffers from, or is at risk of developing a disease, the method comprises: i) calculating one or more sample scores for a biological sample obtained from the subject, according to the method of any one of the preceding statements; and ii) determining whether a subject suffers from, or is at risk of developing a disease based on the one or more calculated sample scores; wherein the one or more calculated sample scores above a pre-determined score is indicative of the subject suffering from or is at risk of developing the disease.
[0087] In some examples, the one or more calculated sample scores above a predetermined score is indicative of the subject suffering from or is at risk of developing the disease.
[0088] Obtaining the nucleic acid sequences from a subject
[0089] In some examples, obtaining the one or more nucleic acid sequences involves obtaining the one or more nucleic acid sequence reads from one or more samples obtained from a subject. The nucleic acid sequence reads may be available in any compatible form or file format, such as but are not limited to FASTQ, FASTQ.ORA, binary base call (BCL) file format, and the like.
[0090] Isolation of one or more nucleic acids (such as RNA) from the sample can be performed using any of the methods known in the art. The one or more nucleic acid (such as RNA) isolated from the sample can be total RNA or mRNA. Nucleic acid (such as RNA) isolation can be performed using a purification kit, a buffer set and protease from commercial manufacturers, such as Qiagen, according to the manufacturer's instructions. In one example, total nucleic acid (such as RNA) is isolated from the sample.
[0091] Commercially available nucleic acid (such as RNA) isolation kits may include such as but is not limited to AllPrep DNA / RNA / miRNA Universal Kit, and the like. Total nucleic acid (such as RNA) from tissue samples can be isolated, for example, using RNA Stat-60 (Tel-Test, Friendswood, Tex.). Nucleic acid (such as RNA) prepared from a tumor can be isolated, for example but is not limited to, by cesium chloride density gradient centrifugation, and the like. Additionally, large numbers of tissue samples can readily be processed using techniques well known to those of skill in the art, such as, for example, but is not limited to, the single-step RNA isolation process of Chomczynski (U.S. Pat. No. 4,843,155, incorporated by reference in its entirety for all purposes), and the like. In one example, total nucleic acid (such as RNA) can be isolated for example but is not limited to from formalin-fixed paraffin-embedded (FFPE) tissues as described by Bibikova et al. (2004) American Journal of Pathology 165:1799-1807, and the like. Likewise, the High Pure RNA Paraffin Kit (Roche) can be used. Paraffin is removed by xylene extraction followed by ethanol wash. RNA can be isolated from sectioned tissue blocks using for example but is not limited to the MasterPure Purification kit (Epicenter, Madison, Wis.); a DNase I treatment step is included. Nucleic acid (such as RNA) can be extracted from frozen samples using an extraction agent (such as Trizol reagent) according to the supplier's instructions (Invitrogen Life Technologies, Carlsbad, Calif). Samples with measurable residual genomic nucleic acid (such as DNA) can be resubjected to a nuclease (such as DNasel) treatment and assayed for nucleic acid (such as DNA) contamination. All purification, nuclease treatment (such as DNase treatment), and other steps can be performed according to the manufacturer's protocol. After total nucleic acid (such as RNA) isolation, samples can be stored at -80 °C until use.
[0092] In some examples, the one or more nucleic acids can be isolated from liquid biopsy samples using methods in the art. In some examples, the liquid biopsy samples are collected through e.g., venipuncture or lumbar puncture followed by a nucleic acid isolation procedure with methods in the art as will be appreciated by a person skilled in the art.
[0093] The total nucleic acid (such as RNA) may undergo small nucleic acid (such as RNA) RNA library preparation workflow using any of the kits or methods known in the art prior to sequencing. Examples of such kits may include but are not limited to NEBNext Multiplex Small RNA Library Prep Set performed according to manufacturer’s protocol. Typically, such library process involves 3’ adaptor ligation, primer hybridisation, primer hybridisation, 5’ adaptor ligation, conversion to complementary DNA (cDNA), PCR amplification, clean up and size selection. Further, conversion of RNA to cDNA can be performed using any of the methods known in the art for such a conversion, such as using reverse transcriptase in a reverse transcription reaction. cDNA does not exist in vivo and therefore is a non-natural molecule. Besides cDNA not existing in vivo, cDNA is necessarily different than RNA (such as transfer RNA (tRNA)), as it includes deoxyribonucleic acid and not ribonucleic acid. The cDNA can then be amplified, for example, by the polymerase chain reaction (PCR) or other amplification method known to those of ordinary skill in the art. For example, other amplification methods that may be employed include the ligase chain reaction (LCR) (Wu and Wallace, Genomics, 4:560 (1989), Landegren et al., Science, 241 :1077 (1988), incorporated by reference in its entirety for all purposes, transcription amplification (Kwoh et al., Proc. Natl. Acad. Sci. USA, 86:1173 (1989), incorporated by reference in its entirety for all purposes), selfsustained sequence replication (Guatelli et al., Proc. Nat. Acad. Sci. USA, 87:1874 (1990), incorporated by reference in its entirety for all purposes), incorporated by reference in its entirety for all purposes, and nucleic acid based sequence amplification (NASBA). Guidelines for selecting primers for PCR amplification are known to those of ordinary skill in the art. See, e.g., McPherson et al., PCR Basics: From Background to Bench, Springer-Verlag, 2000, incorporated by reference in its entirety for all purposes. The product of this amplification reaction, i.e., amplified cDNA is also necessarily a nonnatural product. First, as mentioned above, cDNA is a non-natural molecule. Second, in the case of PCR, the amplification process serves to create hundreds of millions of cDNA copies for every individual cDNA molecule of starting material.
[0094] In some examples, the one or more nucleic acid sequences may be derived from one or more samples obtained from a subject, wherein the one or more samples are processed and subjected to sequencing techniques. In some examples, the sequencing techniques to produce the one or more nucleic acid sequences may include any method known in the art such as but is not limited to Next-Generation Sequencing (NGS), of biological samples, and the like. The NGS system used can be any NGS system known in the art. In one example, the cDNA is amplified with primers that introduce an additional DNA sequence (e.g., adapter) onto the fragments (e.g., with the use of adapter-specific primers) that make the amplified cDNA amendable to an NGS sequencing platform. The methods described herein can be useful for sequencing by the method commercialized by Illumina, as described U.S. Pat. Nos. 5,750,341 ; 6,306,597; and 5,969,119. Complementary DNA (cDNA) products can be prepared as described herein and can then be denatured and can be randomly attached to the inside surface of flow-cell channels. Unlabeled nucleotides can be added to initiate solid-phase bridge amplification to produce dense clusters of double-stranded DNA. To initiate the first base sequencing cycle, four labelled reversible terminators, primers, and DNA polymerase can be added. After laser excitation, fluorescence from each cluster on the flow cell can be imaged. The identity of the first base for each cluster can then be recorded. Cycles of sequencing are performed to determine the fragment sequence one base at a time. In some examples, the methods described herein may involve obtaining nucleic acid sequence reads prepared from sequencing by ligation methods (e.g. commercialised by Applied Biosystems SOLiD sequencing), sequencing by synthesis using the methods (e g. commercialised by 454 / Roche Life Sciences).
[0095] In some examples, the one or more nucleic acid sequences may be obtained through methods in the art such as but are not limited to polymerase chain reaction (PCR) free library preparation or sequencing approaches, and the like. As such, the length of the one or more nucleic acid sequences is dependent on the devices or apparatus used for the method as disclosed herein. In some examples, the one or more nucleic acid sequences are obtained by sequencing a nucleic acid library comprising 10 to 300 nucleotides. In some examples, the nucleic acid library comprises 10 to 300 nucleotides, 10 to 200, 10 to 150, 10 to 100, 10 to 90, 10 to 80, 10 to 70, 20 to 70, 30 to 70, 40 to 60 nucleotides. In some examples, the nucleic acid library may include 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21 , 22, 23, 24, 25, 26, 27, 28, 29, 30, 31 ,
[0096] 32, 33, 34, 35, 36, 37, 38, 39, 40, 41 , 42, 43, 44, 45, 46, 47, 48, 49, 50, 51 , 52, 53,
[0097] 54, 55, 56, 57, 58, 59, 60, 61 , 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75,
[0098] 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91 , 92, 93, 94, 95, 96, 97,
[0099] 98, 99, 100, 101 , 102, 35 103, 104, 105, 106, 107, 108, 109, 110, 111 , 112, 113, 114, 115, 116, 117, 118, 119, 120, 121 , 122, 123, 124, 125, 126, 127, 128, 129, 130,
[0100] 131 , 132, 133, 134, 135, 136, 137, 138, 139, 140, 141 , 142, 143, 144, 145, 146, 147,
[0101] 148, 149, 150, 151 , 152, 153, 154, 155, 156, 157, 158, 159, 160, 161 , 162, 163, 164,
[0102] 165, 166, 167, 168, 169, 170, 171 , 172, 173, 174, 175, 176, 177, 178, 179, 180, 181 ,
[0103] 182, 183, 184, 185, 186, 187, 188, 189, 190, 191 , 192, 193, 194, 195, 196, 197, 198,
[0104] 199, 200, 201 , 202, 203, 204, 205, 206, 207, 208, 209, 210, 211, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221 , 222, 223, 224, 225, 226, 227, 228, 229, 230, 231 ,
[0105] 232, 233, 234, 235, 236, 237, 238, 239, 240, 241 , 242, 243, 244, 245, 246, 247, 248,
[0106] 249, 250, 251 , 252, 253, 254, 255, 256, 257, 258, 259, 260, 261 , 262, 263, 264, 265, 266, 267, 268, 269, 270, 271 , 272, 273, 274, 275, 276, 277, 278, 279, 280, 281 , 282, 283, 284, 285, 286, 287, 288, 289, 290, 291 , 292, 293, 294, 295, 296, 297, 298, 299, 300 nucleotides, and the like.
[0107] Without wishing to be bound by theory, the method of the present disclosure is advantageous in detecting unique nucleic acid sequence reads that may be missed by methods in the art where the nucleic acid library is constructed with smaller size range (such as 22±5 nucleotides) as most nucleic acid libraries in public databases are focused on miRNAs. Methods in the art may miss out on nucleic acid sequences whose lengths are longer 22±5 nucleotides.
[0108] Sequencing read length
[0109] In some examples, the one or more nucleic acid sequences may include lengths of more than 27 bp. In some examples, the one or more nucleic acid sequences may include lengths of 10 to 1000 nucleotides, 20 to 900 nucleotides, 30 to 800 nucleotides, 40 to 700 nucleotides, 50 to 600 nucleotides, 60 to 500 nucleotides, 70 to 400 nucleotides, 80 to 300 nucleotides, 90 to 200 nucleotides, 100 to 150 nucleotides, and the like. In some examples, the one or more nucleic acid sequence may include lengths of 10, 11 , 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41 , 42, 43, 44, 45, 46, 47, 48, 49 or 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950 or 1000 nucleotides. In some examples, the one or more nucleic acid sequence comprises lengths of 10 to 40 nucleotides.
[0110] Without wishing to be bound by theory, the inventors of the present disclosure have found that method of the present disclosure identifies nucleic acid sequence that are longer than 22±5 nucleotides. The nucleic acid sequence that is longer than 22±5 nucleotides may include annotated sequence and unannotated sequence.
[0111] Alignment to reference genome
[0112] In the process of analysing the sequencing data, an alignment process may be performed to determine where particular sequence reads map to a reference sequence to identify regions of the reference sequence that are similar to the sequence reads.
[0113] In some examples, the reference sequence may be selected from a reference genome of a type of organism, such as but is not limited to a human, and the like. In some examples, the reference sequence may be from a synthetic genome for experimental controls.
[0114] In some examples, the reference genome comprises one or more sub- genomic regions of the reference genome. In some examples, the sub-genomic regions comprise annotated genomic RNAs. In some examples, the annotated genomic RNAs may include but are not limited to one or more of transfer RNA (tRNAs), microRNAs (miRNAs), ribosomal RNA (rRNAs), long non-coding RNAs (IncRNAs), small nucleolar RNAs (snoRNAs), small nuclear RNAs (snRNAs), and the like. As such, aligning the sequence reads to a reference sequence may allow for identification of genomic locations or coordinates within the reference sequence for the organism that the sequence reads map to. In addition, alignment of the sequence reads to a reference genome / sequence helps identify the functionality of the sequence (for annotated sequences).
[0115] In some examples, the one or more nucleic acid sequence reads may be aligned to a portion of a reference genome, for example, to specific a tRNA gene, such as but not limited to tRNA-Ala, tRNA-Arg, other tRNA genes, and the like. As will be appreciated by a person skilled in the art, the alignment obtained would be specific to the portion of the reference genome.
[0116] In some examples, the annotated genomes or genomes with annotated regions are used as the reference genome sequence for alignment. Genome annotation is important as the sequencing of the genome or DNA generates sequence information without the sequence identity or its functional role. Therefore, such genome annotations allow deducing the identity or the origin of the DNA sequence along the genome. Based on the identified genomic coordinates and the annotation, one can learn more information in relation to the biological origin of the DNA sequence and, if available, the structural and functional features or information of the DNA sequence.
[0117] Without necessarily bound by theory, an alignment may refer to aligning a nucleic acid sequence reads to another nucleic acid sequence reads a nucleic acid sequence reads to a reference genome (or a portion thereof); aligning a plurality of nucleic acid sequence reads to a reference genome; and / or aligning a nucleic acid sequence reads to more than one reference genome (or more than one portions of a reference genome). In some examples, an alignment may refer to an aligned or partially aligned nucleic acid sequence reads to a reference genome. In some examples, in an end-to-end alignment to a reference sequence, an alignment may include one or more aligned portions and empty portions (no alignment). In some examples, the alignment file may include but is not limited to Sequence Alignment Map (SAM), Binary Alignment Map (BAM) file format, multiple sequence alignment (MSA), and the like.
[0118] In some examples, the aligning of the one or more nucleic acid sequence reads to a reference genome comprises no base mismatch, one base mismatch or two base mismatches. Without wishing to be bound by theory, the inventors of the present disclosure allow up to two mismatches during genome alignment to discover as many nucleic acid sequences (such as tRNA related sequences) as possible. Having one or two base mismatches in the alignment may allow more sequences relevant to the reference genome to be uncovered, thereby allowing a more robust discovery process. Allowing one or two base mismatches also accounts for potential sequencing errors (i.e., false negatives).
[0119] In some examples, aligning the nucleic acid sequence reads to a reference genome to provide an alignment may further comprise converting the alignment to a multiple sequence alignment format at selected portions of the reference genome.
[0120] In some examples, the method comprises converting the alignment to a multiple sequence alignment format. In some examples, multiple sequence alignment is built using algorithms in the art such as but is not limited to progressive algorithm, iterative algorithm, dynamic programming-based algorithms, and the like. In some examples, multiple tools in the art may be used for multiple sequence alignment, such as but is not limited to Clustal Omega, MUSCLE, and the like.
[0121] In some examples, aligning the nucleic acid sequence reads to a reference genome to provide an alignment comprises obtaining alignment to one or more sub- genomic regions of the reference genome / portion of the reference genome. In some examples, the sub-genomic regions comprise annotated genomic RNAs.
[0122] In some examples, the reference genome / the portion of the reference genome may be selected from one or more of tRNAs, miRNAs, rRNAs, IncRNAs, snoRNAs, snRNAs, any other RNAs with genomic annotations, and the like.
[0123] In some examples, the annotated genomic RNAs may be selected from one or more of tRNAs, miRNAs, rRNAs, IncRNAs, snoRNAs, snRNAs, any other RNAs with genomic annotations, and the like.
[0124] Identifying boundaries using a statistical model In some examples, the method comprises identifying a plurality of boundaries / probable boundaries of the one or more nucleic acid sequences using a statistical model, wherein each boundary corresponds to a nucleotide position along the reference genome.
[0125] In some examples, the statistical model comprises Hidden Markov Model and / or Dynamic Bayesians networks.
[0126] In some examples, the statistical model may use algorithms such as but is not limited to a Hidden Markov model, Bayesian networks, and the like.
[0127] In some examples, the statistical model is a Hidden Markov Model.
[0128] In some examples, the statistical model is fitted with a two-state Gaussian Hidden Markov or a two-state Poisson Hidden Markov model.
[0129] In some examples, the statistical model may be fitted with more than two states. In some examples, the two states may be defined as transcriptionally active and inactive, or the presence or absence of authentic sequencing read.
[0130] In some examples, the statistical model of the present disclosure is able to filter nucleic acid sequence reads that are from technical errors (such as “background noise” and to focus on nucleic acid sequence reads with high confidence / that are from the biological target nucleic acid molecule of interest. In some examples, the statistical model of the present disclosure can process nucleic acid reads accurately, and with minimal error as possible to determine the amount / quantity of nucleic acid molecules, which levels indicate a disease status in a subject.
[0131] Selecting boundaries from identified boundaries
[0132] In some examples, the method comprises selecting the boundaries satisfying one or more pre-determined conditions of having:
[0133] (i) a probability of the probable boundaries at or above a pre-determined probability threshold; and
[0134] (ii) a frequency of number of biological samples in which the boundary was observed at or above a pre-determined frequency threshold.
[0135] In some examples, the pre-condition (I) is estimated using the Viterbi algorithm.
[0136] Without wishing to be bound by theory, the Viterbi algorithm is a dynamic programming algorithm used for determining the most probable path (presence of sequence alignment) and probability to arrive at a particular hidden state. Typically, when used with the Hidden Markov Model, it requires knowledge of the parameters of the Hidden Markov Model and a particular output state, and it finds the state sequences that are most likely to have generated that output sequence. It works by finding a maximum over all possible state sequences.
[0137] In some examples, the probability in pre-condition (i) may include one or more types of probability, such as but is not limited to emission, posterior, transition probability, and the like.
[0138] As will be appreciated by the person skilled in the art, the probability is typically associated with the type of statistical models and algorithm used. For example, posterior probability, which relates to updated probability of a state after taking into account new information, is largely associated with Viterbi algorithm.
[0139] In some examples, the selected boundaries / probable boundaries may have a probability at or above a probability threshold / pre-determined probability threshold. Such a frequency threshold may be determined and adjusted accordingly depending on the application of use or output to be achieved. In some examples, the probability threshold / pre-determined probability threshold is selected from 50% to 100% probability. In some examples, the probability threshold / pre-determined probability threshold may include 50, 51 , 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71 , 72, 73, 74, 75, 76, 77, 78, 79, 80, 81 , 82, 83, 84, 85, 86, 87, 88, 89, 90, 91 , 92, 93, 94, 95, 96, 97, 98, 99, or 100% probability.
[0140] In some examples, the probability threshold / pre-determined probability threshold is selected from 50% to 100% probability. In some examples, the probability threshold / pre-determined probability threshold may include 55% to 95%, 60% to 95%, 70% to 95%, 80% to 95%, 80% to 90% probability, and the like.
[0141] In further examples, the boundaries as selected may have a frequency of number of samples / biological samples in which the boundary was observed at or above a frequency threshold / pre-determined frequency threshold. Such a frequency threshold may be determined and adjusted accordingly depending on the application of use or output to be achieved.
[0142] In some examples, the frequency threshold / pre-determined frequency threshold is a frequency of the boundaries observed in 2% to 100% of the biological samples. In some examples, the frequency threshold / pre-determined frequency threshold may include 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14, 15, 16, 17, 18, 19, 20,
[0143] 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31 , 32, 33, 34, 35, 36, 37, 38, 39, 40, 41 , 42,
[0144] 43, 44, 45, 46, 47, 48, 49, 50, 51 , 52, 53, 54, 55, 56, 57, 58, 59, 60, 61 , 62, 63, 64,
[0145] 25 65, 66, 67, 68, 69, 70, 71 , 72, 73, 74, 75, 76, 77, 78, 79, 80, 81 , 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 or 100% of the samples / biological samples.
[0146] In some examples, the pre-determined frequency threshold is a frequency of the boundaries observed in 2% to 100% of the biological samples. In some examples, the frequency threshold / pre-determined frequency threshold may include 5% to 100%, 5% to 90%, 5% to 80%, 5% to 70%, 5% to 60%, 5% to 50%, 5% to 40%, 5% to 30%, 5% to 20%, 5% to 10% of the biological samples.
[0147] Obtaining nucleic acid sequences bounded by the selected boundaries
[0148] In some examples, the nucleic acid sequences are bound by a pair of selected boundaries, wherein the pair of selected boundaries corresponds to the start and end position of the nucleic acid sequence.
[0149] In some examples, the selected boundary is determined to be the start and / or end positions of the nucleic acid sequences based on the proportion of sequence reads having start or end positions occurring at the selected boundaries.
[0150] In some examples, the method of obtaining nucleic acid sequences bounded by the selected boundaries further comprises determining the directions of the nucleic acid sequence reads bounded by the selected boundaries. As used herein, the direction of nucleic acid sequences and sequence reads refers to the arrangement or order of nucleotides in a continuous manner, which is typically denoted in the 5’ to 3’ direction, unless stated otherwise.
[0151] In some examples, the method further comprises allocating the nucleic acid sequences to be an annotated nucleic acid, an unannotated nucleic acid and an irrelevant nucleic acid.
[0152] Quantifying the number of obtained nucleic acid sequences to identify one or more nucleic acid sequences)
[0153] In some examples, the method of quantifying the number of obtained nucleic acid sequences further comprises using the genomic coordinates of the most probable / authentic nucleic acid sequences (as quantified) to identify and quantify said nucleic acid sequences in one or more biological samples.
[0154] In some examples, the most probable / authentic nucleic acids or most probable fragmentation boundary comprises the nucleic acid / boundary within the 10% threshold. Without wishing to be bound by theory, as shown in FIG. 2C, the method showed narrowing down a whole array of detected nucleic acid sequences to a few final nucleic acid sequences. Therefore, the final unique nucleic acid (such as tRF) sequences are the most probable / authentic nucleic acid sequences while the other nucleic acid sequences are experimentally detected derivative sequences of these nucleic acid sequences.
[0155] As such, depending on the sample / biological samples used, this can help identify and quantify nucleic acid sequences that are differentially expressed (overexpressed or under-expressed) between diseased and non-diseased biological samples, or between biological samples obtained from healthy / non-diseased subjects and diseased subjects.
[0156] Obtaining scoring matrices
[0157] In some examples, the method further comprises obtaining one or more scoring matrices based on nucleic acid sequences obtained from diseased or nondiseased biological samples, or biological samples of diseased or non-diseased subjects.
[0158] In some examples, the method comprises obtaining two scoring matrices based on diseased biological samples and non-diseased biological samples, or biological samples of diseased or non-diseased subjects.
[0159] Without wishing to be bound by theory, the scoring matrix of the present disclosure is derived from the study samples (i.e., training cohort), before being applied to a separate set of study samples (i.e., validation cohort). In the present disclosure, two scoring matrices are generated, wherein the first scoring matrix is from a first sample (such as a diseased / cancer subject) and the second scoring matrix is from a second sample (such as non-diseased / healthy subject). For example, reads obtained from cancer subjects will show a high score on the first scoring matrix and a low score for the second scoring matrix. Sample scores summarize all the read scores and hence indicate the likelihood of the sample type.
[0160] In some examples, each scoring matrix is represented by a modified positionspecific scoring matrix (PSSM). In some examples, the modified PSSM comprises likelihood of occurrence of each nucleotide sequence (A, T, G, C) and being empty / missing at different positions along the portion of an annotated genomic locus. In some examples, the modified PSSM may include modifications to the original PSSM with some modifications to some modifications to accommodate sequence variants like indels and to accommodate end-to-end alignment of fragmented sequences with many missing bases. In some examples, the modification to the original PSSM may include such as but is not limited to, (1) incorporating “empty base” as part of the sequence profile to accommodate indels, (2) dealing with the ambiguous sequencing base call “N”, and (3) modifying the position weight probability matrix based on the relative base coverage to accommodate differential coverage along the gene, and the like.
[0161] In some examples, each scoring matrix is represented by a position-specific scoring matrix (PSSM). In some examples, each scoring matrix is represented by a position-specific scoring matrix (PSSM), wherein said PSSM comprises likelihood of occurrence of each nucleotide sequence (A, T, G, C) and being empty at different positions along the portion of an annotated genomic locus (such as one tRNA gene), and wherein the scoring matrices are obtained from a plurality of different types of biological samples such as diseased or non-diseased biological samples, or diseased or non-diseased / healthy subjects. In some examples, the PSSM is built per target (such as one tRNA gene) or with one annotated genomic locus.
[0162] In some examples, the one or more scoring matrices are generated from the following steps: i) obtaining one or more nucleic acid sequence reads from diseased or nondiseased biological samples, or biological samples of diseased or non-diseased subjects or biological samples of different types; ii) aligning the nucleic acid sequence reads to a portion of a reference genome to provide an alignment having one or more aligned and empty portions; iii) determining from the alignment the likelihood of occurrence of each nucleotide (A, T, G, C) and being empty at different positions along the portion of the reference genome; iv) generating a scoring matrix for each portion of the reference sequence using the likelihood of occurrence of each nucleotide (A, T, G, C) and being empty at each nucleotide position of the reference genome.
[0163] Without wishing to be bound by theory, the algorithm of the present disclosure to build the scoring matrices was designed to accommodate sequence variants like insertions and deletions (indels), and to accommodate end-to-end alignment of fragmented sequences with many missing bases. The algorithm includes: (1) incorporating “empty base” as part of the sequence profile to accommodate indels, (2) dealing with the ambiguous sequencing base call “N”, and (3) modifying the position weight probability matrix based on the relative base coverage to accommodate differential coverage along the gene. Typically, PSSM is meant for end-to-end aligned sequences of the same length. Since fragmentation and missing base information are important to distinguish nucleic acid sequence (such as tRF) patterns, the inventors of the present disclosure created the 5th nucleotide base to signify empty or missing base. This requires consideration of what the basal frequency of the missing base is, what the substitution rate would be between the four conventional nucleotides (A,T,C,G) and the missing base, in terms of the relative ratio.
[0164] It was assumed that there was a 1% chance of mismatch (similar to what PAM1 matrix), 0.1% chance of transition between base to no base when building the substitution matrix. This allows a high penalty for indels or introducing a missing base, and vice versa. Additionally, it was assumed that it was 5 times more likely that any given base would be a missing base and not one of the conventional nucleotides since many of the bases along the entire tRNA genes are “empty” from the sequencing data (such as small RNAseq data). The ambiguous sequencing base call, all the “N” bases in sequencing data, was augmented by adjusting the derived position weight matrix with additional split probability. For instance, 2 counts of N base calls were changed to 0.5 (2 divided by 4) counts to each A, T, C, and G base calls. Final position weight probability matrices were adjusted based on the relative base coverage, which created the system to assign higher scores if fragments occurred at coordinates with high base coverage.
[0165] Without wishing to be bound by theory, as methods in the art require end-to-end alignment, fragmented nucleic acids of varying lengths cannot be applied to the method. Advantageously, in the present disclosure, the presence of missing bases allows for the indication of the start and the end of nucleic acid fragments which allows the capture of fragmentation patterns. The method of the present disclosure includes modifying the PSSM as conventional PSSM is limited to A, T, C, G nucleotides. Due to the modification, the present disclosure discloses a PSSM that is different from those in the prior art.
[0166] Calculating sample scores
[0167] In some examples, the method of the present disclosure comprises: calculating the sample score based on the sum of the individual read scores of each nucleic acid sequence read, where in the individual read score of each nucleic acid sequence read is calculated using one or more modified PSSM,
[0168] In some examples, the method of the present disclosure may include calculating a sample score for a plurality of sample groups according to the wherein p1 , [32, ... , pN represent the coefficients; r represents the sequence read;
[0169] R represents the total number of nucleic acid sequence reads in the sample;
[0170] S1, S2, ... SN represent the read scores derived from the first, second, ... Nth modified PSSM, and wherein the sample score represents the condition likelihood.
[0171] In some examples, the sample score is calculated according to the following equation: r is a read of a nucleic acid sequence read;
[0172] R is the total number of nucleic acid sequence reads in the sample;
[0173] 51 is the first read score calculated using a first modified PSSM obtained from the samples of the first sample group;
[0174] 52 is the second read score calculated using a second modified PSSM obtained from the samples of the second sample group;
[0175] SN is the Nth read score calculated using a Nth modified PSSM obtained from the samples of the Nth sample group.; and pi, [32, pN, are the coefficients.
[0176] Without wishing to be bound by theory, If there are more than two sample groups, there will be multiple scoring matrices, and the sample score will be a collection of the sum of the read scores derived from the multiple scoring matrices. That is, the sample score may include a vector (i.e., multiple numbers). Different methods in the art may be used to classify samples using the collection of the sa ple scores, such as but is not limited to clustering, maximum likelihood calculation, or multi-class classification, and the like. The equation of the present disclosure represents a maximum likelihood calculation.
[0177] In some examples, the method of the present disclosure may include calculating a sample score for two sample groups according to the following equation: wherein r represents the sequence read;
[0178] R represents the total number of sequence reads in the sample library;
[0179] 51 represents the first read score derived from the first modified PSSM; and
[0180] 52 represents the second read score derived from the second modified PSSM.
[0181] In some examples, the sample score is calculated according to the following equation: wherein r is a read of a nucleic acid sequence read;
[0182] R is the total number of nucleic acid reads in the sample;
[0183] 51 is the first read score calculated using a first modified PSSM obtained from the samples of the first sample group;
[0184] 52 is the second read score calculated using a second modified PSSM obtained from the samples of the second sample group;
[0185] Without wishing to be bound by theory, if there If there are two sample groups, there will be two scoring matrices, and the sample score will be the sum of the differences of the read scores derived from the two scoring matrices. In this case, the first scoring matrix will be of the diseased group so the sample score will represent how similar the sample is to the diseased group (first group) and dissimilar to the healthy group (second group).
[0186] In some examples, the method may include calculating a sample score for one sample group according to the following equation: wherein r represents the sequence read;
[0187] R represents the total number of sequence reads in the sample library; and
[0188] S represents the read score derived from the modified PSSM.
[0189] In some examples, the sample score is calculated according to the following equation: samp ,le score wherein r is a read of a nucleic acid sequence read; R is the total number of nucleic acid reads in the sample;
[0190] S is the read score calculating using a modified PSSM obtained from the sample.
[0191] Without wishing to be bound by theory, the most basic way to calculate the sample score includes with one sample group. In this case, there will be one scoring matrix (i.e., modified PSSM), and the sample score will be the sum of the read scores derived from the one scoring matrix. The sample score will represent how similar one sample is to the one sample type of our interest.
[0192] In some examples, the method comprises calculating the sample score based on the sum of the individual read score of each nucleic acid sequence reads, wherein the individual read score of each nucleic acid sequence read is calculated based on the difference between: a first read score calculated using a first PSSM; and a second read score calculated using a second PSSM.
[0193] In some examples, the method comprises calculating the one or more sample scores according to the following equation: wherein r is a read of a nucleic acid sequence (such as nucleic acid sequence reads);
[0194] Read Score.Dr is the first read score calculated using a first PSSM obtained from diseased samples or samples of diseased subjects;
[0195] Read Score. Nr is the second read score calculated using a second PSSM obtained from non-diseased samples or samples of non-diseased subjects; and
[0196] R is the total number of reads.
[0197] In some examples, the scoring matrices obtained from samples of the first sample group comprises a first group and the second or more sample comprises a second or more groups, wherein the first and second or more groups are selected from a plurality of distinct sample groups.
[0198] In some examples, the method comprises calculating the sample score based on the sum of the individual read score of each nucleic acid sequence read, wherein the individual read score of each nucleic acid sequence read is calculated based on the difference between: a first read score calculated using a first PSSM obtained from diseased samples or samples of diseased subjects; and a second read score calculated using a second PSSM obtained from nondiseased samples or samples of non-diseased subjects.
[0199] In some examples, when a sample / biological sample is determined or predicted to be a diseased sample, the sample score of said biological sample is higher than that of a non-diseased sample. In some examples, wherein when a biological sample is determined or predicted to be a diseased sample, the first read score (Sir) is higher than the second read score (S2r). tRNA fragments
[0200] In some examples, the one or more nucleic acids comprise transfer ribonucleic acids (tRNAs), microRNAs (miRNAs), ribosomal ribonucleic acids (rRNAs), long noncoding RNAs (IncRNAs), small nucleolar RNAs (snoRNAs), small nuclear RNAs (snRNAs), yRNA, piwi-interacting RNA (piRNA) and / or fragments thereof.
[0201] In some examples, the one or more nucleic acid sequence may include but is not limited to transfer ribonucleic acids (tRNAs), microRNAs (miRNAs), ribosomal ribonucleic acids (rRNAs), long non-coding RNAs (IncRNAs), small nucleolar RNAs (snoRNAs), small nuclear RNAs (snRNAs) and / or fragments thereof, and the like.
[0202] In some examples, the one or more nucleic acid sequence comprises tRNA fragments. In some examples, the one or more nucleic acid sequences may include but is not limited to tRNA-Lys or its fragments thereof, tRNA-Glu or its fragments thereof, tRNA-Pro or its fragments thereof, and the like. In some examples, the one or more nucleic acid sequences may include but is not limited to TCCCCGGCACCTCCA (SEQ ID NO: 1), AGCGCCAAATCCTAACCACTAGACCACAG (SEQ ID NO: 2), GGTAGCGTGGCCGAGC (SEQ ID NO: 3), TCTCTTCGGGGGCGTGG (SEQ ID NO: 4), CTCTTCGGGGGCGTGGGT (SEQ ID NO: 5), GGTTCAAATCCCGGACGAGCCC (SEQ ID NO: 6), GTCACGGTGGCCGAGTGG (SEQ ID NO: 7), GAAAGCGCCGAATCCTAGCCACTAGACCACCAGGGA (SEQ ID NO: 8), or at least 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to the nucleic acid sequences thereof, or 1, or 2, or 3, or 4, substitutions thereof, and the like. tRNA fragments (tRFs) have garnered attention for their potential diagnostic value in cancer, emerging as a noteworthy small RNA-based diagnostic biomarker in liquid biopsy. However, quantifying these fragments poses challenges due to their closely related sequences and the scarcity of databases. tRFs are generated through the cleavage of mature tRNA molecules, and while their biogenesis is becoming clearer, the comprehensive catalogue of tRFs remains less characterised and curated compared to miRNAs. The current list of known tRFs, as documented in databases like tRFdb, is relatively limited. Consequently, many studies detect tRFs by searching for sequences reported in existing databases, which inherently depends on previously published findings. This highlights the need for expanded research and better characterisation of tRFs and fragments of nucleic acid to fully elucidate their role and utility in disease diagnostics
[0203] In the present disclosure, two small RNA sequencing based pipelines were developed that are independent of tRF annotations, aimed at quantitatively characterising tRNA fragmentation patterns. The inventors of the present disclosure evaluated the diagnostic efficacy of the derived tRF-based features using plasma samples from colorectal cancer (CRC) patients.
[0204] Small RNA-sequencing was conducted on 129 CRC tissues (68 tumours; 62 adjacent normal), along with plasma samples from 50 CRC patients and 107 healthy individuals. MicroRNA-seq data from TCGA (The Cancer Genome Atlas), specifically from TCGA-COAD (Colon Adenocarcinoma) and READ (Rectum Adenocarcinoma) cohorts, were also used in this study. Datasets included 616 tumours and 11 adjacent normal tissue samples.
[0205] Frequent tRNA fragmentation boundaries were first identified in CRC tissue data. These coordinates were used to obtain the tRF expression count matrix in plasma samples. The second method of quantifying one or more nucleic acid sequences in the present disclosure measures the similarity of each sample to a collection of CRC tumour tissues using tumour similarity scores (tScores). Scoring matrices were constructed based on tRNA mapped read profiles in tumour tissues and were used to generate scores to each plasma sample. Linear support vector classifier model explored different feature combinations (tRF counts, and tScores). Through 4x5 cross-validation on plasma data, the inventors of the present disclosure assessed the significance of tRF-based features in detecting CRC in plasma samples.
[0206] Annotated vs unannotated sequences / uncovering novel nucleic acid
[0207] In some examples, the method of the present disclosure identifies nucleic acid sequences comprising annotated sequences and unannotated sequences. In some examples, the annotated sequences were obtained from a nucleic acid database (such as a tRF database). In some examples, the one or more nucleic acid sequences may include but are not limited to the fragments of tRNA-Lys, tRNA-Glu, tRNA-Pro, and the like.
[0208] The approach of the present disclosure successfully uncovered novel tRFs within CRC samples. Among the 1254 tRFs incorporated in the feature matrix, 270 tRFs were annotated sequences in a database (such as MINTbase v2.0). Including the nonannotated tRFs broadened the feature space and markedly improved model performance (p-value: 0.0126), as evidenced by an improved Area Under the Curve (AUC) of 0.859 compared to the model comprising solely of annotated tRFs (AUC: 0.781). Notably, the inventors of the present disclosure observed substantial disparities in tRNA read patterns between tissue and plasma samples, as well as across different datasets. This discrepancy was underscored by the presence of numerous samplespecific tRFs and the low tScores observed, indicative of the poor retention of tumor tissue tRNA read profiles in plasma samples. Furthermore, the analysis of the present disclosure revealed weaker model performance by tScores (AUC of 0.752) compared to tRF counts (AUC of 0.859).
[0209] The results of the present disclosure underscore the ability of the method of the present disclosure to identify novel unannotated tRFs, thereby expanding the range of features available for use in cancer detection models.
[0210] Sample
[0211] As used herein, the term "sample / samples" may be used interchangeably with "sample group / sample groups".
[0212] In some examples, the sample may include a biological sample. A biological sample may be from a cell or tissue known to be diseased, non-diseased or suspected of being diseased. Such a sample may be cancerous, non-cancerous or suspected of being cancerous. Blood or peripheral blood sample can, for example, include whole blood, plasma, serum, or any derivative of blood, and the like.
[0213] Examples of such samples / biological samples may include, but are not limited to, biopsies, smears, blood, lymph, urine, saliva, or any other bodily secretion or derivative thereof, and the like. In some examples, the biological sample is a liquid biological sample. In some examples, the biological sample is a non-cellular biological fluid. In some examples, the biological sample may include but is not limited to one or more of a tissue sample (such as a biopsy tissue sample), a bodily fluid sample (such as nasopharyngeal secretion, urine, blood sample), and a non-cellular bodily fluid sample (such as plasma or serum), and the like. In some examples, the sample / biological sample comprises non-cellular bodily fluid sample, such as plasma and / or serum. Samples can be obtained from a subject by a variety of techniques, which are known to those skilled in the art. In some examples, the sample is a liquid biopsy sample. In some examples, the liquid biopsy sample may include but is not limited to serum, plasma, platelet poor plasma, urine, cerebrospinal fluid, and the like. In some examples, the sample is a bodily fluid sample or non-cellular bodily fluid sample. In some examples, the sample is a plasma sample.
[0214] Accordingly, a sample can include without limitation a cryosection of a fresh frozen biopsy, a formalin-fixed paraffin embedded (FFPE) tissue biopsy, a cryopreserved diagnostic cell suspension, a bodily fluid sample (such as nasopharyngeal secretion, urine, peripheral blood, blood sample) and a non-cellular bodily fluid sample (such as plasma or serum), and the like. In some examples, the sample may include a tissue sample and / or a liquid biological sample.
[0215] In some examples, a sample / biological sample may be isolated from a subject having one or more diseases, such as, but is not limited to cancers (such as colorectal cancer, lung cancer, gastric cancer, breast cancer, liver cancer and prostate cancer), and the like. In some examples, the sample / biological sample may be from a healthy subject, a non-diseased subject and / or a cancer-free subject.
[0216] In some examples, the sample is obtained from a plurality of subjects to provide a plurality of samples for the method of the present disclosure.
[0217] In some examples, the plurality of samples may refer to samples obtained from multiple sampling sites of a single individual subject (i.e. samples may be obtained from various types of biological samples). In some examples, the plurality of samples may refer to samples obtained from multiple individual subjects (i.e. samples may be obtained from a population having substantially homogenous biological profile). In some examples, the plurality of samples may refer to samples obtained from multiple variable subject populations (i.e. samples may be obtained from two or more populations having different biological profiles).
[0218] In some examples, the method of the present disclosure can be applied to a plurality of samples from a plurality of biological backgrounds to distinguish each sample type. In some examples, the method of the present disclosure can be applied to disease subtyping e.g., between healthy individuals with predispositions to one or more disease / disease subtypes. In some examples, the method of the present disclosure can be applied to different conditions such as between healthy individuals with one or more physiological states (e.g., athletes who may or may not benefit from further training or supplementation).
[0219] In some examples, the plurality of samples comprises two samples, the samples comprises a first sample representing a first classification group, a second sample representing a second classification group, wherein the first and second groups are selected from a plurality of distinct sample groups.
[0220] In some examples, the plurality of samples may comprise x number of samples, where x is a positive integer (i.e. 1 , 2, 3, etc). Without wishing to be bound by theory, the number of samples (i.e. x, the integer) depends on the sophistication of the apparatus used to perform the calculation or the capacity of the devices employed to perform the methods. Therefore, in some examples, the plurality of samples comprises x samples, wherein the samples comprises a first sample representing a first classification group, an xth sample, representing an nth classification group,
[0221] In some examples, the samples may refer to a first sample with a first group and the second or more sample comprises a second or more group, wherein the first and second or more groups are selected from a plurality of distinct sample groups from a plurality of subjects. In some examples, the sample group may include other condition-specific classifications.
[0222] In some examples, the first sample may include a sample obtained from a subject suffering from a first disease subtype and the second or more sample may include a sample from a subject suffering from a second or more disease subtype. For example, the first sample may be a sample obtained from a person suffering from a subtype of a cancer and the second or more sample may be a sample obtained from a person suffering from another subtype of a cancer.
[0223] In some examples, the first sample comprises a first classification group and the second or more sample comprises a second or more classification group, wherein the first and second or more groups are selected from a plurality of distinct sample groups.
[0224] In some examples, the first sample may include but is not limited to a diseased sample, a cancer sample, a cancer subtype, and the like. In some examples, the first sample group may include but is not limited to a diseased group, a cancer group, a cancer subtype group, an untreated group and the like. In some examples, the second or more sample may include but is not limited to a non-diseased sample, a healthy sample, a sample from a cancer-free subject, another cancer subtype, and the like. In some examples, the second or more sample group may include but is not limited to a non-diseased group, a healthy group, a treated group, and the like.
[0225] Also disclosed herein are apparatuses for performing the methods as disclosed herein.
[0226] Control
[0227] In some examples, the control may include but is not limited to one or more samples / biological samples obtained from a healthy subject, a non-diseased subject, a cancer-free subject, and the like.
[0228] Disease
[0229] In some examples, the disease may include a cancer selected from such as but is not limited to a colorectal cancer, a lung cancer, a gastric cancer, a breast cancer, a liver cancer, a prostate cancer, and the like. In some examples, the disease may include one or more types of cancer, such as but is not limited to colorectal cancer, lung cancer, gastric cancer, breast cancer, liver cancer, prostate cancer, and the like.
[0230] In some examples, the disease is colorectal cancer.
[0231] Computer-implemented method
[0232] In some examples, the method of quantifying nucleic acid sequences may include a computer-implemented method. In some examples, all or some of the steps in the method of quantifying one or more nucleic acid sequences shall be performed by a suitably programmed computer.
[0233] The above-described examples of the present disclosure can be implemented in numerous ways. For example, the examples may be implemented using hardware, software or a combination thereof. When implemented in software, the software code can be executed on any suitable processor or collection of processors, whether provided in a single computer or distributed among multiple computers. It should be appreciated that any component or collection of components that perform the functions described above can be generically considered as one or more controllers that control the abovediscussed functions. The one or more controllers can be implemented in numerous ways, such as with dedicated hardware, or with general purpose hardware (e.g., one or more processors) that is programmed using microcode or software to perform the functions recited above.
[0234] One or more processors may be interconnected by one or more networks in any suitable form, including a local area network or a wide area network, such as an enterprise network or the Internet. Such networks may be based on any suitable technology and may operate according to any suitable protocol and may include wireless networks, wired networks, or fiber optic networks.
[0235] One or more algorithms for controlling methods or processes provided herein may be embodied as a readable storage medium (or multiple readable media) (e.g., a computer memory, one or more floppy discs, compact discs (CD), optical discs, digital video disks (DVD), magnetic tapes, flash memories, circuit configurations in Field Programmable Gate Arrays or other semiconductor devices, or other tangible storage medium) encoded with one or more programs that, when executed on one or more computers or other processors, perform methods that implement the various methods or processes described herein.
[0236] In some examples, a computer readable storage medium may retain information for sufficient time to provide computer-executable instructions in a non-transitory form. Such a computer readable storage medium or media can be transportable, such that the program or programs stored thereon can be loaded onto one or more different computers or other processors to implement various aspects of the methods or processes described herein.
[0237] Executable instructions may be in many forms, such as program modules, executed by one or more computers or other devices. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. Typically, the functionality of the program modules may be combined or distributed as desired in various embodiments.
[0238] Also, data structures may be stored in computer-readable media in any suitable form. Non-limiting examples of data storage include structured, unstructured, localized, distributed, short-term and / or long-term storage. Non-limiting examples of protocols that can be used for communicating data include proprietary and / or industry standard protocols (e.g., HTTP, HTML, XML, JSON, SQL, web services, text, spreadsheets, etc., or any combination thereof). For simplicity of illustration, data structures may be shown to have fields that are related through location in the data structure. Such relationships may likewise be achieved by assigning storage for the fields with locations in a computer- readable medium that conveys relationship between the fields. However, any suitable mechanism may be used to establish a relationship between information in fields of a data structure, including through the use of pointers, tags, or other mechanisms that establish relationship between data elements.
[0239] RESULTS
[0240] Optimising small RNA-seq pipelines to enhance transfer RNA (tRNA) reads detection
[0241] In small RNA sequencing, the predominant focus has been on microRNA (miRNA) studies. However, the study of the present disclosure examines the distinct characteristics of transfer RNA-derived fragments (tRFs), which exhibit transcript lengths that are longer than miRNAs, and sequence similarity to many of the tRNA genes.
[0242] In an effort to optimise the identification of tRF-related sequencing readouts, the inventors of the present disclosure systematically evaluated key factors within small RNA sequencing (RNA-seq) pipelines. Departing from conventional approaches in small RNA-seq which conventionally targets 20-25M single-end reads with the size selection range to include RNA molecules of length 17-35bp with the majority of them at 22bp (which is the length of miRNA), the inventors of the present disclosure performed paired- end sequencing with a larger library size, 40M, and with a wider size selection range to maximise the discovery of small RNA molecules, including the tRFs. The larger library size will support the detection of low abundance tRFs and paired-end sequencing allows a more accurate call on the fragment ends. Following read preprocessing and genome alignment, the inventors of the present disclosure filtered reads mapped to the tRNA gene coordinates. The inventors of the present disclosure designed two approaches to construct features based on these tRNA mapped reads. FIG. 1A and 1B shows the overview of the two approaches.
[0243] Unsupervised identification of tRFs discover novel / unannotated fragments of tRNAs
[0244] The first approach runs the tRF detection pipeline to autonomously identify tRFs without reliance on any pre-existing tRF sequence annotation. FIG. 1A and 1B shows the simplified steps in the tRF detection pipeline, which starts with small RNA sequencing data. To obtain the fragment ranges along each tRNA gene, the inventors of the present disclosure first compute the overall read coverage for each tRNA gene. Based on the coverage, the inventors of the present disclosure identified the boundaries where fragment ends were frequently detected, using Hidden Markov Model (HMM). Once the fragment end positions are identified, the inventors of the present disclosure compute all possible fragment ranges and filter them based on the fragment directionality observed in sequencing reads and based on lengths. For quantification of tRF sequences, identical sequences and those sharing either the same fragment at the start or end were collapsed. Inhouse tissue data was used to detect tRF sequences and their counts were derived from the inhouse tissue data, inhouse plasma data, and The Cancer Genome Atlas Program (TCGA) tissue data.
[0245] The inventors of the present disclosure detected 13,223 tRF coordinates from the colorectal (CRC) tissue samples which contained -13% of 5' tRFs and -16% of 3' tRFs. Many tRFs were sample type- and individual- specific. Many tRFs showed highly correlated expression levels and such counts were merged. The final feature matrix of the present disclosure included 1 ,254 tRFs, of which 270 were annotated sequences in MINTbase v2.0.
[0246] The final tRF count matrix was produced by collapsing the counts from similar tRFs
[0247] The main rationale behind the tRF detection approach of the present disclosure was to start with an exhaustive search followed by evaluation of the candidate sequences by filtering and merging. This is why the inventors of the present disclosure started with all loci, including predicted loci, for tRNA gene coordinates and retained multi-mapping reads.
[0248] FIG. 2A to 2C shows unsupervised identification of tRNA fragments and merging sequences, illustrating the process of how tRF sequences were predicted based on sequence read profile and filtered to obtain the final tRF sequences, demonstrated with tRNA-Pro-AGG as an example. According to the tRNA gene list provided by GtRNAdb, there are 11 distinct genomic coordinates for tRNA-Pro-AGG. These genomic coordinates are based on computational prediction and are highly similar to each other in terms of the sequences. Instead of taking one reference sequence for a tRNA gene, however, this approach allows for a more sensitive and lenient discovery of potential tRF sequences. The different genomic coordinates of the same tRNA gene may also show distinct read coverage profiles (FIG. 2A, 2B). Even though the end-to-end positions from the coordinates are different, the fragmented sequences are often highly similar to each other. For tRNA-Pro-AGG, there were only 177 unique sequences from over 2000 tRF ranges estimated from 10 tRNA gene loci (one gene locus did not have sufficient read coverage and was discarded). Therefore, the inventors of the present disclosure implemented sequence merging to collapse the tRF counts coming from tRFs with either the identical sequences across different coordinates, or from the same tRF group within the same coordinate (FIG. 2C). For the final feature matrix, counts of the tRF sequences that share the same start or end of the fragment were merged. In the case of tRNA-Pro- AGG, 8 distinct tRF groups remained after this process.
[0249] To investigate how the tRF sequences the inventors of the present disclosure discovered compare to those in the tRF related databases, inventors of the present disclosure compared the tRNA-Pro-AGG tRF sequences to those reported in MINTbase v2.0. MINTbase hosts an extensive list of tRFs, which contains 26,531 distinct human tRFs from 11 ,719 datasets (Pliatsika et al., 2017). Out of the 177 unique tRF sequences inventors of the present disclosure discovered for tRNA-Pro-AGG gene, 66 were found in the MINTbase v2.0 (FIG. 2D), demonstrating that the approach of the present disclosure indeed discovers unannotated tRF sequences. tScore describes the sample’s similarity to tumour tissues in terms of tRNA reads profile
[0250] The second approach utilises the tScore pipeline to assign quantitative scores to individual reads, ultimately extending to each sample. FIG. 1 B illustrates how the scoring matrices for tumour samples (PSSM.T) and for normal samples (PSSM.N) are designed to capture the tRNA read profiles. The presence of the 3’-fragment in the tumour samples results in PSSM.T with higher values in the 3’ position compared to those in PSSM.N. Tumour similarity scores (tScore) were designed to numerically describe the similarity of each sample’s tRNA read profiles to those observed in tumour samples. Tumour-like samples, therefore, will show high tScores while normal-like samples will show low tScores.
[0251] The inventors of the present disclosure used the tissue data to build the scoring matrices so that the tScore will reflect the similarity to tRNA read characteristics in tumour. The inventors of the present disclosure first investigated which data would be the most suitable to build the scoring matrix from. The scoring matrices were built from tissue and plasma data separately, and applied to tissue, plasma, and TCGA tissue data to generate respective tScores. FIG. 3A to 3B shows this process. Since TCGA tissue data was strongly unbalanced with a limited number of normal tissues, the inventors of the present disclosure excluded TCGA tissue data when comparing the scoring matrices (FIG. 3A). The inventors of the present disclosure generated three different sets of scores, with the following designs: (1) tScore for tissue samples derived from tissue scoring matrix, (2) tScore for plasma samples derived from tissue scoring matrix, and (3) tScore for plasma samples derived from plasma scoring matrix.
[0252] To compare the choice of dataset for building the scoring matrix, the inventors of the present disclosure plotted FIG. 3B to show the assigned tScores for the samples in each dataset. Tumour and normal tissue samples show clearly distinctive tScore values (FIG. 3B, first plot) / tScores effectively delineate the disparity in the tRNA read profiles between tumour and normal tissues; while cancer and control plasma samples showed no clear separation of tScores (FIG. 3B, second plot). It shows that the tRNA read characteristics are clearly separable between tumour and normal tissue samples but are not distinct among the plasma samples.
[0253] Even though the tissue-derived tRNA read profiles are inadequately reflected in plasma samples, the tScores generated using tissue scoring matrix are able to capture distinct patterns in plasma samples at certain tRNA loci (FIG. 3B, third plot). In other words, tumour-like tRNA read characteristics are better retained in tissue samples and hence tissue samples were selected to build the scoring matrices. Conversely, tScores derived directly from plasma samples did not exhibit sufficiently distinct differences to establish the scoring matrices. FIG. 5A to 5C shows a model comparison between feature matrices built using all discovered tRFs (all tRF), only annotated tRFs (anno tRF), and tScore. The weaker performance of tScores compared to tRF counts in the model of the present disclosure further emphasizes the differences in tRNA read profiles between tissue and plasma samples (FIG. 5A to 5C).
[0254] Plasma samples showed distinct tRNA read characteristics from tumour tissues (FIG. 3B, FIG. 7A to 7C)
[0255] When using the same scoring matrices derived from inhouse tissue data, plasma and tissue samples showed different tScores (FIG. 3B, first and third plots). The plasma tScores generated using the scoring matrices derived from the tissue data showed mostly near-zero values. It indicates that the tRNA read characteristics observed in tissue samples were lost in plasma samples, except at some tRNA loci.
[0256] The type of data influences the tRNA fragment discovery Most small RNA sequencing is intended for miRNA studies. Total library size, proportion of tRNA reads (related to the library preparation method like size selection range and choice of kit used), type of sample (different types of analytes) affects the small RNA composition. For instance, plasma samples have a low proportion of tRNA reads and high proportion of Y RNA (yRNA) (Yeri et al., 2017). FIG. 9 also shows a much higher proportion of tRNA reads in the inhouse tissue data compared to the inhouse plasma data.
[0257] Most small RNA-seq data in public databases are focused on miRNAs. TCGA small RNA-seq libraries were constructed with 22±5 nucleotide (nt) insert (Chu et al., 2016), which will miss out on some tRFs whose lengths could reach 36bp (Kuscu et al., 2018). In the data of the present disclosure, the inventors of the present disclosure observed 36% of tRF sequences with length longer than 27bp (FIG. 10). This is a similar proportion of tRF sequences uniquely detected in the tissue data of the present disclosure and missing in TCGA data (37.8%; 474 out of 1254 sequences) (FIG. 7A). FIG. 9 also shows a much higher proportion of tRNA reads in the inhouse tissue data compared to the TCGA tissue data.
[0258] Due to these differences in data and their characteristics between inhouse generated and TCGA data, inhouse tissue data was used to identify tRF sequences associated with CRC and also to build the scoring matrices. Shared tRFs show strong correlation of overall expression values, but not the fold change values (Table 1, FIG. 4A and 4B, FIG. 7A to 7B). FIG. 4A, 4B, and Table 2 shows differentially expressed tRFs, shared between TCGA, tissue and plasma data.
[0259] Table 1. Trfs differentially expressed in TCGA, tissue and plasma data
[0260] Features asecffn Table 2. Shared trfs differentially expressed tn TCGA, tissue and plasma data trf ID sequence trf_l TCCCCGGCACCTCCA (SEQ ID NO: 1) trf_2 AGCGCC AAATCCT AACCACTAG ACC AC AG (SEQ I D NO: 2) trf_3 GGTAGCGTGGCCGAGC (SEQ ID NO: 3) Irf 4 TCTCTTCGGGGGCGTGG (SEQ ID NO; 4) trf _5 CTCTTCGGGGGCGTGGGTfSEQ ID NO: 5) trf 6 GGTTCAAATCCCGGACGAGCCC (SEQ ID NO: 6)
[0261] GTCACGGTGGCCGAGTGG(SEQ ID NO: 7) trf GAAAGCGCCGAATCCTAGCCACTAGACCACCAGGGA (SEQ ID NO: 8)
[0262] Unannotated tRFs markedly improved the model performance to identify cancer patients (FIG. 5, FIG. 8A to 8B)
[0263] The discovered list of tRFs was categorized into “annotated tRFs” and “unannotated tRFs”, based on to the annotation status in MINTbase. Two models were constructed, one including only the annotated tRFs, and the other including both annotated and unannotated tRFs. The model that incorporated unannotated tRFs in the input feature matrix demonstrated significantly better performance (pval=0.0126) compared to the model with only annotated tRFs. FIG. 5 shows the ROC curve. This finding indicates the value of identifying unreported tRFs and supports the effectiveness of the unsupervised tRF identification method of the present disclosure.
[0264] Discovering unannotated tRFs improves the diagnostic model performance by increasing the feature space, Furthermore, it was observed that including the unannotated tRFs broadened the feature space and markedly improved the model performance. Top performing tRF sequences originate from various tRNA genes, with the strongest biomarker candidates coming from tRNA-Lys, tRNA-Glu, and tRNA-Pro (FIG. 6).
[0265] The present disclosure presents two approaches to quantify nucleic acids. The first method quantifies nucleic acids in an unbiased and unsupervised way and hence discover both annotated and unannotated sequences. Quantification of nucleic acids (such as tRFs) is challenging due to highly conserved and hence similar sequences in large proportions of the reads. As such, the first approach (i.e., the tRF approach) is limited by the heterogeneity of the dataset used in mining tRF coordinates, as it will define the tRF universe. The second method quantifies the relative amounts of nucleic acids that are relevant to the disease status which is reflected by tScore. tScore measures how much of the sequenced reads are similar to those from tumour tissues. The more “tumour-like” the tRNA read patterns are, the higher the tScores will be given to the sample. The second method is complementary to the first method as it is agnostic to knowing precise fragmentation boundaries of the tRFs. The inventors of the present disclosure note that in the presented study, tRF features performed markedly better than tScore features (FIG. 5 and FIG. 8B). However, depending on the dataset and nature of the disease of application, the two methods can be applied independently or together to build the diagnostic model with higher accuracy and sensitivity.
[0266] Tissue and cohort-specific batch effects are present. However, given the strong correlation in the shared tRFs, one could build a dataset comprehensive enough for mining the potential targets. The difference between the inhouse tissue data and TCGA tissue data not only includes the difference in transcript length distribution, but also the type of tissue samples. Inhouse data were generated from fresh frozen tissue samples, while TCGA uses Formalin-Fixed Paraffin-Embedded (FFPE) samples.
[0267] As plasma samples are more heterogeneous than tissue samples, tumour-like read characteristics and directionality of differential expression are poorly retained in the plasma samples. The inventors of the present disclosure presented two approaches to build feature matrices based on tRNA signatures / tRNA fragmentation patterns. The quantification of specific tRF sequences showed a superior performance for tRNA. tRNA signatures present sample- and analyte- specificity. Sample- and analyte-specific heterogeneity, and sequencing-related batch effects may be limiting factors. Results of the present disclosure highlight tRNA fragments as valuable biomarkers in designing targeted panels for cancer diagnostics and the disclosed approaches can be applied to other small RNA fragments and cancer types. Selecting a concise panel of markers is crucial for a cost-effective and targeted diagnostic approach.
[0268] Materials and Methods
[0269] Data sources / Data used in this study
[0270] There are three different types of data used in this study. i) TCGA tissue data refers to the miRNA-seq data from TCGA (The Cancer Genome Atlas) and includes COAD (Colon Adenocarcinoma) and READ (Rectum Adenocarcinoma) cohorts. The dataset includes 616 tumours and 11 adjacent normal samples. Aligned reads in BAM format were downloaded. ii) Inhouse tissue data refers to the small RNA-seq data generated in-house. Tissue (such as colon tissue) samples are collected from colorectal cancer (CRC)) patients. The dataset includes 68 tumours and 62 adjacent normal samples. iii) Inhouse plasma data refers to the small RNA-seq data generated in-house. Plasma samples are collected from colorectal cancer (CRC) patients whose tissues were profiled, and some healthy individuals. The dataset includes 50 cancer patients’ plasma and 107 healthy individuals’ plasma samples.
[0271] Small RNA-sequencing was performed with 1 g of total RNA as the input. The library was sequenced on NovaSeq6000 using PE50.
[0272] Sample processing
[0273] Approximately 30 mg of fresh frozen tissue colorectal cancer (CRC)) sample was excised on dry ice using a scalpel, followed by lysis and homogenization using Lysing Matrix Y and FastPrep-24™ 5G (MP Biomedicals, cat. no. 116960500 and 116005500 respectively). The homogenised lysate was used for the extraction of total RNA using AllPrep DNA / RNA / miRNA Universal Kit (Qiagen, cat. No. 80224). The quality of the total RNA was assessed by using RNA ScreenTape (Agilent Technologies, cat. no. 5067- 5576), Qubit RNA HS Assay Kit (Thermo Fisher, cat. No. Q32852). The processing of plasma samples may be performed by methods known in the art, such as the methods described in: J Mol Diagn 2023, 25: 438e453; WO 2024 / 076307; and Cancers 2019, 11 , 458, which are incorporated herein by reference.
[0274] Small RNA library preparation and sequencing
[0275] The small RNA library was constructed using NEBNext Multiplex Small RNA Library Prep Set (New England Biolabs, cat. no. E7560S) with 1 pg of total RNA as the input following the manufacturer’s protocol. The resulting small RNA library was purified using Monarch® PCR & DNA Cleanup Kit (New England Biolabs, cat. no. T1030L) with a 7:1 buffer to sample ratio. The library size was analysed using High Sensitivity D1000 ScreenTape (Agilent Technologies, cat. no. 5067-5584). A size-selection step was performed on the small RNA library by using HT 3% agarose gel cassette on a PippinHT instrument (Sage Science, cat. no. HTG3010 and HTP0001 respectively) with the size selection range of 135-180nt. The size-selected library was purified using 1.8X SPRIselect beads (Beckman Coulter, cat. no. B23319) and elute in nuclease-free water. The final size-selected small RNA library was quantified using Qubit dsDNA HS Assay Kit (Thermo Fisher, cat. no. Q32853) and the library size was analysed using High Sensitivity D1000 ScreenTape as described above. The final library was sequenced on NovaSeq 6000 (Illumina Inc, cat. no. 20012850) using paired end sequencing with 50 bp reads (PE50), with the target library size of 40M reads.
[0276] Small RNA sequencing data
[0277] For inhouse generated small RNA-sequencing (smRNA-seq) data, adapter sequences were trimmed by Cutadapt (Martin, 2011). Trimmed reads with lengths shorter than 15bp or longer than 40bp were filtered out. Paired-end reads were merged using NGmerge for 5’-end correction (Gaspar, 2018). Reads were aligned to GRCh38 human genome using bwa samse command (Li & Durbin, 2009). GRCh38.d1.vd1 Reference Sequence from Genomic Data Commons (GDC) was used as the reference genome. File was downloaded from https: / / gdc.cancer.gov / about-data / gdc-data-processing / gdc-reference-files. Up to two mismatches were allowed in alignment and the best alignment was reported for multimapping reads. The best alignment is determined by a mapping algorithm that takes into account the quality of mapping and number of matches or mismatches in the alignment. For example, the best alignment may refer to the alignment with the least number of mismatches.
[0278] For TCGA data, TCGA-COAD and TCGA-READ data were downloaded in BAM format. Bedtools (Quinlan & Hall, 2010) was used to filter reads mapped to the tRNA genes. The “hg38-tRNAs.bed” file downloaded from GtRNAdb (Chan & Lowe, 2015), as part of tRNAscan-SE results was used to infer the genomic coordinates for tRNA genes. The alignment files in BAM format were converted to multiple sequence alignment (MSA) format using bam2msa (orangeSi, n.d.).
[0279] Unsupervised identification of tRF sequences
[0280] Reads with double insertion bases were first removed. Multiple sequence alignment format at each tRNA gene coordinate was modified to accommodate frequent insertions. Based on overall read coverage to each tRNA gene coordinate, the inventors of the present disclosure first fitted a Hidden Markov Model (HMM) with 2-state poisson or gaussian models. Fragmentation boundaries were called based on the estimated states by the viterbi algorithm and the fragmentation frequency threshold. Fragmentation frequency indicates how many of the samples included in the analysis shared the same fragmentation position. Fragmentation frequency threshold can be optimised for different datasets. 10% was used in the analysis of the present disclosure. All HMM related analysis was performed using the R package, depmixS4 (Visser & Speekenbrink, 2010).
[0281] After finding the fragmentation boundaries, the directionality was determined for each fragmentation position based on the proportion of positive and negative derivative values from the overall read coverage. All possible tRF ranges were first computed based on the fragmentation positions and the directionality. The ranges were then filtered by length by discarding those shorter than 15bp.
[0282] Justification for using total coverage, keeping duplicates, allowing mismatches in the alignment
[0283] Overall coverage (total read coverage across all the samples included in the analysis) was used when determining the fragmentation boundary, instead of considering coverage in individual samples. This approach was chosen because the inventors of the present disclosure observed that using individual sample coverage made the callouts more sensitive to variations in library depth, particularly problematic for plasma samples given their inherently low tRNA read counts. Discerning genuine fragment boundaries from background noise became more challenging as it mainly relied on the observed frequency fragment ends. Duplicate reads were included in the coverage calculation and the fragmentation boundary estimation for analogous reasons. During the genome alignment, the inventors of the present disclosure allowed up to two mismatches to discover as many tRNA related sequences as possible.
[0284] Quantification of tRF sequences
[0285] Instead of mapping the raw sequencing reads to the identified tRF sequences, quantification relied on the multiple sequence alignment format after insertion modification. This is to distinguish transcripts that only differ by a couple bases in length, and also to disregard sequence variants for a more sensitive detection. The input sequences and the reference tRF sequences were first matched by length and the positions in the multiple sequence alignment formats. For the final matrix of tRF counts, counts from the tRFs with identical sequences and those belonging to the same tRF group within the same tRNA gene were merged.
[0286] Features extraction
[0287] Two approaches were designed to capture tRNA fragmentation patterns: i) Identification of differentially expressed tRFs / unsupervised discovery of tRF coordinates
[0288] DESeq2 (Love et al., 2014) was used to identify differentially expressed tRFs. High confidence fragmentation boundaries are first identified from the overall read coverage and filtered to generate coordinates for tRFs. The coordinates are used in tRF quantification. The tRFs detected in less than 5 samples were removed from the analysis. Adjusted p-value less than 0.05 was used as the threshold.
[0289] II) Construction of scoring matrices for each sequencing read / numeric score assignment to describe a sample's similarity to a tumour tissue based on tRNA read profiles.
[0290] Position specific scoring matrices (PSSM) were generated for each tRNA coordinate and for each sample type (tumour or normal) from the inhouse generated tissue sequencing data. The algorithm to build the scoring matrices was designed based on the original method (Henikoff & Henikoff, 1996), with some modifications to accommodate sequence variants like indels and to accommodate end-to-end alignment of fragmented sequences with many missing bases. The modification include: (1) incorporating “empty base” as part of the sequence profile to accommodate indels, (2) dealing with the ambiguous sequencing base call “N”, and (3) modifying the position weight probability matrix based on the relative base coverage to accommodate differential coverage along the gene.
[0291] Typically, PSSM is meant for end-to-end aligned sequences of the same length. Since fragmentation and missing base information are important to distinguish tRF patterns, the inventors of the present disclosure created the 5th nucleotide base to signify the empty or missing base. This requires consideration of what the basal frequency of the missing base is, what the substitution rate would be between the 4 conventional nucleotides (A, T, C, G) and the missing base, in terms of the relative ratio. It was assumed that there was a 1% chance of mismatch (similar to what PAM1 matrix), 0.1% chance of transition between base to no base when building the substitution matrix. This allows a high penalty for indels. Additionally, it was assumed that it was 5 times more likely that any given base would be a missing base and not one of the conventional nucleotides since many of the bases along the entire tRNA genes are “empty” from small RNAseq data. The ambiguous sequencing base call, all the “N” bases in sequencing data, was augmented by adjusting the derived position weight matrix with additional split probability. For instance, 2 counts of N base calls were changed to 0.5 (2 divided by 4) counts to each A, T, C, and G base calls. Final position weight probability matrices were adjusted based on the relative base coverage, which created the system to assign higher scores if fragments occurred at coordinates with high base coverage.
[0292] The scoring matrices were built based on deduplicated read profiles, but score assignment was done including the duplicate reads. Using coverage adjusted scoring matrices built from deduplicated reads while keeping the duplicate reads in score assignment maximised the difference between tumour-normal scores (data not shown). Two scoring matrices were built, one derived from tumour samples (PSSM.T) and the other derived from normal samples (PSSM.N). Sample-wise tumour similarity score (tScore) is computed based on the dissimilarity between the two read score distributions, each obtained from the tumour scoring matrix and the normal scoring matrix. Reads with profiles that are similar to tumour samples will have a high score when using PSSM.T and a low score with PSSM.N.
[0293] Calculation of tScore per sample
[0294] Each read was given two scores, by using PSSM.N and PSSM.T. For each sample, tScore was derived using the following formula:
[0295] Where R is the number of reads in a sample, tScore. T is the tScore given to the read using PSSM.T and tScore. N is the tScore given to the read using PSSM.N. High values of sample tScore, therefore, represent that the sample contained many reads that were similar to the tumour tissue reads / indicates a greater resemblance of the sample's tRNA read profile to that of a tumour sample..
[0296] Colorectal cancer (CRC) tissue data was used for unsupervised discovery of tRF coordinates, and for building the scoring matrices for tumour and normal samples. Final feature matrices with tRF counts and tScores were built using CRC plasma samples.
[0297] Model building and validation
[0298] Feature matrices were built from tRF counts and tScores. Counts for tRF were normalised using EdgeR TMM normalisation (Robinson et al., 2010). Preprocessing included filtering for features with too many zeroes (more than 10% of the samples) and / or with low expression (10% quantile less than 10), imputation for the remaining zero values, count normalization and global mean standardization. The remaining zero values were replaced with the minimum values. Log2 transformation and global mean standardisation were applied.
[0299] Two-class classification model was built using Support Vector Machine (SVM), linear Ridge regression (Ridge), AdaBoost (adaboost), Random Forest (Random Forest Classifier), and logistic regression (logistic) with scikit-learn (Pedregosa et al., 2011) in Python3. For each type of model framework, three different models were built and tested with the following feature matrices: (1) tRF counts only / all discovered tRFs, (2) annotated tRFs only, (3) tScore only, and (4) tRF counts and tScore combined. Nested 4x4 cross-validation was used for all models. For incremental feature selection, top features were selected by HnearSVC coefficients.
[0300] DETAILED DESCRIPTION OF FIGURES
[0301] Example embodiments of the disclosure will be better understood and readily apparent to one of ordinary skill in the art from the following discussions and if applicable, in conjunction with the figures. Example embodiments are not necessarily mutually exclusive as some may be combined with one or more embodiments to form new exemplary embodiments. The example embodiments should not be construed as limiting the scope of the disclosure.
[0302] FIG. 1A shows an overview of the tRF identification pipeline
[0303] FIG. 1B shows an example of the two PSSM matrices derived from control (normal) samples (top), and from diseased (tumour) samples (bottom). The scoring matrices capture the pattern of the aligned sequence reads.
[0304] FIG. 2A shows the workflow and tRFs identified from an example gene locus, tRNA-Pro-AGG-1-1.
[0305] FIG. 2B shows the workflow and tRFs identified from an example gene locus tRNA-Pro-AGG-2-6.
[0306] FIG. 2C shows graphs where the process of how final tRF sequences are curated from a collection of overlapping and similar tRF sequence reads, with tRNA-Pro-AGG as an example.
[0307] FIG. 2D shows a Venn diagram of the overlap between identified tRFs from Pro- AGG tRNA gene to the annotated tRF sequences from a public database, MINTbase.
[0308] FIG. 3A shows an overview of how different scoring matrices were built and applied to different types of CRC data to generate tScores. FIG. 3B shows a dot plot of generated tScores using the respective scoring matrix and applied to the respective data, as written on the figure. Scoring matrix built from inhouse tissue data generates more distinct tScores.
[0309] FIG. 4A shows a Venn diagram of the overlap between the differentially expressed tRFs identified from tissue, plasma, and TCGA data.
[0310] FIG. 4B shows a box plot of the expression level of the overlapping differentially expressed tRFs in tissue (inhouse_tissue), plasma (inhouse_plasma), and TCGA (tcga_tissue) data in different sample types (normal, tumor, control, cancer)
[0311] FIG. 5 shows line graphs of the model comparison between features matrices built using all discovered tRFs (trf), only annotated tRFs (annotrf), and tscore. ROC curve (top) includes all the selected features (on average, 88 features). The table shows the average AUC values. P-value was calculated using the Venkatraman method with 10,000 permutations.
[0312] FIG. 6 shows a box plot with the top 30 features from the classification model using all trf, ranked by permuted feature importance.
[0313] FIG. 7A shows dot plots of the relationship between the tRF expression level in different datasets (plasma (P2327_CRC_plasma), tissue (P2310_CRC_tissue), and TCGA tissue (TCGA_CRC)). It shows that some tRFs are sample type specific (i.e. uniquely found in plasma and not plasma dataset), and cohort specific (i.e. uniquely found in inhouse generated tissue data and not in TCGA tissue data)
[0314] FIG. 7B shows dot plots of the relationship between the Iog2 fold change in tRF expression level (cancer / control) seen in different datasets.
[0315] FIG. 7C shows dot plots of the relationship between the Iog2 fold change in tScore (tumor / normal) in plasma and tissue data.
[0316] FIG. 8A shows line graphs in various iterations of cross-validation, that all tRF (annotated and unannotated tRFs, “trf) expression demonstrated better performance than annotated (“annotrf’) in diagnosing colorectal cancer.
[0317] FIG. 8B shows line graphs in various iterations of cross-validation, that tRF expression showed better performance than tScore in diagnosis of colorectal cancer.
[0318] FIG. 9 shows the box plots of total sequence reads mapped to the human genome (first plot), reads mapped to the tRNA gene coordinates (second plot), and the proportion of tRNA reads (third plot) in inhouse tissue, inhouse plasma, and TCGA tissue data.
[0319] FIG. 10 shows the histogram of tRF sequence lengths for all discovered tRF sequences in inhouse tissue data. APPLICATIONS
[0320] Embodiments of the methods disclosed herein provide a method of quantifying one or more nucleic acids.
[0321] Advantageously, the present disclosure provides a first approach of quantifying one or more nucleic acids by running a nucleic acid sequence (such as transfer RNA fragment (tRF)) detection pipeline to identify nucleic acid sequences (such as tRFs). The first approach identifies boundaries where the nucleic acid sequence read ends are frequently detected, using a statistical model. All possible fragment ranges were computed and filtered based on fragment directionality and lengths. Identical sequences / those sharing the same fragment were collapsed to quantify the nucleic acid sequences (such as nucleic acid sequence reads) identified.
[0322] Even more advantageously, the statistical model of the present disclosure is able to filter out technical errors (such as “background noise / irrelevant sequence / noise”) to focus on nucleic acid sequence reads that are from biological target nucleic acid molecules of interest.
[0323] Even more advantageously, the present disclosure was able to quantify and / or identify one or more nucleic acids using liquid biopsies from subjects.
[0324] Even more advantageously, the present disclosure provides a method that discovers unannotated nucleic acid (such as tRF) sequences.
[0325] Even more advantageously, the present disclosure provides a second approach of quantifying nucleic acid (such as tRF) sequences using a scoring pipeline (such as tScore pipeline) that compares two sample groups to assign quantitative scores to individual reads that extend to each sample. Tumour-like samples show high tScores while normal samples show low tScores.
[0326] Even more advantageously, the present disclosure shows that the unsupervised models for quantifying / identifying nucleic acid sequences that incorporated both the annotated and unannotated nucleic acids (such as tRFs) demonstrated significantly better performance compared to only annotated nucleic acid (such as tRFs) sequences alone. This highlights the value of identifying unreported nucleic acids to support the methods of the present disclosure.
[0327] Even more advantageously, the present disclosure provides nucleic acid fragments (such as tRFs) as biomarkers for cancer diagnostics.
[0328] Even more advantageously, the present disclosure provides methods that can be applied to a variety of nucleic acid types and cancer types. It will be appreciated by a person skilled in the art that other variations and / or modifications may be made to the embodiments disclosed herein without departing from the spirit or scope of the disclosure as broadly described. For example, in the description herein, features of different exemplary embodiments may be mixed, combined, interchanged, incorporated, adopted, modified, included etc. or the like across different exemplary embodiments. The present embodiments are, therefore, to be considered in all respects to be illustrative and not restrictive.
Claims
CLAIMS1 . A method of quantifying and / or identifying one or more nucleic acids associated with a condition, the method comprising: i) obtaining one or more nucleic acids from a plurality of sample groups and generating one or more nucleic acid sequence reads; ii) identifying a plurality of boundaries of the one or more nucleic acid sequence reads using a statistical model; iii) obtaining a plurality of nucleic acid sequences bound by the selected boundaries; iv) quantifying and / or identifying the annotated nucleic acid and the unannotated nucleic acid that are associated with the condition.
2. The method of claim 1, wherein the identifying one or more nucleic acids comprises comparing the expression level or copy number of the annotated nucleic acid and the unannotated nucleic acid in the plurality of sample groups; wherein an altered expression level or altered copy number as compared between the plurality of sample groups is indicative of the annotated nucleic acid and / or the unannotated nucleic acid to be associated with the condition.
3. A method of determining / predicting a condition in a subject comprising: i) obtaining one or more nucleic acids from a plurality of sample groups and generating one or more nucleic acid sequence reads from the plurality of sample groups; ii) aligning the one or more nucleic acid sequence reads to obtain one or more aligned portion, one or more empty portion, and / or one or more missing portion, iii) generating one or more scoring matrices from the alignment of the plurality of nucleic acid sequences; iv) calculating one or more read scores for the alignment using the generated one or more scoring matrices wherein the read score is determined by calculating the difference between the plurality of nucleic acid sequence read scores; and v) calculating the one or more sample scores based on the calculated one or more read scores;wherein the one or more sample scores provide an indication of the disease condition of the subject.
4. The method according to any one of claim 1 and 3, wherein each scoring matrix is represented by a modified position-specific scoring matrix (PSSM).
5. The method according to any one of claims 1, 3 to 4, wherein the modified PSSM comprises the likelihood of occurrence of each nucleotide sequence (A, T, G, C) and being empty / missing at different positions along the portion of an annotated genomic locus.
6. The method according to any one of claims 1, 3 to 5 comprising calculating the one or more sample scores based on the sum of the individual read scores of one or more nucleic acid sequence reads, wherein the individual read score of the one or more nucleic acid sequence reads is calculated using one or more modified PSSM.
7. The method according to any one of claims 1, 3 to 6 further comprising calculating a sample score for a plurality of sample groups according to the following equation: sample scorewherein [31, [32, ... , pN represent the coefficients; r represents the sequence read;R represents the total number of nucleic acid sequence reads in the sample; S1, S2, ... SN represent the read scores derived from the first, second, ... Nth modified PSSM, and wherein the sample score represents the condition likelihood.
8. The method according to claim 7 further comprising calculating a sample score for two sample groups according to the following equation:wherein r represents the sequence read;R represents the total number of sequence reads in the sample library;51 represents the first read score derived from the first modified PSSM; and52 represents the second read score derived from the second modified PSSM.
9. The method according to any one of claims 1 , 3 to 8 further comprising calculating a sample score for one sample group according to the following equation:wherein r represents the sequence read;R represents the total number of sequence reads in the sample library; and S represents the read score derived from the modified PSSM.
10. A method of determining whether a subject suffers from, or is at risk of developing a condition, the method comprising: i) determining the expression level or copy number of one or more nucleic acid sequence reads identified I quantified comprising: a) obtaining one or more nucleic acids from a plurality of samples and generating one or more nucleic acid sequence reads; b) identifying a plurality of boundaries of the one or more nucleic acid sequence reads using a statistical model; c) obtaining a plurality of nucleic acid sequences bound by the selected boundaries; d) quantifying and / or identifying the annotated nucleic acid and the unannotated nucleic acid that are associated with the condition; comparing the expression level or copy number of the one or more nucleic acid sequence reads in the subject with that in a control; wherein an altered expression level or altered copy number as compared between the plurality of samples is indicative of the subject suffering from or is at risk of developing the condition; and / orii) calculating one or more sample scores for a sample obtained from a subject comprising a) obtaining one or more nucleic acids from a plurality of samples and generating one or more nucleic acid sequence reads from the plurality of samples; b) aligning the plurality of nucleic acid sequence reads to obtain one or more aligned portion, an empty portion, and / or one or more missing portion, c) generating one or more scoring matrices from the alignment of the plurality of nucleic acid sequence reads; d) calculating one or more read scores for the alignment using the generated one or more scoring matrices wherein the read score is determined by calculating the difference between the plurality of nucleic acid sequence read score; and e) calculating the one or more sample scores based on the calculated one or more read scores; wherein the one or more sample scores provide an indication of the condition of the subject, determining whether a subject suffers from, or is at risk of developing a condition based on the one or more calculated sample scores.
11. The method according to any one of the preceding claims further comprising converting the alignment to a multiple sequence alignment format.
12. The method according to any one of the preceding claims, wherein the one or more nucleic acids comprise transfer ribonucleic acids (tRNAs), microRNAs (miRNAs), ribosomal ribonucleic acids (rRNAs), long non-coding RNAs (IncRNAs), small nucleolar RNAs (snoRNAs), small nuclear RNAs (snRNAs), yRNA, piwi-interacting RNA (piRNA) and / or fragments thereof.
13. The method according to any one of the preceding claims, wherein the one or more nucleic acid sequence reads comprise tRNA fragments.
14. The method according to any one of the preceding claims, wherein the one or more nucleic acid sequence reads is selected from the group consisting of tRNA-Lys or its fragments thereof, tRNA-Glu or its fragments thereof, and / or tRNA-Pro or its fragments thereof.
15. The method according to any one of the preceding claims, wherein the condition is a cancer selected from colorectal cancer, lung cancer, gastric cancer, breast cancer, liver cancer and / or prostate cancer.
16. The method according to any one of the preceding claims, wherein the condition is colorectal cancer.
17. The method according to any one of the preceding claims, wherein the sample comprises a tissue sample and / or a liquid biological sample.