Method for distinguishing live and dead microorganisms in a sample

By culturing samples and sequencing RNA in the presence of RNA labeling agents, comparing nucleotide substitution levels, the problem of difficult to distinguish between live and dead microorganisms in the prior art is solved, and an accurate diagnosis of microbial infection and an efficient assessment of the risk of contamination is achieved.

CN112585283BActive Publication Date: 2025-05-23PATHOQUEST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201980054782.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-04-03
Filing Date
2019-06-20
Publication Date
2025-05-23
Estimated Expiration
2039-06-20

AI Technical Summary

Technical Problem

The prior art is difficult to distinguish between live and dead microorganisms. Especially when detecting replicative microorganisms, there is difficulty in distinguishing between contaminated RNA and replicative RNA, which affects detection efficiency and sensitivity.

Method used

The samples were cultured in the presence of an RNA labeling agent, RNA was extracted and sequenced, and the nucleotide substitution levels were compared to distinguish between microbial nucleic acid sequences with transcriptional activity and microbial nucleic acid sequences with transcriptional inertia.

Benefits of technology

Accurate distinction between live and dead microorganisms is achieved, and the accuracy of microbial infection diagnosis and efficiency of contamination risk assessment are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112585283B_ABST
    Figure CN112585283B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for distinguishing live and dead microorganisms in a sample by distinguishing between transcriptionally active microbial nucleic acid sequences and transcriptionally inert microbial nucleic acid sequences in the sample. Specifically, the method according to the present invention is based on the comparison of nucleotide substitution levels in samples cultured in the presence of an RNA marker. The present invention also relates to a method for diagnosing microbial infection in a subject; and a method for evaluating the risk of sample contamination by implementing the differentiation method according to the present invention.
Need to check novelty before this filing date? Find Prior Art

Description

Field of the Invention

[0001] The present invention relates to a method for distinguishing between live and dead microorganisms in a sample by distinguishing between transcriptionally active microbial nucleic acid sequences and transcriptionally inert microbial nucleic acid sequences in the sample. Specifically, the method according to the present invention is based on comparing nucleotide substitution levels in samples cultured in the presence of an RNA marker.

[0002] The present invention also relates to a method for diagnosing a microbial infection in a subject; and a method for assessing the risk of contamination of a sample by implementing the method for differentiation according to the present invention. Background Art

[0003] The ability to detect viruses in cells, and more generally, microorganisms, has many applications in the field of diagnostics, where it proves useful for identifying infectious agents that cause, for example, diseases; in biomedical research, where it determines the interpretation of experimental results; or in the safety assessment of potentially contaminated samples; or in the viability of microorganisms used in biotechnological processes.

[0004] In addition to the current standard detection techniques that rely on amplification of microorganism-specific nucleic acid sequences, several technologies have emerged to circumvent the main limitation of techniques based solely on amplification, namely the inability to distinguish between dead and active (or replicating) microorganisms, including latent viruses. The ability to establish this distinction is crucial in screening biological samples, as the presence of active microbial agents may have different consequences than those associated with the presence of dead and / or inactive microorganisms, especially viruses.

[0005] These techniques can be based on the detection of sequences specific for replicating viruses: the presence of RNA in the case of DNA viruses, the stoichiometry of sense to antisense RNA in the case of negative-sense single-stranded RNA viruses, the presence of antisense RNA (antigenome) in the case of positive-sense single-stranded RNA viruses, and the presence of DNA and positive-sense spliced ​​RNA in the case of retroviruses. These techniques are based on the amplification of the reverse complementary strand, such as reverse transcription polymerase chain reaction (RT-PCR), or RNA sequencing based on high-throughput sequencing methods (RNA sequencing, also called whole transcriptome shotgun sequencing), in particular double-stranded RNA sequencing techniques.

[0006] Current technologies still have several limitations: First, current technologies do not allow for the differentiation of contaminating RNA, such as replication intermediates that exist in the sample independently of active microorganisms, such as active viral particles, from RNA associated with the presence of replicating microorganisms, such as viruses, in the sample. The efficiency of such technologies in detecting replicating double-stranded RNA viruses remains to be verified. In the case of positive-sense single-stranded RNA viruses, the small amount of antisense RNA species produced hinders the sensitivity of existing detection technologies. In the case of negative-sense single-stranded RNA viruses, it is difficult to distinguish between the positive-strand antigenome and specific transcripts, requiring specific bioinformatics analysis for the viral species.

[0007] Another characteristic of RNA species associated with the presence of replicating viruses (and more broadly, replicating microorganisms) is that, in contrast to latent contaminating RNAs, they are in the process of synthesis (i.e., in the case of viruses, in the host cell; or in the case of other microorganisms, directly in bacteria, fungi, or others). Several techniques for labeling and purifying nascent transcripts (also known as RNA metabolic labeling techniques) have been described. For example, the incorporation of 4-thiouridine (4SU) or other types of uridine analogs (BudR) has been used to purify nascent eukaryotic mRNA transcripts and / or to study the dynamics of transcriptional networks (Herzog et al., 2017. Nat Methods. 14(12): 1198-1204; Tani et al., 2012. RNA Biol. 9(10): 1233-8).

[0008] Herein, the inventors have designed, optimized, and validated a method for detecting virally infected cells (including novel viral strains / species) that relies on the use of metabolic labeling to detect viral RNA synthesis within cells of a biological sample. The inventors also show that, when relying on the use of metabolic labeling to detect microbial RNA synthesis in a sample, the method is readily applicable to the detection of any transcriptionally active microorganism. This includes, for example, the detection of viral, bacterial, archaeal, fungal, or protozoal infection or contamination by live microorganisms, as well as the differentiation of live microorganisms from the remnants of dead microorganisms. Summary of the Invention

[0009] The present invention relates to an in vitro method for distinguishing between live and dead microorganisms in a sample, comprising distinguishing between transcriptionally active microbial nucleic acid sequences and transcriptionally inert microbial nucleic acid sequences in the sample, wherein the method comprises the following steps:

[0010] (a) sequencing a first set of RNA extracted from a sample, wherein the first set of RNA is obtained by culturing the sample in the presence of an RNA labeling agent and further by subjecting the extracted RNA to conditions that allow nucleotide substitutions; thereby obtaining a first set of sequence reads;

[0011] (b) comparing the number of substituted nucleotides in the first set of sequence reads that match at least one microbial nucleic acid sequence hit to the number of substituted nucleotides in the control sequence; and

[0012] (c) if the number of substituted nucleotides in the sequence reads that match the at least one microbial nucleic acid sequence hit in the first set of sequence reads is greater than the number of randomly substituted nucleotides in the reference sequence, then inferring that the at least one microbial nucleic acid sequence hit belongs to a viable microorganism.

[0013] In one embodiment, an in vitro method for distinguishing between live and dead microorganisms in a sample comprises distinguishing between transcriptionally active microbial nucleic acid sequences and transcriptionally inert microbial nucleic acid sequences in the sample, wherein the method comprises the following steps:

[0014] (a) sequencing a first set of RNA extracted from a sample, wherein the first set of RNA is obtained by culturing the sample in the presence of an RNA labeling agent and further by subjecting the extracted RNA to conditions that allow nucleotide substitutions; thereby obtaining a first set of sequence reads;

[0015] (b) comparing the number of substituted nucleotides in the first set of sequence reads that match at least one microbial nucleic acid sequence hit to the number of substituted nucleotides in the control sequence; and

[0016] (c) if the number of substituted nucleotides in the sequence reads that match the at least one microbial nucleic acid sequence hit in the first set of sequence reads is greater than the number of randomly substituted nucleotides in the reference sequence, inferring that the at least one microbial nucleic acid sequence hit belongs to a viable microorganism,

[0017] wherein the control sequence is not a second set of sequence reads that match the at least one microbial nucleic acid sequence hit, the second set of sequence reads being obtained by sequencing a second set of RNAs obtained by culturing the sample in the absence of an RNA marker.

[0018] In one embodiment, the control sequence is selected from:

[0019] a second set of sequence reads that match the at least one microbial nucleic acid sequence hit, wherein the second set of sequence reads is obtained by sequencing a second set of RNA obtained by culturing the sample in the absence of an RNA marker;

[0020] a second set of sequence reads that match the at least one microbial nucleic acid sequence hit, wherein the second set of sequence reads is obtained by sequencing a second set of RNA obtained by culturing the sample in the presence of an RNA labeling agent but without subjecting the extracted RNA to conditions that allow nucleotide substitutions;

[0021] - a consensus microbial nucleic acid sequence obtained from sequence reads or contigs of a first set of sequences that match at least one microbial nucleic acid sequence hit;

[0022] - a sequence corresponding to a nucleic acid sequence of the same microorganism found in the closest microorganism strain identified in a nucleic acid sequence database; and / or

[0023] - Similar sequences corresponding to the same microbial nucleic acid sequence hits identified in nucleic acid sequence databases.

[0024] In one embodiment, the control sequence is selected from:

[0025] a second set of sequence reads that match the at least one microbial nucleic acid sequence hit, wherein the second set of sequence reads is obtained by sequencing a second set of RNA obtained by culturing the sample in the presence of an RNA labeling agent but without subjecting the extracted RNA to conditions that allow nucleotide substitutions;

[0026] - a consensus microbial nucleic acid sequence obtained from sequence reads or contigs of a first set of sequences that match at least one microbial nucleic acid sequence hit;

[0027] - a sequence corresponding to a hit of the same microbial nucleic acid sequence found in the closest microbial strain identified in a nucleic acid sequence database; and / or

[0028] - Similar sequences corresponding to the same microbial nucleic acid sequence hits identified in nucleic acid sequence databases.

[0029] In one embodiment, the in vitro method of the invention comprises:

[0030] (a) sequencing the first and second RNA groups extracted from the sample,

[0031] wherein the first set of RNAs and the second set of RNAs are obtained by culturing the sample in the presence of an RNA labeling agent, thereby obtaining labeled RNA, and

[0032] wherein the first set of RNAs is obtained from a first portion of labeled RNA that is subjected to conditions that allow nucleotide substitutions, and the second set of RNAs is obtained from a second portion of labeled RNA that is not subjected to conditions that allow nucleotide substitutions,

[0033] Thus, a first set of sequence reads and a second set of sequence reads are obtained,

[0034] (b) comparing the number of substituted nucleotides in a first set of sequence reads that match at least one microbial nucleic acid sequence hit and the number of substituted nucleotides in a second set of sequence reads that match at least one microbial nucleic acid sequence hit, and

[0035] (c) if the number of substituted nucleotides in the sequence reads that match the at least one microbial nucleic acid sequence hit in the first set of sequence reads is greater than the number of substituted nucleotides in the sequence reads that match the at least one microbial nucleic acid sequence hit in the second set of sequence reads, then it is inferred that the at least one microbial nucleic acid sequence hit belongs to a living microorganism.

[0036] In one embodiment, an in vitro method is used to distinguish between infectious and non-infectious viral nucleic acid sequences in a cell sample and comprises:

[0037] (a) sequencing a first set of RNA and a second set of RNA extracted from a cell sample,

[0038] wherein the first set of RNA is obtained by culturing the cell sample in the presence of an RNA labeling agent, and the second set of RNA is obtained by culturing the cell sample in the absence of an RNA labeling agent,

[0039] Thus, a first set of sequence reads and a second set of sequence reads are obtained,

[0040] (b) identifying at least one viral nucleic acid sequence hit that matches at least one sequence read of the first set of sequence reads,

[0041] (c) comparing the number of substituted nucleotides in sequence reads that match at least one identified viral nucleic acid sequence hit in the first set of sequence reads and the second set of sequence reads, and

[0042] (d) if the number of substituted nucleotides in the sequence reads that match the at least one identified viral nucleic acid sequence hit in the first set of sequence reads is greater than the number of substituted nucleotides in the sequence reads that match the at least one identified viral nucleic acid sequence hit in the second set of sequence reads, then inferring that the at least one viral nucleic acid sequence hit belongs to an infectious virus.

[0043] In one embodiment, an in vitro method for distinguishing between live and dead microorganisms in a sample comprises:

[0044] (a) sequencing the first and second RNA groups extracted from the sample,

[0045] wherein the first set of RNA is obtained by culturing the sample in the presence of an RNA labeling agent, and the second set of RNA is obtained by culturing the sample in the absence of an RNA labeling agent,

[0046] Thus, a first set of sequence reads and a second set of sequence reads are obtained,

[0047] (b) comparing the number of substituted nucleotides in a first set of sequence reads that match at least one microbial nucleic acid sequence hit and the number of substituted nucleotides in a second set of sequence reads that match at least one microbial nucleic acid sequence hit, and

[0048] (c) if the number of substituted nucleotides in the sequence reads that match the at least one microbial nucleic acid sequence hit in the first set of sequence reads is greater than the number of substituted nucleotides in the sequence reads that match the at least one microbial nucleic acid sequence hit in the second set of sequence reads, then it is inferred that the at least one microbial nucleic acid sequence hit belongs to a living microorganism.

[0049] In one embodiment, the first set of RNA is obtained by incubating the sample in the presence of an RNA labeling agent, thereby obtaining labeled RNA, and further subjecting the labeled RNA to conditions that allow nucleotide substitutions.

[0050] In one embodiment, an in vitro method for distinguishing between live and dead microorganisms in a sample comprises:

[0051] (a) sequencing a first set of RNA and a second set of RNA extracted from a sample,

[0052] wherein the first set of RNAs and the second set of RNAs are obtained by culturing the sample in the presence of an RNA labeling agent, thereby obtaining labeled RNA, and

[0053] wherein the first set of RNAs is obtained from a first portion of labeled RNA subjected to conditions of nucleotide substitution, and the second set of RNAs is obtained from a second portion of labeled RNA subjected to conditions of no nucleotide substitution,

[0054] Thus, a first set of sequence reads and a second set of sequence reads are obtained,

[0055] (b) comparing the number of substituted nucleotides in a first set of sequence reads that match at least one microbial nucleic acid sequence hit and the number of substituted nucleotides in a second set of sequence reads that match the at least one microbial nucleic acid sequence hit, and

[0056] (c) if the number of substituted nucleotides in the sequence reads that match the at least one microbial nucleic acid sequence hit in the first set of sequence reads is greater than the number of substituted nucleotides in the sequence reads that match the at least one microbial nucleic acid sequence hit in the second set of sequence reads, then it is inferred that the at least one microbial nucleic acid sequence hit belongs to a living microorganism.

[0057] In one embodiment, the RNA labeling agent is a thiol-labeled RNA precursor.

[0058] In one embodiment, the thiol-labeled RNA precursor is selected from 4-thiouridine, 2-thiouridine, 2,4-dithiouridine, 2-thio-4-deoxyuridine, 5-ethoxyformyl-2-thiouridine, 5-carboxy-2-thiouridine, 5-(n-propyl)-2-thiouridine, 6-methyl-2-thiouridine and 6-(n-propyl)-2-thiouridine, thereby obtaining thiouridine-labeled RNA.

[0059] In one embodiment, the thiol-labeled RNA precursor is preferably 4-thiouridine.

[0060] In one embodiment, the nucleotide substitution comprises chemically modifying the RNA, preferably by alkylation, oxidative-nucleophilic-aromatic substitution or osmium-mediated conversion; more preferably by alkylation; and further reverse transcribing the chemically modified RNA.

[0061] In one embodiment, the second set of RNA is obtained by culturing a cell sample in the presence of an RNA labeling agent to obtain labeled RNA, and further alkylating the labeled RNA.

[0062] In one embodiment, the labeled RNA is alkylated using an alkylating agent selected from the group consisting of iodoacetamide, iodoacetic acid, N-ethylmaleimide, and 4-vinylpyridine.

[0063] In one embodiment, the alkylating agent is preferably iodoacetamide.

[0064] In one embodiment, the step of sequencing RNA extracted from a cell sample comprises:

[0065] (i) reverse transcribing RNA to obtain a cDNA library,

[0066] (ii) optionally, amplifying the cDNA library, and

[0067] (iii) sequencing the cDNA library, preferably sequencing the cDNA library by next generation sequencing (NGS), deep sequencing or targeted sequencing of personalized sequences.

[0068] In one embodiment, the step of sequencing RNA extracted from a cell sample comprises:

[0069] (i) reverse transcribing the total RNA to obtain a total cDNA library,

[0070] (ii) optionally, amplifying the total cDNA library, and

[0071] (iii) sequencing the total cDNA library by next generation sequencing (NGS).

[0072] In one embodiment, when the sample is incubated in the presence of an RNA labeling agent and / or when the labeled RNA undergoes nucleotide conversion, the reverse transcribed RNA converts uridine (U) to cytidine (C) rather than converting uridine (U) to thymidine (T).

[0073] In one embodiment, when a cell sample is cultured in the presence of an RNA labeling agent, reverse transcription of total RNA converts uridine (U) to cytidine (C) rather than converting uridine (U) to thymidine (T).

[0074] In one embodiment, when the sample is incubated in the presence of an RNA labeling agent, preferably a thiol-labeled RNA precursor, and / or when the labeled RNA is placed under conditions that allow nucleotide substitutions, the RNA undergoes substitutions of adenosine (A) to guanosine (G) in the first strand synthesis and thymidine (T) to cytidine (C) in the second strand synthesis during reverse transcription.

[0075] In one embodiment, the step of identifying at least one viral nucleic acid sequence hit that matches at least one sequence read of the first set of sequence reads comprises:

[0076] (i) optionally, filtering the first set of sequence reads,

[0077] (ii) optionally, assembling the sequences into contigs,

[0078] (iii) aligning the sequence reads or contigs to a database comprising viral nucleic acid sequences,

[0079] (iv) identifying at least one viral nucleic acid sequence hit that matches at least one sequence read or contig, and

[0080] (v) optionally, realigning the sequence reads or contigs with the viral nucleic acid sequence hits identified in step (iv) to determine a consensus viral nucleic acid sequence,

[0081] At least one consensus viral nucleic acid sequence is thereby identified.

[0082] In one embodiment, at least one microbial nucleic acid sequence hit is identified by:

[0083] (i) optionally, filtering the first and / or second set of sequence reads,

[0084] (ii) optionally, assembling the sequences into contigs,

[0085] (iii) aligning the sequence reads or contigs to a database comprising microbial nucleic acid sequences,

[0086] (iv) identifying at least one microbial nucleic acid sequence hit that matches the at least one sequence read or contig, and

[0087] (v) optionally, realigning the sequence reads or contigs with the microbial nucleic acid sequence hits identified in step (iv) to determine a consensus microbial nucleic acid sequence,

[0088] Among them, the common microbial nucleic acid sequence corresponds to the microbial nucleic acid sequence hit.

[0089] In one embodiment, if:

[0090] - the number and / or ratio of T→C substitutions in the sequence reads that match at least one microbial nucleic acid sequence hit in the first set of sequence reads is greater than the number and / or ratio of T→C substitutions in the reference sequence; and / or

[0091] - the number and / or ratio of T→C substitutions in sequence reads matching at least one microbial nucleic acid sequence hit in the first set of sequence reads is greater than the number and / or ratio of T→A and / or T→G substitutions in the same sequence reads,

[0092] Then at least one microbial nucleic acid sequence hit belongs to a living microorganism.

[0093] In one embodiment, if the number of T→C substitutions in the sequence reads that match the at least one identified viral nucleic acid sequence hit in the first set of sequence reads is greater than the number of T→C substitutions in the second set of sequence reads, then the at least one viral nucleic acid sequence hit belongs to an infectious virus.

[0094] In one embodiment, if:

[0095] - the number and / or ratio of second-strand synthesis T→C substitutions in the sequence reads that match at least one microbial nucleic acid sequence hit in the first set of sequence reads is greater than the number and / or ratio of second-strand synthesis T→C substitutions in the control sequence; and / or

[0096] - the number and / or ratio of second-strand synthesis T→C substitutions in sequence reads that match at least one microbial nucleic acid sequence hit in the first set of sequence reads is greater than the number and / or ratio of second-strand synthesis T→A and / or second-strand synthesis T→G substitutions in the same sequence reads,

[0097] Then at least one microbial nucleic acid sequence hit belongs to a living microorganism.

[0098] In one embodiment, the in vitro method according to the invention comprises the following steps:

[0099] (1)(i) sequencing unlabeled total RNA extracted from a cell sample, wherein the unlabeled total RNA is obtained by culturing the cell sample in the absence of an RNA labeling agent, thereby obtaining a plurality of sequence reads,

[0100] (ii) identifying at least one viral nucleic acid sequence hit that matches the sequence read, and

[0101] (iii) determining the number of substituted nucleotides in sequence reads that match the at least one viral nucleic acid sequence hit identified; and

[0102] (2)(i) sequencing labeled total RNA extracted from a cell sample, wherein the labeled total RNA is obtained by culturing the cell sample in the presence of a labeling agent, thereby obtaining a plurality of sequence reads,

[0103] (ii) determining the number of substituted nucleotides in sequence reads matching said identified at least one viral nucleic acid sequence hit,

[0104] (3) comparing the number of substituted nucleotides determined in (1)(iii) and (2)(ii), and

[0105] (4) If the number of substituted nucleotides determined in (2)(ii) is greater than the number of substituted nucleotides determined in (1)(iii), it is inferred that the viral nucleic acid sequence hit belongs to an infectious virus.

[0106] In one embodiment, the microorganism is selected from the group consisting of viruses, bacteria, archaea, fungi, and protozoa.

[0107] The present invention also relates to an in vitro method for diagnosing a microbial infection in a subject, comprising:

[0108] (a) providing a sample of the subject,

[0109] (b) performing on said sample an in vitro method for distinguishing between live and dead microorganisms, and

[0110] (c) diagnosing the subject as being infected with a microorganism if at least one of the identified microorganism nucleic acid sequence hits belongs to a live microorganism.

[0111] The present invention also relates to an in vitro method for diagnosing a viral infection in a subject, comprising:

[0112] (a) providing a cell sample from a subject,

[0113] (b) performing an in vitro method for distinguishing between infectious and non-infectious viral nucleic acid sequences in a cell sample according to the present invention on the cell sample, and

[0114] (c) diagnosing the subject as being infected with the virus if at least one of the identified viral nucleic acid sequences is a hit belonging to an infectious virus.

[0115] The present invention also relates to a method of treating a subject affected by a microbial infection, comprising:

[0116] (a) providing a sample of the subject,

[0117] (b) performing on said sample an in vitro method for distinguishing between live and dead microorganisms,

[0118] (c) diagnosing the subject as being infected with a microorganism if at least one of the identified microorganism nucleic acid sequence hits is that of a live microorganism, and

[0119] (d) if the subject is diagnosed as having a microbial infection in step c), treating the subject.

[0120] The present invention also relates to a method for assessing the risk of microbial contamination in a sample, comprising:

[0121] (a) provide samples,

[0122] (b) performing on said sample an in vitro method for distinguishing between live and dead microorganisms, and

[0123] (c) If at least one of the identified microbial nucleic acid sequence hits belongs to a living microorganism, it is inferred that the sample is at risk of being contaminated.

[0124] definition

[0125] In the present invention, the following terms have the following meanings:

[0126] As used herein, "about" or "approximately" can mean within an acceptable error range for a particular value as determined by one skilled in the art, depending in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, according to practice in the art, "about" can mean within 1 or more than 1 standard deviation. "About" before a number means within plus or minus 10% of the numerical value. Alternatively, particularly with respect to biological systems or processes, the term can mean within an order of magnitude, within 5 times, and more preferably within 2 times of the value. Where specific values ​​are described in the application and claims, the term "about" should be assumed to be within an acceptable error range for the specific value unless otherwise indicated.

[0127] As used herein, "amplification" refers to the process of generating multiple copies, ie, at least 2 copies, of a desired template sequence. Techniques for amplifying nucleic acids are well known to those skilled in the art and include specific amplification methods as well as random amplification methods.

[0128] "Biological sample" as used herein refers to any sample obtained from a subject, obtainable from a subject, or otherwise obtained from a subject. "Biological sample" includes "solid tissue sample" and "fluid sample". The term "solid tissue sample" herein refers to a solid tissue sample isolated from anywhere in the body. Tissue samples contain undecomposed cells that are distributed in large clusters. Examples of tissue samples include, but are not limited to, biopsy samples and autopsy samples. The term "fluid sample" herein refers to a fluid sample isolated from anywhere in the body. Examples of fluid samples include, but are not limited to, serum, plasma, whole blood, urine, saliva, breast milk, tears, sweat, joint fluid, cerebrospinal fluid, lymph, sputum, mucus, pelvic fluid, synovial fluid, body cavity flushing fluid, eye brushing, skin scraping, cheek swab, vaginal swab, cervical smear, rectal swab, fluid extraction, semen, vaginal fluid, ascites, and amniotic fluid. In a preferred embodiment, a "biological sample" is a cell sample comprising at least one cell, i.e., any biological sample described herein.

[0129] As used herein, "cDNA library" refers to a library composed of complementary DNA reverse transcribed from mRNA.

[0130] " contig " used herein refers to overlapping sequence reads. Typically, contigs are continuous nucleic acid sequences produced by the recombinant production of small DNA fragments (sequence reads) produced by next generation sequencing. In fact, assembly software will search for overlapping sequence read pairs. Optionally, assembly software can access nucleic acid or amino acid databases with "alignment and inspection", thereby verifying the assembly of sequence reads. The sequence assembled by paired overlapping sequence reads produces longer continuous reads (contigs) of sequencing DNA. By repeating this process multiple times, first using initial short sequence read pairs, then using increasingly longer pairs as a result of previous assembly, longer contigs can be determined.

[0131] As used herein, "deep sequencing" refers to nucleic acid sequencing to a depth that allows each base to be read multiple times from an independent nucleic acid molecule (e.g., sequencing a large number of template molecules relative to the length of the sequence) and allows thousands of molecules to be sequenced simultaneously, thereby allowing the characterization of complex nucleic acid molecule libraries and improving sequencing accuracy. Deep sequencing of the transcriptome, also known as RNA sequencing, provides the sequence and frequency of the species of RNA molecules present at any particular time in a particular sample.

[0132] As used herein, "expected value" or "e-value" refers to a parameter that describes the number of sequence hits that one can expect to see when aligning sequence reads or contigs on a database of a particular size. As the number of matches increases, the e-value decreases exponentially. Essentially, the e-value describes the random background noise. For example, an e-value of 1 assigned to a hit can be interpreted as meaning that in a database of the current size, one might expect to see 1 match with a similarity score only by chance. The lower the e-value, or the closer it is to zero, the more "significant" the match.

[0133] As used herein, "live microorganism" refers to any transcriptionally active microorganism, i.e. a microorganism that is capable of synthesizing RNA by itself (e.g. in the case of bacteria, archaea, fungi or protozoa) or after infecting a host cell (e.g. in the case of viruses). Live microorganisms include latent microorganisms, i.e. dormant microorganisms that can be reactivated. It should be noted that latent microorganisms, although in a dormant state, can exhibit basic transcriptional activity. In contrast, "dead microorganisms" refer to microorganisms that are not transcriptionally active, i.e. microorganisms for which no transcribed genes can be detected. In the context of the present invention, the method aims to distinguish between live microorganisms and inert microbial nucleic acid sequences, which are either free in the sample or contained in so-called dead microorganisms.

[0134] As used herein, "lysate" refers to the collection of liquid or solid material following a lysis process.

[0135] As used herein, "lysis" refers to the destruction of a biological sample (or the act of destroying a biological sample) to obtain material that would otherwise be unavailable. When the biological sample is a cell, lysis refers to breaking the cell membrane of the cell, resulting in the spillage of the cell contents. Lysis methods are well known to those skilled in the art and include, but are not limited to, proteolytic lysis, chemical lysis, thermal lysis, mechanical lysis, and osmotic lysis.

[0136] As used herein, a "nucleic acid sequence primer" or "primer" refers to an oligonucleotide that can hybridize or anneal to a nucleic acid sequence under appropriate conditions and serve as a site for initiation of nucleotide polymerization, such as the presence of nucleoside triphosphates and an enzyme for polymerization, such as DNA polymerase or RNA polymerase or reverse transcriptase, in an appropriate buffer and at an appropriate temperature.

[0137] As used herein, "oligonucleotide" refers to a nucleotide polymer, typically a single-stranded nucleotide polymer. In some embodiments, an oligonucleotide comprises 2 to 500 nucleotides, preferably 10 to 150 nucleotides, and preferably 20 to 100 nucleotides. An oligonucleotide can be synthetic or enzymatic. In some embodiments, an oligonucleotide can comprise nucleotide monomers, deoxynucleotide monomers, or a mixture thereof.

[0138] As used herein, "microorganism" refers to an organism, such as, but not limited to, a virus, bacteria, archaea, fungus, or protozoa, that may infect or contaminate a sample; and / or produce, transmit, or carry disease in a subject.

[0139] As used herein, "polymerase chain reaction" or "PCR" includes, but is not limited to, methods such as allele-specific PCR, asymmetric PCR, hot-start PCR, sequence-specific PCR, methylation-specific PCR, mini-primer PCR, multiplex ligation-dependent probe amplification, multiplex PCR, nested PCR, quantitative PCR, reverse transcription PCR, and / or sequence-descending PCR. DNA polymerases suitable for amplifying nucleic acids include, but are not limited to, Taq polymerase Stoffel fragment, Taq polymerase, Advantage DNA polymerase, AmpliTaq, AmpliTaq Gold, Titanium Taq polymerase, KlenTaq DNA polymerase, Platium Taq polymerase, Accuprime Taq polymerase, Pfu polymerase, Pfu polymerase turbo, Vent polymerase, Ventexo-polymerase, Pwo polymerase, 9Nm DNA polymerase, Therminator, Pfx DNA polymerase, Expand DNA polymerase, rTth DNA polymerase, DyNAzyme-EXT polymerase, Klenow fragment, DNA polymerase I, T7 polymerase, Sequenase™, Tfi polymerase, T4 DNA polymerase, Bst polymerase, Bca polymerase, BSU polymerase, phi-29 DNA polymerase, and DNA polymerase beta, or modified versions thereof. In one embodiment, the DNA polymerase has 3'-5' proofreading activity, i.e., exonuclease activity. In one embodiment, the DNA polymerase has 5'-3' proofreading activity, i.e., exonuclease activity. In one embodiment, the DNA polymerase has strand displacement activity, i.e., the DNA polymerase binds to and approaches template-dependent nucleic acid synthesis, causing the paired nucleic acid to dissociate from its complementary strand in the 5' to 3' direction. For example, Escherichia coli DNA polymerase I, the Klenow fragment of DNA polymerase I, T7 or T5 phage DNA polymerase, and the DNA polymerase of HIV viral reverse transcriptase are all enzymes with both polymerase activity and strand displacement activity. For example, helicase reagents can be used together with inducers that do not have strand displacement activity to produce a strand displacement effect, i.e., nucleic acid displacement is coupled to the synthesis of nucleic acids with the same sequence. Similarly, proteins from Escherichia coli or other organisms, such as Rec A or single-stranded binding proteins, can be used together with other inducers to produce or promote strand displacement (Kornberg & Baker (1992). Chapters 4-6. In DNA replication (Second Edition, 113-225 pages). New York: WH Freeman).

[0140] As used herein, "random amplification techniques" refer to amplifying any nucleic acid present in a biological sample, regardless of its sequence. This includes, but is not limited to, multiple displacement amplification (MDA), random PCR, random amplified polymorphic DNA (RAPD), or multiple annealing and amplification cycles (MALBAC).

[0141] As used herein, "transcriptionally active microbial nucleic acid sequence" refers to a nucleic acid sequence belonging to a living microorganism, i.e., a microorganism that expresses a microbial gene, even if the microorganism is latent. In contrast, as used herein, "inert microbial nucleic acid sequence" refers to a nucleic acid sequence belonging to an inactive microorganism (i.e., a dead microorganism). The term "inert microbial nucleic acid sequence" also refers to a free nucleic acid sequence, i.e., a free nucleic acid sequence outside of a microorganism, whether intact or degraded / fragmented, but in no case active.

[0142] As used herein, a "transcriptionally active viral nucleic acid sequence" refers to a nucleic acid sequence belonging to a live virus, i.e., a nucleic acid sequence of a live virus that expresses viral genes, even if the viral cycle is sterile, i.e., does not result in the formation of viral particles (e.g., in the case of latent viruses). In contrast, an "inert viral nucleic acid sequence" as used herein refers to a nucleic acid sequence belonging to an inactive virus, i.e., a dead virus or a nucleic acid not associated with a viral particle.

[0143] As used herein, "reverse transcription" refers to the use of an RNA-guided DNA polymerase (reverse transcriptase, abbreviated as "RT") to replicate RNA to produce a complementary DNA strand ("cDNA"). Reverse transcription of RNA can be achieved using a reverse transcriptase and a mixture of four deoxyribonucleoside triphosphates (dNTPs), i.e., a mixture of deoxyadenosine triphosphate (dATP), deoxycytidine triphosphate (dCTP), deoxyguanosine triphosphate (dGTP), and (deoxy)thymidine triphosphate (dTTP), by techniques well known to those skilled in the art. In some embodiments, reverse transcription of RNA includes the first step of synthesizing a first-strand cDNA. Methods for synthesizing first-strand cDNA are well known to those skilled in the art. The first-strand cDNA synthesis reaction can use a combination of sequence-specific primers, oligo (dT) primers, or random primers. Examples of reverse transcriptases include, but are not limited to, M-MLV reverse transcriptase, SuperScript II (Invitrogen), SuperScript III (Invitrogen), SuperScript IV (Invitrogen), Maxima (ThermoFisher Science), ProtoScript II (New England Biolabs), PrimeScript (ClonTech).

[0144] As used herein, a "sequence read" refers to a sequence or data representing a nucleotide base sequence, in other words, the order of monomers in a nucleic acid sequence, as determined by a sequencer.

[0145] As used herein, a "sequencer" or "sequencer" refers to a device used to determine the order of the components of a biopolymer (e.g., a nucleic acid or protein). Preferably, a sequencer within the meaning of the present invention refers to a next-generation sequencer. "Next-generation sequencer" can include many different sequencers based on different technologies, such as Illumina sequencing, Roche 454 sequencing, Ion Torrent sequencing, solid-state sequencing, etc.

[0146] As used herein, a "subject" refers to a mammal, preferably a human. In one embodiment, the subject is a pet, including but not limited to a dog, cat, guinea pig, hamster, rat, mouse, ferret, rabbit, bird, or amphibian. In one embodiment, the subject can be a "patient," i.e., female or male, adult or child, who is awaiting medical care, is currently receiving medical care, has been / is currently / will be the subject of a medical procedure, or is being monitored for the development of a disease, disorder, or condition, particularly a viral, bacterial, archaeal, fungal, or protozoal infection.

[0147] As used herein, "template" or "template sequence" refers to a nucleic acid sequence to be amplified. The template may comprise DNA or RNA. In one embodiment, the template sequence is known. In one embodiment, the template sequence is unknown.

[0148] The terms used herein are intended to describe specific situations only and are not intended to be limiting. As used herein, unmodified quantifiers include the singular and the plural, unless the context clearly indicates otherwise. In addition, when the terms "include," "comprising," "containing," "having," or variations thereof are used in the detailed description and / or claims, these terms are intended to be inclusive in a manner similar to the term "comprising." DETAILED DESCRIPTION

[0149] The present invention relates to a method for distinguishing between live and dead microorganisms in a sample, preferably a cell sample. Specifically, the method according to the invention is based on distinguishing between transcriptionally active microbial nucleic acid sequences and inert microbial nucleic acid sequences in a sample, preferably a cell sample.

[0150] The methods of the present invention are particularly suitable for distinguishing between (1) dead microorganisms, such as viruses, bacteria, archaea, fungi or protozoa, and inert microbial sequences; and (2) active (or transcriptionally active) and latent microorganisms, such as viruses, bacteria, archaea, fungi or protozoa.

[0151] Therefore, it can be understood that the present method is easily applicable to the detection of any type of microorganisms, as well as the differentiation of living and dead microorganisms.

[0152] In one embodiment, the microorganism is selected from the group consisting of viruses, bacteria, archaea, fungi, and protozoa.

[0153] In one embodiment, the microorganism is a virus.

[0154] Viruses are small infectious agents that replicate inside living cells and can infect all types of life forms.

[0155] The Baltimore virus classification is based on the mechanism of mRNA production. Viruses must produce mRNA from their genomes to produce proteins and replicate themselves, but each virus family uses a different mechanism to achieve this. The viral genome may be single-stranded (ss) or double-stranded (ds) RNA or DNA and may or may not use reverse transcriptase. Additionally, single-stranded RNA viruses may be sense (+) or negative sense (-). This classification divides viruses into seven groups:

[0156] 1. dsDNA viruses (e.g., adenovirus, herpes virus, or pox virus),

[0157] II. (+) ssDNA viruses (e.g., Anelloviridae, Diviridae, Circoviridae, Geminiviridae, Genoviridae, Boviridae, Microviridae, Nanoviridae, Parvoviridae, Streptoviridae, or Helicoviridae),

[0158] III. dsRNA viruses (e.g. reoviruses),

[0159] IV. (+) ssRNA viruses (e.g., picornaviruses or togaviruses),

[0160] V. (-) ssRNA viruses (e.g., orthomyxoviruses, rhabdoviruses), VI. (+) ssRNA-RT viruses with a DNA intermediate in their life cycle (e.g., retroviruses),

[0161] VII. dsDNA-RT viruses with a DNA-RNA intermediate in their life cycle (eg, hepadnaviruses).

[0162] In one embodiment, the method according to the present invention is used to distinguish between a sample, preferably a cell sample, comprising transcriptionally active viral nucleic acid sequences and inert viral nucleic acid sequences, wherein the transcriptionally active viral nucleic acid sequences and inert viral nucleic acid sequences belong to a virus selected from the group consisting of dsDNA viruses, (+)ssDNA viruses, dsRNA viruses, (+)ssRNA viruses, (-)ssRNA viruses, (+)ssRNA-RT viruses and dsDNA-RT viruses.

[0163] In one embodiment, the method according to the present invention is used to distinguish between samples, preferably cell samples, containing transcriptionally active viral nucleic acid sequences and inert viral nucleic acid sequences, wherein the transcriptionally active viral nucleic acid sequences and inert viral nucleic acid sequences belong to viruses selected from the viruses disclosed in the International Committee on Taxonomy of Viruses (ICTV) database, preferably the viruses disclosed in the ICTV Master Species List 2018b.v2 (MSL#34) on May 31, 2019, which list is incorporated herein by reference in its entirety.

[0164] In one embodiment, the microorganism is a bacterium.

[0165] In one embodiment, the method according to the invention is used for distinguishing between a sample, preferably a cell sample, comprising transcriptionally active bacterial nucleic acid sequences and inert bacterial nucleic acid sequences belonging to bacteria.

[0166] Examples of bacteria include, but are not limited to, bacteria belonging to the phyla Acidobacteria, Actinobacteria, Aquabacteria, Bacteroidetes, Chlamydiae, Chlorobacteria, Chloroflexi, Auricobacteria, Cyanobacteria, Deferrobacteria, Deinococcus-Thermus, Digitococcus, Fibrobacteria, Firmicutes, Fusobacteria, Gemmatimonadetes, Nitrospira, Planctomyces, Proteobacteria, Spirochaetes, Thermodesulfurobacteria, Thermomicrobacteria, Thermotoga, and Verrucomicrobia, including subclasses thereof.

[0167] In a preferred embodiment, the bacterium is of the phylum Firmicutes, preferably of the class Bacillus, more preferably of the subclass Mollicutes, and even more preferably of the genus Mycoplasma.

[0168] In one embodiment, the microorganism is an archaea.

[0169] In one embodiment, the method according to the invention is used for distinguishing between a sample, preferably a cell sample, comprising transcriptionally active archaeal nucleic acid sequences belonging to Archaea and inert archaeal nucleic acid sequences.

[0170] Examples of archaea include, but are not limited to, archaea belonging to the phyla Mysterarchaeota, Eonarchaeota, Altiarchaeia, Archaea, Asgardarchaeota, Deeparchaeota, Crenarchaeota, Prohaloarchaeota, Geoarchaeota, Halophiles, Archaeota, Methanobacteria, Methanococci, Methanobacteria, Methanopyri, Nanohaloarchaea, Parvarchaeota, Thalassoarchaeia, Thaumarchaeota, Thermococci, Thermoplasmas, and Vossarchaeota, including subclasses thereof.

[0171] In one embodiment, the microorganism is a fungus.

[0172] Fungi are eukaryotic organisms, including yeasts and molds, characterized by the presence of chitin in their cell walls.

[0173] In one embodiment, the method according to the invention is used for distinguishing between a sample, preferably a cell sample, comprising transcriptionally active fungal nucleic acid sequences and inert fungal nucleic acid sequences belonging to a fungus.

[0174] Examples of fungi include, but are not limited to, fungi belonging to the phyla Ascomycota, Basidiomycota, Entorrhizomycota, Glomeromycota, Mucormycota, Calcarisporiellomycota, Mortierella, Combretum, Entomomycota, Oleobacteria, Basidiomycota, Neomycota, Chytridiomycota, and Cladosporium; including subphyla thereof.

[0175] In one embodiment, the microorganism is a protozoa.

[0176] In one embodiment, the method according to the invention is used to differentiate between a sample, preferably a cell sample, comprising transcriptionally active and inert protozoan nucleic acid sequences belonging to protozoa.

[0177] Examples of protozoa include, but are not limited to, protozoa belonging to the phyla Euglena, Amoeba, Choanoflagellate, Microsporidia, and Sulcozoa; including subclasses thereof.

[0178] In some embodiments, the sample is a biological sample. Examples of suitable biological samples include, but are not limited to, solid tissue samples and fluid samples.

[0179] In one embodiment, the biological sample is obtained by sampling using minimally invasive or non-invasive methods.

[0180] In one embodiment, the biological sample is previously obtained from the subject, ie the method according to the invention is an in vitro method.

[0181] In some embodiments, the biological sample is a cell sample. "Cell sample" refers to any biological sample described herein that contains at least one cell.

[0182] In one embodiment, the biological sample is cultured. Thus, the term "biological sample" includes cell or tissue cultures, preferably in vitro cell or tissue cultures, such as cultures of cells or tissues isolated from a cytological sample, a tissue sample or a biological fluid sample.

[0183] In one embodiment, the method according to the invention comprises an initial step of culturing a sample, preferably a cell sample, preferably in vitro.

[0184] The cultivation of cell samples, in particular cells or tissues isolated from cytological samples, tissue samples or biological fluid samples, is well known to those skilled in the art.

[0185] In one embodiment, the cell sample is seeded at a density that allows exponential growth. In one embodiment, the biological sample is seeded at a confluence of about 50% to about 80%.

[0186] The initial step of culturing the sample requires (1) allowing latent microorganisms (e.g., viruses, bacteria, archaea, fungi, or protozoa) in the sample to transcribe RNA (this is the key biological process used in the present method to distinguish between living and dead microorganisms) and (2) metabolic labeling. In the case where the microorganism to be detected is not a self-replicating microorganism (e.g., a virus or certain bacteria, such as mycoplasmas), the sample should be a cell sample to allow the latent microorganism to infect the cell and replicate. In the case where the microorganism is a self-replicating microorganism (i.e., the microorganism contains cells or is itself a cell, such as is typically a bacterium, archaea, fungi, or protozoa), it is not mandatory that the sample be a cell sample.

[0187] In some embodiments, the sample is not a biological sample. In this case, the sample can be, for example, an environmental sample such as water, soil, air, etc. Other examples of non-biological samples include food samples. Other examples of non-biological samples include preservation culture media.

[0188] In one embodiment, a method for distinguishing between live and dead microorganisms, preferably viruses, bacteria, archaea, fungi or protozoa, in a sample, preferably a cell sample, comprises distinguishing between transcriptionally active microbial nucleic acid sequences and inert microbial nucleic acid sequences, preferably viruses, bacteria, archaea, fungi or protozoa, in a sample, preferably a cell sample, the method comprising the following steps:

[0189] (a) sequencing a first set of RNA extracted from a sample, preferably a cell sample, wherein the first set of RNA is obtained by culturing the sample, preferably a cell sample, in the presence of an RNA labeling agent and further by subjecting the extracted RNA to conditions that allow nucleotide substitutions; thereby obtaining a first set of sequence reads;

[0190] (b) comparing the number of substituted nucleotides in the first set of sequences that match at least one microbial, preferably viral, bacterial, archaeal, fungal or protozoan nucleic acid sequence hit to a reference sequence;

[0191] (c) if the number of substituted nucleotides in the sequence reads in the first set of sequence reads that match the at least one microbial, preferably viral, bacterial, archaeon, fungal or protozoan nucleic acid sequence hit is greater than the number of nucleotides randomly substituted in the reference sequence, then inferring that the at least one microbial, preferably viral, bacterial, archaeon, fungal or protozoan nucleic acid sequence hit belongs to a living microorganism, preferably a virus, bacterium, archaeon, fungus or protozoan.

[0192] In one embodiment, the method of the present invention is carried out under conditions that cause the reverse transcriptase to produce errors (i.e., incorporation of mismatched nucleotides), which can be detected and compared by reference to a common reverse transcription standard method. These conditions include the presence of a label such as a thiol label in the RNA to be reverse transcribed, and / or the presence of nucleotide modifications by nucleotide substitution techniques. Exemplary conditions will be described in further detail below.

[0193] As used herein, the term "mismatched nucleotide" refers to a nucleotide that binds in a non-Watson-Crick base pairing manner.

[0194] In one embodiment, the error rate of the reverse transcriptase is independent of the fidelity of the reverse transcriptase.

[0195] As used herein, the term "fidelity" with reference to reverse transcriptase refers to the degree of sequence accuracy maintained by the enzyme during the synthesis of DNA from RNA. Fidelity is inversely proportional to the error rate of reverse transcription.

[0196] In one embodiment, the method according to the invention comprises the step of sequencing a first set of RNA extracted from a sample, preferably a cell sample. In one embodiment, the method according to the invention comprises the step of sequencing a first set of total RNA extracted from a sample, preferably a cell sample. In one embodiment, the method according to the invention comprises the step of sequencing a first set of total messenger RNA (mRNA) extracted from a sample, preferably a cell sample.

[0197] In one embodiment, the step of sequencing a first set of RNA extracted from a sample, preferably a cell sample, includes one or more than one or all of the sub-steps of labeling the RNA, lysing the cells, extracting the RNA, replacing nucleotides in the labeled RNA, generating a cDNA library, amplifying the cDNA library, and sequencing the cDNA library.

[0198] During in vitro transcription, RNA is typically labeled in culture by adding a label to the culture medium to be incorporated into the RNA transcript, thereby obtaining labeled RNA. Alternatively or additionally, RNA can also be labeled in culture without adding a label to be incorporated into the RNA transcript, if the sample from which the RNA is extracted already contains such a label, as will be described in further detail below.

[0199] "RNA transcript" refers to any newly synthesized RNA molecule.

[0200] The labeling of RNA transcripts can be performed by techniques well known to those skilled in the art. These techniques include, but are not limited to, those described in Schulz & Rentmeister (2014. Chembiochem. 15(16): 2342-7), Huang & Yu (2013. Curr Protoc Mol Biol. Chapter 4: Unit 4.15), and Liu et al. (2016. Bioessays. 38(2): 192-200).

[0201] Preferably, metabolic labeling of RNA alters Watson-Crick base pairing and results in the reverse transcription of the labeled RNA into substituted nucleotides, i.e., pairing the labeled nucleotide with a non-Watson-Crick nucleotide. For example, during first-strand synthesis, labeled uridine may pair with guanosine instead of adenosine. Thus, cytidine will be incorporated during second-strand synthesis, ultimately resulting in a thymidine (T) to cytosine (C) substitution relative to the initial nucleic acid sequence.

[0202] In one embodiment, the metabolic labeling of RNA transcripts is performed by thiol labeling. Thiol labeling is a technique well known in the art, which includes incorporating a thiol-labeled RNA precursor into a newly synthesized RNA. This technique includes, but is not limited to, Cleary et al. (2005. Nat Biotechnol. 23 (2): 232-7), Miller et al. (2009. Nat Methods. 6 (6): 439-41), Garibaldi et al. (2017. Methods Mol Biol. 1648: 169-176), Russo et al. (2017. Methods. 120: 39-48) and Herzog et al. (2017. Nat Methods. 14 (12): 1198-1204) described techniques.

[0203] Examples of suitable thiol-labeled RNA precursors include, but are not limited to, 4-thiouridine, 2-thiouridine, 2,4-dithiouridine, 2-thio-4-deoxyuridine, 5-ethoxycarbonyl-2-thiouridine, 5-carboxy-2-thiouridine, 5-(n-propyl)-2-thiouridine, 6-methyl-2-thiouridine, 6-(n-propyl)-2-thiouridine, 6-thiguanosine, 6-methylthiguanosine, 6-thiinosine, and 6-methylthiinosine.

[0204] In one embodiment, the thiol-labeled RNA precursor is a thiouridine derivative, preferably selected from 4-thiouridine (4sU), 2-thiouridine (2sU), 2,4-dithiouridine (2,4sU), 2-thio-4-deoxyuridine, 5-ethoxyformyl-2-thiouridine, 5-carboxy-2-thiouridine, 5-(n-propyl)-2-thiouridine, 6-methyl-2-thiouridine and 6-(n-propyl)-2-thiouridine.

[0205] In a preferred embodiment, the thiol-labeled RNA precursor is 4-thiouridine (sometimes abbreviated as "4sU" or "s4U").

[0206] In one embodiment, the thiol-labeled RNA precursor is supplied to the sample, preferably a cell sample, from the culture medium.In one embodiment, the thiol-labeled RNA precursor is added to the culture medium.

[0207] When a thiol-labeled RNA precursor is added to a culture medium, it can be introduced into a sample, preferably a cell sample (e.g., a virally infected cell or a microorganism containing or itself a cell, such as bacteria, fungi, or protozoa) via a specific transporter known as an "equilibrium nucleoside transporter." These transporters are nearly ubiquitous in metazoans. In particular, 4-thiouridine can be introduced into cells via the equilibrative nucleoside transporter 1 (ENT1), encoded by the human SLC29A1 gene.

[0208] In one embodiment, fresh thiol-labeled RNA precursor is added to the culture every 1 hour, 2 hours, 3 hours, 4 hours, 5 hours, 6 hours, or more than 6 hours.

[0209] In one embodiment, the sample, preferably a cell sample, is cultured in the culture medium containing the thiol-labeled RNA precursor for a period of about 2 hours to about 15 hours, preferably about 4 hours to about 12 hours, preferably about 6 hours to about 10 hours.

[0210] In one embodiment, a sample, preferably a cell sample, is cultured in a medium containing a thiol-labeled RNA precursor for a first time period and a second time period, including adding fresh thiol-labeled RNA precursor between the first time period and the second time period. In one embodiment, the first time period is about 1 hour to about 10 hours, preferably about 2 hours to about 8 hours, preferably about 4 hours to about 6 hours, and preferably about 6 hours. In one embodiment, the second time period is about 1 hour to about 6 hours, preferably about 2 hours to about 5 hours, preferably about 3 hours to about 4 hours, and preferably about 3 hours.

[0211] Preferably, the thiol-labeled RNA precursor is non-toxic to the sample, preferably, non-toxic to the cell sample.

[0212] In one embodiment, the thiol-labeled RNA precursor is provided to a sample, preferably a cell sample, at a concentration that does not impair cell viability. In one embodiment, a "concentration that does not impair cell viability" is about 1 μM to about 2 mM final, preferably about 10 μM to about 1.5 mM final, preferably about 100 μM to about 1 mM final, preferably about 250 μM to about 1 mM final, preferably about 500 μM to about 1 mM final, preferably about 700 μM to about 900 μM final, preferably about 800 μM final of the thiol-labeled RNA precursor.

[0213] In one embodiment, the thiol-labeled RNA precursor is provided to a sample, preferably a cell sample, directly from a microorganism, preferably a virus, bacteria, archaea, fungus or protozoa. In one embodiment, the thiol-labeled RNA precursor is not added to the culture medium.

[0214] Certain microorganisms are able to catalyze the biosynthesis of thiol-labeled RNA precursors using enzymes such as, but not limited to, 4-thiouridine synthase (ThiI) (Mueller et al., 1998. Nucleic Acids Res. 26(11):2606-10) or 2-thiouridine synthase (MnmA) (Kambampati & Lauhon, 2003. Biochemistry. 42(4):1109-1; Black & Dos Santos, 2015. J Bacteriol. 197(11):1952-62).

[0215] In the particular embodiment in which the sample comprises a microorganism capable of catalyzing the biosynthesis of a thiol-labeled RNA precursor, it may be advantageous to further supply the thiol-labeled RNA precursor to the sample, preferably a cell sample, from the culture medium.

[0216] In this embodiment, the thiol-labeled RNA precursor further provided from the culture medium may be the same as or different from the thiol-labeled RNA precursor provided by the microorganism.

[0217] In this embodiment, the thiol-labeled RNA precursor further provided from the culture medium can be supplied as described above (with respect to, but not limited to, the addition of fresh thiol-labeled RNA precursor, concentration, time period, etc.).

[0218] Thiol-labeled RNA precursors and thiol-labeled RNA are light-sensitive and easily oxidized. Therefore, in one embodiment, RNA labeling is carried out in the dark, or at least carried out in the absence of light. In one embodiment, RNA labeling is carried out in the presence of a reducing agent. Examples of suitable reducing agents include, but are not limited to, beta-mercaptoethanol, dithiothreitol (DTT), tris (2-carboxyethyl) phosphine (TCEP), cysteine, N-acetylcysteine, cysteamine, sodium 2-mercaptoethanesulfonate, dithioerythritol (DTE) and bis (2-mercaptoethyl) sulfone.

[0219] Typically, the purpose of lysing the cells of a sample, preferably a cell sample, is to release the contents of the cells, in particular their RNA. In one embodiment, lysing the cells of a sample may be optional, for example where the RNA contents of the cells have already been released in the sample.

[0220] In one embodiment, the cells are lysed by chemical lysis, mechanical lysis, proteolytic lysis, thermal lysis and / or osmotic lysis. These cell lysis techniques are well known to those skilled in the art.

[0221] In one embodiment, cells are lysed in a suitable lysis solution. The lysis solution may include various components, including salts, buffers, detergents, reducing agents, protease inhibitors, nuclease inhibitors, glycerol, sugars, etc. Those skilled in the art have knowledge of lysis solutions and can easily design and / or select an appropriate lysis solution based on the type of cells to be lysed.

[0222] In one embodiment, cell lysis is performed in the presence of a ribonuclease (RNase) inhibitor. RNases can sometimes be released from cells during cell lysis or co-purify with the isolated RNA, thereby affecting downstream applications. This RNase contamination can also be introduced through needle tips, tubing, and other reagents used in the procedure. RNase inhibitors are commercially available.

[0223] Thiol-labeled RNA is light-sensitive and easily oxidized. In one embodiment, cell lysis is carried out in the dark, or at least in the dark. In one embodiment, cell lysis is carried out in the presence of a reducing agent. Examples of suitable reducing agents include, but are not limited to, β-mercaptoethanol, dithiothreitol (DTT), tris(2-carboxyethyl)phosphine (TCEP), cysteine, N-acetylcysteine, cysteamine, sodium 2-mercaptoethanesulfonate, dithioerythritol (DTE), and bis(2-mercaptoethyl)sulfone.

[0224] RNA can be extracted by techniques well known to those skilled in the art. These techniques include, but are not limited to, chloroform-isoamyl alcohol extraction, phenol-chloroform extraction, alkaline extraction, guanidine thiocyanate-phenol-chloroform extraction, binding to anion exchange resins, silica matrices, glass particles, diatomaceous earth, magnetic particles made from various synthetic polymers, biopolymers, porous glass, and inorganic magnetic materials.

[0225] Preferably, RNA is extracted by chloroform-isoamyl alcohol extraction using, for example, chloroform:isoamyl alcohol 24:1.

[0226] In one embodiment, the extracted RNA is further precipitated. Precipitating the RNA can be performed using techniques well known to those skilled in the art. These techniques include isopropanol-ethanol precipitation, the triazole method (Chomczynski, 1993. Biotechniques. 15(3): 532-4, 536-7), and the pine tree method (Chang et al., 1993. Plant Mol Biol Report. 11(2): 113-116).

[0227] Preferably, precipitation of RNA is performed by isopropanol-ethanol precipitation.

[0228] Thiol-labeled RNA is light-sensitive and easily oxidized. In one embodiment, RNA extraction is carried out in the dark, or at least in the absence of light. In one embodiment, RNA extraction is carried out in the presence of a reducing agent. Examples of suitable reducing agents include, but are not limited to, β-mercaptoethanol, dithiothreitol (DTT), tris(2-carboxyethyl)phosphine (TCEP), cysteine, N-acetylcysteine, cysteamine, sodium 2-mercaptoethanesulfonate, dithioerythritol (DTE), and bis(2-mercaptoethyl)sulfone.

[0229] In one embodiment, the labeled RNA undergoes nucleotide substitutions.In one embodiment, the labeled RNA is placed under conditions that allow for nucleotide substitutions.

[0230] Hereinafter, the terms "substitution," "conversion," and "transformation" are used interchangeably to refer to the incorporation of a mismatched nucleotide.

[0231] Nucleotide substitution in labeled RNA can be achieved by chemically modifying the labeled RNA and further reverse transcribing the chemically modified labeled RNA. Therefore, conditions that allow nucleotide substitution include chemical modification of the labeled RNA; and reverse transcribing the chemically modified labeled RNA.

[0232] Preferably, the nucleotide conversion method allows for altered Watson-Crick base pairing in the labeled RNA and results in reverse transcription of the labeled RNA incorporating mismatched nucleotides during cDNA synthesis, ie, pairing of the labeled nucleotide with a non-Watson-Crick nucleotide.

[0233] For example, during first-strand cDNA synthesis, labeled uridine (e.g., thiol-labeled uridine) can pair with guanosine (G) instead of adenosine (A). Consequently, cytidine will be incorporated during second-strand synthesis, ultimately resulting in a thymidine (T) to cytidine (C) substitution relative to the original nucleic acid sequence.

[0234] Thus, a nucleotide substitution can be defined as the equivalent first-strand synthesis nucleotide substitution (i.e., a nucleotide substitution that occurs during first-strand synthesis); or as the equivalent second-strand synthesis nucleotide substitution (i.e., a nucleotide substitution that occurs during second-strand synthesis).

[0235] In one embodiment, the labeled RNA undergoes A to G (A→G) substitutions during first strand synthesis. In one embodiment, the labeled RNA undergoes T to C (T→C) substitutions during second strand synthesis. Thus, in these embodiments, labeled uridine (U) in the labeled RNA is converted to cytidine (C) rather than thymidine (T) in the corresponding cDNA.

[0236] Unless explicitly stated otherwise, nucleotide substitutions described herein correspond to second strand synthesis nucleotide substitutions.

[0237] Suitable chemical modifications of the labeled RNA include, but are not limited to, alkylation, oxidation-nucleophilic-aromatic substitution, osmium-mediated conversion, or any other method known to those skilled in the art.

[0238] Alkylation of the labeled RNA can be performed by techniques well known to those skilled in the art, including but not limited to those described by Herzog et al. (2017. Nat Methods. 14(12): 1198-1204).

[0239] Preferably, alkylation of the labeled RNA is performed after extraction of the RNA, as described above.

[0240] In one embodiment, alkylation of the labeled RNA is performed using an alkylating agent. Examples of suitable alkylating agents include, but are not limited to, iodoacetamide, iodoacetic acid, N-ethylmaleimide, and 4-vinylpyridine.

[0241] In a preferred embodiment, the alkylating agent is iodoacetamide.

[0242] Non-limiting examples of alkylation treatments of labeled RNA include adding to the labeled RNA:

[0243] - a 100% ethanol solution of about 1 mM final to about 20 mM final, preferably about 5 mM final to about 15 mM final, preferably about 10 mM final iodoacetamide,

[0244] - a buffer solution of pH 8.0 (e.g. sodium phosphate (NaPO4) buffer) of about 10 mM final to about 100 mM final, preferably about 25 mM final to about 75 mM final, preferably about 50 mM final,

[0245] - about 25 v / v% to about 75 v / v%, preferably about 40 v / v% to about 60 v / v%, preferably about 50 v / v% DMSO.

[0246] Thiol-labeled RNA is light-sensitive, and in one embodiment, RNA alkylation is performed in the dark, or at least protected from light.

[0247] In one embodiment, RNA alkylation is not performed in the presence of a reducing agent.

[0248] In one embodiment, RNA alkylation is quenched, ie, stopped at the end of the alkylation treatment.

[0249] The alkylation treatment quenching can be performed by techniques well known to those skilled in the art.

[0250] In one embodiment, RNA alkylation quenching is performed using a reducing agent.

[0251] Examples of suitable reducing agents include, but are not limited to, β-mercaptoethanol, dithiothreitol (DTT), tris(2-carboxyethyl)phosphine (TCEP), cysteine, N-acetylcysteine, cysteamine, sodium 2-mercaptoethanesulfonate, dithioerythritol (DTE), and bis(2-mercaptoethyl)sulfone.

[0252] A non-limiting example of an RNA alkylation quencher includes adding about 1 mM final to about 100 mM final, preferably about 10 mM final to about 50 mM final, preferably about 10 mM final to about 30 mM final, preferably about 20 mM final dithiothreitol (DTT) to the alkylated RNA.

[0253] Oxidation-nucleophilic-aromatic substitution of labeled RNA can be performed by techniques well known to those skilled in the art. Such techniques include, but are not limited to, those described by Schofield et al. (2018. Nat Methods. 15(3): 221-225).

[0254] Preferably, the oxidative-nucleophilic-aromatic substitution of the labeled RNA is performed after RNA extraction, as described above.

[0255] In one embodiment, an oxidative-nucleophilic-aromatic substitution of the labeled RNA is performed using an oxidizing agent and a nucleophilic agent.

[0256] Examples of suitable oxidizing agents include, but are not limited to, sodium periodate (NaIO 4 ), m-chloroperbenzoic acid (mCPBA), sodium iodate (NaIO 3 ), and hydrogen peroxide (H 2 O 2 ).

[0257] In a preferred embodiment, the alkylating agent is sodium periodate (NaIO4).

[0258] Examples of suitable nucleophiles include, but are not limited to, 2,2,2-trifluoroethylamine (TFEA), hydrazine, benzylamine, ammonia, methoxyamine, 1,1-dimethylethylenediamine, aniline, and 4-(trifluoromethyl)benzylamine.

[0259] In a preferred embodiment, the nucleophile is 2,2,2-trifluoroethylamine (TFEA).

[0260] Thiol-labeled RNA is light-sensitive. In one embodiment, the RNA oxidation-nucleophilic-aromatic substitution is performed in the dark, or at least protected from light.

[0261] In one embodiment, the RNA oxidation-nucleophilic-aromatic substitution is not performed in the presence of a reducing agent.

[0262] In one embodiment, the RNA oxidative-nucleophilic-aromatic substitution is quenched, ie, stopped at the end of the oxidative-nucleophilic-aromatic substitution process.

[0263] The quenching oxidation-nucleophilic-aromatic substitution treatment can be performed by techniques well known to those skilled in the art.

[0264] Osmium-mediated conversion of labeled RNA can be performed by techniques well known to those skilled in the art. Such techniques include, but are not limited to, those described by Riml et al. (2017. Angew Chem Int Ed Engl. 56(43): 13479-13483).

[0265] Preferably, following RNA extraction, osmium-mediated conversion of the labeled RNA is performed as described above.

[0266] In one embodiment, osmium tetroxide (OsO4) and ammonia are used for osmium-mediated conversion of labeled RNA.

[0267] Thiol-labeled RNA is light-sensitive. In one embodiment, the RNA oxidation-nucleophilic-aromatic substitution is performed in the dark, or at least protected from light.

[0268] In one embodiment, the osmium-mediated conversion of RNA is not performed in the presence of a reducing agent.

[0269] In one embodiment, the osmium-mediated conversion of RNA is quenched, ie, stopped at the end of the oxidative-nucleophilic-aromatic substitution process.

[0270] Quenching of the osmium-mediated conversion process can be performed by techniques well known to those skilled in the art.

[0271] The generation of cDNA libraries, in particular for sequencing purposes, is part of the knowledge of those skilled in the art. Kits for cDNA library generation are commercially available and include, but are not limited to, the SMARTerStranged Total RNA Sequencing Kit (ClonTech), the QuantSeq 3' mRNA Sequencing Library Preparation Kit (Lexogen), the NextEra XT DNA Library Preparation Kit (Illumina), the TruSeq Nano DNA Library Preparation Kit (Illumina), the NEBNext DNA Library Preparation Master Mix (New England Biolabs), the NEBNext UltraDNA Library Preparation Kit (New England Biolabs), and the JetSeq DNA Library Preparation Kit (Bioline).

[0272] In one embodiment, generating a cDNA library includes some or all of the following substeps:

[0273] -RNA reverse transcription, including:

[0274] ○ First-strand cDNA synthesis (to obtain a double-stranded mixed RNA-cDNA library),

[0275] o Optionally, RNA template removal (to obtain a single-stranded cDNA library),

[0276] o Second strand cDNA synthesis (thereby obtaining a double stranded cDNA library), and

[0277] -Optional, double-stranded cDNA library purification.

[0278] Reverse transcription of RNA is achieved by techniques well known to those skilled in the art using a reverse transcriptase and a mixture of four deoxyribonucleotides (dNTPs), namely, deoxyadenosine triphosphate (dATP), deoxycytidine triphosphate (dCTP), deoxyguanosine triphosphate (dGTP), and (deoxy)thymidine triphosphate (dTTP).

[0279] In one embodiment, the first-strand cDNA synthesis reaction uses a sequence-specific primer. In one embodiment, the first-strand cDNA synthesis reaction uses a sequence-specific primer. In one embodiment, the first-strand cDNA synthesis reaction uses a random primer.

[0280] In one embodiment, the primers used for first-strand cDNA synthesis include a fixed nucleic acid sequence (including, for example, a linker and / or index for sequencing) and a primed nucleic acid sequence (complementary to the RNA template). In one embodiment, the primers used for synthesizing the first-strand cDNA include a fixed 5' end sequence and a primed 3' end sequence. In one embodiment, the primers used for synthesizing the first-strand cDNA include a fixed 3' end sequence and a primed 5' end sequence.

[0281] In particular, methods for removing RNA templates are well known to those skilled in the art. Removal of RNA templates can be achieved, for example, by incubating the double-stranded mixed RNA-cDNA library with RNase H.

[0282] RNA reverse transcription for generating cDNA libraries can be performed in a random manner, i.e., using random primers, thereby reverse transcribing the entire or most RNA. Alternatively, RNA reverse transcription for generating cDNA libraries can be performed in a targeted manner, i.e., using specific primers, thereby only generating gene libraries with customized sequences.

[0283] In one embodiment, cDNA library is produced, particularly reverse transcription RNA library causes nucleotide replacement.Under the situation that there is no chemical modification, such nucleotide replacement all is random generation on any reverse transcribed RNA.Yet, in the reverse transcription process of the RNA of previous mark, observed the surge of this replacement, and further carried out chemical modification by technology such as alkylation, oxidation-nucleophilic-aromatic substitution, osmium-mediated conversion as mentioned above.The surge of replacement is described in " embodiment " part below.

[0284] Amplification of the cDNA library can be performed by methods well known to those skilled in the art.

[0285] Amplification of a cDNA library can be performed in a random manner, i.e. using random primers, thereby amplifying the entire or a large portion of the cDNA library. Alternatively, amplification of a cDNA library can be performed in a targeted manner, i.e. using specific primers, thereby amplifying only individual sequences in the cDNA library.

[0286] The cDNA library can be sequenced by methods well known to those skilled in the art. In one embodiment, the cDNA library is sequenced by next generation sequencing (NGS), deep sequencing or targeted sequencing of personalized sequences.

[0287] Methods for NGS are well known to those skilled in the art and include, but are not limited to, paired-end sequencing, sequencing by synthesis, and single-read sequencing.

[0288] Platforms that can be used for NGS include, but are not limited to, Illumina MiSeq (Illumina), Ion Torrent PGM (ThermoFisher Science), PacBio RS (PacBio), Illumina GAIIx (Illumina), and Illumina HiSeq 2000 (Illumina).

[0289] The step of sequencing the cDNA library can be performed using a commercially available kit such as the MiSeq kit v2 (Illumina).

[0290] In one embodiment, sequencing a cDNA library generates a set of sequence reads.

[0291] In one embodiment, the method according to the invention comprises the step of comparing the number of substituted nucleotides in a first set of sequences matching at least one microbial, preferably viral, bacterial, archaeal, fungal or protozoan nucleic acid sequence hit with a reference sequence.

[0292] A "substituted nucleotide" is a nucleotide that is substituted with another nucleotide relative to a nucleic acid sequence hit of a microorganism, preferably a virus, bacterium, archaea, fungus, or protozoa. Generally, any nucleotide can be substituted with any other nucleotide, for example, thymidine can be substituted with cytidine (T→C), adenosine (T→A), or guanosine (T→G). The same applies to substitutions of adenosine (A), cytidine (C), and guanosine (G) with any of the other three nucleotides.

[0293] Such substitutions occur randomly in small amounts, particularly during reverse transcription. However, the present invention is based on the surge in such substitutions when RNA is pre-labeled and further subjected to chemical modification methods such as alkylation, oxidative-nucleophilic-aromatic substitution, osmium-mediated conversion, etc.

[0294] In one embodiment, the total number of substituted nucleotides in a first set of sequence reads that match at least one microbial, preferably viral, bacterial, archaeal, fungal, or protozoan nucleic acid sequence hit is compared to the total number of substituted nucleotides in a reference sequence.

[0295] In one embodiment, the number of T→C substitutions in a first set of sequence reads that match at least one microbial, preferably viral, bacterial, archaeal, fungal, or protozoan nucleic acid sequence hit is compared to the number of T→C substitutions in a reference sequence.

[0296] In one embodiment, the nucleotide substitution rate in a first set of sequence reads matching at least one microbial, preferably viral, bacterial, archaeal, fungal or protozoan nucleic acid sequence hit is compared to the nucleotide substitution rate in a reference sequence.

[0297] In one embodiment, the T→C substitution rate in a first set of sequence reads matching at least one microbial, preferably viral, bacterial, archaeal, fungal or protozoan nucleic acid sequence hit is compared to the T→C substitution rate in a reference sequence.

[0298] As used herein, the "substitution rate" is calculated by dividing the number of one or more given nucleotide substitutions (e.g., T→C or any other nucleotide substitution defined above) by the total number of substitutions. Alternatively, the "substitution rate" can be calculated by dividing the number of one or more given nucleotide substitutions by the total number of nucleotides in the sequence reads that match at least one microbial, preferably viral, bacterial, archaeal, fungal, or protozoan nucleic acid sequence hit.

[0299] In one embodiment, the ratio of the T→C substitution rate between the sequence reads in the first set of sequence reads and the reference sequence that match at least one microbial, preferably viral, bacterial, archaeal, fungal or protozoan nucleic acid sequence hit, and the ratio of the average substitution rate of all other nucleotides (i.e., all nucleotides except T→C) between the sequence reads in the first set of sequence reads and the reference sequence that match at least one microbial, preferably viral, bacterial, archaeal, fungal or protozoan nucleic acid sequence hit are compared.

[0300] In one embodiment, the method comprises identifying at least one microbial, preferably viral, bacterial, archaeal, fungal, or protozoan nucleic acid sequence hit that matches at least one sequence read.

[0301] In one embodiment, identifying at least one microorganism, preferably a virus, bacteria, archaea, fungus or protozoa, nucleic acid sequence hit that matches at least one sequence read comprises the substeps of filtering the read set, assembling the sequence reads into contigs, aligning the sequence reads or contigs with a database, identifying at least one microorganism, preferably a virus, bacteria, archaea, fungus or protozoa, nucleic acid sequence hit that matches at least one sequence read or contig, and realigning the sequence reads or contigs with at least one microorganism, preferably a virus, bacteria, archaea, fungus or protozoa, nucleic acid sequence hit.

[0302] Filtering a collection of sequence reads is part of the knowledge of a person skilled in the art.

[0303] In one embodiment, filtering the sequence read set may include, but is not limited to, suppressing sequence read duplications, suppressing low-quality sequence reads, suppressing sequence read homopolymers, removing fixed nucleic acid sequences from sequence reads (e.g., adapters and / or library tags used for sequencing), discarding endogenous sequence reads (i.e., sequence reads that match nucleic acid sequences belonging to the subject cell), discarding unnecessary sequence reads (e.g., rRNA sequence reads, etc.), etc.

[0304] Such filtering can be performed using software readily available to those skilled in the art.

[0305] Assembling sets of sequence reads into contigs is part of the knowledge of a person skilled in the art.

[0306] Sequence reads can be assembled into contigs using software readily available to those skilled in the art.

[0307] Optionally, the sequence reads or contigs can be translated into amino acid sequences.

[0308] Aligning sets of sequence reads or contigs is part of the knowledge of a person skilled in the art.Such alignment of sequence reads or contigs can be performed using software readily available to a person skilled in the art.

[0309] In one embodiment, the sequence reads or contigs are aligned in a microbial database, i.e., a database comprising nucleic acid sequences or amino acid sequences of microorganisms, preferably viruses, bacteria, archaea, fungi or protozoa (where the sequence reads or contigs are translated into amino acid sequences). Such databases can be downloaded, for example, from the EMBL nucleotide sequence database.

[0310] It is part of the knowledge of a person skilled in the art to identify a nucleic acid sequence hit matching at least one microorganism, preferably a virus, bacterium, archaea, fungus or protozoa, to at least one sequence read or contig (or an amino acid sequence hit if the sequence read or contig is translated into an amino acid sequence).

[0311] After aligning a collection of sequence reads or contigs to a database, hit sequences from the database can be identified.

[0312] In one embodiment, at least one hit sequence is identified (and thus selected) based on a threshold expectation value (e-value) obtained when aligned to the sequence reads or contigs. In one embodiment, a hit sequence is selected if the e-value obtained when aligning the sequence hit to the at least one sequence read or contig is less than 10 -2 , preferably less than 5×10 -3 , preferably less than 10 -3 , then the sequence hit is identified (and thus selected).

[0313] It is part of the knowledge of a person skilled in the art to realign a sequence read or contig with previously identified (and therefore selected) nucleic acid sequence hits (or amino acid sequence hits in the case where the sequence read or contig is translated into amino acid sequence) to at least one microorganism, preferably a virus, bacterium, archaea, fungus or protozoa.

[0314] This realignment of sequence reads or contigs can be performed using software readily available to those skilled in the art.

[0315] In one embodiment, upon realignment, at least one final consensus sequence (or amino acid sequence hit if the sequence reads or contigs are translated into amino acid sequences) of at least one previously identified (and therefore selected) microbial, preferably viral, bacterium, archaea, fungal or protozoan nucleic acid sequence hit is determined.

[0316] In one embodiment, the control sequence is selected from:

[0317] a second set of sequence reads matching said at least one microbial, preferably viral, bacterial, archaeal, fungal or protozoan, nucleic acid sequence hits, wherein said second set of sequence reads is obtained by sequencing a second set of RNA obtained by culturing a sample, preferably a cell sample, in the absence of an RNA labeling agent;

[0318] a second set of sequence reads matching said at least one microbial, preferably viral, bacterial, archaeal, fungal or protozoan, nucleic acid sequence hit, wherein said second set of sequence reads is obtained by sequencing a second set of RNA obtained by culturing a sample, preferably a cell sample, in the presence of an RNA labeling agent without subjecting the extracted RNA to conditions that allow nucleotide substitutions;

[0319] a consensus microbial, preferably viral, bacterial, archaeal, fungal or protozoan nucleic acid sequence obtained from sequence reads or contigs of a first set of sequences matching at least one microbial, preferably viral, bacterial, archaeal, fungal or protozoan nucleic acid sequence hit;

[0320] a sequence that corresponds to a nucleic acid sequence hit of the same microorganism, preferably a virus, bacterium, archaea, fungus or protozoa, found in the closest microorganism, preferably a virus, bacterium, archaea, fungus or protozoa strain identified in a nucleic acid sequence database; and / or

[0321] - a similar sequence corresponding to a nucleic acid sequence hit of the same microorganism, preferably a virus, bacterium, archaea, fungus or protozoa, identified in a nucleic acid sequence database.

[0322] In one embodiment, the control sequence is a second set of sequence reads that match the nucleic acid sequence hits of the at least one microorganism, preferably a virus, bacteria, archaea, fungus or protozoa, wherein the second set of sequence reads is obtained by sequencing a second set of RNA obtained by culturing a sample, preferably a cell sample, in the absence of an RNA labeling agent;

[0323] In this embodiment, the method according to the invention comprises the following steps:

[0324] (a) sequencing a first RNA and a second set of RNA extracted from a sample, preferably a cell sample,

[0325] wherein the first set of RNA is obtained by culturing cells, preferably a cell sample, in the presence of an RNA labeling agent, and the second set of RNA is obtained by culturing cells, preferably a cell sample, in the absence of an RNA labeling agent,

[0326] Thus, a first sequence read segment and a second set of sequence read segments are obtained,

[0327] (b) comparing the number of substituted nucleotides in a first set of sequence reads that match at least one microbial, preferably viral, bacterial, archaeal, fungal, or protozoan nucleic acid sequence hit and the number of substituted nucleotides in a second set of sequence reads that match at least one microbial, preferably viral, bacterial, archaeal, fungal, or protozoan nucleic acid sequence hit, and

[0328] (c) if the number of substituted nucleotides in the sequence reads that match the at least one microbial, preferably viral, bacterium, archaea, fungus or protozoan nucleic acid sequence hit in the first set of sequence reads is greater than the number of substituted nucleotides in the second sequence reads, then it is inferred that the at least one microbial, preferably viral, bacterium, archaea, fungus or protozoan nucleic acid sequence hit belongs to a living microorganism, preferably a virus, bacterium, archaea, fungus or protozoan.

[0329] Preferably, the first set of RNA is obtained by culturing a sample, preferably a cell sample, in the presence of an RNA labeling agent, thereby obtaining labeled RNA, and further subjecting said labeled RNA to conditions that allow nucleotide substitutions.

[0330] In one embodiment, the method according to the present invention comprises the following steps:

[0331] (a) sequencing a first set of RNA and a second set of RNA extracted from a sample, preferably a cell sample,

[0332] wherein the first set of RNA is obtained by culturing cells, preferably a cell sample, in the presence of an RNA labeling agent, and the second set of RNA is obtained by culturing cells, preferably a cell sample, in the absence of an RNA labeling agent,

[0333] Thus, a first set of sequence reads and a second set of sequence reads are obtained,

[0334] (b) identifying at least one microbial, preferably viral, bacterial, archaeal, fungal or protozoan, nucleic acid sequence hit that matches at least one sequence read of the first set of sequence reads,

[0335] (c) comparing the number of substituted nucleotides in sequence reads of the first and second sets of sequence reads that match at least one identified microbial, preferably viral, bacterial, archaeal, fungal or protozoan, nucleic acid sequence hit, and

[0336] (d) if the number of substituted nucleotides in the sequence reads that match at least one identified microbial, preferably viral, bacterial, archaeal, fungal or protozoan nucleic acid sequence hit in the first set of sequence reads is greater than the number of substituted nucleotides in the second sequence reads, then it is inferred that the at least one microbial, preferably viral, bacterial, archaeal, fungal or protozoan nucleic acid sequence hit belongs to a living, active microbial, preferably viral, bacterial, archaeal, fungal or protozoan.

[0337] In this embodiment, the method comprises the step of sequencing a first set of RNA extracted from a sample, preferably a cell sample.In this embodiment, the sample, preferably a cell sample, is cultured in the presence of an RNA labeling agent.

[0338] In this embodiment, the step of sequencing a first set of RNA extracted from a sample, preferably a cell sample, includes one or more than one or all of the sub-steps of labeling the RNA, lysing the cells, extracting the RNA, replacing nucleotides in the labeled RNA, generating a cDNA library, amplifying the cDNA library, and sequencing the cDNA library.

[0339] These substeps are defined and described in detail above and applied to the sequencing of the first set of RNAs.

[0340] In this embodiment, the further method comprises the step of sequencing a second set of RNA extracted from the sample, preferably a cell sample.In this embodiment, the sample, preferably a cell sample, is cultured in the absence of an RNA labeling agent.

[0341] In this embodiment, the step of sequencing the second set of RNA extracted from the sample, preferably a cell sample, comprises one or more than one or all of the sub-steps of lysing cells, extracting RNA, generating a cDNA library, amplifying the cDNA library and sequencing the cDNA library.

[0342] These substeps are defined and described in detail above and applied to the sequencing of the second set of RNA.

[0343] In one embodiment, the control sequence is a second set of sequence reads that match the nucleic acid sequence hits of the at least one microorganism, preferably a virus, bacteria, archaea, fungus or protozoa, wherein the second set of sequence reads is obtained by sequencing a second set of RNA, and the second set of RNA is obtained by culturing the sample, preferably a cell sample, in the presence of an RNA labeling agent, but without subjecting the extracted RNA to conditions that allow nucleotide substitutions.

[0344] In this embodiment, the method according to the invention comprises the following steps:

[0345] (a) sequencing a first set of RNA and a second set of RNA extracted from a sample, preferably a cell sample,

[0346] wherein the first set of RNAs and the second set of RNAs are obtained by culturing a sample, preferably a cell sample, in the presence of an RNA labeling agent, thereby obtaining labeled RNA, and

[0347] wherein the first group of RNAs is obtained from a first portion of labeled RNA that has undergone nucleotide substitutions, and the second group of RNAs is obtained from a second portion of labeled RNA that has not undergone nucleotide substitutions,

[0348] Thus, a first set of sequence reads and a second set of sequence reads are obtained,

[0349] (b) comparing the number of substituted nucleotides in a first set of sequence reads that match at least one microbial, preferably viral, bacterial, archaeal, fungal or protozoan nucleic acid sequence with the number of substituted nucleotides in a second set of sequence reads that match at least one microbial, preferably viral, bacterial, archaeal, fungal or protozoan nucleic acid sequence hit, and

[0350] (c) if the number of substituted nucleotides in the sequence reads that match the at least one microbial, preferably viral, bacterium, archaea, fungus or protozoan nucleic acid sequence hit in the first set of sequence reads is greater than the number of substituted nucleotides in the second sequence reads, then it is inferred that the at least one microbial, preferably viral, bacterium, archaea, fungus or protozoan nucleic acid sequence hit belongs to a living microorganism, preferably a virus, bacterium, archaea, fungus or protozoan.

[0351] In one embodiment, the reference sequence can be a consensus microbial, preferably viral, bacterial, archaeal, fungal, or protozoan nucleic acid sequence. In one embodiment, the consensus microbial, preferably viral, bacterial, archaeal, fungal, or protozoan nucleic acid sequence can be obtained from multiple sequence reads of a first set of sequence reads that match at least one microbial, preferably viral, bacterial, archaeal, fungal, or protozoan nucleic acid sequence hit. Since it has been observed that not all target nucleotides are thio-labeled and / or substituted during nucleotide substitution, such a consensus sequence can be readily determined. In practice, according to the methods of the present invention, a sufficient number of target nucleotides are substituted to allow differentiation; however, this number remains sufficiently low to establish a consensus sequence.

[0352] In one embodiment, the control sequence can be a nucleic acid sequence hit corresponding to the same microorganism, preferably a virus, bacteria, archaea, fungus or protozoa nucleic acid sequence found in the closest microorganism, preferably a virus, bacteria, archaea, fungus or protozoa strain identified in a nucleic acid sequence database.

[0353] In one embodiment, the control sequence can be a similar sequence corresponding to a nucleic acid sequence hit of the same microorganism, preferably a virus, bacteria, archaea, fungus or protozoa, identified in a nucleic acid sequence database.

[0354] In one embodiment, the method according to the invention comprises the step of inferring whether at least one microbial, preferably viral, bacterial, archaeal, fungal or protozoan nucleic acid sequence hit belongs to a living microorganism, preferably a virus, bacterium, archaea, fungus or protozoan.

[0355] In one embodiment, the living microorganism, preferably a virus, bacterium, archaea, fungus or protozoa, is characterized by the taxonomic assignment of at least one microorganism, preferably a virus, bacterium, archaea, fungus or protozoa nucleic acid sequence hit.

[0356] In one embodiment, if the total number of nucleotide substitutions in the sequence reads that match at least one microbial, preferably viral, bacterial, archaeon, fungal or protozoan nucleic acid sequence hit in the first set of sequence reads is greater than the total number of nucleotide substitutions randomly substituted in the reference sequence, then it is inferred that the at least one microbial, preferably viral, bacterial, archaeon, fungal or protozoan nucleic acid sequence hit belongs to a living microbial, preferably viral, bacterial, archaeon, fungal or protozoan.

[0357] In one embodiment, if the number of T→C substitutions in the sequence reads that match at least one microbial, preferably viral, bacterial, archaeon, fungal or protozoan nucleic acid sequence hit in the first set of sequence reads is greater than the number of T→C substitutions randomly substituted in the reference sequence, then it is inferred that the at least one microbial, preferably viral, bacterial, archaeon, fungal or protozoan nucleic acid sequence hit belongs to a living microbial, preferably viral, bacterial, archaeon, fungal or protozoan.

[0358] In one embodiment, "the number of substitutions [...] in the first set of sequence reads is greater than the number of substitutions [...] in the reference sequence" when the number of substitutions in the sequence reads that match at least one microbial, preferably viral, bacterial, archaeal, fungal or protozoan nucleic acid sequence hit in the first set of sequence reads is two times, preferably three times, more preferably four times, five times, six times, seven times, eight times, nine times, ten times, fifteen times, twenty times, fifty times, or even one hundred times greater than the number of substitutions in the reference sequence.

[0359] In one embodiment, if the nucleotide substitution rate in the sequence reads that match at least one microbial, preferably viral, bacterial, archaeon, fungal or protozoan nucleic acid sequence hit in the first set of sequence reads is greater than the nucleotide substitution rate of random substitution in the reference sequence, then it is inferred that the at least one microbial, preferably viral, bacterial, archaeon, fungal or protozoan nucleic acid sequence hit belongs to a living microbial, preferably viral, bacterial, archaeon, fungal or protozoan.

[0360] In one embodiment, if the T→C substitution rate in the sequence reads that match at least one microbial, preferably viral, bacterial, archaeon, fungal or protozoan nucleic acid sequence hit in the first set of sequence reads is greater than the T→C substitution rate of random substitution in the reference sequence, then it is inferred that the at least one microbial, preferably viral, bacterial, archaeon, fungal or protozoan nucleic acid sequence hit belongs to a living microbial, preferably viral, bacterial, archaeon, fungal or protozoan.

[0361] As used herein, the term "T→C substitution rate" can be defined by the following formula:

[0362]

[0363] In one embodiment, "the substitution rate [...] in the first set of sequence reads is greater than the substitution rate [...] in the reference sequence" when the substitution rate of sequence reads matching at least one microbial, preferably viral, bacterial, archaeal, fungal or protozoan nucleic acid sequence hit in the first set of sequence reads is two times, preferably three times, more preferably four times, five times, six times, seven times, eight times, nine times, ten times, fifteen times, twenty times, fifty times, or even one hundred times greater than the substitution rate of the reference sequence.

[0364] In one embodiment, if the T→C substitution rate is greater than the average substitution rate of all other nucleotides in the sequence reads that match at least one microbial, preferably viral, bacterial, archaeon, fungal or protozoan nucleic acid sequence hit in the first set of sequence reads, then it is inferred that the at least one microbial, preferably viral, bacterial, archaeon, fungal or protozoan nucleic acid sequence hit belongs to a living microorganism, preferably a virus, bacterium, archaeon, fungus or protozoan.

[0365] In one embodiment, the "T→C substitution rate is greater than the average substitution rate of all other nucleotides" when the T→C substitution rate is 2 times, preferably 3 times, more preferably 4 times, 5 times, 6 times, 7 times, 8 times, 9 times, 10 times, 15 times, 20 times, 50 times, or 100 times the average substitution rate of all other nucleotides.

[0366] The “average substitution rate for all other nucleotides” refers to the average of the A→C, A→G, A→T, C→A, C→G, C→T, T→A, T→G, G→A, G→C, and G→T substitution rates.

[0367] In one embodiment, if the T→C substitution rate is greater than the average T→A and T→G substitution rates in the sequence reads that match at least one microbial, preferably viral, bacterial, archaeon, fungal or protozoan nucleic acid sequence hit in the first set of sequence reads, it is inferred that the at least one microbial, preferably viral, bacterial, archaeon, fungal or protozoan nucleic acid sequence hit belongs to a living microorganism, preferably a virus, bacterium, archaeon, fungus or protozoan.

[0368] In one embodiment, the "T→C substitution rate is greater than the average substitution rate of T→A and T→G" when the T→C substitution rate is 2 times, preferably 3 times, more preferably 4 times, 5 times, 6 times, 7 times, 8 times, 9 times, 10 times, 15 times, 20 times, 50 times, or 100 times greater than the average substitution rate of all other nucleotides. In one embodiment, if the ratio of the T→C substitution rates between sequence reads matching at least one microbial, preferably viral, bacterial, archaeal, fungal, or protozoan nucleic acid sequence hit in the first set of sequence reads and in the reference sequence is greater than the ratio of the average substitution rates of all other nucleotides between sequence reads matching at least one microbial, preferably viral, bacterial, archaeal, fungal, or protozoan nucleic acid sequence hit in the first set of sequence reads and in the reference sequence, then the at least one microbial, preferably viral, bacterial, archaeal, fungal, or protozoan nucleic acid sequence hit is inferred to belong to a viable microorganism, preferably a virus, bacteria, archaeal, fungal, or protozoan.

[0369] In one embodiment, if the ratio of the T→C substitution rates between the sequence reads matching at least one microorganism, preferably a virus, bacteria, archaea, fungus or protozoa nucleic acid sequence hit in the first set of sequence reads and in the reference sequence is greater than the ratio of the average T→A and T→G substitution rates between the sequence reads matching at least one microorganism, preferably a virus, bacteria, archaea, fungus or protozoa nucleic acid sequence hit in the first set of sequence reads and in the reference sequence, then it is inferred that the at least one microorganism, preferably a virus, bacteria, archaea, fungus or protozoa nucleic acid sequence hit belongs to a living microorganism, preferably a virus, bacteria, archaea, fungus or protozoa.

[0370] In one embodiment, if the T→C substitution index in the sequence reads that match at least one microorganism, preferably a virus, bacteria, archaea, fungus, or protozoa nucleic acid sequence hit in the first set of sequence reads is greater than a threshold, then the at least one microorganism, preferably a virus, bacteria, archaea, fungus, or protozoa nucleic acid sequence hit is inferred to belong to a living microorganism, preferably a virus, bacteria, archaea, fungus, or protozoa. In this embodiment, the threshold value can be determined experimentally. In one embodiment, the threshold value is greater than the T→C substitution index in the sequence reads that match at least one microorganism, preferably a virus, bacteria, archaea, fungus, or protozoa nucleic acid sequence hit in the second set of sequence reads. In one embodiment, the threshold value is at least 2, preferably at least 2.5, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more than 20.

[0371] As used herein, the term "T→C substitution index" can be defined by the following formula:

[0372]

[0373] According to the present invention, the method for distinguishing between living and dead microorganisms, preferably viruses, bacteria, archaea, fungi or protozoa, in a sample, preferably a cell sample, comprises distinguishing between nucleic acid sequences of transcriptionally active microorganisms and inert microorganisms, preferably viruses, bacteria, archaea, fungi or protozoa, in a sample, preferably a cell sample, and the method is useful in many different applications.

[0374] Indeed, the risk of contamination by microorganisms, particularly viruses, bacteria, archaea, fungi, or protozoa, is a major concern for biological products. This includes the risk of contamination of both Good Manufacturing Practice (GMP) facilities and the final drug product. Viral testing of raw materials, cells, viral seeds, master / working libraries, vaccine serum batches, and more is crucial for drug safety. This is particularly critical for live vaccines, gene therapy viral vectors, and cell therapy drug products, as their production does not include downstream viral elimination steps. Therefore, the safety of these products relies heavily on in-process viral testing.

[0375] All previously reported contamination of cell culture-based products has been due to unpredictable animal viruses that were not detected during virus testing of the raw materials or production cells. In fact, routine virus testing is limited because many viruses will not grow in the cell lines used for in vitro testing or in rodents or eggs used for in vivo testing.

[0376] The methods of the present invention provide another way to accurately test sample contamination and distinguish harmless contamination from live microorganisms, including latent microorganisms, and inert microbial nucleic acid fragments (e.g., microbial nucleic acid fragments after gamma-irradiation inactivation).

[0377] On this basis, a large number of industrial applications are foreseeable.

[0378] In the field of vaccines, inactivated vaccines are currently usually controlled by culturing the vaccine, which is believed to be inactivated, and then investigating whether live microorganisms are present. The method of the present invention will allow live microorganisms to be distinguished from the background noise of inactive microbial nucleic acid sequences, which are inactivated and therefore harmless.

[0379] Similarly, the method according to the present invention can be easily implemented to detect contamination with live microorganisms in biological samples, but can also be detected in blood cultures and other types of biological samples used for diagnosis, such as raw materials (e.g., serum batches in the case of vaccines), cells, master / working libraries, etc. For example, after a subject has been treated with antibiotics, it is possible to consider the presence of remaining live microorganisms in the subject and thus identify latent drug-resistant microorganisms.

[0380] In the field of viral therapy, the method according to the invention can be implemented to test viral vectors (eg viral vectors used for gene therapy) for the presence of replication-competent revertant viruses.

[0381] Preservation media may also be contaminated with microorganisms, and the method according to the invention can be easily used to test for such contamination before contacting the sample for preservation.

[0382] The realm of possibilities also extends to non-biological samples. For example, food safety is a major concern. The rise of health scandals and the demand for food that focuses on quality and safety have resonated with the need to test food for microbial contamination. The method of the present invention can address this issue by providing a practical and definitive answer as to whether a food sample is contaminated with live microorganisms.

[0383] Environmental samples can also be tested. For example, it is known that water and / or air conditioning circuits may carry microorganisms. The method according to the present invention can be implemented to confirm the presence or absence of these living microorganisms.

[0384] Another object of the present invention is a diagnostic method, preferably an in vitro diagnostic method, for diagnosing a microbial infection, preferably a viral, bacterial, archaeal, fungal or protozoal infection, in a subject.

[0385] In one embodiment, the diagnostic method according to the present invention comprises the step of providing a sample, preferably a cell sample, from a subject.

[0386] In one embodiment, the diagnostic method according to the invention further comprises the step of performing any method for distinguishing between nucleic acid sequences of transcriptionally active microorganisms and inert microorganisms, preferably viruses, bacteria, archaea, fungi or protozoa, in a sample, preferably a cell sample, according to the invention.

[0387] In one embodiment, the diagnostic method according to the present invention further comprises the following step: if at least one identified microorganism, preferably a virus, bacterium, archaea, fungus or protozoa nucleic acid sequence hit belongs to a living microorganism, preferably a virus, bacterium, archaea, fungus or protozoa, then the diagnosed subject is infected with a microorganism, preferably a virus, bacterium, archaea, fungus or protozoa.

[0388] Another object of the present invention is a method of treating a microbial, preferably a viral, bacterial, archaeal, fungal or protozoal, infection in a subject.

[0389] In one embodiment, the method for treating a microbial infection according to the present invention comprises the steps of performing a diagnostic method according to the present invention; and treating the subject if the subject is diagnosed as being infected by a microorganism, preferably a virus, bacteria, archaea, fungus or protozoa.

[0390] Means and methods for treating microbial infections are well known to those skilled in the art and include, but are not limited to, administering to a subject at least one antiviral, antibacterial, antifungal, or antiprotozoal drug.

[0391] Suitable examples of antiviral drugs include, but are not limited to, those classified in therapeutic subgroup J05 of the Anatomical Therapeutic Chemical Classification System. Other examples include, but are not limited to, acemannan, acyclovir, acyclovir sodium, amantadine, adefovir, adenine arabinoside, alovudine, avesutol, amantadine hydrochloride, alanodine, aridone, ativiridine mesylate, avridine, cidofovir, sipanthelline, cytarabine hydrochloride, BMS 806, C31G, carrageenan, zinc salts, cellulose sulfate, cyclodextrin, dapivirine, delavirdine mesylate, deciclovir, dextrin 2-sulfate, didanosine, di Thali, dolutegravir, edoxuridine, enviradin, enviradin, etravirine, famciclovir, famotine hydrochloride, fecitabine, fealuridine, foscarnet, sodium fosfoacetate, FTC, ganciclovir, ganciclovir sodium, GSK 1265744, 9-2-hydroxy-ethoxymethylguanine, ibalizumab, idoxuridine, interferon, 5-iodo-2'-deoxyuridine, IQP-0528, ketosar, lamivudine, lobcavir, maraviroc, memotil, pirodavir, penciclovir, raltegravir, ribavirin, rimantadine hydrochloride, rilpivirine (TMC-278), saquinavir mesylate, SCH-C, SCH-D, somantonine hydrochloride, somanton Levudine, vilstorone, stavudine, T20, tilorone hydrochloride, TMC120, TMC125, trifluridine, trifluridine, tenofovir, tenofovir alafenamide, tenofovir disoproxil fumarate, prodrugs of tenofovir, UC-781, UK-427, UK-857, valacyclovir, valacyclovir hydrochloride, vidarabine, vidarabine phosphate, vidarabine sodium phosphate, viroxime, zalcitabine, zidovudine, viroxime, and combinations thereof.

[0392] Suitable examples of antibacterial agents include, but are not limited to, those classified in therapeutic subgroup J01 of the Anatomical Therapeutic Chemical Classification System. Other examples include, but are not limited to, aminoglycosides (e.g., amikacin, gentamicin, kanamycin, neomycin, netilmicin, streptomycin, tobramycin, bacomycin, etc.), ansamycins (e.g., geldanamycin, herbimycin, etc.), carbacephems (e.g., loracarbef, etc.), carbapenems (e.g., ertapenem, doripenem, imipenem, cilastatin, meropenem, etc.), first generation cephalosporins (e.g., Cefuroxime, cefadroxil, cefuroxime, cefazolin, cefalotin, cefalexin, etc.), second-generation cephalosporins (e.g., cefaclor, cefamandole, cefoxitin, cefprozil, cefuroxime, etc.), third-generation cephalosporins (e.g., cefixime, cefdinir, cefditoren, cefoperazone, cefotaxime, cefpodoxime, ceftazidime, ceftibuten, ceftizoxime, ceftriaxone, etc.), fourth-generation cephalosporins (e.g., cefepime, etc.), ), fifth-generation cephalosporins (e.g., ceftobiprole), glycopeptides (e.g., teicoplanin, vancomycin), macrolides (e.g., azithromycin, clarithromycin, dirithromycin, erythromycin, roxithromycin, telithromycin, telithromycin, spectinomycin, etc.), monolactamases (e.g., aztreonam), penicillins (e.g., amoxicillin, ampicillin, azlocillin, carbenicillin, o-cloxacillin, dicloxacillin, flucloxacillin, mezlocillin, methoxythromycin), (e.g., nafcillin, oxacillin, penicillin, piperacillin, ticarcillin, etc.), antimicrobial peptides (e.g., bacitracin, colistin, polymyxin B, etc.), quinolones (e.g., ciprofloxacin, enoxacin, gatifloxacin, levofloxacin, lomefloxacin, moxifloxacin, norfloxacin, ofloxacin, trovafloxacin, etc.), sulfonamides (e.g., mafenide, prontosil, sulfacetamide, sulfamethizole, sulfonamides, sulfasalazine, sulfaiso Trimethoprim, trimethoprim-sulfamethoxazole azoles, tetracyclines (e.g., demeclocycline, doxycycline, minocycline, oxytetracycline, tetracycline, etc.), other antibiotics (e.g., arsphenamine, chloramphenicol, clindamycin, lincomycin, ethambutol, fosfomycin, fusidic acid, furazolidone, isoniazid, linezolid, metronidazole, mupirocin, nitrofurantoin, plate mycin, pyrazinamide, quinupristin / dalfopristin, rifampicin, tinidazole, etc.), and combinations thereof.

[0393] Suitable examples of antifungal drugs include, but are not limited to, those classified in therapeutic subgroup J02 of the Anatomical Therapeutic Chemical Classification System. Other examples include, but are not limited to, abafungin, albaconazole, amorolfine, amphotericin B, anidulafungin, atovaquone, biafungin, bifonazole, bromochlorosalicylate, butenafine, butoconazole, caspofungin, clobenzyl, chlorophetanol, chlorphenesin, ciclopirox, cilofungin, citronella oil, clotrimazole, fluconazole, crystal violet, dapsone, demazole, eberconazole, econazole, efinaconazole, ethylparaben, fenticonazole, fluconazole, flucytosine, flutrimazole, fluconazole, griseofulvin, halopronin, hafnazole, fluconazole, griseofulvin, hafnazole, flucon ... ampicillin, hexaconazole, isavuconazole, isoconazole, itraconazole, ketoconazole, citronella, lemon myrtle, luliconazole, micafungin, miconazole, naftifine, natamycin, netaconazole, nystatin, omoconazole, orange oil, oxiconazole, patchouli, pentamidine, polyoxil, posaconazole, potassium iodide, ravuconazole, salicylic acid, selenium sulfide, sertaconazole, sodium thiosulfate, sulpentine, sulconazole, taurolidine, tavaborole, tea tree oil, terbinafine, terconazole, ticlazone, tioconazole, tolcitramate, tolnaftate, tribromo-metacresol, undecylenic acid, voriconazole, Whitefield's ointment, and combinations thereof.

[0394] Suitable examples of antiprotozoal drugs include, but are not limited to, those classified in therapeutic subgroup P01 of the Anatomical Therapeutic Chemical Classification System. Other examples include, but are not limited to, albendazole, amodiaquine, amphotericin B, arsothiol, artemether, artemisinin, arteether, artemisinin, artesunate, atovaquone, azanidazole, benznidazole, bromoquinol, carnidazole, quiniodine, chlorhexidine, chloroquine, chlorproguanil, chlorquinaldol, clindamycin, clindamycin, clindamycin, dehydroemetine, dihydroartemisinin, diiodohydroxyquinoline, diloxanide, doxycycline, elonisole, emetine, etofamide, fexinidazole, fumigatus chlorpyrifos, furazolidone, bismuth arsine, halofantrine, hydroxychloroquine, iodoquinol, lumefantrine, mefloquine, meglumine antimoniate, melarsoprol, mepacrine, metronidazole, miltefosine, nifurtimox, nimorazole, nitazoxanide, ornidazole, pyrimethamine, paromomycin, pentamidine, ubiquinone, piperaquine, primaquine, proguanil, propamidine, propranidazole, pyrimethamine, bisquinacrine, quinacrine, quinidine, quinine, secnidazole, sodium stibogluconate, sulfadiazine, sulfadoxine, sulfamethoxazole, sulfamethoxazole azole, suramin, tafenoquine, ticlorazepam, tenonidazole, tetracycline, bromoquinol, tinidazole, trimethoprim, trimesacil, and combinations thereof.

[0395] Another object of the present invention is a method for assessing the risk of contamination by microorganisms, preferably viruses, bacteria, archaea, fungi or protozoa, in a sample.

[0396] In one embodiment, the method for assessing the risk of contamination by microorganisms, preferably viruses, bacteria, archaea, fungi or protozoa, according to the present invention comprises the step of providing a sample. In one embodiment, the sample may be a biological sample or a non-biological sample.

[0397] In one embodiment, the method according to the present invention for assessing the risk of contamination by microorganisms, preferably viruses, bacteria, archaea, fungi or protozoa, comprises the following step: performing any method for distinguishing between nucleic acid sequences of transcriptionally active microorganisms and inert microorganisms, preferably viruses, bacteria, archaea, fungi or protozoa, in a sample according to the present invention.

[0398] In one embodiment, the method according to the present invention for assessing the risk of contamination by microorganisms, preferably viruses, bacteria, archaea, fungi or protozoa, comprises the following steps: if at least one identified microorganism, preferably a virus, bacteria, archaea, fungi or protozoa nucleic acid sequence hit belongs to a living microorganism, preferably a virus, bacteria, archaea, fungi or protozoa, it is inferred that the sample is at risk of contamination. BRIEF DESCRIPTION OF THE DRAWINGS

[0399] FIG1 is a set of two graphs illustrating the substitution rate and substitution index of T nucleotides. Figure 1A : The substitution rate of T nucleotides is expressed as the ratio of substituted T to the total number of Ts. Figure 1B : The substitution index is expressed as the ratio of the "T to C" substitution rate to the average of the "T to A" + "T to G" substitution rates [T: TBEV; S: SMRV; C: cellular RNA].

[0400] Figure 2 is a graph illustrating the substitution rate (in %) of T nucleotides to C, G, or A in samples treated with 4sU and alkylated for microbial nucleic acid sequence hits used as the TBEV consensus sequence constructed from data for the current conditions.

[0401] Figure 3 It is a graph illustrating the substitution rate (in %) of T nucleotides to C, G or A in samples treated with 4sU and alkylated with microbial nucleic acid sequence hits used as the SMRV consensus sequence constructed from data of the current conditions.

[0402] Figure 4 is a graph showing the GC content distribution of 662 contigs selected as candidates for reconstruction of the LC5_ALAID_CNS reference genome.

[0403] Figure 5Figure 1 shows a tiled image of the A. laidlawii strain PG8A genome and selected reads from 662 contigs initially assembled under LC5 experimental conditions. Matching contigs (black forward, gray reverse) are reported at their true similarity percentage (top) and normalized to 10% similarity to flatten coverage and facilitate visualization.

[0404] Figure 6 is a set of seven graphs illustrating the substitution rate (or conversion rate) (in %) of T nucleotides to C, G, or A along the LC5_ALAID_CNS reference sequence under different test conditions. Only events with high confidence (≥20X depth) were selected for analysis. Figure 6A :CTRL5 label condition; Figure 6B : LC5 condition; Figure 6C : LC5 labeling conditions; Figure 6D : 40-fold dilution LC5 tag condition; Figure 6E : HC_HK5 tag condition; Figure 6F : HC_HK5 tag condition; Figure 6G : HC_G5 tag conditions

[0405] Example

[0406] The above and other aspects and features of the present invention will be further illustrated by the following examples. These examples are merely illustrative and not restrictive.

[0407] Example 1: Detection of replicative tick-borne encephalitis virus (TBEV) in cultured Vero cells

[0408] Materials and methods

[0409] Material

[0410] Vero cells were grown in minimum essential medium (MEM) supplemented with 2% fetal bovine serum (FBS). The virus used for infection was tick-borne encephalitis virus (TBEV), which belongs to the Flaviviridae family and consists of a ssRNA(+) genome with an average size of 10 kb.

[0411] method

[0412] Viral infection

[0413] Vero cells were seeded at 400,000 cells / well in 3 wells of MW6 plates and the cells reached 10 6 Pieces / hole.

[0414] The cells were then infected with TBEV virus at an MOI (multiplicity of infection) of 1 and incubated on ice for 1 hour with agitation.

[0415] For one well, the medium was removed immediately after incubation, and the cells were lysed with 1 mL of TRIzol and stored at -80°C until RNA extraction (condition 1).

[0416] For the other two wells (conditions 2 and 3), the culture medium was removed, replaced with MEM + 2% FBS, and incubated overnight at 37°C.

[0417] 4sU mark

[0418] This step was performed using the SLAMseq Kinetics Kit—Synthetic Metabolic Kinetics Module) (Lexogen, Cat. No. 061).

[0419] During cell culture, 4-thiouridine (4sU) is incorporated into the cell culture medium, and the 4sU nucleotide is incorporated into newly synthesized RNA.

[0420] Prepare a medium containing 800 μM 4sU by adding 8 μL of a 100 nM 4sU solution in 992 μL MEM.

[0421] On the second day after viral infection, the medium was removed and replaced with medium without 4sU in one well (Condition 2) or with medium containing 4sU (800 μM) in the last well (Condition 3). Six hours later, the medium was removed and replaced with fresh medium without 4sU in Condition 2 or with fresh medium containing 4sU (800 μM) in Condition 3.

[0422] Three hours later, the culture medium was removed from three wells, and the cells were lysed with 1 mL of TRIzol and stored at -80°C until RNA extraction.

[0423] RNA sampling

[0424] This step was performed using the SLAMseq Kinetics Kit—Synthetic Metabolic Kinetics Module) (Lexogen, catalog number 061).

[0425] RNA extraction was performed in the dark using a 24:1 mixture of chloroform and isoamyl alcohol (Sigma Aldrich, catalog number 25666) followed by isopropanol / ethanol precipitation. During the extraction process, reducing agent (RA) was used to maintain 4sU-treated samples under reducing conditions.

[0426] The isolated total RNA contains both existing (unlabeled) RNA and newly synthesized (labeled) RNA.

[0427] Alkylation

[0428] This step was performed using the SLAMseq Kinetics Kit—Synthetic Metabolic Kinetics Module) (Lexogen, catalog number 061).

[0429] Total RNA extracted from condition 3 was mixed with iodoacetamide (IAA) to modify the 4-thiol groups of 4sU-containing nucleotides by adding carboxamidomethyl groups. RNA was then purified by ethanol precipitation prior to library preparation.

[0430] Library preparation

[0431] The SMARTer Stranded Total RNA-Sequencing Kit - Pico Input Mammalian (ClonTech) was used to construct libraries directly from 10 ng of RNA. The workflow used in this kit incorporates a proprietary technology (PathoQuest, Paris, France) that depletes ribosomal cDNA using probes specific for mammalian rRNA and some mitochondrial RNAs.

[0432] Sequencing

[0433] Sequencing was performed on a NextSeq instrument (Illumina) using the NextSeq 500 / 550 High Output Reagent Kit v2 (FC-404-2002, Illumina).

[0434] Sequencing was single-read with a read length of 150 nucleotides, generating approximately 125 million reads per sample.

[0435] Overview

[0436] Table 1 below summarizes the protocols used in three different situations (1, 2, and 3).

[0437] Table 1: Overview of the program

[0438]

[0439] Bioinformatics analysis—TBEV genome analysis

[0440] The first goal of this study was to obtain the complete TBEV genome sequence of this isolate for reference. This analysis was performed on samples infected at day 0 without 4sU treatment.

[0441] Raw read filtering

[0442] First, the raw data reads are filtered to select high-quality and relevant reads.

[0443] Classify the raw data to suppress or reduce duplicate data, low-quality read data and homopolymers (dedicated software). Sequences introduced during library preparation (adapters, primers) were removed using Skewer (Jiang et al., 2014. BMC Bioinformatics. 15: 182).

[0444] Finally, endogenous primate reads (from Vero cells) that aligned to the human genome (reference GRCh37 / hg19) or reads that aligned to bacterial rRNA were discarded.

[0445] Local alignment was performed using BWA (Li et al., 2009. Bioinformatics. 25(14): 1754-60).

[0446] The human genome was downloaded from the UCSC Genome Browser (2002. Genome Res. 12(6):996-1006). The bacterial rRNA database was downloaded from the EMBL-EBI ENA rRNA database with additional in-house sequence cleaning and clustering processes.

[0447] These filtered reads are considered as target sequences.

[0448] Assemble from scratch

[0449] The remaining and related sets of reads were then assembled into longer sequences, called contigs. This de novo assembly step was performed using the CLC Assembly Cell Solution (Qiagen).

[0450] Agnostic virus identification

[0451] The resulting contigs were aligned with unassembled reads (singletons) using BLAST alignment (Altschul et al., 1990. J Mol Biol. 215(3):403-10). Contigs and singletons were first aligned against a viral nucleotide database. -3 The hits were compared against a comprehensive nucleotide database. If their best hit was still a virus taxon, that hit was reported.

[0452] The nucleotide virus and comprehensive databases were downloaded from the EMBL-EBI nucleotide sequence database STD in November 2017. A dedicated software was developed to remove duplicate and low-confidence sequences (due to being too short, having multiple classifications, low-quality associated keywords, etc.).

[0453] Contigs without any viral nucleotide hits were similarly aligned consecutively against the viral and comprehensive protein databases to check for more distant viral hits.

[0454] Protein, virus, and comprehensive databases were downloaded from the Uniref100 database in November 2017. Although the Uniref100 database was already non-redundant, a taxonomic cleaning process was performed to generate the final database.

[0455] The best hit was reported for the classification task. Contigs that were not assigned after these two rounds of alignment were classified as unknown or non-viral species.

[0456] Table 2 below shows the results of the analysis.

[0457]

[0458] Table 2: Final consensus set of agnostic virus identification TBEV

[0459] This process identified contigs containing the full TBEV genome sequence.

[0460] All reads were then realigned on this sequence using the CLC assembly cell solution (Qiagen) to extract the final consensus sequence, which was labeled the "TBEV reference" for this study.

[0461] Bioinformatics analysis - nucleotide substitution rate study

[0462] The purpose was to compare the reads of different samples to check whether the "T to C" substitution rate was significantly higher in the 4sU+alkylated sample (condition 3).

[0463] Creation of a tick-borne encephalitis virus library

[0464] The "TBEV reference" sequence was used to create the blast library.

[0465] To detect potential sequences with a very high rate of "T to C" substitutions, the reference sequence was also modified to replace every T with a C. This sequence was named "TBEV TC Reference". The "TBEV Reference" and "TBEV TC Reference" sequences were merged together to form a single "TBEV BLAST" library.

[0466] Raw read filtering

[0467] First, a quality filtering process was performed to remove or trim low-quality reads (dedicated software).

[0468] Skewer was then used to remove sequences introduced during Illumina library preparation (adapters, primers) (Jiang et al., 2014. BMC Bioinformatics. 15: 182).

[0469] To avoid any analytical bias, duplicate reads were not removed.

[0470] Filtered read blasts on TBEV blast library

[0471] The remaining relevant reads were then aligned by BLAST (Altschul et al., 1990. J Mol Biol. 215(3):403-10) on a previously designed “TBEV BLAST” library. The maximum e-value was set to 10 -8 .

[0472] All aligned reads were considered TBEV positive and selected for further analysis.

[0473] Matching of selected reads on the TBEV whole genome

[0474] TBEV positive reads were then realigned against the “TBEV reference” sequence by alignment using the CLC assembly cell solution (Qiagen). Quality controls were established to ensure that at least 99% of blast-selected reads were positively realigned against the reference.

[0475] Table 3 below summarizes the number of matched sense and antisense reads obtained in each case and the resulting sequence coverage.

[0476]

[0477] Table 3: Read matches, orientations, and coverage.

[0478] Substitution rate estimation

[0479] Each mismatch at each position in the TBEV study reference was detected using the CLC program "CLC_FIND_VARIANIES." The global variation profile was then analyzed by a dedicated script to define the substitution rate for each nucleotide. The proportion of substituted nucleotides was compared to the total number of aligned nucleotides. Typically, the "T to C" substitution rate was calculated using the following formula:

[0480]

[0481] TBEV chain analysis

[0482] Target and strand analysis was performed on TBEV-identified reads. This analysis performed a more stringent alignment of filtered reads. This alignment provided detailed horizontal genome coverage and depth profiles.

[0483] Local alignment was performed using BWA (Li et al., 2009. Bioinformatics. 25(14): 1754-60).

[0484] Because the sample libraries were prepared using the SMARTer Stranded RNA Sequencing Kit, RNA strand information was preserved. Therefore, the alignment analysis provided information about the parent strand of each read (forward or reverse orientation relative to the parent strand).

[0485] Transcript coverage allows conclusions to be drawn about the characteristics of viral replication in cell samples.

[0486] Results and Conclusions

[0487] The ratio of the “T to C” substitution rate calculated based on mismatches on the TBEV reference genome in condition 3 to the T to C substitution rate calculated based on mismatches on the TBEV reference genome in condition 1 was equal to 7.86:1 (Table 4).

[0488] This indicates that the proportion of TBEV RNA species increased with the addition of 4sU, thereby newly synthesizing viral RNA within 9 h of Vero cell culture with 4sU.

[0489] Thus, the method presented here allows the detection of the replicative (+)ssRNA virus, TBEV, using metabolic labeling.

[0490]

[0491] Table 4: T to C substitution rates.

[0492] Example 2: Detection of replicative squirrel monkey retrovirus (SMRV) in cultured Vero cells

[0493] Following the identification of agnostic viruses in Example 1, the best hits showed contigs assigned to squirrel monkey retroviruses. This virus is known to be endogenous and fully integrated in some monkey species. In particular, the Vero cells used in this study have been described as containing multiple monkey endogenous type D retrovirus sequences, particularly SMRV sequences (Sakuma et al., 2018. Sci Rep. 8(1):644).

[0494] Based on this knowledge and in view of the results shown in Table 2 above, the same bioinformatics program was performed to identify SMRV sequence hits and to study the nucleotide substitution rates of the sequence hits.

[0495] Table 5 below summarizes the number of matched sense and antisense reads obtained in each case and the resulting sequence coverage.

[0496]

[0497] Table 5: Read matches, orientations, and coverage.

[0498] The ratio of the "T to C" substitution rate calculated based on mismatches on the SMRV reference genome in condition 3 to the T to C substitution rate calculated based on mismatches on the SMRV reference genome in condition 1 was equal to 41.86:1 (Table 6).

[0499] This indicates that the proportion of SMRV RNA species increased with the addition of 4sU, thereby newly synthesizing viral RNA within 9 h of Vero cell culture with 4sU.

[0500] Thus, the method presented here allows the detection of the replicative (+)ssRNA-RT virus, SMRV, using metabolic labeling.

[0501]

[0502] Table 6: T to C substitution rates.

[0503] Example 3:

[0504] Materials and methods

[0505] Cells and viruses

[0506] A vial of Vero cells (ATCC-CCL-81, lot number 62488537, Molsheim, France) was frozen at passage 3 and then thawed in a BSL-3 laboratory and grown in MEM supplemented with 10% FBS. Cells were used at passage 18.

[0507] A second bottle of Vero cells (lot number 70005907) was purchased from the same source and used directly for PCR testing.

[0508] Tick-borne encephalitis virus (TBEV) is a member of the Flaviviridae family and consists of a ssRNA(+) genome. The Hypr strain (Wallner et al., 1996. J Gen Virol. 77(Pt 5):1035-42) was kindly provided by Sarah Moutailler, ANSES, Maisons-Alfort, France.

[0509] TBEV infection of Vero cells

[0510] Vero cells were seeded at 400,000 cells / well in 3 wells of MW6 plates so that the cells reached 10 6 The cells were then infected with Hypr TBEV at a multiplicity of infection (MOI) of 1 and incubated on ice for 1 hour with agitation.

[0511] Immediately after incubation, the culture medium was removed and the cells were lysed with 1 mL of Trizol and stored at -80°C until RNA extraction (condition "D0-no 4SU").

[0512] For the other two wells (conditions 2, 3, and 4), the culture medium was removed and replaced with MEM + 10% fetal bovine serum, and incubated overnight at 37°C.

[0513] 4sU labeling and RNA extraction

[0514] Adding 4-thiouridine (4sU) to the cell culture medium allows the 4sU nucleotide to be incorporated into newly synthesized RNA. Reverse transcription of 4sU shows a certain proportion of misincorporation resulting in T>C transitions in cDNA, which can be identified by sequencing (Herzog et al., 2017. Nat Methods. 14(12): 1198-1204).

[0515] A medium containing 800 μM 4sU was prepared by adding 8 μL of 100 nM 4sU to 992 μL of MEM. The day after viral infection, the medium was removed and replaced with medium without 4sU in one well (condition "D1-no 4SU") or with medium containing 4sU (800 μM) in another well (condition "D1-with 4SU").

[0516] Six hours later, the culture medium was removed and replaced with fresh medium without 4sU, the condition being "D1-without 4sU", or with fresh medium containing 4sU (800 μM), the condition being "D1-with 4sU".

[0517] Three hours later, the culture medium was removed from the three wells, and the cells were lysed with 1 mL of Trizol and stored at −80°C until RNA extraction.

[0518] RNA was extracted in the dark using a 24:1 mixture of chloroform and isoamyl alcohol (Sigma Aldrich, Cat. No. 25666, St. Louis, USA) followed by isopropanol / ethanol precipitation. During the extraction process, a reducing agent was used to keep the 4sU-treated samples under reducing conditions.

[0519] Alkylation reactions were performed using the SLAMseq Kinetics Kit - Synthome Kinetics Module (Lexogen, catalog number 061, Vienna, Austria) with the "D1-containing 4sU" fraction as the condition. Extracted total RNA was mixed with iodoacetamide (IAA) to modify the 4-thiol groups of 4sU-containing nucleotides by adding a carboxamidomethyl group, resulting in the "D1-containing 4sU + alkylation" condition. This alkylation amplifies the frequency of T>C mismatches during reverse transcription. The other fraction was labeled "D1-containing 4sU without alkylation."

[0520] RNA was then purified using ethanol precipitation prior to library preparation.

[0521] Library preparation and sequencing

[0522] The SMARTer Stranded Total RNA-Sequencing Kit - Pico Input Mammalian (ClonTech, Mountain View, USA) was used to construct libraries directly from 10 ng of RNA. The workflow used with this kit incorporates a proprietary technology (PathoQuest, Paris, France) that uses probes specific for mammalian rRNA and some mitochondrial RNA to deplete ribosomal cDNA. Sequencing was performed on a NextSeq instrument (Illumina, San Diego, USA) using the NextSeq500 / 550 High Output Kit v2 (FC-404-2002, Illumina). Sequencing was single-read with a read length of 150 nucleotides, generating approximately 125 million reads per sample.

[0523] Agnostic bioinformatics analysis

[0524] The raw data reads were filtered to select high-quality and relevant reads and classified to suppress or reduce duplicates, low-quality reads, and homopolymers (PathoQuest proprietary software).

[0525] Skewer was used to remove sequences introduced during the preparation of Illumina libraries (adapters, primers) (Jiang et al., 2014. BMC Bioinformatics. 15: 182).

[0526] Discard primate reads (from Vero cells) aligned with the human genome (reference GRCH37 / HG19) or reads aligned with bacterial rRNA. Local alignment was performed with BWA (Li et al., 2009. Bioinformatics. 25 (14): 1754-60). The human genome was downloaded from the UCSC genome browser (Kent et al., 2002. Genome Res. 12 (6): 996-1006). The bacterial rRNA database was initially downloaded from the EMBL-EBI ENA rRNA database (ebi.ac.uk / pub / Databases / ena / rRNA / release), followed by additional internal sequence cleaning and clustering processes. These filtered reads were considered to be target sequences and assembled into longer sequences, referred to as "overlapping groups," using CLC assembly cell solutions (Qiagen Hilden, Germany). The resulting contigs were aligned with the unassembled reads (monads) using BLAST alignment (Altschul et al., 1990. J Mol Biol. 215(3):403-10). Contigs and monoids were first aligned against a viral nucleotide database. -3 The hits are compared against a comprehensive nucleotide database. If the best hit is still a virus, it is reported.

[0527] The nucleotide viral and comprehensive databases were downloaded from the EMBL-EBI nucleotide sequence database STD in November 2017. A dedicated software (PathoQuest, Paris, France) was developed to remove duplicates and low-confidence sequences (e.g., too short, multiple taxa, low-quality associated keywords, etc.). Contigs without any viral nucleotide hits were similarly aligned consecutively against the viral and comprehensive protein databases to check for more distant viral hits. In November 2017, the protein, viral, and comprehensive databases were downloaded from the Uniref100 database (https: / / www.uniprot.org). Although the Uniref100 database is already non-redundant, a taxonomic cleaning process was used to generate the final databases. Taxonomic assignments were reported for the best hits, and contigs without assignments after these two rounds of alignment were classified as unknown or non-viral species.

[0528] The above process identified contigs containing the full TBEV genome sequence (see "Results"). All reads were then realigned on this sequence using the CLC assembly cell solution (Qiagen, Hilden, Germany) to extract the final consensus sequence. Data retrieved from the "D0-no 4sU" condition allowed the identification of contigs covering the entire TBEV and SMRV genomes, and the resulting sequences were labeled "TBEV reference" and "SMRV reference," respectively.

[0529] Estimation of T>C substitution rate

[0530] To enable detection of viral sequences with a high rate of “T to C” substitutions, each reference sequence was also modified by replacing each T with a C. These sequences were named “TBEV TC Reference” and “SMRV-TC Reference”. The “TBEV Reference” and “TBEV TC Reference” as well as the “SMRV Reference” and “SMRV TC Reference” were merged together to form two libraries named “TBEVBLAST” and “SMRV BLAST”. The quality-filtered read sets were then aligned by BLAST using these previously designed “BLAST” libraries. The maximum e-value was set to 10 -8 Only aligned reads were selected for the next step of analysis.

[0531] Each mismatch at each position in the TBEV study reference was detected using the CLC program "CLC_FIND_VARIANIES". The global variation map was then analyzed by a dedicated script (PathoQuest, Paris, France) to define the substitution rate for each nucleotide. The proportion of substituted nucleotides was compared to the total number of aligned nucleotides. For example, the "T to C" substitution rate was calculated using the following formula:

[0532]

[0533] The substitution rate at each time point was normalized using the following substitution index:

[0534]

[0535] As a quality control for labeling, the average substitution index of a panel of exons was examined using unlabeled cells as a reference. The exons of the following human genes described by Eisenberg and Levanon (2013. Trends Genet. 29(10):569-74) (reference sequence accession numbers) were used: C1orf43 (NM_015449), CHMP2A (NM_014453), EMC7 (NM_020154), and GPI (NM_000175).

[0536] These human exons were used to identify the corresponding exons in the green monkey genome, from which Vero cells are derived. The complete assembly of the green monkey (accession number GCF_000409795.2) was retrieved from the NCBI assembly database (https: / / www.ncbi.nlm.nih.gov / assembly / ). The selected human exons were mapped to the green monkey assembly using minimap2 (Li, 2018. Bioinformatics. 34(18):3094-3100), and the resulting .bam files were converted to .bam files using the bamtobed ​​module in the BEDTool utility (Quinlan & Hall, 2010. Bioinformatics. 26(6):841-2). Only hits with a match quality higher than 30 (41 exons) were retained, and the corresponding sequences were extracted from the green monkey assembly using the getfasta module in the BEDTool utility and tagged for further analysis. If the substitution index is better than 10, the label is considered satisfactory.

[0537] Chain Analysis

[0538] Target and strand analysis was performed on the identified TBEV reads. This analysis is based on a matched alignment of the more stringently filtered reads, which provides detailed horizontal genome coverage and deep distribution. Local alignment was performed using BWA. Because the sample library was prepared using the SMARTer Stranded RNA Sequencing Kit, RNA strand information was also preserved. As a result, the map alignment analysis can provide information about the parent strand of each read (forward or reverse relative to the parent strand).

[0539] result

[0540] Identification of exogenous viruses by agnostic RNA sequencing in Vero cells

[0541] Vero cells were first exposed to a high dose of TBEV at +4°C (D0). At this temperature, only the virus binds to cell receptors, blocking viral entry. Therefore, this experimental condition mimics carriage by a non-replicating virus. RNA was extracted and sequenced as a marker for DNA or RNA viral infection. The results of the agnostic analysis and the matches of the reads to the two main viral hits found by the agnostic analysis (TBEV and SMRV) are shown in Tables 7 and 8, respectively.

[0542]

[0543] Table 7: Reads (antisense / total reads) and genome level coverage (% genome) of the TBEV and SMRV genomes. Reads on the TBV and SMRV genomes found by the agnostic program were matched (Table 8).

[0544]

[0545] Table 8: Agnostic analysis - results of reads, de novo assembly and BLAST analysis after each step of the filtering process.

[0546] As expected, the main virus species detected at D0 was TBEV, but also unexpectedly SMRV (Table 7). Of a total of approximately 150 million raw reads (Table 8), more than 160,000 TBEV reads were identified, covering the entire genome. The Vero cells were then moved to 37°C to allow viral entry and then incubated for one day before harvesting. The number of reads increased significantly, with 5.2 million to 6.4 million TBEV reads recorded. In addition, 1.6 million to 1.8 million reads matched SMRV-H (SMRV isolated from a human lymphocyte cell line) (Oda et al., 1988. Virology. 167 (2): 468-76)) and were not related to the harvest day. This means that SMRV transcripts are expressed by cells and have nothing to do with experimental TBEV infection.

[0547] Some other hits were also identified (Table 8). The main other hits were baboon endogenous viruses, known Vero cell endogenous viruses (Ma et al., 2011. J Virol. 85(13): 6579-88). Hundreds of reads matching endogenous human retroviruses were also recorded. Based on experience, this finding is common in primate / human cell lines. Some typical BVDV reads related to the use of gamma-irradiated bovine serum were also found. Some reads for different herpes viruses (<50) were also identified, which were considered background noise.

[0548] Distinguishing between cell infection and inert sequence carriage

[0549] Since the main goal was to simulate challenging conditions that distinguish between cell infection and carriage and to test the ability of HTS to detect early infection of cells, the results of blocking viral replication in cells exposed to a high dose of TBEV virus at +4°C were compared with the results of cells infected with the same dose of virus 24 hours after infection. The former simulated cell-inactivated virus or free nucleic acid, while the latter simulated cells infected before library preparation. Since TBEV is a positive-sense ssRNA virus, antisense RNA was used as a marker of viral replication. The three conditions tested at D1 (no 4sU; with 4sU + alkylation; with 4sU without alkylation) showed that 0.32% to 0.36% of the reads were antisense, compared to 0.27% at D0, a very small but very significant difference (chi-square test, p < 0.0001). This type of comparative analysis is not relevant to chronic infection of cells by SMRV, which is a retrovirus whose transcription uses a DNA provirus as a substrate and mainly produces positive-sense RNA but also antisense RNA (Manghera et al., 2017. Virol J. 14(1):9).

[0550] Then, the newly synthesized RNA was metabolically labeled by 4sU, and the TBEV ratio of "T to C" substitutions was subsequently examined (Table 9 and Figure 1).

[0551]

[0552] Table 9: T nucleotide substitution rate and substitution index

[0553] At D1, in the absence of metabolic labeling, the "T to C" ratio was very low (0.13%), similar to the "T to A" or "T to G" ratios (0.04-0.13%), resulting in a calculated background substitution index of 1.68. Similar results were obtained at D0, indicating good reproducibility of the substitution background.

[0554] In stark contrast, the T to C substitution rate of labeled and alkylated TBEV RNA was much higher (0.79%) at D1, resulting in a substitution index of 6.4, 3.8 times the background value. The substitution index of labeled and alkylated SMRV cells at D1 was 24.16, 10.7 times the background value.

[0555] Comparison between metabolically labeled RNA and non-labeled RNA requires two culture conditions. Therefore, the TBEV and SMRV substitution indices at D1 were also compared in 4sU-labeled cultures before and after RNA alkylation. This only requires one culture condition, followed by RNA extraction and alkylation, or no treatment. The low level of substitution in RNA 4sU-labeled non-alkylated cells did not affect the detection of latent virus hits by BLAST analysis (Table 8). As shown in Tables 9 and Figure 1BAs shown, the substitution index of 4sU-labeled non-alkylated RNA remains low, close to the substitution index under non-labeling conditions (the substitution index of TBEV and SMRV is 1.71 and 2.27, respectively, which increases to 4.0 and 10.6 times under alkylation conditions). This indicates that non-alkylated RNA extracted from the same cell culture can be used to establish a reference consensus viral library for calculating substitution rates. Therefore, the results show that after 4sU labeling of cells, RNA sequencing can specifically identify newly synthesized viral RNA with a high signal-to-noise ratio.

[0556] Finally, for TBEV( Figure 2 ) and SMRV( Figure 3 ), and also compared the ratio between the T→C substitution rate in 4sU-labeled alkylated cells and the average T→A and T→G substitutions observed in the same cells.

[0557] These ratios are given in Table 10. A substitution ratio greater than 1 indicates transcriptional activity in the sample. Therefore, these results clearly demonstrate that the method of the present invention is capable of distinguishing and detecting viable TBEV and SMRV by comparing the substitution rates of different nucleotides under a single condition (D1-containing 4SU + alkylation).

[0558] TBEV SMRV T→A 0.10 0.10 T→C 0.83 3.54 T→G 0.18 0.12 T→C / average(T→A, T→G) 5.87 32.41

[0559] Table 10: Ratio of T→C substitution rate to the average T→A / T→G substitution rate

[0560] Example 4:

[0561] Materials and methods

[0562] cells and molluscs

[0563] Before contamination, A549 (ATCC_CCL-185) cells were cultured in DMEM-Dulbecco's Modified Eagle Medium in 6-well plates to approximately 70% confluence.

[0564] Acholeplasma reinhardtii was the representative of the family Mollicutes that was selected to infect A549 cells.

[0565] Infection of A549 cells with Acholesterola

[0566] When about 70% confluence, the culture medium of A549 cells was changed into MEM-Earle culture medium supplemented with 7% fetal bovine serum and 1% L-glutamine and without antibiotics. Cells were infected with different infection doses of A. leucoderma (Table 11) at day 0. On day 5, 4-thiouridine (4sU) (800 μM) was added to the culture medium 9 hours, 6 hours and 3 hours before the supernatant was harvested. 2 mL of culture medium was taken out after incubation at 37 ° C for 5 days, and 200 g was centrifuged for 5 minutes to clarify it. 1 mL of clarified supernatant was centrifuged at 15000 g to 20000 g for 10 minutes, 900 μ L of supernatant was removed, and the precipitation was homogenized in the remaining 100 μ L of supernatant. The sample was then frozen before the nucleic acid was extracted.

[0567] Adding 4-thiouridine (4sU) to cell culture media allows the 4sU nucleotide to be incorporated into newly synthesized RNA. Reverse transcription of 4sU shows a certain proportion of misincorporation, resulting in T>C transitions in cDNA, which can be identified by sequencing (Herzog et al., 2017. Nat Methods. 14(12): 1198-1204).

[0568]

[0569] Table 11: Test item description

[0570] * : This assay will be evaluated with and without dilution; sample LC5 tags will be diluted prior to inactivation and infection of cells to obtain similar cholesterogen-free counts.

[0571] ** : Before inactivation.

[0572] CTRL5 label is a control sample that was not infected with A. reinhardtii and was labeled with 4sU on the 5th day.

[0573] LC5 refers to the sample infected with a low concentration of A. reinhardtii on day 5.

[0574] The LC5 tag was a sample infected with a low concentration of A. reinhardtii and labeled with 4sU on day 5.

[0575] HC_HK5 label is a sample that was heat-killed before infection and infected with a high concentration of A. reinhardtii and labeled with 4sU on day 5.

[0576] HC_G5 label is a sample infected with a large dose of A. reinhardtii and treated with gentamicin before infection and labeled with 4sU on day 5.

[0577] RNA extraction

[0578] RNA extraction was performed in the dark using a 24:1 mixture of chloroform and isoamyl alcohol (Sigma Aldrich, Cat. No. 25666, St. Louis, USA) followed by isopropanol / ethanol precipitation. During the extraction process, a reducing agent was used to maintain 4sU-treated samples under reducing conditions.

[0579] Alkylation reactions were performed using the SLAMseq Kinetics Kit - Synthome Kinetics Module (Lexogen, catalog number 061, Vienna, Austria) using only the "D1-containing 4sU" fraction. Extracted total RNA was mixed with iodoacetamide (IAA) to modify the 4-thiol groups of 4sU-containing nucleotides by adding a carboxamidomethyl group, resulting in the "D1-containing 4sU + alkylation" condition. This alkylation amplifies the frequency of T>C mismatches during reverse transcription. The remaining fraction is labeled "D1-containing 4sU without alkylation."

[0580] RNA was then purified using ethanol precipitation prior to library preparation.

[0581] Library preparation and sequencing

[0582] SMARTer Stranded Total RNA-Sequencing Kit-Pico Input Mammalian (ClonTech, Mountain View, USA) was used to construct libraries directly starting from 10 ng of RNA. Depletion of bacterial-derived ribosomal RNA (16S and 23S) was performed on total RNA using the Ribominus Bacterial Transcriptome Analysis Kit (thermoFisher). Prior to library preparation using the manufacturer's recommendations (ClonTech), ribosomal cDNA (included in the SMARTer Stranded Total RNA Sequencing Kit) was depleted using probes specific for mammalian rRNA and certain mitochondrial RNAs. Sequencing was performed on a NextSeq instrument (Illumina, San Diego, USA) using a NextSeq MID output flow unit (FC-404-1001, Illumina). Sequencing was single-read with a read length of 150 nucleotides, generating approximately 125 million reads per sample.

[0583] Agnostic bioinformatics analysis

[0584] The raw data reads were filtered to select high-quality and relevant reads and classified to suppress or reduce duplicates, low-quality reads, and homopolymers (PathoQuest proprietary software).

[0585] Skewer was used to remove sequences introduced during the preparation of Illumina libraries (adapters, primers) (Jiang et al., 2014. BMC Bioinformatics. 15: 182).

[0586] The filtered reads of the LC5 condition are first considered as the target sequence. Since this condition is likely to include a large number of unlabeled sequences of the target organism, this will allow the genome of the target organism (Acholesteroplasma lei) to be reconstructed. Therefore, the LC5 reads are assembled into longer sequences, called "overlap groups" (Li et al., 2015.Bioinformatics.31(10):1674-1676) using MegaHit. Then, the resulting overlap groups are matched back to the Acholesteroplasma lei strain PG8A genome (reference sequence accession number CP000896.1) using minimap2 (Li, 2018.Bioinformatics.34(18):3094-3100). Mummer 3 (Kurtz et al., 2004.Genome Biol.5(2):R12) is then used to tile the positive hits on the Acholesteroplasma lei strain PG8A genome to:

[0587] 1. Confirm the identity of the contig potentially detected as A. leishmanii,

[0588] 2. Ensure the integrity of the newly created sequence.

[0589] After the identities of the contigs were assessed and verified by tiling, the contigs were merged into .fasta files and used as a reference genome (hereafter referred to as LC5_ALAID_CNS) for further analysis.

[0590] Estimation of T>C substitution rate

[0591] To detect A. leishmanii sequences with a very high number of T→C substitutions, the quality-filtered read set was matched back to LC5_ALAID_CNS (Li, 2018. Bioinformatics. 34(18): 3094-3100) using minimap2 in non-multiple match mode. All mismatches (base quality at least equal to 30) at each position of the LC5_ALAID_CNS sequence were then detected using the stacking module of the htsbox software (https: / / github.com / lh3 / htsbox). The global variation map was then analyzed using a dedicated script (PathoQuest, Paris, France) to define the substitution rate for each nucleotide. The proportion of substituted nucleotides was compared to the total number of aligned nucleotides. For example, the "T→C" substitution rate was calculated using the following formula:

[0592]

[0593] The substitution rate at each time point was normalized using the following substitution index:

[0594]

[0595] result

[0596] Sequencing amount

[0597] The sequencing run yields are shown in Table 12. Over 10 million single-end reads were generated in almost all cases. For each condition, over 90% of reads were retained after the filtering step, indicating that the sequencing results were of high quality and therefore suitable for subsequent analysis.

[0598] Test items Raw reads Filtered reads ratio CTRL5 tag 14 092 086 12 698 272 0.90 LC5 14 922 171 14 766 414 0.99 LC5 tag* 17 690 706 17 519 522 0.99 Diluted LC5 tag* 17 124 090 16 894 486 0.99 HC HK5 Label 20 410 793 20 076 010 0.98 4°HC HK5 label 20 410 793 20 076 010 0.98 HC_G5 label 16 591 955 16 060 004 0.97

[0599] Table 12: Sequencing yields under all experimental conditions

[0600] Reference genome reconstruction

[0601] The LC5 read assembly process allowed the generation of a set of 877 contigs (cumulative length 1374213; minimum length = 201; average length = 1566.9; maximum length = 15043). Realigning them to the Acholera reinhardtii PG8A (CP000896.1) genome sequence allowed the unambiguous selection of 662 contigs (cumulative length 1287020; minimum length = 301; average length = 1944.1; maximum length = 15043) as candidates for the reconstruction of LC5_ALAID_CNS. As a first check, the GC content distribution and statistics were investigated to demonstrate the possible mixture of organisms in the contig set ( Figure 4 ).

[0602] like Figure 4 As shown, the GC content distribution is unimodal, which indicates that there is a low probability that contigs representing several organisms exist in the contig set. In addition, the average GC content of this set of samples is not significantly different from the expected value (32.01% and 31.93% for A. reinhardtii strain PG8A).

[0603] To ensure that the entire genome (or at least a large part) of the close relative of A. reinhardtii strain PG8A could be reconstructed, the latter was “tiled” with contigs selected from the initial assembly of reads from the LC5 experimental condition. Figure 5 .

[0604] like Figure 5As shown, the set of 662 contigs covered almost the entire A. reinhardtii PG8A sequence with high similarity (>99% in all cases; data not shown), which strongly suggested that the constructed LC5_ALAID_CNS was a very close relative of A. reinhardtii PG8A.

[0605] In summary, it is possible to:

[0606] 1. Select the cleaned contig set corresponding to A. reinhardtii, and

[0607] 2. Covering the whole genome of a close relative (Acholesteroplasma leijkeri strain PG8A).

[0608] Therefore, this process validated the reference sequence LC5_ALAID_CNS for further analysis.

[0609] Substitution rate and index

[0610] Low coverage positions can lead to biased coverage estimates because they are weighted the same as positions that are fairly well covered. In fact, if a position is only covered three times and there is a T→C substitution, then the T→C substitution rate at that position will be 33%, regardless of whether it may be a true substitution or a sequencing / assembly error. Therefore, to avoid overestimating the substitution rate and thus the substitution index, an analysis was performed, first selecting all detected events (i.e., covered at least once (1X)) and then selecting events that were covered at least 20 times (i.e., 20X), the latter being considered high-confidence events.

[0611] The substitution rates and indices are shown in Table 13. Overall, it is shown here that the T→C transition rate is always higher than the transversion T→A and T→G rates, which is expected because the classical mutation pattern favors transitions over transversions.

[0612] Furthermore, the T→C substitution rates were significantly higher in the LC5 tag and 40-fold diluted LC5 tag conditions compared to all other conditions, regardless of the event selection level, including the high-load inactivated sample (HC_HK5 tag). Furthermore, including low-coverage positions in this analysis had little impact on the results, as there was no significant difference in the ratios observed at 1X and 20X thresholds, although the latter would have limited background noise. The substitution index showed the same trend.

[0613]

[0614] Table 13: Substitution rate and substitution index (SI) for each experimental condition for all detected events (1X threshold) and high confidence events (20X threshold).

[0615] In summary, the reported results indicate that both the experiments with added A. leishmanii and the experiments with 4sU labeling were detected as expected.

[0616] Location Analysis

[0617] Increases in global substitution rates and substitution indices have been reported. To investigate whether these increases were due to substitution hotspots, a positional analysis was performed along the LC5_ALAID_CNS reference sequence, assessing the substitution rate ( Figures 6A to 6G ).

[0618] The presence of A. reinhardtii reads was noted in the CTRL5 tag experimental condition, as some peaks were visible, even though A. reinhardtii was not added to this condition ( Figure 6A Most frequencies reached 100%, indicating that these substitutions were actually true SNPs. This observation suggests either contamination at the experimental level or cross-index contamination during sequencing when multiplexing samples (so-called index jumping). Nevertheless, due to the rather limited number of positions involved, this did not affect the analysis.

[0619] Figure 6B The substitution background for the LC5 condition was shown to be quite low, with all genomic positions well covered (i.e., no coverage gaps). No truly dominant substitution types were observed, consistent with the global scaling analysis results. In contrast, for both the LC5 tag and the 40-fold diluted LC5 tag conditions, a large T→C peak emerging from the background could be distinguished, indicating successful labeling in A. reinhardtii, leading to transcriptional activation ( Figure 6C and 6D ).

[0620] Figure 6E and Figure 6F Results are shown for 4sU-labeled experimental conditions in which the A. reinhardtii cells were heat-killed (HC_HK5 label and 4°HC_HK5 label). In both cases, a lower number of peaks (particularly the T→C peak) was observed compared to the LC5 label and 40X_LC5 label conditions, confirming that the amount of extracted RNA was lower due to the lower number of viable bacterial cells remaining in the culture medium after heating.

[0621] Likewise, gentamicin treatment had the same effect (HC_G5 condition; Figure 6G ), but compared with the thermal effect under the HC_HK5 labeling experimental conditions, it is obviously much milder ( Figure 6E and 6F However, 4sU labeling was still visible and confirmed the global analysis with a moderate substitution index (Table 13).

[0622] Example 5

[0623] Materials and methods

[0624] cells and molluscs

[0625] Before contamination, A549 (ATCC_CCL-185) cells were cultured in DMEM-Dulbecco's Modified Eagle Medium in 6-well plates to approximately 70% confluence.

[0626] No choleplasma or mycoplasma infection of A549 cells

[0627] At approximately 70% confluence, the culture medium of A549 cells was changed to MEM-Earle's medium supplemented with 7% fetal bovine serum and 1% L-glutamine without antibiotics. Cells were infected with either cholesteroplasma or mycoplasma at different infection doses.

[0628] Several scenarios were tested, including:

[0629] -CTRL5 label: control sample, not infected with cholesteroplasma or mycoplasma, labeled with 4-sU on day 5.

[0630] - LC5: infection with low concentration of cholesteroplasma- or mycoplasma-free samples on day 5.

[0631] - LC5 label: Samples infected with low concentrations of cholesteroplasma or mycoplasma-free and labeled with 4sU on day 5.

[0632] - HC_HK5 label: Samples infected with high concentrations of Cholesterol- or Mycoplasma-free bacteria that were heat-killed before infection and labeled with 4sU on day 5.

[0633] - HC_G5 label: samples infected with a large dose of cholesteroplasma or mycoplasma free treated with gentamicin before infection and labeled with 4sU on day 5.

[0634] On day 5, 4-thiouridine (4sU) (800 μM) was added to the culture medium 9, 6, and 3 hours before cell harvest. After incubation at 37°C for 5 days, the culture medium was removed and the cells were pelleted and frozen before RNA extraction.

[0635] Adding 4-thiouridine (4sU) to cell culture media allows the 4sU nucleotide to be incorporated into newly synthesized RNA. Reverse transcription of 4sU shows a certain proportion of misincorporation, resulting in T>C transitions in cDNA, which can be identified by sequencing (Herzog et al., 2017. Nat Methods. 14(12): 1198-1204).

[0636] RNA extraction

[0637] RNA extraction was performed in the dark using a chloroform:isoamyl alcohol mixture 24:1 (Sigma Aldrich, Cat. No. 25666, St. Louis, USA) followed by isopropanol / ethanol precipitation. During the extraction process, 4sU-treated samples were kept under reducing conditions using a reducing agent.

[0638] Alkylation reactions were performed using the SLAMseq Kinetics Kit - Synthalom Kinetics Module (Lexogen, catalog number 061, Vienna, Austria) using the "D1-containing 4sU" condition. Extracted total RNA was mixed with iodoacetamide (IAA) to modify the 4-thiol groups of 4sU-containing nucleotides by adding a carboxamidomethyl group, using the "D1-containing 4sU + alkylation" condition. This alkylation increased the frequency of T>C mismatches during reverse transcription.

[0639] RNA was then purified using ethanol precipitation prior to library preparation.

[0640] Library preparation and sequencing

[0641] The SMARTer Stranded Total RNA-Sequencing Kit-Pico Input Mammalian (ClonTech, Mountain View, USA) was used to construct libraries directly starting from 10 ng of RNA. Depletion of bacterial-derived ribosomal RNA (16S and 23S) was performed on total RNA using the RiboMinus Bacterial Transcriptome Analysis Kit (ThermoFisher). Prior to library preparation using the manufacturer's recommendations (ClonTech), ribosomal cDNA (included in the SMARTer Stranded Total RNA Sequencing Kit) was depleted using probes specific for mammalian rRNA and certain mitochondrial RNAs. Sequencing was performed on an Illumina instrument (Illumina, San Diego, USA) using the NextSeq 500 / 550 High Output Kit v2 (FC-404-2002, Illumina). Sequencing was paired with a read length of 150 nucleotides, generating approximately 100 million reads per sample.

[0642] Agnostic bioinformatics analysis

[0643] The raw data reads were filtered to select high-quality and relevant reads and classified to suppress or reduce duplicates, low-quality reads, and homopolymers (PathoQuest proprietary software).

[0644] Skewer was used to remove sequences introduced during the preparation of Illumina libraries (adapters, primers) (Jiang et al., 2014. BMC Bioinformatics. 15: 182).

[0645] The filtered reads of the negative control conditions (unlabeled, unactivated or both) are first considered as the target sequence. Since these conditions are likely to include a large number of sequences of the target organism, this allows the reconstruction of the genome of the target organism (no choleplasma or mycoplasma). Therefore, these reads are assembled into longer sequences, called "contigs" (Li et al., 2015. Bioinformatics. 31 (10): 1674-1676) using MegaHit. Then, the resulting contigs are matched back to the choleplasma or mycoplasma strain PG8A genome (reference sequence accession number CP000896.1) using minimap2 (Li, 2018. Bioinformatics. 34 (18): 3094-3100). Mummer 3 (Kurtz et al., 2004. Genome Biol. 5 (2): R12) is then used to tile the positive hits onto the choleplasma or mycoplasma strain PG8A genome so that:

[0646] 1. Confirm the identity of the contigs potentially tested as free of Cholesterolemia or Mycoplasma, and

[0647] 2. Ensure the integrity of the newly created sequence (hereinafter referred to as ALAID_CNS).

[0648] Estimation of T>C substitution rate

[0649] To detect acholestolapa or mycoplasma sequence with a very high number of T→C substitutions, the quality-filtered read set was matched back to ALAID_CNS (Li, 2018. Bioinformatics. 34(18): 3094-3100) using minimap2 in non-multiple match mode. All mismatches (base quality at least equal to 30) at each position of the ALAID_CNS sequence were then detected using the stacking module of the htsbox software (https: / / github.com / lh3 / htsbox). The global variation map was then analyzed using a dedicated script (PathoQuest, Paris, France) to define the substitution rate of each nucleotide. The ratio of the substituted nucleotides was compared to the total number of aligned nucleotides. For example, the "T→C" substitution rate was calculated using the following formula:

[0650]

[0651] The substitution rate at each time point was normalized using the following substitution index:

[0652]

Claims

1. An in vitro method for differentiating live microorganisms and dead microorganisms in a sample, which comprises differentiating microbial nucleic acid sequences having transcriptional activity and microbial nucleic acid sequences having transcriptional inertia in the sample, wherein the sample is selected from environmental samples, food samples or preservation media, and wherein the method comprises the following steps: (a) Sequencing an RNA pool extracted from the sample, wherein the RNA pool is obtained by culturing the sample in the presence of an RNA labeling agent and further by placing the extracted RNA under conditions allowing nucleotide substitution; thereby obtaining a set of sequence reads; (b) Identifying at least one microbial nucleic acid sequence hit by: (i) Aligning the sequence reads or their contigs with a database comprising microbial nucleic acid sequences, (ii) Identifying at least one microbial nucleic acid sequence hit that matches at least one sequence read or its contig, and (iii) Determining a consensus microbial nucleic acid sequence by the sequence reads that match the microbial nucleic acid sequences in the database; (c) Determining the number and / or ratio of substituted nucleotides in the set of sequence reads or contigs that match the microbial nucleic acid sequence hit of step (b)(ii) compared to the consensus microbial nucleic acid sequence determined in step (b)(iii); and (d) In the set of sequence reads or contigs that match at least one microbial nucleic acid sequence hit of step (b)(ii), if the number and / or ratio of nucleotides substituted from thymidine (T) to cytidine (C) in the set of sequence reads is greater than the average number and / or ratio of substitutions of all other nucleotides in the same sequence reads, then infer that the consensus microbial nucleic acid sequence hit belongs to live microorganisms; wherein the conditions allowing nucleotide substitution consist of chemically modified RNA with a chemical modification label; and further reverse transcribing the chemically modified RNA.

2. The in vitro method according to claim 1, wherein the RNA labeling agent is a thiol-labeled RNA precursor.

3. The in vitro method according to claim 2, wherein the thiol-labeled RNA precursor is selected from 4-thiouridine, 2-thiouridine, 2,4-dithiouridine, 2-thio-4-deoxyuridine, 5-ethoxycarbonyl-2-thiouridine, 5-carboxy-2-thiouridine, 5-(n-propyl)-2-thiouridine, 6-methyl-2-thiouridine and 6-(n-propyl)-2-thiouridine, thereby obtaining thiouridine-labeled RNA.

4. The in vitro method according to claim 1, wherein the RNA labeling agent is a thiol-labeled RNA precursor, and the thiol-labeled RNA precursor is 4-thiouridine.

5. The in vitro method according to claim 1, wherein the conditions allowing nucleotide substitution consist of chemically modifying RNA by alkylation, oxidative-nucleophilic-aromatic substitution or osmium-mediated transformation; and further reverse transcribing the chemically modified RNA.

6. The in vitro method according to claim 1, wherein the RNA labeling agent is a thiol-labeled RNA precursor, and the conditions allowing nucleotide substitution consist of alkylating RNA; and further reverse transcribing the alkylated RNA.

7. The in vitro method according to claim 6, wherein the thiol-labeled RNA precursor is 4-thiouridine.

8. The in vitro method according to claim 1, wherein the conditions allowing nucleotide substitution consist of alkylation using an alkylating agent selected from iodoacetamide, iodoacetic acid, N-ethylmaleimide and 4-vinylpyridine.

9. The in vitro method according to claim 1, wherein the step of sequencing RNA include: (i) reverse transcribing the RNA to obtain a cDNA library, and (ii) sequencing the cDNA library.

10. An in vitro method according to claim 1, wherein when the sample is cultured in the presence of an RNA labeling agent, and / or when the labeled RNA is placed under conditions that allow nucleotide substitutions, the RNA undergoes substitutions of adenine (A) to guanosine (G) in the first strand synthesis and thymidine (T) to cytidine (C) in the second strand synthesis during reverse transcription.

11. The in vitro method of claim 1, wherein the at least one microbial nucleic acid sequence hit is identified by: (i) filtering the sequence read groups, (ii) assembling sequence reads into contigs, (iii) aligning the sequence reads or contigs with a database comprising microbial nucleic acid sequences, (iv) identifying at least one microbial nucleic acid sequence hit that matches the at least one sequence read or contig, and (v) realigning the sequence reads or contigs with the microbial nucleic acid sequence hits identified in step (iv) to thereby determine a consensus microbial nucleic acid sequence, wherein the consensus microbial nucleic acid sequence corresponds to the microbial nucleic acid sequence hits.

12. The in vitro method according to claim 1, wherein the microorganism is selected from the group consisting of viruses, bacteria, archaea, fungi and protozoa.

13. Use of an RNA marker in the preparation of a product for detecting a microorganism in a sample from a subject, wherein the product is used by: (a) provide a sample of the subject, (b) performing an in vitro method on the sample, the method The following steps are involved: (a) sequencing an RNA set extracted from a sample, wherein the RNA set is obtained by culturing the sample in the presence of an RNA labeling agent and further by subjecting the extracted RNA to conditions that allow nucleotide substitution; thereby obtaining a sequence read set; (b) identifying at least one microbial nucleic acid sequence hit by: (i) aligning the sequence reads or contigs thereof with a database comprising microbial nucleic acid sequences, (ii) identifying at least one microbial nucleic acid sequence hit that matches the at least one sequence read or contig thereof, and (iii) determining a consensus microbial nucleic acid sequence by matching sequence reads with microbial nucleic acid sequences in a database; (c) determining the number and / or ratio of substituted nucleotides in the set of sequence reads or contigs matching the microbial nucleic acid sequence hits of step (b)(ii) compared to the consensus microbial nucleic acid sequence determined in step (b)(iii); and (d) in the set of sequence reads or contigs matching at least one microbial nucleic acid sequence hit of step (b)(ii), if the number and / or ratio of nucleotides substituted from thymidine (T) to cytidine (C) in the set of sequence reads is greater than the average number and / or ratio of substitutions of all other nucleotides in the same sequence reads, then inferring that the consensus microbial nucleic acid sequence hit belongs to a live microorganism; The conditions allowing nucleotide substitutions consist of chemically modifying the labeled RNA; and further reverse transcribing the chemically modified RNA.

14. A method for assessing the risk of microbial contamination in a sample, wherein include: (a) provide samples, (b) subjecting the sample to an in vitro method for distinguishing between live and dead microorganisms, and (c) if at least one of the identified microbial nucleic acid sequence hits belongs to a live microorganism, it is inferred that the sample is at risk of contamination; The in vitro method comprises the following steps: (a) sequencing an RNA set extracted from a sample, wherein the RNA set is obtained by culturing the sample in the presence of an RNA labeling agent and further by subjecting the extracted RNA to conditions that allow nucleotide substitution; thereby obtaining a sequence read set; (b) identifying at least one microbial nucleic acid sequence hit by: (i) aligning the sequence reads or contigs thereof with a database comprising microbial nucleic acid sequences, (ii) identifying at least one microbial nucleic acid sequence hit that matches the at least one sequence read or contig thereof, and (iii) determining a consensus microbial nucleic acid sequence by matching sequence reads with microbial nucleic acid sequences in a database; (c) determining the number and / or ratio of substituted nucleotides in the set of sequence reads or contigs matching the microbial nucleic acid sequence hits of step (b)(ii) compared to the consensus microbial nucleic acid sequence determined in step (b)(iii); and (d) in the set of sequence reads or contigs matching at least one microbial nucleic acid sequence hit of step (b)(ii), if the number and / or ratio of nucleotides substituted from thymidine (T) to cytidine (C) in the set of sequence reads is greater than the average number and / or ratio of substitutions of all other nucleotides in the same sequence reads, then inferring that the consensus microbial nucleic acid sequence hit belongs to a live microorganism; The conditions allowing nucleotide substitutions consist of chemically modifying the labeled RNA; and further reverse transcribing the chemically modified RNA.

15. Use of an RNA marker in the preparation of a product for distinguishing between live and dead microorganisms in a sample, wherein the product is used by the following steps: (a) sequencing an RNA set extracted from a sample, wherein the RNA set is obtained by culturing the sample in the presence of an RNA labeling agent and further by subjecting the extracted RNA to conditions that allow nucleotide substitution; thereby obtaining a sequence read set; (b) identifying at least one microbial nucleic acid sequence hit by: (i) aligning the sequence reads or contigs thereof with a database comprising microbial nucleic acid sequences, (ii) identifying at least one microbial nucleic acid sequence hit that matches the at least one sequence read or contig thereof, and (iii) determining a consensus microbial nucleic acid sequence by matching sequence reads with microbial nucleic acid sequences in a database; (c) determining the number and / or ratio of substituted nucleotides in the set of sequence reads or contigs matching the microbial nucleic acid sequence hits of step (b)(ii) compared to the consensus microbial nucleic acid sequence determined in step (b)(iii); and (d) in the set of sequence reads or contigs matching at least one microbial nucleic acid sequence hit of step (b)(ii), if the number and / or ratio of nucleotides substituted from thymidine (T) to cytidine (C) in the set of sequence reads is greater than the average number and / or ratio of substitutions of all other nucleotides in the same sequence reads, then inferring that the consensus microbial nucleic acid sequence hit belongs to a live microorganism; The conditions allowing nucleotide substitutions consist of chemically modifying the labeled RNA; and further reverse transcribing the chemically modified RNA.

Citation Information

Patent Citations

  • Synthesis of double-stranded nucleic acids

    CN106460052A

  • Attenuated reoviruses for selection of cell populations

    US20090104162A1