Commensal bacterial and viral nucleic acid detection

By using commensal bacterial and viral nucleic acid sequences as proxies, the method addresses the insensitivity of existing pathogen detection in indoor environments, enhancing the reliability of pathogen detection and enabling effective disinfection measures.

WO2026039519A1PCT designated stage Publication Date: 2026-02-19THE RGT UNIV OF MICHIGAN
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/041800
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-13
Filing Date
2025-08-13
Publication Date
2026-02-19

AI Technical Summary

Technical Problem

Existing methods for detecting respiratory pathogens in indoor environments are often insensitive and yield inconsistent results due to the low concentrations of pathogens, making it difficult to interpret the presence or absence of pathogens in air and surface samples.

Method used

Utilizing commensal bacterial and viral nucleic acid sequences normally present in the respiratory tract and oral cavity of humans as proxies to detect respiratory pathogens, employing assays and methods to sample air and surfaces in enclosed spaces, and taking actions based on the detection of these sequences.

Benefits of technology

Provides a reliable method to determine the presence of respiratory pathogens by detecting commensal biomarkers, allowing for informed decisions on patient entry or environmental disinfection measures, improving the accuracy of pathogen detection and intervention effectiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000036_0001
    Figure IMGF000036_0001
  • Figure IMGF000036_0002
    Figure IMGF000036_0002
  • Figure IMGF000021_0001
    Figure IMGF000021_0001
Patent Text Reader

Abstract

Provided herein are devices, systems, kits, and methods for sampling air and / or a surface in an enclosed space (e.g., hospital, clinic room, daycare room, classroom, workspace, or factory) with an assay for at least one human commensal bacterial and / or viral nucleic acid sequence normally present in the respiratory track and / or oral cavity of human, such that such assay serves as a proxy for detecting at least one human respiratory pathogen (which can be very hard to detect in such spaces as they can be at very low levels and still be relevant). Such commensal bacterial and / or viral nucleic acid detection can inform actions such as moving a patient into the enclosed space or taking actions to kill or remove microorganisms in the enclosed space (e.g., replacing HVAC filter that feeds into the enclosed space). Such commensal bacterial and or viral nucleic acid detection can serve as a proxy for how well the action to kill or remove pathogens is performing.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Attorney Docket Number: UM-43622.601

[0002] COMMENSAL BACTERIAL AND VIRAL NUCLEIC ACID DETECTION

[0003] CROSS REFERENCE TO RELATED APPLICATION

[0004] The present application claims priority to U.S. provisional application serial number 63 / 682,548 filed August 13, 2024, which is herein incorporated by reference in its entirety.

[0005] SEQUENCE LISTING PARAGRAPH

[0006] The text of the computer readable sequence listing filed herewith, titled “UM_43622_601_SequenceListing.xml”, created August 13, 2025, having a file size of 1,316,698 bytes, is hereby incorporated by reference in its entirety.

[0007] FIELD OF THE INVENTION

[0008] Provided herein are devices, systems, kits, and methods for sampling air and / or a surface in an enclosed space (e.g., hospital, clinic, classroom, daycare, workspace or factory) with an assay for at least one human commensal bacterial and / or viral nucleic acid sequence normally present in the respiratory track and / or oral cavity of human, such that such assay serves as a proxy for detecting at least one human respiratory pathogen (which can be very hard to detect in such spaces as they can be at very low levels and still be relevant). Such commensal bacterial and / or viral nucleic acid detection can inform actions such as moving a patient into the enclosed space or taking actions to kill or remove microorganisms in the enclosed space (e.g., replacing HVAC filter that feeds into the enclosed space).

[0009] BACKGROUND

[0010] Indoor air and surfaces are increasingly sampled for the presence of respiratory pathogens. The purpose of these measurements is to indicate the presence of a risk of infection and sometimes to test the effectiveness of non-pharmaceutical interventions (e.g., filters, increased ventilation). These measurements are rarely positive for pathogens due to the inherent low concentrations of pathogens in indoor air and on surfaces, and sometimes due to the poor sensitivity of the methods used to detect the pathogens. There are major issues with interpreting the inconsistent results from indoor air and surface measurements.

[0011] Biomarkers have been widely implemented for monitoring fecal contamination and for interpreting wastewater epidemiology data. These arc typically prokaryotes and viruses with Attorney Docket Number: UM-43622.601 high abundance and prevalence in human populations, such as coliforms (8), enterococci (8), crAssphagc (9, 10), and Pepper Mild Mottle Vims (11, 12). For example, the United States Environmental Protection Agency bases its fecal pollution indicator policies on the concentration of coliforms and enterococci (13). In wastewater-based epidemiology studies, Pepper Mild Mottle Vims and crAssphage normalize to the proportion of fecal matter in wastewater by accounting for transient changes in population (14). What is needed are biomarkers for human respiratory emissions.

[0012] SUMMARY OF THE INVENTION

[0013] Provided herein are devices, systems, kits, and methods for sampling air and / or a surface in an enclosed space (e.g., hospital, clinic, or classroom) with an assay for at least one human commensal bacterial and / or viral nucleic acid sequence normally present in the respiratory track and / or oral cavity of human, such that such assay serves as a proxy for detecting at least one human respiratory pathogen (which can be very hard to detect in such spaces as they can be at very low levels and still be relevant). Such commensal bacterial and / or viral nucleic acid detection can inform actions such as moving a patient into the enclosed space or taking actions to kill or remove microorganisms in the enclosed space (e.g., replacing HVAC filter that feeds into the enclosed space).

[0014] In some embodiments, provided herein are methods comprising: a) sampling air and / or a surface in an enclosed space with an assay for at least one human commensal bacteria or virus present in the respiratory track and / or oral cavity of humans such that either: i) a negative result is detected with the assay, or ii) a first positive level is detected with the assay, and b) performing at least one of the following: i) if the negative result is detect, then moving a patient into the enclosed space, wherein the enclosed space is a hospital room, clinic room, operating room, or medical procedure room or area, and / or allowing medical personal into the enclosed space, or ii) if the first positive level is detected, then treating the enclosed space air or surface with at least one of the following: A) replacing an air filter if an air filter present in the enclosed space, B) replacing an air filter in an HVAC system that provides air to the enclosed space, C) increasing air exchange in the enclosed space, D) exposing the enclosed space to UV radiation, E) wiping down surfaces in the enclosed space, and F) increasing humidity in the enclosed space.

[0015] In certain embodiments, the enclosed space is a hospital room, clinic room, operating room, medical procedure room or area, childcare room, or classroom. In other embodiments, the Attorney Docket Number: UM-43622.601 methods further comprise if the positive level is detected, repeating the sampling to generate cither a negative result or second positive level. In other embodiments, the methods further comprise: moving a patient into the enclosed space, wherein the enclosed space is a hospital room, clinic room, operating room, or medical procedure room or area; and wherein if the second positive level is detected it is at a level of less than 10% or less than 5%, of the first positive level; or allowing children to enter a classroom or common space. In other embodiments, the negative result is detected, and the negative result is employed as a proxy for very low or no detectable levels of pathogens. In additional embodiments, the first positive level is detected, and the first positive results is employed as a proxy for at least low levels, and possibly high levels, of pathogens.

[0016] In some embodiments, provided herein are methods comprising: a) sampling air and / or a surface in an enclosed space with a first assay for at least one human respiratory pathogen such that either: i) a negative result with the first assay is detected, or ii) a first positive level is detected with the first assay, and b) sampling the air and / or the surface in the enclosed space with a second assay for at least one human commensal bacteria or virus present in the respiratory track and / or oral cavity of humans such that either: i) a negative result is detected with the second assay, or ii) a first positive level is detected with the second assay, and wherein the sampling in b) is carried at, or about the same time, as the sampling in a).

[0017] In certain embodiments, the about the same time is within about 1 hour or less. In other embodiments, the sampling in a) produces the negative result in the first assay for the at least one human respiratory pathogen, and wherein the sampling in b) produces the negative result in the second assay for the at least one human commensal bacteria or virus, and wherein the method further comprises performing an action in the enclosed space that requires, or is best, when the human respiratory pathogens are not present, such as, but not limited to, i) generating a report that the enclosed space is free from the at least one human respiratory pathogen, ii) moving patient into the enclosed space, wherein the enclosed space is a hospital or clinic room; iii) repeating the sampling in a), optionally wherein the first assay is re-calibrated prior to repeating the sampling in a); iv) allowing medical personal into the enclosed space; and v) allowing children and / or students to enter a daycare room or classroom, or vi) allowing workers back into their workspace or factory.

[0018] In particular embodiments, the sampling in a) produces the negative result in the first assay for the at least one human respiratory pathogen, and wherein the sampling in b) produces Attorney Docket Number: UM-43622.601 the positive level in the second assay for the at least one human commensal bacteria or virus, and wherein the method further comprises: c) treating the enclosed space air or surfaces with at least one of the following: i) replacing an air filter in an air filter present in the enclosed space, ii) replacing an air filter in an HVAC system that provides air to the enclosed space, iii) increasing air exchange in the enclosed space, iv) exposing the enclosed space to UV radiation, v) wiping down surfaces in the enclosed space, and vi) increasing humidity level in the enclosed space.

[0019] In some embodiments, the methods further comprise: d) sampling the air and / or the surface in the enclosed space with the second assay for the at least one human commensal bacteria or virus such that: i) a negative result is detected with the second assay, or ii) a second positive level is detected with the second assay. In other embodiments, the sampling in d) produces the negative result, or produces the second positive level which is at least 2x, 3x, 4x, 5x, 6x, 7x, or 8x lower than the first positive level, and wherein the method optionally further comprises: e) moving a patient into the enclosed space, wherein the enclosed space is a hospital room or clinic room, or allowing children and / or student to enter a daycare room or classroom. In additional embodiments, the sampling in a) produces the first positive result in the first assay for the at least one human respiratory pathogen, and wherein the sampling in b) produces the negative result in the second assay for the at least one human commensal bacteria or virus, and wherein the method further comprises: c) repeating the sampling in a), optionally wherein the first assay is re-calibrated prior to repeating the sampling in a). In other embodiments, the sampling in a) produces the first positive result in the first assay for the at least one human respiratory pathogen, and wherein the sampling in b) produces the first positive level result in the second assay for the at least one human commensal bacteria or virus, and wherein the method further comprises: c) treating the enclosed space air or surfaces with at least one of the following: i) replacing an air filter in an air filter present in the enclosed space, ii) replacing an air filter in an HVAC system that provides air to the enclosed space, iii) increasing air exchange in the enclosed space, iv) exposing the enclosed space to UV radiation, and v) wiping down surfaces in the enclosed space.

[0020] In certain embodiments, the at least one human respiratory pathogen is selected from: Sars-CoV2, seasonal coronaviruses, RSV, parainfluenza virus, adenovirus, rhinovirus, and influenza virus. In additional embodiments, the at least one human commensal bacteria or vims is selected from: Strep, salivarius, Strep, subflava, Neisseria subflava, and H. influenzae phage HPlcl. In particular embodiments, the enclosed space is a workspace or factory. In other Attorney Docket Number: UM-43622.601 embodiments, the enclosed space is a hospital room or clinic room or other medical area. In certain embodiments, the first and / or second assay is performed in a laboratory. In other embodiments, the first and / or second assay is incorporated into a portable detection device.

[0021] In some embodiments, provided herein are systems and kits comprising: a) a first assay for sampling air and / or a surface to detect at least one human respiratory pathogen, and b) a second assay for sampling air and / or a surface to detect at least one human commensal bacteria or virus present in the respiratory track and / or oral cavity of humans. In further embodiments, the first and / or second assay is incorporated into a portable detection device. In other embodiments, the first and second assay are incorporated into the same portable detection device.

[0022] In some embodiments, provided herein are methods comprising: a) sampling air and / or a surface in an enclosed space with an assay for at least one human commensal viral or bacterial nucleic acid sequence present in the respiratory track and / or oral cavity of humans such that either: i) a negative result is detected with the assay, or ii) a first positive level is detected with the assay, and b) performing at least one of the following: i) if the negative result is detected, then moving a patient into the enclosed space, and / or allowing medical personal into the enclosed space, or ii) if the first positive level is detected, then treating the enclosed space air or surface with at least one of the following: A) replacing an air filter if an air filter present in the enclosed space, B) replacing an air filter in an HVAC system that provides air to the enclosed space, C) increasing air exchange in the enclosed space, D) exposing the enclosed space to UV radiation, E) wiping down surfaces in the enclosed space, and F) increasing humidity in the enclosed space.

[0023] In certain embodiments, the enclosed space is a hospital room, clinic room, operating room, medical procedure room or area, childcare room, classroom, workspace or factory. In other embodiments, the methods further comprises, if the positive level is detected, repeating the sampling to generate either a negative result or second positive level. In certain embodiments, the methods further comprise: moving a patient into the enclosed space, wherein the enclosed space is a hospital room, clinic room, operating room, or medical procedure room or area; and wherein if the second positive level is detected it is at a level of less than 10% of the first positive level; or allowing children to enter a classroom or common space; or allowing workers into their workspace or factory. In additional embodiments, when the negative result is detected, and the negative result is employed as a proxy for very low or no detectable levels of pathogens. In Attorney Docket Number: UM-43622.601 other embodiments, the first positive level is detected, and the first positive results is employed as a proxy for at least low levels, and possibly high levels, of pathogens.

[0024] In certain embodiments, the at least one human commensal viral nucleic acid sequence is selected from any one of SEQ ID NOs:l-124, or sequences with 95%, 96%, 97%, 98, or 99% identify over at least 15, 16, 17, 18, 19, or 20 consecutive nucleotides. In additional embodiments, the at least one human commensal bacterial nucleic acid sequence is selected from any one of SEQ ID NOs: 125-145. In further embodiments, the sampling comprises a detection method that employs at least one set of forward and reverse primers, and optionally a corresponding probe, found in SEQ ID NOs: 146-751 (e.g., the pairs of primers, and probes in Tables 3 and 4), or such sequence that are share at least 90%, or 95% identify with SEQ ID NOs 146-751 (over at least 15-20 nucleotides), or that are one or two nucleotide 5' and / or 3' truncations of SEQ ID NOs: 146-751. In further embodiments, the sampling comprises a detection method selected from PCR, quantitative PCR and nucleic acid sequencing.

[0025] In some embodiments, the at least one human commensal viral nucleic acid sequence comprises at least one NOVB_1 nucleic acid sequence. In additional embodiments, the at least one human commensal viral nucleic acid sequence further comprises a second viral nucleic acid sequence that comprises at least one OVB_1 nucleic acid sequence. In other embodiments, the at least one human commensal viral nucleic acid sequence comprises a third viral nucleic acid sequence that comprises at least one NVB_1 nucleic acid sequence. In further embodiments, the at least one human commensal viral nucleic acid sequence comprises at least one nucleic acid sequence from: NOVB_1, NOVB_2, NVB_1, NVB_2, OVB_1, OVB_2, OVB_3, OVB_4, OVB_5, OVB_6, OVB_7, and OVB_8, or any combination of two, three, four, five, six, seven, eight, nine, ten, eleven, or twelve nuclei acid sequences therefrom. In certain embodiments, the at least one human commensal bacterial nucleic acid sequence comprises a first bacterial nucleic acid sequence from: N. subflava, S. salivarius, and S. sanguinis, or multiple nucleic acid sequences from any combination thereof, or any combination with the 1 - 12 of the viral sequences immediately above.

[0026] In some embodiments, provided herein are methods comprising: sampling air and / or a surface, and / or water, and / or biological sample, which is optionally in an enclosed space, with an assay for at least one human commensal viral or bacterial nucleic acid sequence present in the respiratory track and / or oral cavity of humans, wherein the at least one human commensal viral nucleic acid sequence is selected from SEQ ID NOs:l-124 (and / or at least part of a sequence Attorney Docket Number: UM-43622.601 selected from SEQ ID NOs: 1-124 is detected as present in the sample), or a fragment thereof that is at least 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides, and wherein the at least one human commensal bacterial nucleic acid sequence comprises a sequence selected from SEQ ID NOs: 125- 145 (and / or at least part of a sequence selected from SEQ ID NOs: 125-145) and is detected as present in the sample), or a fragment thereof that is at least 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides. In other embodiments, the sampling comprises a detection method that employs at least one (or two, three, four, five, six, seven, or eight) set of forward and reverse primers, and optionally a corresponding probe, selected from the sets in SEQ ID NOs: 146-751 (see Tables 3 and 4). In other embodiments, the sampling comprises a detection method selected from PCR, quantitative PCR, and nucleic acid sequencing.

[0027] In particular embodiments, the sampling comprises a detection method that employs at least one(or two, three, four, five, six, seven, or eight) set of forward and reverse primers, and optionally a corresponding probe, found in SEQ ID NOs: 146-751, or such sequences with 95- 99% identity over at least 15-20 nucleotides (see sets in Tables 3 and 4). In other embodiments, the sampling comprises a detection method selected from: PCR, quantitative PCR, and nucleic acid sequencing.

[0028] In certain embodiments, the at least one human commensal viral nucleic acid sequence comprises at least one NOVB_1 nucleic acid sequence. In other embodiments, the at least one human commensal viral nucleic acid sequence further comprises a second viral nucleic acid sequence that comprises at least one OVB_1 nucleic acid sequence. In other embodiments, the at least one human commensal viral nucleic acid sequence further comprises a third nucleic acid sequence that comprises at least one NVB_1 nucleic acid sequence. In further embodiments, the at least one human commensal viral nucleic acid sequence comprises a first viral nucleic acid sequence that comprises at least one nucleic acid sequence from: NOVB_1, NOVB_2, NVB_1, NVB 2, OVB 1, OVB 2, OVB 3, OVB 4, OVB 5, OVB 6, OVB 7, and OVB 8. In other embodiments, the at least one human commensal bacterial nucleic acid sequence comprises a first bacterial nucleic acid sequence from: N. subflava, S. salivarius, and S. sanguinis.

[0029] In certain embodiments, provided herein are composition comprising: a nucleic acid molecule, wherein nucleic acid molecule comprises a sequence of at least 12 nucleotides in length that hybridizes under stringent conditions to any of SEQ ID NOS: 1-124 or complement thereof, and wherein the nucleic acid molecule: i) comprises a detectable label; ii) comprises at least one modified base; iii) is linked to a heterologous nucleic acid sequence, and / or iv) is linked Attorney Docket Number: UM-43622.601 to a purification tag (e.g., biotin). In particular embodiments, the sequence is at least 14, 15, 16, 17, 18, 19, 20, 21, or 22 nucleotides in length. In other embodiments, the nucleic acid molecule comprises RNA or DNA. In particular embodiments, the at least one modified base is selected from: 5-methylcytosine (5mC) and / or 5-hydroxylmethylcytosine (5hmC). In other embodiments, the detectable label comprises a fluorophore, and wherein the nucleic acid molecules optionally further comprises a quencher. In some embodiments, the heterologous nucleic acid sequence is selected from: a sequencing adapter, a sequencing sample index sequence, a universal primer binding site, a sequencing UMI barcode, and a flow cell tag sequence.

[0030] In certain embodiments, provided herein are compositions or kits comprising: a) a fist forward primer, and b) a corresponding first reverse primer, and c) optionally a corresponding first probe sequence, wherein the forward primer and corresponding reverse primer, and optionally the probe sequence, are found in SEQ ID NOs: 146-169 and 176-751, or such sequence that are share at least 90%, or 95% identify with SEQ ID NOs 146-169 and 176-751, or that are one or two nucleotide 5' and / or 3' truncations of SEQ ID NOs: 146-149 or 176-751. In other embodiments, the first forward primer and corresponding first reverse primer are configured to amplify a portion of: NOVB.l, NOVB_2, NVB_1, NVB_2, OVB_1, OVB_2, OVB_3, OVB_4, OVB_5, OVB_6, OVB_7, and OVB_8. In other embodiments, the first forward primer and corresponding first reverse primer are configured to amplify a portion of NOVB_1. In embodiments, the composition further comprises a second forward primer and corresponding second reverse primer configured to amplify a portion of OVB_1. In certain embodiments, the composition further comprises a third forward primer and corresponding third reverse primer configured to amplify a portion of NVB_1. In particular embodiments, the first forward primer comprises or consists of the sequence in SEQ ID NOs: 146, 177, 180, 183, or 186 (or such sequences that are share at least 90%, or 95% identify with SEQ ID NOs: 146, 177, 180, 183, or 186, or that are one or two nucleotide 5' and / or 3' truncations of SEQ ID NOs: 146, 177, 180, 183, or 186) and the corresponding first reverse primer comprises or consists of the sequence in SEQ ID NOs: 147, 178, 181, 184, and 187 (or such sequences that are share at least 90%, or 95% identify with SEQ ID NOs: 147, 178, 181, 184, and 187 or that are one or two nucleotide 5' and / or 3' truncations of SEQ ID NOs: 147, 178, 181, 184, and 187), and the optional probes are selected from SEQ ID NOs: 176, 179, 182, 185, and 188 (or such sequences that arc share at least 90%, or 95% identify with SEQ ID NOs: 176, 179, 182, 185, and 188 or Attorney Docket Number: UM-43622.601 that are one or two nucleotide 5' and / or 3' truncations of SEQ ID NOs: 176, 179, 182, 185, and 188).

[0031] DESCRIPTION OF THE FIGURES

[0032] Figure 1. This plot shows the proportion of measurements above limit of detection across sample types for pathogens and biomarkers. The numbers on each bar indicate the total number of measurements for context. The main patterns seen here are that in all sample types, the salivary biomarkers are detected significantly more frequently than pathogens. It is also seen that biomarkers are observed in near 20% of surface samples. This number is closer to 5% for NIOSH samples. Samples were more likely to have a pathogen present when a biomarker was detected. It is noted that NIOSH detection (National Institute for Occupational Safety and Health) generally refers to methods and procedures used to collect and analyze workplace air samples for chemical and physical hazards. These samples are used in health hazard evaluations and workplace safety assessments. NIOSH methods are designed to accurately measure substances in the air and provide data to help identify and control workplace risks.

[0033] Figure 2. This plot compares the concentration relative to the limit of detection between pathogens and commensal biomarkers, and across samples. In this figure, the air samples collected with a NIOSH were more frequently positive for commensal biomarkers than pathogens, while surface sample biomarker positives have more observations an order of magnitude higher.

[0034] Figure 3. Interesting patterns of biomarker positivity and concentrations are observed in NIOSH and surface samples. Positivity rate and the concentration of positives remained relatively stable throughout the seasons in surface samples. While NIOSH samples were slightly more variable. Notably, we saw a lower overall positivity of commensal biomarkers in the air during the winter. And we saw significantly much higher concentrations of commensal biomarkers in the spring.

[0035] Figure 4. We observe that in some cases the commensal biomarker positivity on surfaces and in air samples correlates with CO2 concentrations measured in the room over similar’ time periods. In some cases, these do not correlate. CO2 concentrations are often used for assessing air quality. These results suggest that CO2 concentrations will not always be good indicators of the presence of respiratory organisms in air and on surface samples.

[0036] Figure 5. We did not observe correlations between commensal biomarker concentrations in air and on surfaces. We anticipate that in general, concentrations of biomarkers arc a better Attorney Docket Number: UM-43622.601 indicator of respiratory pathogen exposure than CO2 due to the different transport mechanisms and detection methods of microorganisms (i.c. particulates) compared to gaseous compounds.

[0037] Figure 6 shows a summary of the process to identify viral biomarker candidates from healthy human metagenomes. (A) A random selection of metagenomes from the Human Microbiome Project was used to generate vOTUs (viral operational taxonomic units) and identify biomarker candidates. (B) A subset of the metagenomes had their assemblies downloaded from IMG to generate vOTUs. (C) Biomarker candidates were identified using machine learning modeling with the relative abundance of vOTUs across the metagenomes from the Human Microbiome Project.

[0038] Figure 7 shows a relative abundance of vOTUs selected as key features in machine learning models across the HMP metagenomes. (A) vOTUs are organized by hierarchical clustering with euclidean distances based on relative abundances across all HMP metagenomes. The vertical color bar denotes clusters of vOTUs indicative of particular sample origins where nasal, oral, stool, and skin are represented by purple, blue, yellow, and green, respectively. The vOTUs identified by more than one model were selected as viral biomarker candidates, indicated with a star and bolded name. (B) Logarithmic relative abundance of vOTUs in each sample, where dark purple indicates a vOTU was not present in a Human Microbiome Project metagenome. The metagenome sample origin is indicated by the x-axis color bar. (C) Heatmap indicating which vOTUs were selected during feature selection in each of the kernelized support vector machine learning models. The models had two classification options to determine sample origins: two classes (respiratory or non-respiratory) and three classes (nasal, oral, or non- respiratory). There were three sets of features used: all 1,232 vOTUs, 105 non-specific vOTUs removed, and 529 rare or non-specific vOTUs removed (Figure 6C).

[0039] Figure 8 shows the prevalence of viral biomarker candidates, best viral biomarker candidate cocktails, three saliva bacteria from Jung et al. (2018), and crAssphage are compared across the different sample origins. The best viral biomarker candidate cocktails included a two- vOTU cocktail (NOVB_1 and OVB_1) and a three-vOTU cocktail (NOVB_1, OVB_1, and NVB_1). The points indicate the percent of samples with a target present in non-respiratory compared to respiratory samples. The dashed 1:1 line divides which targets are more prevalent in respiratory or non-respiratory samples. Points are shaped by viral biomarker candidates, two- vOTU cocktail, three-vOTU cocktail, saliva bacterial biomarkers, crAssphage CPQ_056 amplicon, and all crAss-likc viruses as circles, triangles, squares, plus sign, crossed box, and Attorney Docket Number: UM-43622.601 asterisk, respectively. Points are colored based on the ratio of the prevalence in nasal compared to oral samples.

[0040] Figure 9 shows the abundance of viral biomarker candidates, best viral biomarker candidate cocktails, three saliva bacteria from Jung et al. (2018), and crAssphage are compared across the different HMP metagenome sample origins. The best viral biomarker candidate cocktails included a two-vOTU cocktail (NOVB_1 and OVB_1) and three-vOTU cocktail (NOVB_1, OVB_1, and NVB_1). Boxplot of logarithmic relative abundances for targets across the different sample origins where nasal, oral (buccal mucosa, saliva, and throat), skin, and stool are represented in purple, blue, green, and yellow, respectively. Any instances where a target was not present in a sample were not included.

[0041] Figure 10 shows the presence and abundance of human respiratory emission biomarker candidates in nasal swab (n=10) and saliva (n=10) samples. (A) Percent of samples present for each biomarker candidate where present is defined as a greater than one-fold increase from the negative controls. The dashed line denotes when at least half of the samples had a biomarker present. (B) Boxplot of the fold-increase from the negative controls amongst the samples with the biomarker present. The dashed line indicates a 10-fold increase in biomarker abundance compared to the negative controls. The gray regions contain results for the bacterial saliva biomarkers from Jung et al. (2018).

[0042] Figure 11 shows the prevalence results for all of the two-vOTU and three-vOTU combinations for cocktails of viral biomarker candidates. The points indicate the percent of samples with a target present in non-respiratory compared to respiratory samples. A dashed 1:1 line divides which targets are more prevalent in respiratory or non-respiratory samples with points above the line having a higher prevalence in respiratory samples. Points are shaped by the number of vOTUs in each cocktail with two-vOTU cocktails and three-vOTU cocktails as circles and triangles, respectively. Points are colored based on the ratio of the prevalence in nasal compared to oral samples. The red oval highlights the “best” two- and three-vOTU cocktails that are shown in Figure 8 and 9.

[0043] Figure 12 shows statistics from Wilcoxon’s test comparing biomarker abundances in respiratory samples. Resulting statistics of comparing the relative abundances of each biomarker in respiratory samples compared to non-respiratory samples (A). The statistics from comparing viral biomarker candidate relative abundances to mean crAss-like viruses relative abundance and mean saliva bacteria relative abundance in nasal (B) and oral (C) samples. Attorney Docket Number: UM-43622.601

[0044] DETAILED DESCRIPTION

[0045] Provided herein arc devices, systems, kits, and methods for sampling air and / or a surface in an enclosed space (e.g., hospital, clinic, or classroom) with an assay for at least one human commensal bacterial and / or viral nucleic acid normally present in the respiratory track and / or oral cavity of human, such that such assay serves as a proxy for detecting at least one human respiratory pathogen (which can be very hard to detect in such spaces as they can be at very low levels and still be relevant). Such commensal bacterial and / or viral nucleic acid detection can inform actions such as moving a patient into the enclosed space or taking actions to kill or remove microorganisms in the enclosed space (e.g., replacing HVAC filter that feeds into the enclosed space).

[0046] Indoor air and surfaces are increasingly sampled for the presence of respiratory pathogens. The purpose of these measurements is to indicate the presence of a risk of infection and sometimes to test the effectiveness of non-pharmaceutical interventions (e.g., filters, increased ventilation). These measurements are rarely positive for pathogens due to the inherent low concentrations of pathogens in indoor air and on surfaces, and sometimes due to the poor sensitivity of the methods used to detect the pathogens. There are major issues with interpreting the results from indoor air and surface measurements. For example, does a “non-detect” measurement (i.e., below detection limits) in the air of an inhabited room may reflect a scenario where: (1) no pathogens were released into the environment, (2) pathogens were released to the environment, but were diluted by diffusion, dispersion, or deposition, or (3) pathogens were released to the environment, but there were issues with the detection method. Measurements of the pathogen signal alone cannot differentiate these scenarios. To date, viral biomarkers for human or animal model respiratory material have not been developed. The human respiratory tract has a unique microbial community containing an abundance of bacteria and bacteriophages, as does the oral microbial community. In certain embodiments, provided herein are methods of using highly abundant bacteria and viruses nucleic acid sequence (e.g., at least one bacteria or vims) found in human respiratory and oral systems as biomarkers to inform air and surface measurements. For example, one may employ Strep, salivarius, Strep, subflava, Neisseria subflava, and H. influenzae phage HPlcl for this purpose. Measuring these biomarkers can help in improving air and surface sampling methods (e.g., changing methods leads to an increase in biomarkers), in developing nonpharmaceutical interventions (e.g., installing filters leads to fewer biomarkers in the air or on surfaces), as well as in interpreting human pathogen measurements Attorney Docket Number: UM-43622.601

[0047] (e.g., does a sample negative for human pathogens actually have any human respiratory material in it?).

[0048] In certain embodiments, the bacterial and viral nucleic acid sequence herein are detected by various biomolecules associated with these microorganisms. For example, one can detect the RNA / DNA genomes of the biomarkers or can detect proteins specific to the biomarkers (e.g., by sequencing or by antibody, or antibody fragment, testing). The signal of the biomarkers is an indicator of how much human commensal respiratory material is present in a samples and also on how well the human respiratory pathogen detection methods are working or how well a nonpharmaceutical intervention is working. Often, the pathogen signals are too low and inconsistent to test how well the sampling / detection methods are working (e.g. method 1 = concentration of 50 + / - 20 and method 2 = concentration of 75 + / - 25). But the commensal biomarkers signals are high enough and consistent enough to show this. Likewise, the human respiratory pathogen signals aren't consistent enough to show that say an air purifier is working (e.g. without purification concentration = 50 + / - 25 and with purification concentration = 40 + / - 20), but the abundant human oral and respiratory track biomarkers are feasibly high enough to show us that the purification is removing bacteria and viruses.

[0049] One other example of how they are useful when you are monitoring pathogens in air. Air monitoring for pathogens results in many negative results. One often wonders if this is because 1) there is nobody in the room exhaling the pathogen, or 2) there is someone in the room exhaling the pathogen and our sampling method wasn't effective at collecting viruses and bacteria exhaled from a person. If we detect the abundant human oral and respiratory track commensal bacterial and viral nucleic acid biomarkers, we have more confidence that the method collected human respiratory material and that there is no pathogen being exhaled.

[0050] Abundant Commensal Oral Bacteria are known in the art such as in Jung et al., Scientific Reports volume 8, Article number: 10852 (2018) (herein incorporated by reference). Examples of such bacterial include: Neisseria subflava, Streptococcus sanguinis, and Streptococcus salivarius. There are previously developed assays to detect the presence of saliva for forensic applications. It is noted that oral bacteria are more prevalent in the environment than pathogens, and pathogens are more frequently detected in biomarker positive samples.

[0051] The present disclosure is not limited with regard to how the human commensal viral and bacterial nucleic acid sequences (e.g., in Tables 3 and 4) are detected. In some embodiments, detection involves measurement or detection of a characteristic of a non-amplified nucleic acid, Attorney Docket Number: UM-43622.601 amplified nucleic acid, a component comprising amplified nucleic acid, or a byproduct of the amplification process, such as a physical, chemical, luminescence, or electrical aspect, which correlates with amplification (e.g. fluorescence, pH change, heat change, etc.). In some embodiments, fluorescence detection methods are provided for detection of amplified or nonamplified nucleic acid.

[0052] In certain embodiments, various detection reagents, such as fluorescent and non- fluorescent dyes and probes are employed. For example, the protocols may employ reagents suitable for use in a TaqMan reaction, such as a TaqMan probe; reagents suitable for use in a SYBR Green fluorescence detection; reagents suitable for use in a molecular beacon reaction, such as molecular beacon probes; reagents suitable for use in a scorpion reaction, such as a scorpion probe; reagents suitable for use in a fluorescent DNA-binding dye-type reaction, such as a fluorescent probe; and / or reagents for use in a LightUp protocol, such as a LightUp probe. In some embodiments, provided herein are methods and compositions for detecting and / or quantifying a detectable signal (e.g. fluorescence) from the human commensal viral and bacterial nucleic acid. Thus, for example, methods may employ labeling (e.g. during amplification, postamplification) amplified nucleic acids with a detectable label, exposing partitions to a light source at a wavelength selected to cause the detectable label to fluoresce, and detecting and / or measuring the resulting fluorescence. Fluorescence emitted from label can be tracked during amplification reaction to permit monitoring of the reaction (e.g., using a SYBR Green-type compound), or fluorescence can be measure post-amplification.

[0053] In some embodiments, detection of human commensal viral and bacterial nucleic acid sequences employs one or more of fluorescent labeling, fluorescent intercalation dyes, FRET- based detection methods (U.S. Pat. No. 5,945,283; PCT Publication WO 97 / 22719; both of which are incorporated by reference in their entireties), quantitative PCR, real-time fluorogenic methods (U.S. Pat. Nos. 5,210,015 to Gelfand, 5,538,848 to Livak, et al., and 5,863,736 to Haaland, as well as Heid, C. A., et al., Genome Research, 6:986-994 (1996); Gibson, U. E. M, et al., Genome Research 6:995-1001 (1996); Holland, P. M., et al., Proc. Natl. Acad. Sci. USA 88:7276-7280, (1991); and Livak, K. J., et al., PCR Methods and Applications 357-362 (1995), each of which is incorporated by reference in its entirety), molecular beacons (Piatek, A. S., et al., Nat. Biotechnol. 16:359-63 (1998); Tyagi, S. and Kramer, F. R., Nature Biotechnology 14:303-308 (1996); and Tyagi, S. et al., Nat. Biotechnol. 16:49-53 (1998); herein incorporated by reference in their entireties), Invader assays, (Neri, B. P., ct al., Advances in Nucleic Acid and Attorney Docket Number: UM-43622.601

[0054] Protein Analysis 3826:1 17- 125, 2000; herein incorporated by reference in its entirety), nucleic acid scqucncc-bascd amplification (NASBA; (See, e.g., Compton, J. Nucleic Acid Sequencebased Amplification, Nature 350: 91-91, 1991.; herein incorporated by reference in its entirety), Scorpion probes (Thelwell, et al. Nucleic Acids Research, 28:3752-3761, 2000; herein incorporated by reference in its entirety), partially double- stranded linear probes (Luk, K.-C., et al, J. Virological Methods 144:1-11, 2007; herein incorporated by reference in its entirety), capacitive DNA detection (See, e.g., Sohn, et al. (2000) Proc. Natl. Acad. Sei. U.S.A. 97:10687- 10690; herein incorporated by reference in its entirety), etc.

[0055] Target human commensal viral and bacterial nucleic acids sequences may be analyzed by any number of techniques to determine the presence of, amount of, or identity of the molecule. Non-limiting examples include sequencing, mass determination, and base composition determination. The analysis may identify the sequence of all or a part of the amplified nucleic acid or one or more of its properties or characteristics to reveal the desired information.

[0056] Any number of DNA sequencing techniques are suitable, including fluorescence-based sequencing methodologies (See, e.g., Birren et al., Genome Analysis: Analyzing DNA, 1, Cold Spring Harbor, N.Y.; herein incorporated by reference in its entirety). In some embodiments, the present disclosure finds use in automated sequencing techniques understood in that art. In some embodiments, the present technology finds use in parallel sequencing of partitioned amplicons (PCT Publication No: W02006084132, herein incorporated by reference in its entirety). In some embodiments, the technology finds use in DNA sequencing by parallel oligonucleotide extension (See, e.g., U.S. Pat. No. 5,750,341, and U.S. Pat. No. 6,306,597, both of which are herein incorporated by reference in their entireties). Additional examples of sequencing techniques in which the technology finds use include the Church polony technology (Mitra et al., 2003, Analytical Biochemistry 320, 55-65; Shendure et al., 2005 Science 309, 1728-1732; U.S. Pat. No. 6,432,360, U.S. Pat. No. 6,485,944, U.S. Pat. No. 6,511,803; all of which are herein incorporated by reference in their entireties), the 454 picotiter pyro sequencing technology (Margulies et al., 2005 Nature 437, 376-380; US 20050130173; herein incorporated by reference in their entireties), the Solexa single base addition technology (Bennett et al., 2005, Pharmacogenomics, 6, 373-382; U.S. Pat. No. 6,787,308; U.S. Pat. No. 6,833,246; herein incorporated by reference in their entireties), the Lynx massively parallel signature sequencing technology (Brenner et al. (2000). Nat. Biotechnol. 18:630-634; U.S. Pat. No. 5,695,934; U.S. Pat. No. 5,714,330; all of which arc herein incorporated by reference in their entireties), and the Attorney Docket Number: UM-43622.601

[0057] Adessi PCR colony technology (Adessi et al. (2000). Nucleic Acid Res. 28, E87; WO 00018957; herein incorporated by reference in its entirety). In certain embodiments, the library preparation and sequencing technologies are as described in any of the following U.S. patents, each of which is herein incorporated by reference: 9,752,188; 10,570,451; 11,479,807; 8,383,345; 10,876,172; 9,598,731; 9,902,992; 10,801,063; 11,091,797; 8,532,930; 9,639,657; and 10,011,870.

[0058] EXAMPLES

[0059] EXAMPLE 1 Establishing Viral Biomarkers of Human Respiratory Emissions from Oral and Nasal Metagenomes

[0060] Humans spend approximately 90% of their lives in built environments making virus transmission indoors a key determinant of health. Environmental sampling of respiratory viral pathogens is often challenging because of frequent non-detect measurements. Non-detect measurements do not differentiate between samples containing low or no pathogens from samples that simply lack respiratory expulsions altogether. This ambiguity can be resolved by scanning samples for a biomarker of human respiratory emissions. To do so, reliable biomarkers for environmental monitoring need to be identified. Ideal biomarkers are prevalent across individuals, abundant, and unique to the human respiratory tract. Here, we present viral biomarker sequences (e.g., viral SEQ ID NOs 1-124, and bacterial SEQ ID NOs: 125-145) identified from oral and nasal metagenomes of healthy individuals. Twelve viral biomarker candidates were selected for further analysis with a machine learning technique from 1,232 curated viral operational taxonomic units. The viral biomarker candidates had as much as 63% prevalence across respiratory metagenomes and prevalence was further increased to 77-81% by combining two or three biomarkers. Quantitative PCR confirmed that these viral biomarkers were prevalent and abundant in nasal swabs and saliva samples.

[0061] Developing non-pharmaceutical interventions to reduce virus transmission indoors relies on robust environmental monitoring methods. Monitoring viral pathogens is challenging because of frequent non-detect measurements that introduce uncertainty. For instance, a non- detect measurement could indicate either the absence of the pathogen or simply the lack of human respiratory activity and thus exposure. To aid in distinguishing these scenarios, this Example identifies viruses from the human respiratory tract that can be incorporated into Attorney Docket Number: UM-43622.601 environmental monitoring as biomarkers of human respiratory activity. These viral biomarkers will improve indoor monitoring to help enact interventions to mitigate vims transmission.

[0062] Some of most effective respiratory virus biomarker would be bacteriophages commonly found in the human respiratory tract, as their dissemination, persistence, and sampling recovery in air and surfaces would mirror those of viral pathogens. Many pathogens, including SARS- CoV-2 and influenza, have distinct tissue tropisms that result in different spatial patterns of shedding within the respiratory system (15). Therefore, it is helpful to identify biomarkers that represent different respiratory expulsion fluids, such as nasal mucus and saliva. Further, it is preferable that viral biomarkers are unique to the human respiratory tract and not found in other human samples, such as in stool or on skin. Lastly, as is the case for crAssphage and Pepper Mild Mottle Vims in stool samples, preferable viral respiratory biomarkers will be highly prevalent and abundant in humans, so they are able to be identified consistently in environmental samples.

[0063] This Examples describes an approach to identify viral biomarker candidates from existing healthy human metagenomes collected as part of the Human Microbiome Project (16). We created viral operational taxonomic units (vOTUs) from nasal and oral metagenomes, then applied various machine learning approaches that balanced biomarker candidate uniqueness to the human respiratory tract with prevalence and abundance in human oral and nasal metagenomes. In this Example, twelve viral biomarker candidates were then specifically quantified in saliva and nasal swabs with quantitative PCR (qPCR) and compared to the abundance of three prevalent and abundant saliva bacteria populations previously proposed for use in forensics (17). These viral biomarkers can be applied in environmental sampling studies and other studies.

[0064] MATERIALS AND METHODS

[0065] Healthy Human Metagenome Dataset Download and Quality Control

[0066] A collection of healthy human metagenomes from the Human Microbiome Project were downloaded from NCBI DbGaP (Project Accession phs000228). A subset of available metagenomes was randomly selected using the sample_n function (dplyr, vl.1.4). Samples representing the respiratory tract included 108 anterior nares (i.e., nasal), 108 buccal mucosa, 8 saliva, and 19 throat metagenomes non-rcspiratory tract samples included 97 stool and 61 Attorney Docket Number: UM-43622.601 retroauricular crease skin metagenomes (Figure 6A). Of the metagenomes selected, samples were collected from 164 healthy individuals with 48% (n = 78 / 164) of individuals identifying as female and ages ranging 18-40 years old. Quality control of reads was performed by trimming Illumina adaptors, then reads were decontaminated of PhiX174 with BBDuk (BBTools, v37.64). Bases with quality scores less than 10 were trimmed from reads, then trimmed reads with quality scores less than 10 or lengths less than 100 bp were removed with BBDuk. Human reads were removed using the Human Host Filtration Pipeline (18).

[0067] Generation of Viral Operational Taxonomic Units (vOTUs)

[0068] Assemblies from 49 random respiratory metagenomes (19 nasal, 27 buccal mucosa, 2 throat, and 1 saliva) were downloaded from IMG to curate viral operational taxonomic units (vOTUs) (Figure 6B). To assess likelihood of each contig being viral, six viral detection methods were run in sequence on contigs longer than 2,000 bp: VirSorter (vl.0.6) (19), VirSorter2 (v2.2.2) (20), VIBRANT (vl.2.1) (21), DeepVirFinder (vl.0) (22), Kaiju (vl.8.0) (23), and CheckV (vl.0.3) (24). Potential viral contigs were identified from these results using previously established rules (25). The viral contigs were binned using vRhyme (v 1.1.0) (26) with a minimum contig length of 2,000 bp. The vRhyme bins labeled as “best” bins or circular viral contigs were retained. Viral bins from all of the samples were pooled and dereplicated with dRep (v3.4.3) (27) if bins shared greater than 95% average nucleotide identity and greater than 85% coverage of the shortest contig to form vOTUs (28). Viral sorting, binning, and dereplicating resulted in 1,232 vOTUs (Figure 6B).

[0069] Read Mapping to vOTUs, Bacterial Saliva Biomarkers, and Stool Biomarkers

[0070] Read mapping was performed to assess the prevalence and uniqueness of vOTUs and other targets amongst the selected healthy human metagenomes (Figure 6C). We compared the prevalence and abundance of viral biomarker candidates to three commensal oral bacteria and a commonly used fecal biomarker, crAssphage. The three saliva bacteria and crAssphage are known to be highly prevalent and abundant in their respective matrices. The three saliva bacteria were Streptococcus salivarius, Streptococcus sanguinis, and Neisseria subflava (17). To map reads to saliva bacteria, the 3,000 bp region of their genomes including and surrounding previously developed PCR amplicon sequences were downloaded from NCBI (see bacterial SEQ ID NOs: 125-145). The RcfScq crAss-likc vims database were downloaded to have a reference Attorney Docket Number: UM-43622.601 set of dsDNA viruses commonly used as stool biomarkers (9). Often, a PCR primer assay, CPQ_056 (29), is used in environmental sampling that captures a subset of crAssphagc (i.e., NCBI Accessions MW063138.1, MW067003.1, MW067002.1, MW067001.1, MW067000.1, MT006214.1, MK415410.1, MK415408.1, MK415404.1, MK415403.1, MK238400.1, NC_024711.1, BK049789.1, MZ130481.1, and MK415399.1). Five Bowtie2 (v2.4.2) indexes were built with default parameters for vOTUs, N. subflava, S. salivarius, S. sanguinis, and crAss- like virus targets. Read mapping with deinterleaved fastq formatted files of QC short reads was performed using Bowtie2 with the default mapping parameters. The number of basepairs mapping to each target was determined from Bowtie2 sam file output using flagstat (samtools, vl.l l). The relative abundance of each target was calculated as the ratio of basepairs mapping to each target divided by basepairs in a sample.

[0071] Biomarker Candidate Selection Using Machine Learning

[0072] To identify vOTUs as candidates for human respiratory emission biomarkers, we used machine learning to select vOTUs that best distinguished respiratory samples from stool and skin samples (Figure 6C). We used supervised classification machine learning with kemelized support vector machines (Python v3.11) to handle our high dimensional data (30). In total, we created six different models by varying classification structure and logically excluding vOTU features. We utilized two classification structures: (1) respiratory (i.e., nasal, buccal mucosa, saliva, and throat) or non-respiratory (i.e., skin and stool) samples; and (2) nasal, oral (i.e., buccal mucosa, saliva, and throat), and non-respiratory samples. There were three sets of vOTU features used in our machine learning models: (1) all 1,232 vOTUs created; (2) removing nonspecific vOTUs that excluded 105 vOTUs present in more than 80% of stool or skin samples; and (3) removing non-specific vOTUs and rare vOTUs that excluded 529 vOTUs present in more than 80% of stool or skin samples and present in less than 20% of respiratory samples. Models were created based on the relative abundances of vOTUs in each sample. The dataset was divided into training and test sets composed of 279 and 94 samples, respectively. Model accuracy measured the frequency that metagenome origin was correctly predicted. Ten vOTU features that best predicted sample origin classification were identified with forward feature selection. Model accuracy was measured again after feature selection. vOTUs selected as features by more than one model were chosen as viral biomarker candidates of human respiratory emissions. The biomarker candidates were combined to form “cocktails” to assess if having Attorney Docket Number: UM-43622.601 mixes of multiple targets would improve prevalence and abundance in respiratory samples.

[0073] Every combination of two and three viral biomarkcr candidates was assessed for prevalence and abundance in respiratory and non-respiratory samples (Figure 11). Confirming Biomarker Candidates are Viruses

[0074] To confirm that all of the selected biomarker candidates are viruses, the genomic content was assessed for each contig comprising the vOTUs (Table 1) and biomarkers were assessed in pooled saliva purified for viral particles.

[0075] TABLE 2 Attorney Docket Number: UM-43622.601

[0076] Two methods were utilized to assign taxonomy and functional potential. Taxonomy and functional potential were assigned with geNomad (vl.l l with database vl.9) (31) run end-to-end that performs assignments by combining a neural network-based classifier and protein markerbased classifier. An alignment approach to taxonomic assignment was also run with megaBLASTn to viruses in the nucleotide database. Functional annotation was also performed with GhostKoala (v3.1 ) (32) to the KEGG database using amino acid sequences generated with prodigal (v2.6.3) using the meta flag. Viruses are known to have compact, efficient genomes Attorney Docket Number: UM-43622.601 with high coding density (33). The coding density of vOTUs was calculated as the total sequence length within open reading frames divided by the total sequence length. If taxonomy, functional potential annotation, and coding density resulted in uncertainty, further alignments were performed with BLASTn to the core nucleotide, viruses in the whole genome shotgun sequencing contigs, and bacteria in the RefSeq Genome databases. Once the biomarker candidates were confirmed as viruses, the vOTUs were renamed based on where their prevalence is highest in respiratory environments (O = Oral; N = Nasal) and ranked based on prevalence and abundance in the metagenomic and qPCR datasets.

[0077] Saliva and Nasal Swab Collection and DNA Extraction

[0078] Saliva and nasal swabs were collected in the evening from ten individuals with no selfidentified chronic respiratory diseases over the age of 18. Saliva samples were collected with SDNA-1000 saliva collection kits (Spectrum Solutions) and nasal swabs were collected with Quickvue Influenza Nasal Swab Tubes (Quidel). The study and all associated documents and protocols were approved by the Institutional Review Board (IRB-HSBS) at the University of Michigan (HUM00241431). Samples were stored at -80°C for a maximum of 36 days until DNA extraction. Duplicate extractions were performed on each sample. Immediately prior to DNA extraction, samples were thawed on ice and 1 mL of PBS with 0.5% bovine serum albumin added to nasal swabs. DNA was extracted from duplicate 200 L saliva and 200 pL nasal PBS solution with a Kingfisher Flex instrument equipped with a 96-well attachment. The Applied BiosystemsTM MagMAXTM Viral / Pathogen II Nucleic Acid Isolation Kit (Fisher Scientific Cat. No. A48383) with two wash cycles was performed with 50 pL elutions. DNA extracts were stored at -20°C for a maximum of one week until qPCR was performed.

[0079] Biomarker Candidate PCR Primer Design and qPCR Reaction Conditions

[0080] The abundance of viral biomarker candidates and bacterial saliva biomarkers were assessed in saliva and nasal swab DNA extracts using qPCR. Primers for the twelve vOTUs selected as viral biomarker candidates were designed using the IDT PrimerQuestTM tool. The best primer assay for each vOTU was selected to maximize specificity. Specificity of each primer assay was assessed with NCBI Primer BLAST to the nr database where no or few hits indicated high specificity. Selected viral biomarker candidates’ primers and previously Attorney Docket Number: UM-43622.601 developed saliva bacterial biomarkers (17) primers are provided in Table 3.

[0081] TABLE 3

[0082] The qPCR threshold cycle (Ct) for each biomarker in the DNA extracts were compared to ddH20 negative controls (NTC). qPCR reactions of duplicate saliva and nasal swab sample DNA extracts and triplicate virus purification experimental replicates were performed with the QuantStudio 3 thermocycler (Thermo Fisher Scientific, Inc.). For each plate, one saliva and nasal sample were analyzed with two ddH2O negative controls for each primer set. The 20 pL reactions were prepared with 10 pL of Luna Universal qPCR mastermix (New England Biolabs, Cat. No. M3003), 0.5 pM of forward and reverse primers, and 5 pL of 1:10 diluted template. The reactions consisted of initial denaturation at 95°C for 60 seconds followed by 45 cycles of denaturation for 15 seconds at 95°C, then annealing and extension for 30 seconds at 60°C. The mean Ct value for duplicate reactions was divided by the mean Ct value of the duplicate NTC to determine the fold increase from NTC. Attorney Docket Number: UM-43622.601

[0083] Statistical Analysis

[0084] All statistical analysis and figure creation was performed with R (v4.4.0) using ggplot2 (v3.5.1) and plotly (v4.10.4). Normality was tested with the Shapiro-Wilkes test. Wilcoxon’s test with Benjamini-Hochberg’s correction was performed on non-normal datasets.

[0085] RESULTS

[0086] Oral metagenomes generated more vOTUs than nasal metagenomes

[0087] To identify viruses pervasive in human respiratory emissions, we curated 1,232 vOTUs from 49 healthy human oral and nasal metagenomes (Figure 6B). The oral samples (i.e., buccal mucosa, saliva, and throat) had significantly more assembled viruses than nasal samples with oral and nasal assemblies having an average of 45 viral genomes and two viral genomes, respectively (p-value = 1.4xl0‘7). A previous analysis of viruses in samples from the Human Microbiome Project identified significantly more vOTUs in oral than nasal samples (34). This observation may be due to fewer viruses comprising the nasal microbiome or issues with the sequencing data quality. Here, sequencing depth after quality control and human read filtration was significantly greater in oral metagenomes than nasal metagenomes (p-value = 2.6x10-5). The oral assemblies also contained significantly more contigs longer than 1,000 bp than nasal assemblies (p-value = 4.9x10-7). Despite the nasal assemblies resulting in fewer contigs, the median contig lengths were not significantly different between oral and nasal assemblies (p- value = 0.26), indicating similar assembly quality between oral and nasal metagenomes. Previous reports comparing oral and nasal microbiomes are inconsistent regarding if the oral microbiome is more diverse than the nasal microbiome (35-38). Previous studies have found diverse viral communities in the oral cavity (34, 39-42). However, the limited prior work on viruses in nasal cavities resulted in few vOTUs suggesting low viral richness (34, 43). vOTUs were selected as viral biomarker candidates for deeper analysis

[0088] Viral biomarker candidates of human respiratory emissions were identified from the vOTUs, the sequences of which are shown in Table 4 below. Attorney Docket Number: UM-43622.601

[0089] TABLE 4 Attorney Docket Number: UM-43622.601 Attorney Docket Number: UM-43622.601 Attorney Docket Number: UM-43622.601 Attorney Docket Number: UM-43622.601 Attorney Docket Number: UM-43622.601 Attorney Docket Number: UM-43622.601

[0090] Preferable viral biomarkers are highly abundant and prevalent in respiratory samples and not found in other human environments, such as in stool or on skin. The identification of preferable candidates was based on: (1) identification of vOTUs that were indicators of oral, nasal, skin, or stool samples with machine learning approaches, (2) evaluation of the abundance and prevalence of vOTUs across target (i.e., oral and nasal) and non-target (i.e., stool and skin) body sites and the consistency of model vOTU selections, and (3) agreement across models as to the best choices of respiratory vOTU biomarker candidates.

[0091] Six machine learning models were applied to predict the origin of the 400 metagenomes based on vOTUs relative abundances across respiratory (i.e., oral and nasal) and non-respiratory (i.e. stool and skin) metagenomes. We performed modeling using three sets of features: all 1,232 vOTUs, removal of non-specific vOTUs (i.e., vOTUs present in more than 80% of stool or skin samples), and removal of non-specific vOTUs and rare vOTUs (i.e., vOTUs present in less than 20% of oral and nasal samples). The accuracy of the machine learning models ranged from 0.49- 0.85, which indicated that our models were not overfitted to the training data. With this result, we proceeded to next apply feature selection to identify the ten vOTUs that best predicted sample origin from each model. The models were rerun with the selected ten vOTUs and the models’ accuracy range improved to 0.70-0.93. Across the six models, 40 vOTUs (see Table 2) were selected by at least one model as key features (Figure 7 A).

[0092] We next evaluated the abundance and prevalence of the 40 potential biomarkers in respiratory (i.e., nasal and oral cavities) and non-respiratory (i.e., stool and skin) locations to Attorney Docket Number: UM-43622.601 characterize which environments each vOTU inhabits. Of the 40 candidate vOTUs, 24 were highly abundant and prevalent in the oral environment exclusively (Figure 7A; Table 5).

[0093] TABLE 5 Attorney Docket Number: UM-43622.601

[0094] Although published studies on the viral component of these human microbiomes are sparse, previous research on bacterial communities has similarly shown that the oral microbiome is distinct from the other human microbiomes (35, 44). We identified three candidate vOTUs that were prevalent in nasal and oral samples (Figure 7A and B). Here, three vOTUs were highly unique to nasal (i.e., 79422_vRhyme_3) or oral samples (i.e., 0VB_5 and 0VB_6), but were not considered prevalent in human environment (Figure 7A and B). However, a few of the candidate vOTUs were poor candidates for biomarkers of respiratory emissions given their prevalence on skin or in stool: five vOTUs were prevalent in nasal and skin samples with five other candidate vOTUs prevalent and abundant in oral and stool samples (Figure 7 A and B). Although these skin and stool vOTUs aided models in determining sample origin, they are not generally preferred respiratory biomarker candidates. The identification of these vOTUs prevalent in stool reinforced the decision to incorporate logical exclusion criteria in four of the models to exclude vOTUs that were highly prevalent in the non-target stool or skin samples (Figure 7C). The overlap of vOTUs in skin and nasal metagenomes aligns with previous findings of similar microbiota in skin and nasal samples (35). Some correspondence in the distribution of stool and oral vOTUs is not surprising, given prior observations of overlap between the oral and distal colon microbial communities (45) that has given rise to hypotheses that the oral microbiome may seed the gastrointestinal tract (46). Our data suggest this also may hold for human virome dispersal, as some vOTUs were prevalent and abundant in both oral and stool samples. Attorney Docket Number: UM-43622.601

[0095] In our pursuit of candidate respiratory emission viral biomarkers, we next looked for agreement between the machine learning models. Across the six models, 12 vOTUs were selected as key features by more than one model (Figure 7C). The two most often selected vOTUs, NOVB_1 and OVB_2, were identified by five and four models, respectively. NOVB_1 and OVB_2 were commonly found in oral samples with NOVB_1 also found in nasal samples. The remaining vOTUs were selected by two (n = 7) or three (n = 3) models with greatest prevalence in oral samples (n = 4), nasal and skin samples (n = 2), nasal and oral samples (n = 1), or oral and stool samples (n = 1). Two vOTUs, OVB_5 and OVB_6, were highly specific to oral samples although were rarely present in the metagenomes. Given that these 12 vOTUs were identified by more than one model, all were further evaluated as viral biomarker candidates in this example.

[0096] Biomarker candidate genomes demonstrate evidence of viral origins

[0097] We next evaluated whether the 12 biomarker candidates identified were of viral, rather than bacterial or human, origins. All contigs that comprised the 12 candidate vOTUs were confidently identified as being of viral origin based on a previously developed and rigorously evaluated viral contig sorting algorithm (25). To increase our confidence in this viral assignment, we sought additional evidence, including (1) viral taxonomy assignment, (2) identification of known viral proteins on the contigs, or (3) a high coding density characteristic of viral, rather than cellular, genomes. Based on our taxonomy assignment approach, all biomarker candidates were assigned as viruses by at least one method, except for NOVB_1 and OVB_1. Viral proteins, or those known to be associated with viral functions, were identified on the contigs of most of the biomarker candidates, with the exception of NOVB_1, OVB_4, and OVB_6. The coding density of the vOTUs was high (83-99%) compared to that expected of the human genome (1-2%) (47) for eleven of the biomarker candidates, with the exception of NOVB_1, which had a coding density of 16%. These observations supported the hypothesis that these contigs were of non-human origin, given that viruses and bacteria have compact, efficient genomes with high coding densities (33, 48). With the exception of NOVB_1, all candidate biomarkers fit expectations of viral origins.

[0098] We further evaluated the contigs comprising the NOVB_1 vOTU to determine whether the putative vOTU was misassigned as viral, or if instead it represented novel, not yet characterized viral diversity, referred to as “viral dark matter” (49, 50). Methods for identifying Attorney Docket Number: UM-43622.601 viral dark matter rely on deep learning methods, such as DeepVirFinder (51), and alignments to viral contigs from shotgun sequencing (52). DeepVirFinder indicated that both NOVB_1 contigs were of viral origin (p-values = 0.032 and 0.033). When NOVB_1 was aligned to the core nucleotide database (BLASTn), it aligned best to the Homo sapiens chromosome 5 (NCBI accession OZ171101.1; bit score = 10,761). However, we also searched for homology in the whole genome sequencing database (NCBI WGS) and identified that the NOVB_1 contigs aligned best to two assembled phage contigs (bit scores = 73.4 and 66.2). Based on this evidence, combined with its high coding density (16%) relative to that expected of the human genome (1-2%) (47), we concluded that NOVB_1 is a novel phage sequence that represents undescribed and uncharacterized sequence space of “viral dark matter”. We next evaluated the prevalence and abundance of the 12 vOTUs biomarker candidates across the metagenomes to simulate a quantitative survey across body sites.

[0099] Most viral biomarker candidates were more prevalent and abundant in respiratory samples than non-respiratory samples

[0100] Most of the viral biomarker candidates had a greater prevalence in respiratory samples (nasal and oral) than non-respiratory samples (stool and skin) (n = 10 / 12; Figure 8; Table 6).

[0101] TABLE 6

[0102] OVB_5 0.06 8.7 3.2

[0103] OVB_6 0.00 5.8 0.0

[0104] NVB_1 2.00 5.8 2.5

[0105] NVB_2 2.00 6.2 22.2

[0106] NOVB_1 0.72 55.0 28.5

[0107] NOVB_2 0.32 63.2 39.2

[0108] OVB_8 0.12 52.1 36.1

[0109] OVB_2 0.16 54.5 18.4

[0110] OVB_3 0.10 50.0 10.8

[0111] OVB_1 0.06 48.3 5.7

[0112] OVB_7 0.04 41.3 8.9

[0113] OVB_4 0.06 34.3 49.4

[0114] 2-vOTU Cocktail (NOVB 1, OVB_1) Attorney Docket Number: UM-43622.601

[0115] 3-vOTU Cockta i NOVB 1, „ „„ „ „„ „

[0116] - 0.59 81.0 32.9

[0117] OVB_1, NVB_1)

[0118] W. subflava 0.13 47.9 6.3

[0119] S. salivarius 0.11 56.2 31.6

[0120] S. sanguinis 0.10 54.1 13.9 crAssphage CPQ_056 0.80 47.5 66.5 crAss-Uke viruses 0.72 66.5 90.5

[0121] Seven of the viral biomarkers were present in more than 40% of respiratory samples (range = 41- 63%) and were less commonly found in non-respiratory samples (range = 6-39%). Of these biomarker candidates, five were specific to oral samples and two were found in both oral and nasal samples (Table 5). Both of the viral biomarkers not observed in oral samples were relatively rare (less than 10% prevalence) across respiratory samples. In total, four of the viral biomarkers (OVB_5, OVB_6, NVB_1, NVB_2) were present in less than 10% of respiratory samples (Figure 8) and were selected by the machine learning models without the rare vOTUs roughly excluded (Figure 7C; Table 5).

[0122] All of the viral biomarker candidates had a greater mean relative abundance in respiratory samples than non-respiratory samples (Figure 9; Figure 12A). For ten of the viral biomarkers, the mean relative abundance in respiratory samples was significantly greater than non-respiratory samples (p-values < 7xl0'4). For two of the relatively rare viral biomarkers, OVB_5 and NVB_2, the mean relative abundance was not significantly greater in respiratory samples (p-values = 0.71 and 0.27, respectively). The best performing viral biomarker (in terms of abundance), NOVB_1, had the greatest mean relative abundance in oral (0.0106%; n = 84 / 108) and nasal (0.0585%; n = 49 / 108) samples (Figure 9; Figure 12A).

[0123] Viral biomarker candidates had similar prevalence and abundance to accepted standards In order to benchmark our observations with accepted standards in the field, the prevalences of viral biomarkcr candidates were compared to three bacterial species saliva biomarkers (17) and a commonly used viral fecal biomarker, cr Assphage (9). The prevalence of the five viral biomarkers specific to oral samples was similar to that of the saliva bacteria biomarkers (Figure 8-9). CrAssphage was highly prevalent in respiratory and non-respiratory samples, with the prevalence of all crAssphage exceeding the most prevalent viral biomarker candidate, NOVB_2, in respiratory metagenomes (Figure 8). crAssphage CPQ_056 primers (29) and all crAss-like viruses were not specific to stool with high prevalence across all human microbiome samples. Attorney Docket Number: UM-43622.601

[0124] Previously, Tisza et al. (2021 ) found that most viruses are prevalent at only one body site with approximately 5% viruses prevalent in at least two sites, including some crAss-likc viruses (34). To assess whether the observed relative abundances of viral biomarker candidates were ample in oral or nasal environments, we compared their abundance to the saliva bacteria biomarkers and crAssphage. The three bacteria species are known to be abundant in saliva (17, 35, 46) and crAssphage is the most abundant phage in stool (53). In comparison to saliva bacteria biomarkers and crAssphage in respiratory samples, the viral biomarkers were of similar relative abundances (Figure 9).

[0125] When combined, the saliva bacteria biomarkers were more prevalent in oral (n=133 / 134) than nasal (n=23 / 108) samples. However, when the saliva bacteria biomarkers were present, their mean relative abundances were similar in oral and nasal samples (0.00311% and 0.00207%, respectively). These results indicate that the saliva bacteria are abundant in both oral and nasal environments when present. Two of the viral biomarkers were significantly more abundant in nasal samples than the mean abundance of the three saliva bacteria biomarkers (p-values = 8.7xl0‘9and 1.3xl0’5for NOVB_1 and NVB_1, respectively; Figure 12B) and the other nine viral biomarkers present in nasal samples were not statistically different in abundance. Whereas two of the viral biomarkers were significantly more abundant than the bacterial biomarkers in oral samples (p-values = 0.0032 and 5.5xl0'18for NOVB_1 and NOVB_2, respectively; Figure 12B). In summary, several viral biomarker candidates were highly abundant in either nasal or oral cavities, with NOVB_1 abundant in both environments, indicating it is a particularly good marker for respiratory monitoring.

[0126] CrAss-like viruses were highly prevalent and abundant in oral (n=102 / 134; 0.00992%) and nasal (n=59 / 108; 0.0742%) samples (Figure 9; Figure 12B and C). Only NOVB_1 was significantly more abundant than crAss-like viruses in oral samples (p-values = 2.1xl0‘4); however, NOVB_1 was significantly less than crAss-like viruses in nasal samples (p-values = 0.0013). All other viral biomarker candidates were either significantly less than crAss-like viruses or similar in relative abundance in nasal and oral samples. CrAss-like viruses were previously shown to be present in multiple body sites (34). However, their high abundance may be partially due to their abundance comprising multiple populations as opposed to a specific single vOTU, as is the case for the viral biomarker candidates. Attorney Docket Number: UM-43622.601

[0127] Combining viral biomarkers into cocktails increases their prevalence in respiratory samples

[0128] While single viral biomarker candidates showed high prevalence and abundance, we also explored the potential benefits of combining multiple viral biomarker candidates into cocktail mixtures. In practice, these cocktails could be measured by performing separate assays for each biomarker using qPCR and then summing the results. Alternatively, a single assay could be conducted using a mixture of the primers, allowing the combined detection of multiple biomarkers within the qPCR reaction, thereby reducing reagent and supply costs. To facilitate this approach, we designed the primer sets for the viral biomarker candidates to have uniform reaction conditions and similar amplicon lengths.

[0129] We conducted an in silico assessment of all combinations of two- and three-vOTU cocktails (Figure 11). Based on their prevalence across all of the metagenomes, we identified a two-vOTU cocktail and a three-vOTU cocktail that showed high prevalence in respiratory samples (77.3% and 81.0%, respectively) and low prevalence in non-respiratory samples (31.6% and 32.9%, respectively; Figure 8). Both cocktails included NOVB_1, the most abundant biomarker in nasal and oral metagenomes, and OVB_1, a highly prevalent biomarker in oral metagenomes (Figure 7B). The three-vOTU cocktail also included NVB_1, which was unique to and highly abundant in the nasal cavity (Figure 7B). NOVB_1 alone was found in 62.7% of oral metagenomes, whereas the two-vOTU cocktail was found in 99.3% of oral metagenomes. In the nasal metagenomes, the three vOTUs cocktail was present in 58.3% of samples, compared to 45.5% for NOVB_1 alone. Both cocktails had minimal impact on prevalence in non-respiratory samples, with the two-vOTU and three-vOTU cocktail being present in 3.2% and 4.4% more of non-respiratory metagenomes, respectively. Although the cocktails increased prevalence in respiratory samples across individuals, the combined abundances of vOTUs in the cocktails was not significantly greater than NOVB_1 alone in any sample type (p-values = 0.53-0.94).

[0130] Overall, assessing multiple viral biomarker candidates increases the robustness of biomarker assays by reducing the likelihood of missing specific target populations and minimizing effects of sample-to-sample fluctuations. Due to variations in the human microbiome across individuals (45, 54), it is unlikely that everyone in a shared space over time will have a common vOTU in their respiratory fluids. Others have demonstrated this challenge in fecal monitoring, whereby natural variations in fecal microbial composition between individuals impacts the utility of crAssphagc and Pepper Mild Mottle Virus (PMMoV) biomarkers. For Attorney Docket Number: UM-43622.601 example, PMMoV is consistently shed across individuals, but its concentration in a single individual’s fcccs can be highly variable (12), plausibly due to fluctuations in diet (55, 56). On the other hand, crAssphage concentrations have been found to be stable within individuals, can differ by six orders of magnitude between individuals (12). Such variations in PMMoV and crAssphage present challenges in environmental monitoring, particularly as community size decreases, because natural variations within and between individuals will result in large fluctuations in biomarker concentrations (14). Assessing multiple biomarkers with cocktail mixtures will reduce false negatives when assessing a sample for human respiratory emissions. By utilizing multiple biomarkers, natural variations in single vOTUs amongst and within individuals will not decrease the reliability of the assay as greatly.

[0131] Viral biomarkers were found in saliva and nasal swab samples with qPCR

[0132] To assess the efficacy of these viral biomarker candidates, we assessed their abundance in recently collected saliva and nasal swab samples with qPCR (Figure 10). In nasal swabs, any saliva bacteria biomarkers were found, at most, in half of the samples. There were five viral biomarker candidates that were more prevalent than the three saliva bacteria in nasal swabs. Notably, NOVB_1 was found in all of the nasal swabs. The bacterial biomarkers were highly prevalent in saliva samples, as previously reported (17). There were five viral biomarkers highly prevalent amongst saliva samples, with NOVB_1 found in all of the saliva samples. The most abundant viral biomarker in nasal swabs, NOVB_1, aligned with the in silico observations from the Human Microbiome Project metagenomes. Only one other viral biomarker, NVB_2, had a mean fold increase above the NTC greater than an order of magnitude in nasal swabs. Among the saliva bacteria biomarkers, only S. salivarius had a mean abundance greater than one order of magnitude above the NTC when present. Overall, the targets were more abundant in saliva samples, which may be attributed to lower overall biomass collected given the methodology. There were seven viral biomarkers with a mean fold increase from the NTC greater than an order of magnitude in saliva. The three saliva bacteria biomarkers were highly abundant in saliva with mean abundances more than two orders of magnitude above the NTC. The viral biomarkers were generally less abundant with only two viral biomarkers, NOVB_1 and OVB_1, having mean abundances greater than two orders of magnitude. In this Example, we did not use probes for the saliva bacteria targets in our qPCR assays. This reduces the specificity of the assays based on NCBI primer BLAST results and may have inflated the qPCR abundances. Attorney Docket Number: UM-43622.601

[0133] By assessing the prevalence and abundance of viruses in existing human metagenomes, we identified several viral biomarkcr candidates with high abundances and prevalences in newly collected nasal swabs and saliva. Three of the viral biomarkers, NVB_1, NVB_2, and NOVB_1, were highly prevalent in nasal and oral samples. There were four viral biomarkers, OVB_3, OVB_2, OVB_1, and OVB_4, that were prevalent and abundant in only oral samples. None of the viral biomarker candidates were unique to the nasal environment based on qPCR. The metagenomic results indicated that NVB_1 was unique to nasal samples; however, we observed an equally high prevalence in nasal swabs and saliva samples with qPCR.

[0134] REFERENCES

[0135] 1. Klepeis et al., 2001. The National Human Activity Pattern Survey (NHAPS): a resource for assessing exposure to environmental pollutants. Journal of Exposure Analysis and Environmental Epidemiology 11:231-252.

[0136] 2. Mao et al., 2020. Transmission risk of infectious droplets in physical spreading process at different times: A review. Build Environ 185:107307.

[0137] 3. Wang et al., 2021. Airborne transmission of respiratory viruses. Science 373.

[0138] 4. Stephens et al., 2019. Microbial Exchange via Fomites and Implications for Human Health. Curr Pollut Rep 5:198-213.

[0139] 5. Yang W, Elankumaran S, Marr LC. 2011. Concentrations and size distributions of airborne influenza A viruses measured indoors at a health centre, a day-care centre and on aeroplanes. J R Soc Interface 8:1176-84.

[0140] 6. Chamseddine et al., 2021. Detection of influenza virus in air samples of patient rooms. J Hosp Infect 108:33-42.

[0141] 7. Borges et al., 2021. SARS-CoV-2: a systematic review of indoor air sampling for virus detection. Environ Sci Pollut Res Int 28:40460-40473.

[0142] 8. Noble et al., 2003. Comparison of total coliform, fecal coliform, and enterococcus bacterial indicator response for ocean recreational water quality testing. Water Res 37:1637-43.

[0143] 9. Stachler E, Bibby K. 2014. Metagenomic Evaluation of the Highly Abundant Human Gut Bacteriophage CrAssphage for Source Tracking of Human Fecal Pollution. Environmental Science & Technology Letters 1:405-409. Attorney Docket Number: UM-43622.601

[0144] 10. Sabar MA, Honda R, Haramoto E. 2022. CrAssphage as an indicator of human-fecal contamination in water environment and virus reduction in wastewater treatment. Water Res 221:118827.

[0145] 11. Kitajima M, Sassi HP, Torrey JR. 2018. Pepper mild mottle virus as a water quality indicator, npj Clean Water 1.

[0146] 12. Arts et al., 2023. Longitudinal and quantitative fecal shedding dynamics of SARS-CoV- 2, pepper mild mottle virus, and crAssphage. mSphere 8.

[0147] 13. Li E, Saleem F, Edge TA, Schellhorn HE. 2021. Biological Indicators for Fecal Pollution Detection and Source Tracking: A Review. Processes 9.

[0148] 14. Holm RH, Mukherjee A, Rai JP, Yeager RA, Talley D, Rai SN, Bhatnagar A, Smith T.

[0149] 2022. SARS-CoV-2 RNA abundance in waste water as a function of distinct urban sewershed size. Environmental Science: Water Research & Technology 8:807-819.

[0150] 15. Leung NHL. 2021. Transmissibility and transmission of respiratory viruses. Nat Rev Microbiol 19:528-545.

[0151] 16. Turnbaugh PJ, Ley RE, Hamady M, Fraser-Liggett CM, Knight R, Gordon JI. 2007. The human microbiome project. Nature 449:804-10.

[0152] 17. Jung JY, Yoon HK, An S, Lee JW, Ahn ER, Kim YJ, Park HC, Lee K, Hwang JH, Lim SK. 2018. Rapid oral bacteria detection based on real-time PCR for the forensic identification of saliva. Sci Rep 8:10852.

[0153] 18. Guccione C, Patel L, Tomofuji Y, McDonald D, Gonzalez A, Sepich-Poore GD, Sonehara K, Zakeri M, Chen Y, Dilmore AH, Damle N, Baranzini SE, Hightower G, Nakatsuji T, Gallo RL, Langmead B, Okada Y, Curtius K, Knight R. 2025. Incomplete human reference genomes can drive false sex biases and expose patient-identifying information in metagenomic data. Nat Commun 16:825.

[0154] 19. Roux S, Enault F, Hurwitz BL, Sullivan MB. 2015. VirSorter: mining viral signal from microbial genomic data. PeerJ 3:e985.

[0155] 20. Guo J, Bolduc B, Zayed AA, Varsani A, Dominguez-Huerta G, Delmont TO, Pratama AA, Gazitua MC, Vik D, Sullivan MB, Roux S. 2021. VirSorter2: a multi-classifier, expert- guided approach to detect diverse DNA and RNA viruses. Microbiome 9:37.

[0156] 21. Kieft K, Zhou Z, Anantharaman K. 2020. VIBRANT: automated recovery, annotation and curation of microbial viruses, and evaluation of viral community function from genomic sequences. Microbiomc 8:90. Attorney Docket Number: UM-43622.601

[0157] 22. Ren J, Song K, Deng C, Ahlgren NA, Fuhrman JA, Li Y, Xie X, Poplin R, Sun F. 2020. Identifying viruses from mctagcnomic data using deep learning. Quant Biol 8:64-77.

[0158] 23. Menzel P, Ng KL, Krogh A. 2016. Fast and sensitive taxonomic classification for metagenomics with Kaiju. Nat Commun 7:11257.

[0159] 24. Nayfach S, Camargo AP, Schulz F, Eloe-Fadrosh E, Roux S, Kyrpides NC. 2021. CheckV assesses the quality and completeness of metagenome- assembled viral genomes. Nat Biotechnol 39:578-585.

[0160] 25. Hegarty B, Riddell J, Bastien E, Langenfeld K, Lindback M, Saini J, Wing A, Zhang J, Duhaime M. 2024. Benchmarking informatics approaches for virus discovery: caution is needed when combining in silico identification methods. mSystems 9:e01105-23.

[0161] 26. Kieft K, Adams A, Salamzade R, Kalan L, Anantharaman K. 2022. vRhyme enables binning of viral genomes from metagenomes. Nucleic Acids Res 5O:e83.

[0162] 27. Olm MR, Brown CT, Brooks B, Banfield JF. 2017. dRep: a tool for fast and accurate genomic comparisons that enables improved genome recovery from metagenomes through dereplication. ISME J 11:2864-2868.

[0163] 28. Roux et al. 2019. Minimum Information about an Uncultivated Virus Genome (MIUViG). Nat Biotechnol 37:29-37.

[0164] 29. Stachler E, Kelty C, Sivaganesan M, Li X, Bibby K, Shanks OC. 2017. Quantitative CrAssphage PCR Assays for Human Fecal Pollution Measurement. Environ Sci Technol 51:9146-9154.

[0165] 30. Ghaddar B, Naoum-Sawaya J. 2018. High dimensional data classification and feature selection using support vector machines. European Journal of Operational Research 265:993- 1004.

[0166] 31. Camargo AP, Roux S, Schulz F, Babinski M, Xu Y, Hu B, Chain PSG, Nayfach S, Kyrpides NC. 2024. Identification of mobile genetic elements with geNomad. Nat Biotechnol 42:1303-1312.

[0167] 32. Kanehisa M, Sato Y, Morishima K. 2016. BlastKOALA and GhostKOALA: KEGG Tools for Functional Characterization of Genome and Metagenome Sequences. J Mol Biol 428:726-731.

[0168] 33. DiMaio D. 2012. Viruses, masters at downsizing. Cell Host Microbe 11:560-1.

[0169] 34. Tisza MJ, Buck CB. 2021. A catalog of tens of thousands of viruses from human metagenomes reveals hidden associations with chronic diseases. Proc Natl Acad Sci U S A 118. Attorney Docket Number: UM-43622.601

[0170] 35. Wang J, Feng J, Zhu Y, Li D, Wang J, Chi W. 2022. Diversity and Biogeography of Human Oral Saliva Microbial Communities Revealed by the Earth Microbiomc Project. Front Microbiol 13:931065.

[0171] 36. De Boeck I, Wittouck S, Wuyts S, Oerlemans EFM, van den Broek MFL, Vandenheuvel D, Vanderveken O, Lebeer S. 2017. Comparing the Healthy Nose and Nasopharynx Microbiota Reveals Continuity As Well As Niche-Specificity. Front Microbiol 8:2372.

[0172] 37. Biswas K, Hoggard M, Jain R, Taylor MW, Douglas RG. 2015. The nasal microbiota in health and disease: variation within and between subjects. Front Microbiol 9:134.

[0173] 38. Vogtmann E, Chaturvedi AK, Blaser MJ, Bokulich NA, Caporaso JG, Gillison ML, Hua X, Hullings AG, Knight R, Purandare V, Shi J, Wan Y, Freedman ND, Abnet CC. 2023.

[0174] Representative oral microbiome data for the US population: the National Health and Nutrition Examination Survey. Lancet Microbe 4:e60-e61.

[0175] 39. Yahara K, Suzuki M, Hirabayashi A, Suda W, Hattori M, Suzuki Y, Okazaki Y. 2021. Long-read metagenomics using PromethlON uncovers oral bacteriophages and their interaction with host bacteria. Nat Commun 12:27.

[0176] 40. Paietta EN, Kraberger S, Custer JM, Vargas KL, Espy C, Ehmke E, Yoder AD, Varsani A. 2023. Characterization of Diverse Anelloviruses, Cressdnaviruses, and Bacteriophages in the Human Oral DNA Virome from North Carolina (USA). Viruses 15.

[0177] 41. Li S, Guo R, Zhang Y, Li P, Chen F, Wang X, Li J, Jie Z, Lv Q, Jin H, Wang G, Yan Q. 2022. A catalog of 48,425 nonredundant viruses from oral metagenomes expands the horizon of the human oral virome. iScience 25:104418.

[0178] 42. Pride DT, Salzman J, Haynes M, Rohwer F, Davis-Long C, White RA, 3rd, Loomer P, Armitage GC, Reiman DA. 2012. Evidence of a robust resident bacteriophage population revealed through analysis of the human salivary virome. ISME J 6:915-26.

[0179] 43. Wylie KM, Mihindukulasuriya KA, Sodergren E, Weinstock GM, Storch GA. 2012. Sequence analysis of the human virome in febrile and afebrile children. PLoS One 7:e27735.

[0180] 44. Bassis C, Tang A, Young V, Pynnonen M. 2014. The nasal cavity microbiota of healthy adults. Microbiome 2.

[0181] 45. Ding T, Schloss PD. 2014. Dynamics and associations of microbial community types across the human body. Nature 509:357-60.

[0182] 46. Proctor DM, Reiman DA. 2017. The Landscape Ecology and Microbiota of the Human Nose, Mouth, and Throat. Cell Host Microbe 21:421-432. Attorney Docket Number: UM-43622.601

[0183] 47. Comfort N. 2015. Genetics: We are the 98%. Nature 520:615-616.

[0184] 48. Land M, Hauser L, Jun SR, Nookacw I, Lcuzc MR, Ahn TH, Karpincts T, Lund O, Kora G, Wassenaar T, Poudel S, Ussery DW. 2015. Insights from 20 years of bacterial genome sequencing. Funct Integr Genomics 15:141-61.

[0185] 49. Roux S, Hallam SJ, Woyke T, Sullivan MB. 2015. Viral dark matter and virus-host interactions resolved from publicly available microbial genomes. Elite 4.

[0186] 50. Brum J, Ignacio-Espinoza JC, Kim E, Sullivan MB. 2016. Illuminating structural proteins in viral “dark matter” with metaproteomics. PNAS 113:2436-2441.

[0187] 51. Santiago-Rodriguez TM, Hollister EB. 2022. Unraveling the viral dark matter through viral metagenomics. Front Immunol 13:1005107.

[0188] 52. Kiishnamurthy SR, Wang D. 2017. Origins and challenges of viral dark matter. Virus Res 239:136-142.

[0189] 53. Dutilh BE, Cassman N, McNair K, Sanchez SE, Silva GG, Boling L, Barr JJ, Speth DR, Seguritan V, Aziz RK, Felts B, Dinsdale EA, Mokili JL, Edwards RA. 2014. A highly abundant bacteriophage discovered in the unknown sequences of human faecal metagenomes. Nat Commun 5:4498.

[0190] 54. Kolde R, Franzosa EA, Rahnavard G, Hall AB, Vlamakis H, Stevens C, Daly MJ, Xavier RJ, Huttenhower C. 2018. Host genetic variation and its microbiome interactions within the Human Microbiome Project. Genome Med 10:6.

[0191] 55. Zhang et al. 2006. RNA viral community in human feces: prevalence of plant pathogenic viruses. PLoS Biol 4:e3.

[0192] 56. Symonds EM, Nguyen KH, Harwood VJ, Breitbart M. 2018. Pepper mild mottle virus: A plant pathogen with a greater purpose in (waste)water treatment development and public health management. Water Res 144:1-12.

[0193] 57. Rockey et al., 2024. Seasonal influenza viruses decay more rapidly at intermediate humidity in droplets containing saliva compared to respiratory mucus. Applied and Environmental Microbiology 90.

[0194] 58. Morawska et al., 2009. Size distribution and sites of origin of droplets expelled from the human respiratory tract during expiratory activities. Journal of Aerosol Science 40:256-269.

[0195] 59. Gao NP, Niu JL. 2007. Modeling particle dispersion and deposition in indoor environments. Atmos Environ (1994) 41:3862-3876.

[0196] 60. Tandukar ct al., Scientific Reports volume 10, Article number: 3616 (2020) Attorney Docket Number: UM-43622.601

[0197] 61 . Letourneau et al., Journal of Environmental Sciences, Volume 148, February 2025, Pages 69-78.

[0198] 62. Jung et al., Scientific Reports volume 8, Article number: 10852 (2018).

[0199] All publications and patents mentioned in the present application are herein incorporated by reference. Various modification and variation of the described methods and compositions of the invention will be apparent to those skilled in the ail without departing from the scope and spirit of the invention. Although the invention has been described in connection with specific preferred embodiments, it should be understood that the invention as claimed should not be unduly limited to such specific embodiments. Indeed, various modifications of the described modes for carrying out the invention that are obvious to those skilled in the relevant fields are intended to be within the scope of the following claims.

Claims

Attorney Docket Number: UM-43622.601CLAIMS l / wc claim:

1. A method comprising: a) sampling air and / or a surface in an enclosed space with an assay for at least one human commensal viral or bacterial nucleic acid sequence present in the respiratory track and / or oral cavity of humans such that either: i) a negative result is detected with said assay, or ii) a first positive level is detected with said assay, and b) performing at least one of the following: i) if said negative result is detected, then moving a patient into said enclosed space, and / or allowing medical personal into said enclosed space, or ii) if said first positive level is detected, then treating said enclosed space air or surface with at least one of the following:A) replacing an air filter if an air filter present in said enclosed space,B) replacing an air filter in an HVAC system that provides air to said enclosed space,C) increasing air exchange in said enclosed space,D) exposing said enclosed space to UV radiation,E) wiping down surfaces in said enclosed space, andF) increasing humidity in said enclosed space.

2. The method of claim 1, wherein said enclosed space is a hospital room, clinic room, operating room, medical procedure room or area, childcare room, classroom, workspace or factory.

3. The method of claim 1, further comprising, if said positive level is detected, repeating said sampling to generate either a negative result or second positive level.

4. The method of claim 3, further comprising: moving a patient into said enclosed space, wherein said enclosed space is a hospital room, clinic room, operating room, or medical procedure room or area; and wherein if said second positive level is detected it is at a level ofAttorney Docket Number: UM-43622.601 less than 10% of said first positive level; or allowing children to enter a classroom or common space; or allowing workers into their workspace or factory.

5. The method of claim 1, wherein said negative result is detected, and said negative result is employed as a proxy for very low or no detectable levels of pathogens.

6. The method of claim 1, wherein said first positive level is detected, and said first positive results is employed as a proxy for at least low levels, and possibly high levels, of pathogens.

7. The method of claim 1, wherein said at least one human commensal viral nucleic acid sequence is selected from SEQ ID NOs: 1-124, and / or at least pail of a sequence selected from SEQ ID NOs: 1-124 is detected as present in said sample.

8. The method of claim 1, wherein said at least one human commensal bacterial nucleic acid sequence is selected from SEQ ID NOs: 125-145, and / or at least part of a sequence selected from SEQ ID NOs: 1-124 is detected as present in said sample.

9. The method of claim 1, wherein said sampling comprises a detection method that employs at least one set of forward and reverse primers, and optionally a corresponding probe, found in SEQ ID NOs: 146-751, or at least one set of forward and reverse primers that amplify a detectable portion of any of SEQ ID NOs: 1-145.

10. The method of claim 1, wherein said sampling comprises a detection method selected from PCR, quantitative PCR and nucleic acid sequencing.

11. The method of claim 1, wherein said at least one human commensal viral nucleic acid sequence comprises at least one NOVB_1 nucleic acid sequence.

12. The method of claim 11, wherein said at least one human commensal viral nucleic acid sequence further comprises a second viral nucleic acid sequence that comprises at least one OVB_1 nucleic acid sequence.Attorney Docket Number: UM-43622.60113. The method of claim 12, wherein said at least one human commensal viral nucleic acid sequence further comprises a third viral nucleic acid sequence that comprises at least one NVB_1 nucleic acid sequence.

14. The method of claim 1, wherein said at least one human commensal viral nucleic acid sequence comprises a first viral nucleic acid sequence that comprises at least one nucleic acid sequence from: NOVB_1, NOVB_2, NVB_1, NVB_2, OVB_1, OVB_2, OVB_3, OVB_4, OVB_5, OVB_6, OVB_7, and OVB_8.

15. The method of claim 1, wherein said at least one human commensal bacterial nucleic acid sequence comprises a first bacterial nucleic acid sequence from: N. subflava, S. salivarius, and S. sanguinis.

16. A method comprising: sampling air and / or a surface, and / or water, and / or biological sample, which is optionally in an enclosed space, with an assay for at least one human commensal viral or bacterial nucleic acid sequence present in the respiratory track and / or oral cavity of humans, wherein said at least one human commensal viral nucleic acid sequence is selected from SEQ ID NOs: 1-124, and / or at least part of a sequence selected from SEQ ID NOs: 1-124 is detected, and wherein said at least one human commensal bacterial nucleic acid sequence is selected from SEQ ID NOs: 125-145, and / or at least part of a sequence selected from SEQ ID NOs: 125- 145 is detected.

17. The method of claim 16, wherein said sampling comprises a detection method that employs at least one set of forward and reverse primers, and optionally a corresponding probe, selected from the sets in SEQ ID NOs: 146-751.

18. The method of claim 16, wherein said sampling comprises a detection method selected from PCR, quantitative PCR, and nucleic acid sequencing.Attorney Docket Number: UM-43622.60119. The method of claim 16, wherein said sampling comprises a detection method that employs at least one set of forward and reverse primers, and optionally a corresponding probe, found in SEQ ID NOs: 146-751, or at least one set of forward and reverse primers that amplify a detectable portion of any of SEQ ID NOs: 1-145.

20. The method of claim 16, wherein said sampling comprises a detection method selected from: PCR, quantitative PCR, and nucleic acid sequencing.

21. The method of claim 16, wherein said at least one human commensal viral nucleic acid sequence comprises at least one NOVB_1 nucleic acid sequence.

22. The method of claim 21, wherein said at least one human commensal viral nucleic acid sequence further comprises a second viral nucleic acid sequence that comprises at least one OVB_1 nucleic acid sequence.

23. The method of claim 22, wherein said at least one human commensal viral nucleic acid sequence further comprises a third viral nucleic acid sequence that comprises at least one NVB_1 nucleic acid sequence.

24. The method of claim 16, wherein said at least one human commensal viral nucleic acid sequence from: NOVB_1, NOVB_2, NVB_1, NVB_2, OVB_1, OVB_2, OVB_3, OVB_4, OVB_5, OVB_6, OVB_7, and OVB_8.

25. The method of claim 16, wherein said at least one human commensal bacterial nucleic acid sequence is selected from: N. subflava, S. salivarius, and S. sanguinis.

26. A method comprising: a) sampling air and / or a surface in an enclosed space with a first assay for at least one human respiratory pathogen such that either: i) a negative result with said first assay is detected, or ii) a first positive level is detected with said first assay, and b) sampling said air and / or said surface in said enclosed space with a second assay for at least one human commensal bacterial or viral nucleic acid sequence is present in theAttorney Docket Number: UM-43622.601 respiratory track and / or oral cavity of humans such that either: i) a negative result is detected with said second assay, or ii) a first positive level is detected with said second assay, and wherein said sampling in b) is carried at, or about the same time, as said sampling in a).

27. The method of claim 26, wherein said about the same time is within about 1 hour or less.

28. The method of claim 26, wherein said sampling in a) produces said negative result in said first assay for said at least one human respiratory pathogen, and wherein said sampling in b) produces said negative result in said second assay for said at least one human commensal bacteria or virus, and wherein the method further comprises performing an action in said enclosed space that requires, or is best, when said human respiratory pathogens are not present, such as, but not limited to, i) generating a report that the enclosed space is free from said at least one human respiratory pathogen, ii) moving patient into said enclosed space, wherein said enclosed space is a hospital or clinic room; iii) repeating said sampling in a), optionally wherein said first assay is re-calibrated prior to repeating said sampling in a); iv) allowing medical personal into said enclosed space; v) allowing children and / or students to enter a daycare room or classroom; and vi) allowing workers back into their workspace or factory.

29. The method of claim 26, wherein said sampling in a) produces said negative result in said first assay for said at least one human respiratory pathogen, and wherein said sampling in b) produces said positive level in said second assay for said at least one human commensal bacterial or viral nucleic acid sequence, and wherein the method further comprises: c) treating said enclosed space air or surfaces with at least one of the following: i) replacing an air filter in an air filter present in said enclosed space, ii) replacing an air filter in an HVAC system that provides air to said enclosed space, iii) increasing air exchange in said enclosed space, iv) exposing said enclosed space to UV radiation, v) wiping down surfaces in said enclosed space, andAttorney Docket Number: UM-43622.601 vi) increasing humidity level in said enclosed space.

30. The method of claim 29, further comprising: d) sampling said air and / or said surface in said enclosed space with said second assay for said at least one human commensal bacterial or viral nucleic acid sequence such that: i) a negative result is detected with said second assay, or ii) a second positive level is detected with said second assay.

31. A composition comprising: a nucleic acid molecule, wherein nucleic acid molecule comprises a sequence of at least 12 nucleotides in length that hybridizes under stringent conditions to any of SEQ ID NOS: 1-124 or complement thereof, and wherein said nucleic acid molecule: i) comprises a detectable label; ii) comprises at least one modified base; and / or iii) is linked to a heterologous nucleic acid sequence.

32. The composition of claim 31, wherein said sequence is at least 14, 15, 16, 17, 18, 19, 20, 21, or 22 nucleotides in length.

33. The composition of claim 31, wherein said nucleic acid molecule comprises RNA or DNA.

34. The composition of claim 31, wherein said at least one modified base is selected from: 5- methylcytosine (5mC) and / or 5-hydroxylmethylcytosine (5hmC).

35. The composition of claim 31, wherein said detectable label comprises a fluorophore, and wherein said nucleic acid molecules further comprises a quencher.

36. The composition of claim 31, wherein said heterologous nucleic acid sequence is selected from: a sequencing adapter, a sequencing sample index sequence, a universal primer binding site, a sequencing UMI barcode, and a flow cell tag sequence.Attorney Docket Number: UM-43622.60137. A composition or kit comprising: a) a fist forward primer, and b) a corresponding first reverse primer, and c) optionally a corresponding first probe sequence, wherein said forward primer and corresponding reverse primer, and optionally said probe sequence, are found in SEQ ID NOs: 146-169 and 176-751.

38. The composition or kit of claim 33, wherein said first forward primer and corresponding first reverse primer are configured to amplify a portion of: NOVB_1, N0VB_2, NVB_1, NVB_2, OVB_1, 0VB_2, 0VB_3, 0VB_4, 0VB_5, 0VB_6, 0VB_7, and 0VB_8.

39. The composition or kit of claim 33, wherein said first forward primer and corresponding first reverse primer are configured to amplify a portion of NOVB_1.

40. The composition or kit of claim 39, wherein said composition further comprises a second forward primer and corresponding second reverse primer configured to amplify a portion of OVB_1.

41. The composition or kit of claim 41, wherein said composition further comprises a third forward primer and corresponding third reverse primer configured to amplify a portion of NVB_1.

Citation Information

Patent Citations

  • Pathogen mitigation

    US20230270911A1

  • Signal normalization of nucleic acid amplification reaction products

    WO2023104384A1