Method for estimating molecular complexity
The joint molecular assembly method addresses the limitations of individual molecular complexity estimation by accounting for shared precursors, enhancing the representation and comparison of sample complexity and phylogenetic relationships.
Patent Information
- Application Number
- PCT/EP2025/065922
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-07
- Filing Date
- 2025-06-06
- Publication Date
- 2025-12-11
AI Technical Summary
Existing methods for estimating molecular complexity focus on individual molecules, neglecting the relationships between molecules in a sample, which limits the representation of the sample's overall complexity and hinders comparative analysis.
A method using MS/MS data to calculate a joint molecular assembly (JMA) that accounts for the shared complexity among components in a sample, allowing for a more comprehensive estimation of molecular complexity and enabling comparisons between samples.
JMA provides a more faithful representation of sample complexity, facilitating the comparison of samples and inferring phylogenetic relationships, distinguishing between closely related species, and identifying biological markers.
Smart Images

Figure EP2025065922_11122025_PF_FP_ABST
Abstract
Description
[0001] METHOD FOR ESTIMATING MOLECULAR COMPLEXITY
[0002] Related Application
[0003] The present case claims the benefit of, and priority to, GB 2408156.4 filed on 07 June 2024 (07.06.2024), the contents of which are hereby incorporated by reference in their entirety.
[0004] Field of the Invention
[0005] The present invention relates to a method for estimating the molecular complexity of one sample or a plurality of samples, wherein each sample includes two or more components, a method for estimating a phylogenetic relationship of two samples and a method for estimating a phylogenetic character of a sample.
[0006] Background
[0007] Early taxonomical classifications focused on practical medicinal properties of plants and broad phenetic organization of animals (Mayr). These attempts at categorization were based on available data of morphological similarities and shared characteristics across organisms.2. Later, Darwin’s realization that different taxonomic groups might share a common ancestry by evolutionary descent with modification was corroborated by close examination of phenotypical characteristics across similar species, revealing how these species persisted and / or changed over time (Darwin et al.). His findings lead to the advent of phylogeny and the formulation of the modern conception of the “tree of life” as a unified model capturing evolutionary relationships across life on Earth (Hug et al.', Delsuc et al.', Mukherjee et al.).
[0008] After the advent of evolutionary theory, expansive discoveries in molecular biology eventually led to genomic and proteomic sequence information as the definitive cornerstone of modern phylogenetics. Exploring evolutionary relationships in sequence data has significantly enhanced our understanding of the history and evolution of life on Earth (Delsuc et al.', Mukherjee et al.', Parks et al. (2018); Ciccarelli et al.). However, relying on gene sequence information alone comes with substantial limitations that have been difficult to overcome. There are incongruences between organismal genealogies based on genes’ selection and their coverage. Fast evolving entities like viruses have defied being rooted in a consensus evolutionary tree, and no viral species has yet been classified at every taxonomic level (Gorbalenya et al.). Genomic data is often sparse and inconsistent, leaving much of Earth’s extant biodiversity unmapped (Mills et al.). Generating high genetic coverage for the diverse range of organisms in the modern biosphere is a huge technical challenge (Mukherjee et al.', Parks et al. (2017); Kapil et al.). The identification of evolutionary relationships has so far relied on pre-existing knowledge of taxonomy or biochemistry, heavily depending on genomic data. This limits analytical capacities as we investigate more complex cases where sequence information is scarce or non-existent.
[0009] One of the present inventors has previously proposed a measure of molecular complexity known as object assembly (also known as pathway complexity) which is derived from the theory of assembly pathways (Marshall (2017)). Assembly pathways are sequences of joining operations that start with basic building blocks (e.g. bonds) and end with a final product. In these sequences, sub-units generated within the sequence can combine with other basic or compound sub-units later in the sequence to recursively generate larger structures (see Figure 19A and 19B). The inventor has also previously shown that life could be detected and differentiated using Assembly Theory (AT) and spectroscopic measurements, such as NMR, IR or mass spectrometry (Marshall et al. (2021 )), as described in WO 2021 / 186193.
[0010] In this previous work, a molecule in a sample was given a molecular complexity value. According to AT, every molecule has an inherent complexity (Molecular Assembly, MA) corresponding to the shortest construction pathway of its molecular graph (see Figure 19A and 19B). Central to AT is the concept of assembly space, composed of a molecule’s precursors along its shortest construction pathway, where each fragment is recursively built up from simpler building blocks. The smallest possible cardinality of this space is indicated by the molecule’s MA index and quantifies the minimal number of constraints within chemical space necessary for its construction.
[0011] However, the MA index is limited to providing a MA for individual molecules. Practical inferences from this work rely on the MA for each molecule taken in isolation. For example, a single molecule with a high MA may be taken as an indication of life. The analysis focuses on the molecule in a sample having the highest MA, and does not leverage the information provided by the MA for all other molecules in a sample. Moreover, the relationship between molecules in a sample - such as common molecules or common precursors - is not taken into account in the MA calculation.
[0012] The MA described in this earlier work focuses on the complexity of each molecule in a sample individually, which may not be representative of the complexity of the sample as a whole and makes the useful comparison of complexity between samples difficult.
[0013] The present invention has been devised in the light of these considerations. Summary of the Invention
[0014] The present invention provides an experimental method for estimating molecular complexity of a sample including two or more components. The method uses simple experimental techniques, such as MS / MS, to estimate the molecular complexity of the multicomponent sample, using a joint molecular assembly (JMA).
[0015] In general, the invention provides a method for estimating the molecular complexity of one sample or a plurality of samples, the method comprising obtaining MS / MS data for the sample or the plurality of samples, identifying the ions of the components in the MS1and the ion fragments in the MSndata, for the ions identified in the MS1data, and calculating the molecular complexity for the one sample or the plurality of samples from the ions of the components and the ion fragments.
[0016] Typically, estimating the molecular complexity comprises calculating a joint molecular assembly (JMA) for the one sample or the plurality of samples. JMA is a measure of the complexity of the sample or the plurality of samples, which represents the shared complexity amongst the components which make up the sample(s).
[0017] In a general aspect of the invention, there is provided a method for estimating the molecular complexity of one sample or a plurality of samples, wherein each sample includes two or more components, the method comprising obtaining MS / MS data for the sample or the plurality of samples; identifying ions of the components and ion fragments in MSndata; and calculating a joint molecular assembly (JMA) for the one sample or the plurality of samples from the ions of the components identified and the ion fragments identified in the MSndata.
[0018] In a first aspect of the invention, there is provided a method for estimating the molecular complexity of one sample or a plurality of samples, wherein each sample includes two or more components, the method comprising:
[0019] (a) obtaining MS / MS data for the sample or the plurality of samples, the MS data comprising MS1data for ions of the components, and MSndata for the nthgeneration of ion fragments formed from the ions of the components, wherein n is 2 or more;
[0020] (b) identifying the ion fragments in the MSndata for the ions of the components identified in the MS1data, and;
[0021] (c) calculating a joint molecular assembly (JMA) for the one sample or the plurality of samples from the ions of the components identified in the MS1data and the ion fragments identified in the MSndata, wherein JMA is a measure of the complexity of the sample or the plurality of samples, which represents the complexity of the components and the shared complexity amongst the components.
[0022] The present invention allows the concept of assembly space to apply to multiple components of a sample. The components may be a set of distinct molecules, but which can have a joint assembly space. The shortest and simultaneous construction pathway for the entire set of molecules together is called a Joint Assembly Space (JAS) (Figure 19C). Compared to the assembly space shown in Figure 19A and 19B, the Joint Assembly Space allows for the existence of common precursors between components. The pathways in a JAS represent the minimal set of constraints required to construct a given ensemble of molecules simultaneously, consisting of the smallest number of construction steps required to generate all the target molecules.
[0023] This number is then represented by the Joint Molecular Assembly (Joint MA). Within the framework of a generalized AT, JAS represents the Assembly Observed (subset of Assembly Contingent) measured at a single point in time. This can be used as a powerful tool to reveal contingent relationships between individual molecules or entire chemical systems. This is particularly important when the exact chemical reaction pathways are too complex or the temporal formation histories of molecules are unknown. The JAS does not have explicit knowledge of geological time, but for a set of complex molecules it represents the underlying causal relationships that have a necessary temporal ordering for that set of molecules to be co-constructed through an evolutionary process.
[0024] In this way, JMA provides a more faithful estimate of complexity for a sample including multiple components, as it includes contributions from multiples components of the sample and accounts for their common precursors. Where a sample includes multiple similar components (which share many common precursors), the JMA will be lower than in a sample which includes multiple diverse components (which share few common precursors). A more diverse range of components may be an indication of life. This information is not captured when analysing a sample using an MA index only.
[0025] Accordingly, the invention allows for a generalised method of inferring the joint assembly of complex, multi-component samples, such as environmental samples, and preferably whole cell samples.
[0026] JMA can also be obtained for a plurality of samples, which allows for the comparison of the JMA for two samples. As the JMA includes contributions from multiples components of the sample and accounts for their common precursors, the JMA can be used to estimate the joint assembly overlap (JAO) between samples. This data rich comparison can be used to estimate the similarity of two samples - which in turn can be used to infer the phylogenetic relationship between the samples. This comparison is significantly more informative than a comparison using the MA index of a single molecule in a sample.
[0027] Accordingly, in a further general aspect of the invention there is provided a method for estimating a phylogenetic relationship of a first sample and a second sample, the method comprising estimating the molecular complexity of the samples, and estimating a phylogenetic relationship of the first sample and second sample from the similarity of the molecular complexity.
[0028] In a second aspect of the invention there is provided a method for estimating a phylogenetic relationship of a first sample and a second sample, the method comprising: estimating the molecular complexity of the samples according to the method of the first aspect, to determine the JMA of the first sample and the second sample; and estimating a phylogenetic relationship of the first sample and second sample from the similarity of the JMA of the samples.
[0029] The first sample and the second sample referred to in the second aspect correspond to the plurality of samples is the first aspect.
[0030] In a third aspect of the invention there is provided a method for estimating a phylogenetic character of a first sample, the method comprising:
[0031] (a) estimating the molecular complexity of the first sample and a second sample according to the method of the first aspect, to determine the JMA of the first sample and second sample; and
[0032] (b) estimating the phylogenetic character of the first sample from the similarity of the JMA of the first and second samples.
[0033] The first sample and the second sample referred to in the third aspect correspond to the plurality of samples is the first aspect.
[0034] Preferably, the phylogenetic relationship of the samples is determined based on the JMA of the first and second samples, such as using joint assembly overlap (JAO).
[0035] In some embodiments the method is for estimating the phylogenetic clade of two or more samples. The method may be used to generate a phylogenetic tree from the estimated phylogenetic clades of the samples.
[0036] In some embodiments the method is for identifying a candidate pharmaceutical or agrochemical in a first sample.
[0037] In some embodiments the method is for identifying diseased cells in a first sample.
[0038] In some embodiments the method is for identifying treatment-resistance pathogens in a first sample. ln some embodiments the method is for the detection of life in a sample. The comparison of JMA for samples has practical applications in life detection, but also can be used for life classification. It can be used to determine the relationship between species, and to generate a phylogenetic tree of life for multiple samples. These phylogenetic trees have been shown to conform with current biological knowledge, offering a comparable alternative to other phylogenetic methods, see Figure 3. This illustrates that the distinct molecular networks of biological cells carry information about their evolutionary history, which can be deciphered and employed as a universal apparatus for life’s detection and classification using JMA. The inventors have also shown that JMA can be used to infer minor morphological and genetic differences, such as between a series of bacterial generations. JMA has been shown to identify evolutionary differences on a short time scale, where the differences may not be genomically encoded. This has further applications in identification of candidate pharmaceutical or agrochemical, identifying diseased cells and treatmentresistance pathogens.
[0039] These and other aspects and embodiments of the invention are described in further detail below.
[0040] Summary of the Figures
[0041] The present invention is described herein with reference to the figures listed below.
[0042] Figure 1 shows an overview of the experimental and analytical design. A) Sample preparation. Fresh samples are freeze-dried then ground in a mortar and pestle to produce a fine powder (< 1 mm). B) Extraction. Four falcon tubes are prepared- three replicates containing 0.1 g of sample and one control. 10 ml_ of 80% methanol extraction solvent is added before sonication for 20 minutes at 30 °C. Solids are separated by centrifuging for 8 minutes at 4400 rpm. The supernatant is passed through a syringe filter (Nylon, 0.22 pm) before analysing the extract using HPLC-MS / MS (Orbitrap Fusion Lumos). C) MS Analysis, i) Chromatograms of sample and control illustrate the high level of contamination and noise (red) that need to be omitted before analysis, while capturing the sample analytes (green), ii) A background subtraction is executed, providing a list of analytes native to the sample, iii) Due to the presence of contamination and noise, not all sample analytes have been selected for fragmentation, while contamination analytes have been erroneously selected. A second round of MS analysis is performed inputting a list of sample analytes, resulting in the fragmentation of all sample analytes. This data is subsequently used in the later MA algorithms.
[0043] Figure 2 shows the extraction workflow for animal, fungi and plant samples.
[0044] Figure 3 shows the workflow outlining the methodology for bacteria colony survival (blue) and the preparation of bacteria cultures for extraction (green). Figure 4 shows the bacterial extraction workflow.
[0045] Figure 5 shows chromatograms for three sample repeats and a control. The area highlighted in red illustrates the high contamination present in both the sample and control, while the green area displays analytes unique only to the sample.
[0046] Figure 6 shows scenarios for the retention or removal of analytes found in the sample and control. First row shows the nominal detection of an analyte (which would be retained in row V).
[0047] Figure 7 shows raw data containing analytes that are a product of contamination (top) and filtered data, after background subtraction that contains a peaklist of analytes extracted from the sample (bottom).
[0048] Figure 8 shows an optimised MS2workflow. Metabolites are extracted from a sample and analysed in the Orbitrap MS. A filtered peaklist is then fed into an Orbitrap method (unique for each sample) and samples are analysed a second time to specifically fragment ions from the peak list. This greatly improves the quantity of MS2spectra and thus builds more accurate MS2consensus spectra.
[0049] Figure 9 shows a single MS2scan for an analyte. Peaks that lie below the intensity threshold (red) are removed, and peaks from isotopic patterns are collapsed into a single peak (zoom).
[0050] Figures 10A to 10F show scatterplots of the detected analytes and the number of peaks in their associated MS2spectra for each sample in the dataset.
[0051] Figures 11A and 11B show scatterplots (left side) and KDE maps (right side) of the detected analytes and the number of peaks in their associated MS2spectra for each cladal domains in the dataset.
[0052] Figure 12 is a diagram showing how the molecular assembly may inferred from tandem MS data using the recursive MA algorithm and its generalisation to calculating joint assembly. Arrows between nodes at the same layer indicate measured or deduced fragmentation event and thus a possible construction operation.
[0053] Figure 13 is a diagram showing how the joint assembly index (Joint MA(A, B)) may be calculated from two samples as part of defining joint assembly overlap JAO. Arrows between nodes at the same layer indicate measured or deduced fragmentation event and thus a possible construction operation.
[0054] Figure 14 shows a matrix of JAO values as computed by the Recursive MA algorithm between all biological samples in the dataset. Figure 15 shows KDE plots of the OOC values for the different predictions attributed by the life discrimination classifier to samples of biological and non-biological sources. The OOC values are a result of averaging the values of the MS1and MS2prediction models. A) For non-biological samples, the OOC values for their predictions. B) For biological samples, the OOC values for their predictions.
[0055] Figure 16 shows KDE plots of the OOC values attributed by the domain classifier to samples of each cladal domains. The distributions show the OOC values of each predicted label, and for each set of samples belonging to their “true” group in separate figures. The OOC values are a result of averaging the values of the MS1and MS2prediction models. A) Archaea samples. B) Fungi samples. C) Bacteria samples. D) Plant samples. E) Animal samples.
[0056] Figure 17 shows spectral clustering of assembly-generated JAO between samples in the dataset. A) Discriminating between biological and non-biological samples. B) Classification of cladal domains, showing the clustering decision by the algorithm for varying number of clusters
[0057] Figure 18 shows spectral clustering of OOC values from MS2data between samples in the dataset. A) Discriminating between biological and non-biological samples. B) Classification of cladal domains, showing the clustering decision by the algorithm for varying number of clusters.
[0058] Figure 19 shows an illustration of assembly pathways. A) Basic building blocks used in the assembly processes. B) The assembly pathway for caffeine, consisting of the shortest sequence of joining operations to provide the Molecular Assembly (MA) index (here, MA = 9). C) The Joint Assembly Space (JAS) of Cytosine and Thymine, consisting of their combined assembly pathways to give the shortest sequence of joining operations for their simultaneous construction. Both target molecules have MA of 6 on their own, but when assembled together their Joint MA is 9.
[0059] Figure 20 shows an illustration of assembly spaces of biological systems across evolution. A) The cellular composome and its informational architecture. The flow of information is portrayed, from the genome to the phenome (e.g. proteins and ribozymes) and the metabolome. Observable analytes from these levels serve as molecular assembly markers, which give information about the assembly space underlying the cellular composome, and which can inform about the unique identity and state of the cell. B) An assembly representation of evolution as a tree of molecular networks, from an initial network composed of an assortment of building blocks towards more defined complex organismal networks. The molecular networks evolve over time and become more distinctly associated with recent clades. The observed assembly spaces of molecular networks are shown at the top, defined by the diversity and complexity of their molecular structures and used for phylogenetic inference. Colours indicate association to specific phylogenetic groupings, and their idiosyncratic traces through their inferred evolutionary past.
[0060] Figure 21 shows histograms and Kernel Density Estimates used for fingerprinting molecular networks. A) Histograms of the intensity of peaks detected for all samples within the labeled groups. B) Histograms of the number of peaks detected for samples within each labeled group. C) Analytes from a Cod sample (purple) layed over Kernel Density Estimate (KDE) maps of the Animals and Inorganics (minerals) groups (green), according to the m / z and number of MS2peaks of their associated analytes. The KDE bandwidth estimation was conducted with a scaling factor of 0.55. Scatterplots and KDE maps for all samples and groups can be found in the ‘experimental methods’ section.
[0061] Figure 22 shows assembly-based analyses and classification of composite MS information. A) Schematic representation of the methodology of the Recursive MA algorithm emulating the Joint Fragment Assembly Space of multimolecular samples. B). Boxenplots of all MS2fragments observed in the analysed samples, showing their m / z and their occurrence in the samples dataset. Each box represents a bin of 100 Da. C) Boxplot of the maximal estimated MA of each sample for each group, calculated with Recursive MA and normalized across the set. D) Spectral clustering over the JAO and OOC matrices for cladal domain classification, presenting Silhouette scores for each number of clusters in the model. The Silhouette score indicates how similar the objects within the clusters are to each other compared to objects in other clusters, with high values indicating better clustering scheme.
[0062] Figure 23 shows a schematic for phylogenetic inference using Assembly Theory. A) Algorithmic pipeline for the generation of a phylogenetic tree from MS data using Assembly theory. Recursive MA is employed to calculate Joint MA indices for single and paired samples, which are used in JAO calculations (see experimental methods). The JAO matrix represents the overlap of assembly-spaces between samples, which is then transformed into a tree through an inference algorithm (WPGMA). B) Comprison between an assembly-based (JAO) and a non-assembly-based (OOC) methods for phologenetic tree inference. Colors refer to the predefined cladal domains, as in Figures 4 and 8, and indicate the cladal cohesiveness of the clustering. The OOC tree was constructed from MS2information, and its full annotation is included in the SI section 8.
[0063] Figure 24 shows an example of an Assembly-based phylogenetic Tree of Life from nontargeted metabolite information. A) The Assembly tree with species annotation. The position of samples within this tree is not pre-determined by a priori cladistic knowledge and no prior knowledge of genomic, proteomic or metabolic chemistry of the samples was used. Colours refer to cladal domains. Lengths of the branches correspond to 2 minus JAO of the clustered clades. This observation highlights the relatedness of samples, as it points to a shared contingent chemical history, where exploration in their assembly spaces occurred in a similarly directed trajectory (Sharma et al.). B) Kernel Density Estimate (KDE) plots of the m / z values of all shared MS2fragments between samples of different geodesic distances in the assembly tree model. Arrows indicate the trends in the data. C) Histograms of randomly- generated phylogenetic trees models showing their Generalized Robinson-Foulds (GRF) similarity scores to the consensual genome-based tree. The different distributions refer to higher degree of constraints in the randomity of the tree generation, depicted in the inset, where randomness was allowed within predefined cladal groups. Purple - completely random, blue - separating prokeryotes and eukaryotes, light blue - separating the bacterial, archaeal and eukaryotic domains, and green - separating the cladal domains of bacteria, archaea, plants, fungi and animals. Colored circles in the inset refers to the cladal groups as appeared in Figures 7 and 8. The vertical dashed lines represent the score of the OOC and JAO tree models, which are 0.36 and 0.62 respectively. T-tests of the assembly-based model against the random-trees distributions all results in essentially p-value = 0.0.
[0064] Figure 25 shows an example of an assembly-based phylogenetic Tree of Life from nontargeted metabolite information, presenting all samples including multiple samples belonging to individual species. Colours refer to cladal domains. Lengths of the branches correspond to 2 minus JAO of the clustered clades.
[0065] Figure 26 shows an example of a Phylogenetic Tree of Life generated using the OOC metric over MS1data, presenting all samples including multiple samples belonging to individual species. Colours refer to cladal domains. Lengths of the branches correspond to 2 minus OOC of the clustered clades.
[0066] Figure 27 shows an example of a Phylogenetic Tree of Life generated using the OOC metric over MS2data, presenting all samples including multiple samples belonging to individual species. Colours refer to cladal domains. Lengths of the branches correspond to 2 minus OOC of the clustered clades.
[0067] Figure 28 shows an example of a Phylogenetic Tree of Life generated using the OOC metric over MS2data, presenting collapsed samples belonging to the same species. Colours refer to cladal domains. Lengths of the branches correspond to 2 minus OOC of the clustered clades.
[0068] Figure 29 shows histograms and Kernel Density Estimate (KDE) maps of overlapping MS2fragments between samples clustered at different geodesic distances (GD) on the assembly tree. A) GD 1 , B) GD 2, C) GD 3, D) GD 4, E) GD 5, F) GD 6, G) GD 7, H) GD 8.
[0069] Figure 30 shows an example of a consensual phylogenomic Tree of Life based on genomic data and generated from the TimeTree database. Colours refer to cladal domains as in previous figures. Lengths of the branches in the tree correspond to time lengths recorded in the database. Figure 31 shows histograms of randomly-generated phylogenetic trees models showing their Quartet similarity scores to the consensual genome-based tree. The different distributions refer to higher degree of constraints in the randomity of the tree generation, depicted in the inset, where randomness was allowed within predefined cladal groups. Purple - completely random, blue - separating prokaryotes and eukaryotes, light blue - separating the bacterial, archaeal and eukaryotic domains, and green - separating the cladal domains of bacteria, archaea, plants, fungi and animals. Coloured circles in the inset refers to the cladal groups as appeared in previous tree figures. The vertical dashed lines represent the scores of the tree models generated by the MS1Overlap, MS1&2Overlap and JAO phylogenetic metrics, which are 0.817, 0.839 and 0.926 respectively. T-tests of the assembly-based model against the random-trees distributions all result in essentially p-value=0.0.
[0070] Figure 32 shows phylogenetic and morphological inferences in multicellular species based on multiple tissue samples. A) A schematic comparison of clustering choices guided by phylogeny (shading) or morphology (shape). B) An assembly-based JAO phylogenetic model of the botanical cohort, consisting of seven species (indicated by shading) and three tissue types (leaves, flowers and branches, specified by icons). Parentheses indicate different strains (v - variegata, sdd - souvenir de desio).
[0071] Figure 33 shows an assembly-based similarity matrix between all samples from the tissue heterogeneity cohort (based on multiple tissue samples), presenting their JAO values between each pair of samples as calculated by the Recursive MA algorithm. The diagonal values are set to zero for visual clarity.
[0072] Figure 34 shows a phylogenetic tree of the tissue heterogeneity cohort, constructed with the MS1Overlap metric.
[0073] Figure 35. shows a phylogenetic tree of the tissue heterogeneity cohort, constructed with the MS1&2Overlap metric.
[0074] Figure 36 shows how bacterial lineages may be deduced using assembly-based phylogenetics. A) A schematic representation of a recursive step in the experiment, consisting of plating a culture of E. coll and re-culturing two picked colonies from the plate. B) The tree-like structure of the recursive experiment, spanning four generations of colony culturing, and resulting in eight leaf node samples indicated by bold letters. C) A scatterplot detailing the capacity of the phylogenetic algorithms to predict a tree that closely resembles the actual experimental tree structure, assessed by GRF and Quartet tree similarity metrics. The black dots represent every possible tree structure comprised of eight leaf nodes (135,135 trees), and their tree similarity values with normally-distributed errors for better visual dispersion. The dashed lines represent percentiles of trees included within the encapsulated areas. D) A plot depicting the prediction accuracy of the phylogenetic algorithms in relation to the number of samples included in the tree. The calculations were done through jackknife cross-validation, in which every possible set permutation of n nodes was tested, comparing the resultant predicted tree to the actual experimental tree structure. The tree similarity used was an average of the GRF and Quartet values obtained for each prediction. Error bars represent standard deviations.
[0075] Figure 37 shows a heatmap of the JAO values between all E. coll samples in the recursive culturing experiment. The letters correspond to the notation in Figure 36. The diagonal in the matrix is set to zero for clarity.
[0076] Figure 38 shows a histogram of the GRF tree similarity values between all possible phylogenetic trees constructed from eight leaf nodes (135,135), and the true experimental tree structure of the recursive culturing experiment of E. coll. The dashed lines indicate the GRF values calculated for the tree predictions of the MS1Overlap, MS1&2Overlap and JAO m etrices.
[0077] Figure 39 shows a histogram of the Quartet tree similarity values between all possible phylogenetic trees constructed from eight leaf nodes (135,135), and the true experimental tree structure of the recursive culturing experiment of E. coll. The dashed lines indicate the Quartet values calculated for the tree predictions of the MS1Overlap, MS1&2Overlap and JAO m etrices.
[0078] Figure 40 shows a jackknife cross-validation analysis of the bacterial recursive culturing cohort. All sets were generated that include 5-8 samples, and their corresponding phylogenetic trees were constructed with the controls MS1and MS1&2Overlap metrics, as well as the assembly-based JAO metric. The trees were compared to the real experimental tree structure, and the similarity between them was quantified with A) GRF (General Robinson- Foulds) and B) Quartet tree similarity heuristics.
[0079] Detailed Description of the Invention
[0080] The present invention provides an experimental method for estimating molecular complexity of a sample including two or more components. The method estimates the complexity of the sample based on the two or more components and their fragments, accounting for the similarity of the components and fragments. The method uses simple experimental techniques such as MS / MS to estimate the molecular complexity of the multicomponent sample, using a joint molecular assembly (JMA).
[0081] In general, the invention provides a method for estimating the molecular complexity of one sample or a plurality of samples, the method comprising obtaining MS / MS data for the sample or the plurality of samples, identifying the ion fragments and ions of components of the samples, identified in the MS1data, and estimating the molecular complexity for the one sample or the plurality of samples from the ions of the components and the ion fragments. Typically, estimating the molecular complexity comprises calculating a joint molecular assembly (JMA).
[0082] In WO 2021 / 186193, an experimental method for estimating molecular complexity is described. The method uses experimental techniques to assess the structural heterogeneity of a molecule within a sample, and provides an estimate of the molecular assembly index (MA) of a molecule.
[0083] In this earlier work, a molecule in a sample was given a molecular complexity value (MA). In a sample with multiple molecules, each molecule was given an individual molecular complexity value (MA). In contrast, the present invention estimates a joint molecular assembly (JMA) which represents the joint molecular complexity of all molecules in a sample.
[0084] The JMA allows information about all molecules in the sample to be used for characterisation, as opposed to focussing on a molecule having the highest MA value as being determinative, as in WO 2021 / 186193. Although this earlier work may determine a collection of MA values, the values only provide information on each component. The JMA considers the relationship between each molecule in a sample, specifically whether the molecules have shared precursors. This provides an indication of the complexity between molecules in a sample. For example, two similar molecules having many shared precursors will provide a smaller contribution to the JMA than two different molecules having few shared precursors. This information is not captured when estimating the MA of individual molecules only, as described in WO 2021 / 186193.
[0085] In this way, JMA provides a more representative estimation of complexity for a sample as a whole - which not only accounts for complexity of each component of the sample, but also accounts for the complexity attributable to the diversity of the components within that sample.
[0086] The JMA effectively summarises the overall casual graph that is estimated to connect the components using their shared precursor history. The JMA captures information on the casual links between different components, rather than looking at isolated MA values for a collection of individual components.
[0087] WO 2021 / 186193 also describes comparison of the MA of a molecule with a threshold value, in order to estimate if the molecule is of biotic or abiotic origin. This threshold-based comparison takes the MA for each molecule in isolation. In contrast, the present invention allows for the comparison of a multiple component sample using JMA.
[0088] The comparison of JMA for two samples may be used to estimate their phylogenetic relationship or estimate their phylogenetic character. This comparison is more sophisticated than a threshold based comparison, and allows for discrimination between closely related samples including similar components. The examples in the present invention illustrate that closely related animal, bacteria, fungi, plant and archaea can be distinguished. The examples show that different taxonomic families genus and even species can be distinguished using JMA. For example, the examples show how Staphylococcus aureus ‘ H’ , Staphylococcus epidermidis 8558, Staphylococcus hominis and Staphylococcus saprophyticus, bacteria species all in the same genus, may be distinguished using JMA (see Figure 24A). In contrast, WO 2021 / 186193 only demonstrates that baker’s yeast (Eukaryota) can be distinguished from Escherichia coll MG1655 (Bacteria).
[0089] While the MA described in this earlier work focuses on the complexity of what is typically the most complex molecule in a sample, the present invention makes use of the complexity of multiple components in the sample and their interrelation. The JMA is representative of the complexity of the sample as a whole, as it accounts for the diversity of the components as well as the components inherent complexity.
[0090] Thus, the present invention allows for useful and meaningful comparison of different samples - which has applications in identification of candidate pharmaceutical or agrochemical compounds, identifying diseased cells, identifying treatment-resistance pathogens and the detection of life.
[0091] Method for Estimating Molecular Complexity
[0092] In a first aspect of the invention, there is provided a method for estimating the molecular complexity of one sample or a plurality of samples, wherein each sample includes two or more components, the method comprising:
[0093] (a) obtaining MS / MS data for the sample or the plurality of samples, the MS data comprising MS1data for ions of the components, and MSndata for the nthgeneration of ion fragments formed from the ions of the components, wherein n is 2 or more;
[0094] (b) identifying the ion fragments in the MSndata for the ions of the components identified in the MS1data, and;
[0095] (c) calculating a joint molecular assembly (JMA) for the one sample or the plurality of samples from the ions of the components identified in the MS1 data and the ion fragments identified in the MSndata, wherein JMA is a measure of the complexity of the sample or the plurality of samples, which represents the complexity of the components and the shared complexity amongst the components.
[0096] The method may be for estimating the molecular complexity of one sample.
[0097] In some embodiments, the method is for estimating the molecular complexity of a sample, wherein the sample includes two or more components, the method comprising: (a) obtaining MS / MS data for the sample, the MS data comprising MS1data for ions of the components, and MSndata for the nthgeneration of ion fragments formed from the ions of the components, wherein n is 2 or more;
[0098] (b) identifying the ion fragments in the MSndata for the ions of the components identified in the MS1data, and;
[0099] (c) calculating a joint molecular assembly (JMA) of the sample from the ions of the components identified in the MS1data and the ion fragments identified in the MSndata.
[0100] The method may be for estimating the molecular complexity of a plurality of samples. The method may be for estimating the molecular complexity of two samples.
[0101] In some embodiments, the method is for estimating the molecular complexity of two samples, wherein each sample includes two or more components, the method comprising:
[0102] (a) obtaining MS / MS data for each sample, the MS data comprising MS1data for ions of the components, and MSndata for the nthgeneration of ion fragments formed from the ions of the components, wherein n is 2 or more;
[0103] (b) identifying the ion fragments in the MSndata for the ions of the components identified in the MS1data, and;
[0104] (c) calculating a joint molecular assembly (JMA) of the two samples from the ions of the components identified in the MS1data and the ion fragments identified in the MSndata.
[0105] The method may also consider neutral fragments. The neutral fragments are typically formed during formation of an ion fragment. An ion of the component (e.g., identified in the MS1data) may be fragmented to produce an ion fragment (e.g., identified in the MS2data) and a neutral fragment. An ion fragment (e.g., identified in the MS2data) may be fragmented to produce an ion fragment (e.g., identified in the MS3data) and a neutral fragment. Herein, the discussion of a ‘fragment’ may refer to a ion fragment and optionally a neutral fragment.
[0106] In some embodiments, step (b) further comprises determining neutral fragments formed during MSnfragmentation, and wherein in step (c) the JMA is calculated from the components identified in the MS1 data, the ion fragments identified in the MSndata and the determined neutral fragments.
[0107] In some embodiments the neutral fragments are determined by calculating the difference between a parent ion and a child ion fragment. That is, the mass of the neutral fragment is equal to the mass of the parent ion less the mass of the fragment ion.
[0108] In some embodiments the method is for estimating a phylogenetic relationship of a sample or the phylogenetic relationship of a plurality of samples.
[0109] In some embodiments the method is for estimating a phylogenetic distance between a plurality of samples. The phylogenetic distance between two samples refers to the distance between the two samples on a phylogenetic tree. The distance may refer to the number of branch points on the phylogenetic tree between the two samples.
[0110] In some embodiments the method is for estimating a phylogenetic characteristic of a sample or the phylogenetic characteristic of a plurality of samples.
[0111] In some embodiments the method is for estimating the phylogenetic clade of a sample or the phylogenetic clade of a plurality of samples.
[0112] In some embodiments, the phylogenetic clade is a taxonomic domain, phylum, class, order, family, genus or species. Preferably, the phylogenetic clade is a taxonomic class, order, family, genus or species. More preferably, the phylogenetic clade is a taxonomic family, genus or species. Even more preferably, the phylogenetic clade is a taxonomic genus or species. Yet more preferably, the phylogenetic clade is a taxonomic species.
[0113] For example, the method of the invention is for distinguishing samples of the same taxonomic domain, phylum, class, order, family, or genus. Preferably, the invention is for distinguishing samples of the same taxonomic genus. More preferably, the invention is for distinguishing samples of the same taxonomic species. That is, the method may distinguish between samples having a different genotype.
[0114] Sample
[0115] The method of the invention may be carried out on one sample. The method may be carried out on a plurality of samples, such as two samples.
[0116] Each sample comprises two or more components. That is, the sample comprises two or more different components. The sample may be known as a complex or multicomponent sample.
[0117] The sample may comprise 2 or more components, such as 10 or more components, 20 or more components, 30 or more components, 50 or more components, such as 100 or more components. Typically, the sample comprises from 2 to 10,000 components, such as 10 to 1,000 components, 20 to 500 components, 30 to 200 components, such as 50 to 200 components. In general, the components in the sample are taken to exclude solvent, such as an extraction solvent used to extract the sample. Thus, the components are those present in the pre-extracted sample. In general, the components in the sample are taken as the components having a molecular weight of 100 Da or more, as determined by MS / MS.
[0118] The sample may comprise 2 or more components which are identifiable in MS1data, such as 10 or more components, 20 or more components, 30 or more components, 50 or more components, such as 100 or more components which are identifiable in MS1data. Typically, the sample comprises from 2 to 10,000 components which are identifiable in MS1data, such as 10 to 1 ,000 components, 20 to 500 components, 30 to 200 components, such as 50 to 200 components which are identifiable in MS1data. In the examples, at least 30 ions of components are identified in the MS1data. An ion of a component is identifiable in the MS1data if it has sufficient intensity in the MS1spectrum for the peak to be identified.
[0119] The number of components may be determined using any suitable method. For example, the number of components may be determined using MS / MS, such as LC-MS / MS, for example by identification of the ions of the components in the MS1data. In addition, or alternatively, the number of components may be determined using chromatography, such as liquid chromatography (e.g., LC-MS / MS). In addition, or alternatively, the number of components may be determined using NMR.
[0120] The sample may be an environmental sample. That is, a sample collected from the environment. The number of components described above are typical for an environmental sample.
[0121] The sample may be an extract from a biological source (e.g. an extract from a fermentation broth, or a plant extract).
[0122] The sample may be from an extra-terrestrial source such as a sample of extra-terrestrial water, carbon dioxide, rock or other minerals. In such cases, the method may be used to detect extra-terrestrial life.
[0123] In some embodiments, the sample is a biological sample, such as extracts from a biological source.
[0124] In some embodiments each sample comprises two or more components, such as three or more, or four or more, selected from the group consisting of polysaccharide, polypeptide, lipid co-factor, steroid and polynucleotide. In some embodiments each sample comprises two or more components, such as three or more, or four or more, selected from the group consisting of polysaccharide, polypeptide, lipid and polynucleotide.
[0125] In some embodiments each sample is a whole cell sample.
[0126] A whole cell sample typically comprises cell membrane, cell nucleus, cytoplasm and organelles. The organelles may comprise mitochondrion, ribosomes, endoplasmic reticulum, golgi apparatus, and lysosomes. Typically a whole cell sample is lysed prior to obtaining MS / MS data on the sample.
[0127] In some embodiments each sample is a tissue sample. A tissue is a functional group of cells. That is, the sample is taken from the tissue of a biological multicellular organism. In some embodiments, the organism is a plant, and the tissue sample may be an epidermal tissue, a vascular tissue, or a ground tissue. In some embodiments, multiple samples of different tissues from the same organism are used, such as two or more tissues, such as three or more tissues.
[0128] In some embodiments each sample is an organ sample. An organ is a functional group of tissues. That is, the sample is taken from the organ of a multicellular organism. In some embodiments, the organism is a plant, and the organ sample may be a leaf, a stem, a branch, a flower or a root. In some embodiments, multiple samples of different organs from the same multicellular organism are used, such as two or more organs, such as three or more organs.
[0129] In some embodiments each sample is an organ system. An organ system is a functional group of organs. That is, the sample is taken from the organ system of a multicellular organism. In some embodiments, the organism is a plant, and the organ system is a root system or a shoot system. In some embodiments, multiple samples of different organs from the same multicellular organism are used, such as two or more organs, such as three or more organs.
[0130] By providing samples taken from the same tissue, organ or organ system of different organisms, it has been found that the phylogenetic relationship between the organisms can be more reliably determined.
[0131] In some embodiments the preparation step does not comprise removing components of the cell.
[0132] In some embodiments the preparation step does not comprise removing components of the tissue. In some embodiments the preparation step does not comprise removing components of the organ. In some embodiments the preparation step does not comprise removing components of the organ system. In some embodiments the preparation step does not comprise removing components of the multicellular organism.
[0133] In some embodiments, the sample is a whole tissue sample. In some embodiments, the sample is a whole organ sample. In some embodiments, the sample is a whole organ system sample. In some embodiments, the sample is a whole organism.
[0134] In some embodiments a sample is not purified. In some embodiments, a sample is not substantially purified. Thus, a sample may not be exposed to one or more of extraction, filtering, chromatography, concentration, and washing.
[0135] In some embodiments a sample is not purified beyond that required for MS / MS analysis. The MS / MS analysis may require some purification for practical purposes - e.g., for the sample to be introduced to the MS / MS machine. However, the sample is not purified to isolate particular components of the sample, such as polynucleotide or polypeptide. The samples may be whole cell samples. That is, the cells have not been treated to extract polynucleotide and / or polypeptide material. The samples may include a plurality of cells
[0136] In some embodiments each sample does not consist of a polynucleotide, such as wherein the samples do not consist essentially of a polynucleotide.
[0137] In some embodiments each sample does not substantially consist of a polypeptide, such as wherein the samples do not consist essentially of a polypeptide.
[0138] In this context the ‘sample’ only refers to components present in the original sample (e.g., the original biological sample). The ‘sample’ is not intended to refer to components added to the sample during sample preparation, such as extraction solvents, buffers or other additives, such as enzymes.
[0139] The sample may be extracted using a solvent. The sample may be dissolved or suspended in a suitable solvent. The sample may be filtered prior to analysis. Typically, the sample does not require further purification, for example chromatographic separation. The filtration may be required for introduction of the sample to a MS / MS machine.
[0140] Suitable preparation steps for the sample are described herein.
[0141] The use of whole cell, or an otherwise unpurified sample, simplifies the analytical procedure, and allows for the use of environmental samples directly. The methods of the invention do not require a high quality, purified sample, but can be carried out on a lower quality environmental sample without requiring extensive purification, and often only a single purification, such as filtration, merely for practical purposes for the introduction of the sample into a mass spectrometer.
[0142] MS / MS data
[0143] The methods of the invention comprise a step of obtaining MS / MS data for the sample or the plurality of samples. This may be known as the ‘MS / MS obtaining step’.
[0144] The MS / MS data obtaining step is typically carried out before the identifying step. The MS / MS data obtaining step may be after the preparation step, where present.
[0145] The method comprises a step of obtaining MS / MS data for the sample or the plurality of samples, the MS data comprising MS1data for ions of the components, and MSndata for the nthgeneration of ion fragments formed from the ions of the components, wherein n is 2 or more. MS / MS refers to tandem mass spectrometry, such as multistage MSn. In MS / MS, a first mass spectrum (MS1) of a sample is recorded. The ions in the MS1correspond to ions of the components of the sample. Then, an ion of the component (such as the parent ion) from the first mass spectrum is selected for fragmentation. After fragmentation, a second mass spectrum (MS2) of the ion fragments is recorded.
[0146] The MS / MS may include further generations of fragmentation. The generations of fragmentation are denoted as MSn, for the nthgeneration of ion fragments, where n is 2 or more. In some embodiments, n is 2 to 5, preferably 2 to 4, more preferably 2 to 3. Typically n is 2 or 3, such as 2. The greater number of generations (n) of fragmentation the greater the accuracy in the estimation of the molecular complexity, as the more information is obtained about the fragmentation (and precursor molecules) which make up the two or more components of the sample.
[0147] An advantage of tandem mass spectrometry is that it generates separate signals for different ions within a complex mixture. Thus, the method can be performed on a complex mixture of components of interest.
[0148] Suitable tandem mass spectrometry configurations include triple quadrupole (TQ), quadrupole time-of-flight (Q-TOF), ion trap time-of-flight (IT-TOF), quadrupole ion trap (Q-IT), quadrupole ion-cyclotron-resonance (Q-ICR), ion trap ion-cyclotron resonance (IT-ICR), ion trap orbitrap (IT-Orbitrap), double time-of-flight (TOF-TOF) and multistage MS (MSn).
[0149] Suitable fragmentation methods include in-source fragmentation, collision-induced dissociation (CID), electron transfer dissociation (ETD), electron capture dissociation (ECD), photodissociation and surface- induced dissociation. Preferably, CID is used. Suitable CID methods include low-energy CID and high-energy CID (HECID), for example higher-energy collisional dissociation (HCD) also known as higher-energy C-trap dissociation. Preferably, HCD is used as it is able to resolve ions of lower molecular mass, such as ions in the range 200 to 1 ,000 Da.
[0150] In some embodiments the MS / MS data for a sample is obtained in the same way for each sample.
[0151] In some embodiments obtaining the MS / MS data comprises performing MS / MS on each of the sample or the plurality of samples.
[0152] Preferably, the MS / MS data may be obtained by performing MS / MS on each sample separately.
[0153] In some embodiments the MS / MS data is obtained from a database for the sample or the plurality of samples. The MS / MS data may be retrieved from a database using any suitable means of communication. The MS / MS data obtained from the database typically includes information on the peak intensity and molecular weight for each MSnspectra. This may be known as a ‘peak list’.
[0154] The database may include data on the MS / MS spectra and the conditions under which the MS / MS spectra were obtained.
[0155] In some embodiments, the MS / MS data obtained from the database also includes information on the MS / MS conditions, such the instrument used, the method and / or parameters used for acquisition, and / or the methods and / or parameters and post-processing. In some embodiments, the MS / MS data obtained from the database also includes information on the sample, such as the identity of the sample. In some embodiments, the MS / MS data obtained from the database also includes information on the sample preparation, such as the extraction and / or purification method.
[0156] In some embodiments obtaining the MS / MS data comprises performing MS / MS on one or more samples and obtaining MS / MS data from a database for one or more samples.
[0157] Preferably, obtaining the MS / MS data comprises performing MS / MS on one sample and obtaining MS / MS data from a database for one or more samples.
[0158] Identifying Fragments
[0159] The methods of the invention comprise a step of identifying components and fragments in the MS / MS data. This may be known as the ‘identifying step’.
[0160] The identifying step is typically after the MS / MS obtaining step. The identifying step is typically undertaken before the complexity calculation step.
[0161] The method comprises a step of identifying the ion fragments in the MSndata for the ions of the components identified in the MS1data.
[0162] The identification step comprises identifying ion fragments in the MSndata and identifying ions of components in the MS1data. The identification step comprises identifying ion fragments in the MS2data and MS3data, and identifying ions of components in the MS1data.
[0163] The ions of the components identified in the MS1data may be filtered. The filtering is used to reduce the noise in the data and / or to allow peaks from impurities to be discounted.
[0164] In some embodiments the ions of the components identified in the MS1data are represented by peaks present within 0.01 Da across multiple MS1spectra and within 25% of the median peak intensity calculated across multiple MS1spectra. The sample contains multiple components, and so multiple component ions are selected in the MS1spectrum.
[0165] Preferably, the method comprises selecting 30 or fewer parent ions in the MS1spectrum, more preferably 25 or fewer parent ions, even more preferably 20 parent ions or fewer and most preferably 15 or fewer parent ions are selected in the MS1spectrum. Limiting the number of parent ions selected for fragmentation and analysis reduces the time required for analysis and improves the efficiency of the method.
[0166] The ions of the components identified in the MS1data may only correspond to the most intense peaks. Typically, the most intense component ions in the MS1spectrum are selected for fragmentation. That is, the parent ions having the largest intensity (counts per second, cps) are selected.
[0167] In some embodiments the ions of the components identified in the MS1data have from 5 to 30 of the most intense peaks in the MS1data, such as from 15 to 25 of the most intense peaks in the MS1 data, such as the 20 most intense peaks in the MS1data. The peak intensity may be determined as the median peak intensity calculated across multiple MS1spectra.
[0168] Typically, the ion fragments identified in the MSndata are formed from fragmentation of the ions of the components identified in the MS1data.
[0169] For example, ion fragments in the MS2data are typically formed from fragmentation of the ions identified in the MS1data. For example, ion fragments in MS3data are typically formed from fragmentation of the ions identified in MS2data. For example, ion fragments in MS4data are typically formed from fragmentation of the ions identified in MS3data.
[0170] The ion fragments identified in the MSndata may be filtered. The filtering is used to reduce the noise in the data and / or to allow peaks from impurities to be discounted. The filtering may be achieved by obtaining multiple MSnspectra (e.g., multiple MS2sepctra), such as three or more spectra, five or more spectra or ten or more spectra. The filtering may be achieved by obtaining multiple MSnspectra, such as from two to ten spectra.
[0171] In some embodiments the ion fragments identified in the MSndata are represented by peaks which are present in 35% or more of the MSnspectra. The ion fragments identified in the MS2data are represented by peaks which are present in 35% or more of the MS2spectra.
[0172] In some embodiments the ion fragments identified in the MSndata are represented by peaks which are merged if they are within ± 0.01 Da in the MS2 spectrum and / or merged if they are within ± 1.0 Da of an adjacent peak in the MS2 spectrum. The ion of the component identified in the MS1data may be selecting for or fragmentation.
[0173] The ion fragment identified in the MS2data may be selecting for or fragmentation.
[0174] Where the sample comprises a known molecule, the component ion can be appropriately selected in the MS1spectrum. That is, the parent ion can be selected to correspond to the mass of the molecular ion. Typical positive molecular ions take the form M+, [M+H]+or [M+X]+, where X is a cationic species, such as an alkali metal cation. Typical negative molecular ions take the form M-, [M-H]- or [M+Y]-, where Y is an anionic species, such as a formate anion.
[0175] Where the sample comprises an unknown component, the component ion can be selected in the MS1spectrum based on the observed mass.
[0176] Preferably, the method comprises selecting a component ion in the MS1spectrum having a minimum mass of 200 Da or more, more preferably 250 Da or more, even more preferably 275 Da or more and most preferably 300 Da or more. The inventor has found that molecules having this minimum mass provide a suitable amount of information for the JMA calculation.
[0177] Preferably, the method comprises selecting a component ion in the MS1spectrum having a maximum mass of 1 ,000 Da or less, more preferably 800 Da or less, even more preferably 600 Da or less and most preferably 500 Da or less.
[0178] The component ion may have a mass that is in a range selected from the upper and lower amounts given above, for example in the range 200 to 1 ,000 Da, such as 250 to 800 Da or 300 to 500 Da.
[0179] Where multistage MSnis employed, the fragment ion in the MS2is also selected for fragmentation. The fragment ion in the MS2data is then fragmented to form a fragment ion in an MS3spectra. The description of the selection of the component ion for fragmentation given above also applies to the ion fragment in the MS2data which may be selected for further fragmentation.
[0180] Optionally, the method comprises temporarily excluding from analysis those component ions which have recently been fragmented and analysed to provide an MS2spectrum. This maximises the number of component ions that can be selected and analysed. Methods of temporarily excluding component ions from analysis are known, and include Data Dependent Acquisition (DDA). In such methods, a dynamic exclusion window is used to ensure that component ions appearing multiple times in a set time interval are excluded from the analysis for a certain time period. The relevant time periods can be appropriately adjusted. Typically, the component ions appearing twice in a 5 to 30 second interval are excluded, such as component ions appearing twice in 10 seconds. Typically, such component ions are excluded from analysis for a period ranging from 10 to 90 seconds, for example 20 seconds, 30 seconds, or 45 seconds.
[0181] The method also comprises identifying the ion fragments in the MSndata. The ion fragments may be unique peaks, that is peaks having a unique mass and which correspond to distinct fragments of the component ion.
[0182] Typically, all peaks in the MSnspectrum are counted. Then, the peaks are filtered to arrive at the number of unique peaks. That is, certain peaks in the MSnspectrum are either disregarded or merged (combined) with adjacent peaks.
[0183] Typically, the method comprises disregarding peaks in the MSnspectrum having a maximum intensity below a certain threshold.
[0184] The threshold can be determined based on the spectrometer. For example, the method may comprise disregarding all peaks having a maximum intensity of 10,000 cps or less, more preferable 20,000 cps or less, even more preferably 40,000 cps or less and most preferably 50,000 cps or less. The threshold can be determined based on the observed MSnspectrum. For example, the method may comprise disregarding all peaks having a maximum intensity of 0.5% or less relative to the highest recorded intensity in the MSnspectrum, such as 1 .0% or less or 2.0% or less.
[0185] Typically, the method comprises merging all peaks having an observed mass within a certain distance of an adjacent peak in the MSnspectrum. Merging of peaks can be appropriately done based on standard methods. For example, the largest local peak (local maximum) can be chosen and the peaks within a given distance merged (discounted) with that peak. The process can be repeated until suitable spectra have been recorded. Merging of peaks can be appropriately selected based on the resolution of the mass spectrometer. Preferably, the method comprises merging all peaks within ± 0.005 Da of an adjacent peak in the MS2spectrum, more preferably the method comprises merging all peaks within ± 0.01 Da of an adjacent peak in the MS2spectrum.
[0186] One MSnspectra per component ion may be recorded. Typically, the method comprises repeating the measurement to generate multiple MS1and MSnspectra. In such cases, peaks not appearing in a certain proportion of the MS2spectra corresponding to the same parent ion can be disregarded. Thus, the analysis removes inconsistent peaks. This improves reproducibility.
[0187] Preferably, the method comprises disregarding all peaks present in fewer than 10% of the MS2spectra for a selected parent ion, more preferably fewer than 15%, even more preferably fewer than 20% and most preferably fewer than 25% of the MSnspectra for a selected parent ion. The number of peaks in the MSnspectrum may be adjusted based on the observed peaks in the MS1spectrum. Optionally, the number of peaks in the MS2spectrum may be divided by the number of peaks found within 0.5 Da of the parent mass in the MS1spectrum. This adjustment may compensate for co-fragmentation and merging of different MS1parent ions into the same MS2spectra.
[0188] Additional peak filtering may be used to account for excessive number of ions in the spectra. For example, the method comprises disregarding all peaks having a relative intensity below a certain fraction of the most intense peak in the MSnspectrum. Typically, the method comprises disregarding all peaks having a relative intensity below 2% of the highest recorded intensity in the MS2spectrum, such as below 5% or below 10%.
[0189] Joint Molecular Assembly (JMA)
[0190] The method comprises a step of calculating the molecular complexity for the one sample or the plurality of samples from the ions of the components and the ion fragments. This may be known as the ‘complexity calculation step’.
[0191] The complexity calculation step is typically undertaken after the identifying step.
[0192] The method comprises a step of calculating a joint molecular assembly (JMA) for the one sample or the plurality of samples from the ions of the components identified in the MS1data and the ion fragments identified in the MSndata.
[0193] Generally, JMA is a measure of the complexity of the sample or the plurality of samples, which represents the complexity of the components and the shared complexity amongst the components. That is, JMA is a measure of the complexity of the sample or the plurality of samples, which represents the complexity of each component and the complexity amongst the two or more components of the sample or the plurality of samples. The complexity amongst the components is indicated based on the commonality of precursors (fragments) amongst the components.
[0194] JMA may be a numerical measure.
[0195] JMA is calculated from the ions of the components of each sample and the ion fragments of the components. Preferably, JMA is calculated based on the ions of the components of each sample, and the ion fragments and neutral fragments of the components.
[0196] The JMA contribution from each ion fragments takes into account if that ion fragment is a precursor to multiple components. Thus, if the two or more components share precursors, the contribution from the common ion fragments to the JMA is lower (on a per component basis), indicative of the lower complexity attributable to structurally similar components. In the same way, unique ion fragments (which may only be present for one component, for example) provide a higher contribution to the JMA (on a per component basis), indicative of the higher complexity attributable to more structurally diverse components.
[0197] In some embodiments the JMA is calculated by:
[0198] (II) determining a shortest construction pathway for the ions of the components within the sample or plurality of samples from the fragments, optionally together with neutral fragments;
[0199] (II) calculating a joint molecular assembly (JMA) based on the location of the ions of the components and the fragments within the shortest construction pathway.
[0200] A construction pathway is the pathway from the fragments (ion fragments and neutral fragments) to the ions of the components. The MS / MS data reveals how the ions of the components may be ‘fragmented’ to give the fragments. The construction pathway is the reverse of the fragmentation - it illustrates how the ions of the components can be constructed from the fragments as building blocks.
[0201] Multiple ions of the components are present, and so the construction pathway illustrates how the multiple ions of the components may be constructed from the fragments. Typically, one fragment may be used in the construction pathway of two or more ions of the components.
[0202] The construction pathway typically takes the form of a branched network, connecting the ions of the components (at the top) with the fragments (below). Connections within the network indicate a route of construction.
[0203] A terminus of the construction pathway may be known as a ‘leaf node. The terminus is an ion of a component or a fragment (ion fragment, or neural fragment where present) which does not have a precursor on the construction pathway. That is, the terminus only includes a single connection to the construction pathway. The connection is typically to a lower generation MSnlevel. For example, where the intermediate is in the MS2generation, the connections are typically to the MS1generation.
[0204] The terminus may be able to undergo further fragmentation, but has not been selected for further fragmentation under the MS / MS method used. Alternatively, the terminus may not be able to undergo further fragmentation under the MS / MS method used (e.g., because the ion fragment is sufficiently stable that it cannot undergo further fragmentation).
[0205] In some embodiments the fragments are ion fragments only. Preferably, the fragments are ion fragments and neutral fragments.
[0206] Where neutral fragments are determined in step (b), the neutral fragments are assigned to the MS generation in which the corresponding ion fragment is identified. For example, if a neutral fragment is formed during MS2fragmentation, the neutral fragment is assigned to the MS2data with its corresponding MS2ion fragment.
[0207] In some embodiments, determining a shortest construction pathway for the ions of the components within the sample or plurality of samples further comprises stratifying the ions of the components and fragments into groups based on the MSndata in which they were identified.
[0208] In some embodiments, the shortest construction pathway is determined based on the fragmentation of the components and fragments.
[0209] The fragmentation may be the actual fragmentation determined from the MSnspectra. For example, the actual fragmentation may be determined by comparing the MS1spectra with the MS2spectra, and / or the MS2spectra and the MS3spectra. Preferably, the fragmentation for the MS2ion fragments are identified using the actual fragmentation.
[0210] In addition, or alternatively, the fragmentation may be a calculated fragmentation. The calculated fragmentation may be determined by any suitable means, such as computer simulation or using a database. For example, the calculated fragmentation may be determined for the ion fragments identified in the MSnspectra, wherein n is 2, or more using the calculated fragmentation. Preferably, theoretical fragmentation is used where n is 3 or more.
[0211] In some embodiments, the shortest construction pathway is determined by enumerating all possible construction pathways, and selecting the shortest pathway. The shortest pathway is typically the pathway with the fewest number of steps from the fragments to the ions of the components.
[0212] In some embodiments, in step (ii) the JMA is calculated by determining if each of the ions of the components or the fragments are at a terminus of the shortest construction pathway, and determining if each of the ions of the components or the fragments are intermediates on the shortest construction pathway. Here, terminus refers to a starting point which does not have a precursor on the construction pathway.
[0213] In some embodiments, in step (ii) the JMA is calculated by: estimating a molecular assembly (MAT) for each of the ions of the components and each fragment, which is a terminus of the construction pathway; calculating a sum of the MAT(Z MAT); calculating the total number of the ions of components and fragments which are intermediates on the construction pathway (Nj); and calculating a joint molecular assembly (JMA) as the sum of Z MATand Nj. The step of calculating a sum of the MAT(Z MAT) is a numerical sum of the MATvalues. The sum of the molecular assembly for each terminus (MAT) is determined.
[0214] The step of calculating the total number of ions of components and fragments which are intermediates on the construction pathway (Nj) is a numerical sum of the number of components and fragments (ion fragment and optionally neutral fragments). Preferably, the total number of components and fragments is the total number of unique components and fragments. Each component and fragment is given a standard value - e.g., one. The value of Nj is thus the total number of unique components and fragments which are intermediates on the construction pathway.
[0215] An intermediate on the construction pathway is an ion of the component or fragment (ion fragment, or neural fragment where present) which has a precursor on the construction pathway. That is, the intermediate includes two connections to the construction pathway. The two connections are typically to a precursor and a successor. The connections are typically to a lower generation MSnlevel (e.g., for the successor) and to a higher generation MSnlevel (e.g., for the precursor). For example, where the intermediate is in the MS2generation, the connections are typically to the MS1and the MS3generations.
[0216] In some embodiments, the MATis estimated based on the molecular weight of the component and / or fragment. In some embodiments, the MATis estimated based on the number of unique peaks for the component and / or the fragment. This is described herein.
[0217] Calculating Molecular Assembly
[0218] The MATmay be estimated based on the molecular weight of a component and / or a fragment.
[0219] Typically, the molecular assembly number is directly proportional with the molecular weight of the component or the fragment. Thus, the molecular assembly number (MA) can be calculated by scaling the molecular weight (z) by a certain magnitude (n). That is, the molecular number can be calculated using the equation:
[0220] MA = n(z)
[0221] Typically, the molecular assembly number is linearly correlated with the molecular weight of the component or the fragment. Thus, the molecular assembly number (MA) can be calculated by adding an off-set (b) to the molecular weight (z) after scaling by a magnitude (n). That is, the molecular assembly number can be calculated using the equation:
[0222] MA = n(z) + b The MATmay be estimated using on the method described in WO 2021 / 186193 at pages 14-16, the contents of which are incorporated by reference in their entirety.
[0223] The MAT is an integer measure that describes the complexity of a molecule. It indicates the minimum number of steps required to construct the molecule from basic building blocks (Marshall (2019)).
[0224] The MAT number is proportional to the number of unique peaks in the MSnspectrum. Thus, by scaling the number of unique peaks in the MSnspectrum, it is possible to determine the MA index of the molecule. Scaling (normalizing) the number of unique peaks to obtain the MA index of the molecule allows the values obtained from different experimental methods to be appropriately compared.
[0225] Typically, the molecular assembly number is directly proportional with the number of unique peaks in the MSnspectrum. Thus, the molecular assembly number (MA) can be calculated by scaling the number of unique peaks (x) by a certain magnitude (m). That is, the molecular number can be calculated using the equation:
[0226] MA = m(x)
[0227] Typically, the molecular assembly number is linearly correlated with the number of unique peaks in the MSnspectrum. Thus, the molecular assembly number (MA) can be calculated by adding an off-set (c) to the number of unique peaks (x) after scaling by a magnitude (m). That is, the molecular assembly number can be calculated using the equation:
[0228] MA = m(x) + c
[0229] Where the molecular assembly number (MA) is calculated based on the number of unique peaks in the MSnspectrum (xMSn), the value of m is typically in the range 0.3 to 0.7. Preferably, m is in the range 0.4 to 0.6, more preferably, 0.45 to 0.55.
[0230] Where the molecular assembly number (MA) is calculated based on the number of unique peaks in the MSnspectrum (xMSn), the value of c is typically in the range 4 to 9. Preferably, c is in the range 5 to 8, more preferably 6 to 7.
[0231] Preferably, where the molecular assembly number (MA) is calculated based on the number of unique peaks in the MSnspectrum (xMSn), the molecular assembly number (MA) is calculated using the equation:
[0232] MA = 0.48(xMSn) + 6.58 Most preferably, where the molecular assembly number (MA) is calculated based on the number of unique peaks in the MS2spectrum (xMS2), the molecular assembly number (MA) is calculated using the equation:
[0233] MA = 0.48(XMS2) + 6.58
[0234] Preparation Step
[0235] In some embodiments, the method further comprises a step of preparing the sample or the plurality of samples, wherein the preparation step precedes step (a). This may be known as a ‘preparation step’.
[0236] In some embodiments the preparation step is the same for each of the plurality of samples.
[0237] In some embodiments the preparation step comprises separating the two or more components by chromatography, such as liquid chromatography. The preparation step and MS / MS data obtaining step may be carried out using LC-MS / MS.
[0238] In some embodiments the preparation step comprising extracting the samples with a solvent, preferably using sonication, such as ultra sonification. The ultra sonication may occur at about 37 kHz.
[0239] The extraction typically includes mixing the sample with an extraction solvent, and sonicating the mixture. The extraction may take place at elevated temperature, such as about 30 °C.
[0240] Any suitable solvent may be used, such as an aqueous solvent or an organic solvent. Preferably the solvent is methanol or water, such as methanol.
[0241] Typically, the ratio of sample to solvent is about 1 to 100. For example, 0.1 g of sample is extracted in 10 g of solvent.
[0242] In some embodiments, the preparation step comprises freezing the samples, such as freeze drying the samples.
[0243] In some embodiments, the preparation step comprises grinding the samples. For example, the samples may be ground using a mortar and pestle. Preferably, the samples are ground to a size such that the particles have a cross-sectional area of 1 mm2or less.
[0244] In some embodiments, the preparation step comprises size sorting the samples, such as sieving the samples. Preferably, the samples are sieved to provide particles having a cross- sectional area of 1 mm2or less. Sieving may remove particles having a cross-sectional area of more than 1 mm2. In some embodiments, the preparation step further comprises freeze drying the sample and dividing the sample.
[0245] Method for Estimating a Phylogenetic Relationship
[0246] The method for estimating molecular complexity has several technical applications.
[0247] In a general aspect of the invention, there is provided a method for estimating a phylogenetic relationship of a first sample and a second sample, the method comprising estimating the molecular complexity of the samples, and estimating the phylogenetic relationship of the first sample and the second sample from the similarity of the molecular complexity.
[0248] A phylogenetic relationship is a relationship between two samples based on their evolutionary history. For example, two samples having a close phylogenetic relationship would share a high proportion of evolutionary history, while two samples having a distant phylogenetic relationship would share a low proportion of evolutionary history. The phylogenetic relationship is indicative of the closest common ancestor for the two samples. For example, two samples having a close phylogenetic relationship would have a close common ancestor, while two samples having a distant phylogenetic relationship would have a distant common ancestor.
[0249] In a second aspect of the invention there is provided a method for estimating a phylogenetic relationship of a first sample and a second sample, the method comprising: estimating the molecular complexity of the samples according to the method of the first aspect, to determine the JMA of the first sample and the second sample; and estimating the phylogenetic relationship of the first sample and the second sample from the similarity of the JMA of the samples.
[0250] In some embodiments the similarity of the JMA of the samples is calculated from the ratio of common elements to unique elements of the JMA of the first sample and the second sample.
[0251] In some embodiments the similarity of the JMA of the samples is calculated from the JMA of the first sample, the JMA of the second sample and the JMA of the first and the second samples.
[0252] JMA of the first and second samples refers to the combined JMA of the first and the second samples.
[0253] Preferably, the similarity of the JMA of the first and second samples is determined using joint assembly overlap (JAO).
[0254] JAO is a measure of the similarity of the JMA. The JAO is a numerical measure. The JAO compares the JMA of a first sample and the JMA of a second sample to the combined JMA of the first and second samples. The JAO is scaled based on the maximum possible value of the JMA for the first sample and second sample - to give a JAO value between 0 and 1 , where 0 represents no overlap between the JMA of the first and second samples and 1 represents complete overlap between the JMA of the first and second samples. ln some embodiments the similarity of the JMA of the samples is a joint assembly overlap (JAO), wherein JAO is calculated based on equation 1 :
[0255] Equation 1. where JMA(A) is the JMA for a first sample, JMA(B) is the JMA for a second sample JMA(A+B) is the JMA for the first and second samples and max(JMA(A),JMA(B)) is the maximum value for JMA(A) and JMA(B).
[0256] Typically, MS / MS data for the samples is determined for each sample (e.g., the first and second samples) independently.
[0257] The JMA(A) is the JMA for a first sample. The JMA(B) is the JMA for a second sample. The JMA for each sample is determined based on the ion components identified in the MS1data and the ion fragments identified in the MSndata for the sample. The JMA takes into account the different components present in each of the samples.
[0258] Typically, the MS / MS data used for determining JMA(A) and JMA(B) is MS / MS data on the first and second samples, respectively.
[0259] The JMA(A+B) is the JMA for the first and second samples. The JMA for the two samples is determine based on the ion components identified in the MS1data and the ion fragments identified in the MSndata for both of the samples. The JMA(A+B) takes into account the ion components and fragments present in both of the samples (i.e. , the common elements) and also takes into account the unique ion components and fragments present in only one of the samples (i.e., the unique elements).
[0260] As a result, the value of JMA(A+B) is usually smaller than JMA(A) + JMA(B). This is because the value of JMA(A+B) takes into account potential shared precursors (i.e., common elements) between the two samples, which are not accounted for in the calculation of JMA(A) and JMA(B). Thus, the contribution for the common elements appears in both JMA(A) and JMA(B), but is only accounted for once in JMA(A+B). Hypothetically, for two samples with no overlap, JMA(A) + JMA(B) would equal JMA(A+B); and for two samples which are identical JMA(A+B) would equal JMA(A) or JMA(B). Typically, the MS / MS data used for determining JMA(A+B) is MS / MS data on the first and second samples. The MS / MS data is typically obtained for the first and second sample independently. The ion components and ion fragments identified in the MSndata is typically identified in silico. In this way, it can be determined if the ion components and / or ion fragments are attributable to the first sample, second sample, or both the first and second samples.
[0261] The max(JMA(A), JMA(B)) is the maximum value for JMA(A) and JMA(B). The max(JMA(A),JMA(B)) is equal to the maximum value of JMA(A) + the maximum value of JMA(B). The max(JMA(A),JMA(B)) is used as a scaler, in order to allow for comparison using JAO. The max(JMA(A),JMA(B)) scales the JAO based on the maximum hypothetical value of JMA(A) and JMA(B).
[0262] Method for Estimating Phylogenetic Character
[0263] In a general aspect of the invention, there is also provided a method for estimating a phylogenetic characteristic of a first sample.
[0264] A phylogenetic characteristic is a characteristic associated with the evolutionary history of a sample. A phylogenetic characteristic of a first sample may be inferred by comparison of the first sample to a second sample. The second sample will typically have known characteristics. Where the first sample and second sample have a close phylogenetic relationship, the characteristics of the second sample may be inferred onto the first sample. However, where the first sample and second sample do not have a close phylogenetic relationship, the characteristics of the second sample may not be present for the first sample.
[0265] For example, if a second sample is a diseased cell, (e.g., a cancer cell), a first sample having a close phylogenetic relationship to the second sample is more likely to be diseased (e.g., cancerous) than a first sample having a distant phylogenetic relationship to the second sample.
[0266] In a third aspect of the invention there is provided a method for estimating a phylogenetic character of a first sample, the method comprising:
[0267] (a) estimating the molecular complexity of the first sample and a second sample according to the method of the first aspect, to determine the JMA of the first sample and second sample; and
[0268] (b) estimating the phylogenetic character of the first sample from the similarity of the JMA of the first and second samples.
[0269] Step (a) is as described above.
[0270] In some embodiments, step (b) comprises estimating the phylogenetic character of the first sample from the phylogenetic relationship of the first and second samples. In some embodiments the method is for estimating the phylogenetic clade of a first sample, wherein step (a) is repeated with one or more different second samples, each second sample being a reference sample for a different clade; and in step (b) the phylogenetic clade is estimated for the first sample from the similarity of the JMA of the first and second samples. A phylogenetic clade is a grouping based on evolutionary history.
[0271] The phylogenetic clade of the second sample is typically known. The phylogenetic clade of the first sample may be informed by the clade of a particular second sample.
[0272] In some embodiments, step (b) comprises assigning a phylogenetic clade to the first sample based on the phylogenetic clade of a second sample with which it has the greatest similarity.
[0273] In some embodiments the method is for estimating the phylogenetic clade of a two or more samples, wherein step (a) is repeated for each pair of two or more samples; and in step (b) the phylogenetic clade is estimated for the two or more samples from the similarity of the JMA of each pair of samples.
[0274] In some embodiments the phylogenetic clade is estimated from the pairs of samples having the greatest similarity.
[0275] In some embodiments the pairs of samples having the greatest similarity are calculated using a weighted pair group with arithmetic mean (WPGMA) algorithm for the similarity.
[0276] In some embodiments the method further comprises a step of generating a phylogenetic tree from the estimated phylogenetic clades of the samples. A phylogenetic tree is a graphical representation which shows the evolutionary history between phylogenetic clades.
[0277] The similarity is used to constructs the phylogenetic tree based on the pairwise similarity of each two samples. A distance matrix is formed to determine the similarity between all possible pairs of samples. The pairs of samples which are most similar are assigned adjacent branches on the tree.
[0278] Any suitable algorithm for determining the greatest similarity may be used. Preferably, similarity is calculated using a weighted pair group with arithmetic mean (WPGMA) algorithm.
[0279] In some embodiments the method is for estimating a characteristic of a first sample, the method comprising: comparing the similarity of the JMA of the first sample and second sample to a threshold value, wherein the threshold value is indicative of a characteristic. The complexity of a molecule may be correlated with the biological activity of a molecule. Thus, identifying highly complex samples is useful in providing candidate molecules for pharmaceutical or agrochemical development.
[0280] In some embodiments the method is for identifying a candidate pharmaceutical or agrochemical; the method further comprising: identifying a first sample having a similarity of the JMA of the first sample and second sample which is greater than a threshold value as a candidate agrochemical or pharmaceutical.
[0281] In some embodiments the second sample is a therapeutic or agrochemical target and / or a confirmed candidate agrochemical or pharmaceutical.
[0282] Thus, where the similarity is above a threshold value, the first sample may be identified as a being sufficiently similar to the second sample (typically, a known therapeutic or agrochemical target and / or a confirmed candidate agrochemical or pharmaceutical) that the first sample may also be a suitable candidate pharmaceutical or agrochemical.
[0283] In some embodiments the method is for identifying diseased cells in a first sample, wherein the first sample contains a suspected diseased cell; the method further comprising: identifying a sample having a similarity of the JMA of the first sample and second sample greater than a threshold value as diseased cells.
[0284] In some embodiments the second sample is confirmed healthy cells and / or confirmed diseased cells. The confirmed diseased cells may be cancer cells.
[0285] Thus, where the similarity is above a threshold value, the first sample may be identified as a being sufficiently similar or dissimilar to the second sample (typically, confirmed healthy cells or confirmed diseased cells) that the first sample may also be a healthy cell or diseased cell.
[0286] In some embodiments the method is for identifying treatment-resistance pathogens in a first sample, wherein the first sample is suspected treatment-resistance pathogens; the method further comprising: identifying a sample having a similarity of the JMA of the first sample and second sample greater that a threshold value as a treatment-resistance pathogen.
[0287] In some embodiments the second sample is confirmed non-treatment-resistance pathogen and / or confirmed treatment-resistance pathogen.
[0288] Thus, where the similarity is above a threshold value, the first sample may be identified as a being sufficiently similar or dissimilar to the second sample (typically, confirmed non- treatment-resistance pathogen or confirmed treatment-resistance pathogen) that the first sample may also be a non-treatment-resistance pathogen or treatment-resistance pathogen.
[0289] In some embodiments the method is for the detection of life in a sample, wherein the first sample is a suspected biotic sample and the second sample is a confirmed biotic and / or confirmed abiotic sample; the method further comprising: identifying a first sample having a similarity of the JMA of the first sample and second sample greater that a threshold value as biotic.
[0290] Thus, where the similarity is above a threshold value, the first sample may be identified as a being sufficiently similar or dissimilar to the second sample (typically, a biotic or confirmed abiotic sample) that the first sample may also be a biotic or abiotic sample.
[0291] In some embodiments the sample is a sample of extra-terrestrial material or terrestrial material, such as terrestrial soil, water, ice or rock.
[0292] Computer Program
[0293] The invention comprises computer-implemented methods and computer programs for estimating the molecular complexity of a sample, according to the method of the first aspect.
[0294] The invention provides a computer-implemented method for estimating the molecular complexity of one sample or a plurality of samples, the method comprising:
[0295] (a) obtaining MS / MS data for the sample or the plurality of samples, the MS data comprising MS1data for ions of the components, and MSndata for the nthgeneration of ion fragments formed from the ions of the components, wherein n is 2 or more;
[0296] (b) identifying the ion fragments in the MSndata for the ions of the components identified in the MS1data, and;
[0297] (c) calculating a joint molecular assembly (JMA) for the one sample or the plurality of samples from the ions of the components identified in the MS1data and the ion fragments identified in the MSndata, wherein JMA is a measure of the complexity of the sample or the plurality of samples, which represents the complexity of the components and the shared complexity amongst the components.
[0298] The invention also provides a data processing device comprising:
[0299] (a) means for obtaining MS / MS data for the sample or the plurality of samples, the MS data comprising MS1data for ions of the components, and MSndata for the nthgeneration of ion fragments formed from the ions of the components, wherein n is 2 or more;
[0300] (b) means for identifying the ion fragments in the MSndata for the ions of the components identified in the MS1data, and;
[0301] (c) means for calculating a joint molecular assembly (JMA) for the one sample or the plurality of samples from the ions of the components identified in the MS1 data and the ion fragments identified in the MSndata, wherein JMA is a measure of the complexity of the sample or the plurality of samples, which represents the complexity of the components and the shared complexity amongst the components.
[0302] The invention also provides a computer program comprising instruction which, when the program is executed on a computer, cause the computer to carry out the steps of:
[0303] (a) obtaining MS / MS data for the sample or the plurality of samples, the MS data comprising MS1data for ions of the components, and MSndata for the nthgeneration of ion fragments formed from the ions of the components, wherein n is 2 or more;
[0304] (b) identifying the ion fragments in the MSndata for the ions of the components identified in the MS1data, and;
[0305] (c) calculating a joint molecular assembly (JMA) for the one sample or the plurality of samples from the ions of the components identified in the MS1 data and the ion fragments identified in the MSndata, wherein JMA is a measure of the complexity of the sample or the plurality of samples, which represents the complexity of the components and the shared complexity amongst the components.
[0306] The invention also provides a computer-readable storage medium comprising instructions which, when executed by computer, cause the computer to carry out the steps of:
[0307] (a) obtaining MS / MS data for the sample or the plurality of samples, the MS data comprising MS1data for ions of the components, and MSndata for the nthgeneration of ion fragments formed from the ions of the components, wherein n is 2 or more;
[0308] (b) identifying the ion fragments in the MSndata for the ions of the components identified in the MS1data, and;
[0309] (c) calculating a joint molecular assembly (JMA) for the one sample or the plurality of samples from the ions of the components identified in the MS1 data and the ion fragments identified in the MSndata, wherein JMA is a measure of the complexity of the sample or the plurality of samples, which represents the complexity of the components and the shared complexity amongst the components.
[0310] That is, the invention also provides a computer-readable storage medium having stored thereon any computer program disclosed herein.
[0311] The invention further comprises computer-implemented methods and computer programs for estimating the phylogenetic relationship of a sample, according to the method of the second aspect.
[0312] The invention also comprises computer-implemented methods and computer programs for estimating the phylogenetic characteristic of a sample, according to the method of the third aspect.
[0313] Other Embodiments Each and every compatible combination of the embodiments described above is explicitly disclosed herein, as if each and every combination was individually and explicitly recited. Various further aspects and embodiments of the present invention will be apparent to those skilled in the art in view of the present disclosure.
[0314] “and / or” where used herein is to be taken as specific disclosure of each of the two specified features or components with or without the other. For example “A and / or B” is to be taken as specific disclosure of each of (i) A, (ii) B and (iii) A and B, just as if each is set out individually herein.
[0315] Unless context dictates otherwise, the descriptions and definitions of the features set out above are not limited to any particular aspect or embodiment of the invention and apply equally to all aspects and embodiments which are described.
[0316] Certain aspects and embodiments of the invention will now be illustrated by way of example and with reference to the figures described above.
[0317] Examples
[0318] Experimental Methods
[0319] To build a Tree of Life using Assembly algorithms, we required accurate MS2spectra for each analyte in the chosen samples. This necessitated a sample preparation stage that removed water and ground the sample to leave a powdered product (Figure 1A). The powdered sample was then passed through a sonication extraction procedure that provided an analyte mixture ready for MS analysis (Figure 1B). Comparisons between the samples and their controls (Figure 1 C-i) highlighted a high level of contamination and noise that must be omitted before the Tree of Life analysis. A background subtraction is therefore executed, providing a list of analytes native to the sample (Figure 1 C-ii). Due to the presence of contamination and noise, not all sample analytes were selected for fragmentation, while contamination analytes have been chosen erroneously (Figure 1 C-iii). A second round of MS analysis was performed with an input list of sample analytes, resulting in the fragmentation of all sample analytes. This data is subsequently used in the Assembly algorithms. Each section will be described in detail in the following sections.
[0320] Samples Database
[0321] To experimentally determine the MA of molecules within complex mixtures, a variety of animal (Table 1 ), bacteria (Table 2), fungi (Table 3), plant (Table 4), archaea (Table 5) and inorganic (Table 6) samples were prepared and analysed. For each sample, three sample repeats and a control were prepared via the methods detailed below. Table 1. Animal Samples
[0322] Table 2. Bacteria Samples
[0323] Table 3. Fungi Samples
[0324] Table 4. Plant Samples. Table 4A. Additional Plant Samples
[0325] Table 5. Archaea Samples.
[0326] Table 6. Inorganic Samples.
[0327] Sample Preparation
[0328] A variety of animal, bacteria, fungi, plant, and inorganic samples consisting of complex analyte mixtures were prepared and analysed to garner information on molecular assembly. For each sample, three sample replicates and a control were investigated.
[0329] All animal, fungi and plant samples were frozen overnight at -80 °C, within two hours of purchasing or picking. Samples were then placed in -48 °C freeze-drier at 10 mbar for 48 hours (Oyinloye et al.). A mortar and pestle were used to grind samples to a fine powder before being sieved to remove particles >1 mm2. Samples were then stored in a 4 °C fridge until extraction.
[0330] Ultrasound assisted extraction offers many advantages over conventional extraction techniques (e.g. Soxhlet or maceration) such as shorter extraction durations and lower temperatures. The lower temperatures required for extraction are particularly beneficial towards retaining the structural integrity of extracted compounds (Azwanida, Kumar et al. (2021 )). Extractions were conducted in triplicate using 0.1 g of sample and 10 ml_ of extraction solvent which were placed in a 37 kHz sonic bath for 20 minutes at 30 °C (see Figure 2). A control was also prepared that follows the same procedure without the addition of 0.1 g of sample. The resulting mixture was then removed and centrifuged at 4400 rpm for 6 minutes. Using a syringe, the supernatant (containing the extract) was carefully removed and then passed through a 0.2 pm Agilent syringe filter into a clean glass vial. 0.4 pL of the extraction solution was pipetted into an HPLC vial in preparation for HPLC-MS analysis while the remaining solution was stored in a 4 °C fridge.
[0331] Bacteria samples were cultured overnight in nutrient broth, at 37 °C, then centrifuged to separate the broth supernatant from the bacteria pellet. Inorganic samples were crushed and sieved to produce particle sizes < 0.5 mm. To procure a wide range of metabolites, ultrasound-assisted extraction (UAE, 20 mins / 30 °C) was employed with 10 mL of 80% methanol extraction solvent.
[0332] Culturing was carried out using aseptic technique to sterilize equipment (see Figure 3). Surfaces were first cleaned with ethanol, then 5% Chemgene. A colony from each strain was stored in a cryovial containing glycerol solution to ensure long-term survival and stored in a -80 °C freezer.
[0333] Bacteria were streaked on nutrient agar weekly to maintain healthy bacteria. When needed for extraction, a single colony was used to inoculate 50 mL of nutrient broth in a 250 mL sterile conical flask and incubated for 24 hours at 37 °C with shaking to promote growth.
[0334] Extraction requires separation of the nutrient broth from the bacteria via centrifugation (see Figure 4). Although knowledge is scarce regarding bacterial cell destruction due to centrifugal damage, the viability of E. coll has been proven to be significantly reduced at high rpm, indicating damage to the cell surface (Peterson et al.). Metabolite leakage into the nutrient broth due to this damage must be minimised so centrifugation was carried out at low rpm (3,000) for 10 minutes to prevent cell lysis. After removal of the supernatant, bacteria were separated into three centrifuge tubes each containing 0.1 g of bacteria. These were each quenched in 10 mL of -40 °C 80% methanol to stop the metabolism (Chen et al.). This mixture was then placed in a 37 kHz sonic bath for 20 minutes at 30 °C. The supernatant was carefully removed and filtered using a 0.2 pm Agilent syringe filter. 0.4 pL of the filtered solution was placed in an HPLC vial for HPLC-MS analysis while the remainder was stored in a 4 °C fridge. A control was present alongside each strain.
[0335] Archaea Samples were delivered in double-walled vacuum-sealed glass vials. In a Class II biosafety cabinet, the outer layer of glass was heated at one end until almost melting, then water was dripped onto the vial causing it to shatter. After removing the glass shards, the inner vial was extricated, and the plug removed to expose the freeze-dried archaea pellet. These pellets were reconstituted in 80% methanol before being divided equally between three sample replicates. The samples then followed the same extraction procedure as for the animal, fungi and plant samples as shown in Figure 2. Inorganic samples were crushed using the Mad Mining rock crusher before sieving to gather particles < 1 mm. The samples then followed the same extraction procedure as for the animal, fungi and plant samples as shown in Figure 2.
[0336] LC-MS / MS Analysis
[0337] Samples were analysed by HPLC-MS / MS.
[0338] Samples were reconstituted in 250 pL of H2O before LC-MSMS analysis. Chromatographic separation was achieved with a 5 pm, 200 A, 150 x 4.6 mm, SeQuant® ZIC-HILIC™ column on a Thermo U300 HPLC system with a flow rate of 0.400 ml / min. Mobile phases for the separation were Buffer A: H2O / 0.1 % Formic acid and Buffer B: Methanol / 0.1 % Formic Acid. A gradient was applied which consisted of a start at 95% B falling to 60% B at 30 minutes. The column was then washed at 55% B for 5 minutes, followed by an equilibration step for 5 minutes at 95% B. Total run time was 40 minutes per sample.
[0339] Samples were analysed in line by a Thermo Fusion Lumos Tribrid Orbitrap using a HESI ion source with a +4 kV voltage applied. Samples were analysed via a DDA method where the 20 most abundant ions were selected for MS2fragmentation (Defossez; Broeckling et al.). There was dynamic exclusion of ions for 90 seconds after the ion had been selected eight times in 30 secs. Ion fragmentation was achieved via HCD set to 35%, with an isolation window of 5 ppm. The resolution of the MS1full scan was 240,000 and the MS2fragmentation spectra set at 30,000.
[0340] Additionally, for the study of Single-Species Lineages (recursive culturing experiment of E. coll) the same settings were used with slight changes. Separation was achieved with a 7 pm, 150 x 4.6 mm Poroshell 120 EC-C18 column, run with a flow rate of 0.600 mL / min, where the mobile phase was ramped up to 95% B in 26 minutes. The acquisition was conducted via the automated DDA Thermo software AcquireX with a blank, an exclusion list run, an inclusion list run (sample), and two sample runs MS / MS acquisition with dynamic inclusion of fragmented ions. A mass range for analytes detection of 120-1200 Da was used.
[0341] Analytical Optimisation
[0342] The complex nature of the unprocessed samples necessitated an adaptive analysis protocol, implemented via data-dependent MS / MS acquisition (DDA), whereby the most intense analytes from MS1scans are selected for MS2fragmentation (Broeckling et al.). However, DDA can encounter limitations that result in the incomplete sampling of detectible analytes, such as constraints on the number of analytes available for MS / MS due to scan speed, masking of analytes by contamination and noise (Koelmel et al.). Since it was crucial to improve the abundance and quality of MS2spectra for these overlooked analytes (Davies et al.), a modified DDA approach was developed whereby an initial MS experiment was conducted to obtain a list of analytes (an inclusion list) unique to each sample, devoid of contamination and noise, before a secondary MS experiment was conducted specifically fragmenting analytes from the inclusion list.
[0343] Background Subtraction
[0344] Data analysis requires the .raw files originating from the Orbitrap to firstly be converted to mzML using MSConvert (Chambers et al.), then subsequently converted into .json using an in-house script mzmlripper. These .json files were then analysed using python.
[0345] The excellent sensitivity of the Orbitrap allows for the detection of trace quantities of analytes, however, this can also expose contamination from the column, the spectrometer, or the extraction procedure (Hecht). Figure 5 illustrates contamination that appears in both the sample and the control (highlighted in red) which must be removed before data analysis of the extracted analytes (highlighted in green). A python script was written to achieve this background subtraction (removal of control analytes, contamination, and noise) and works as follows.
[0346] We take unfiltered peaklists from the three sample repeats and control that consist of analyte mass-to-charge ratios and their corresponding intensities. For an analyte to be retained after background subtraction the analyte must firstly appear across all three sample repeats within a 0.005 mass window and within 25% of the median intensity (Figure 6(i)). If it only appears in one or two sample repeats, the analyte is discarded (Figure 6(H)). Likewise, if the analyte appears across all three sample repeats but at significantly different intensities, the analyte is discarded (Figure 6(iii)). Analytes that have successfully passed this stage are then compared against the control peaklist. If the analyte appears in the control above an intensity threshold (10% of the median intensity of the analyte in the sample repeats), the analyte is discarded (Figure 6(iv)). If the analyte does not appear in the control (Figure 6(v)) or appears at an intensity lower than the threshold (Figure 6(vi)), the analyte is retained. Thus, the only analytes that remain will be (I) seen across all 3 sample repeats, (II) at similar intensity levels, (ill) either not in the control or at very low intensities (this could just be slight column contamination from the previous run). The remaining analytes are now a set of compounds that we can confidently say are from the sample (Figure 7). AcquireX was used for the samples of the single-species lineages example (recursive culturing of E. coll experiment), in which all these considerations were implemented without a need for further runs or analysis.
[0347] However, after consulting this filtered peaklist, it became evident that the MS2data for sample analytes was poor- the analytes of interest were not being fragmented resulting from the contamination and noise in the raw data. As only the most intense ions from each MS1scan were selected for MS2fragmentation, the limited sampling period afforded by LC-MS analysis (compounds elute over a narrow time-period) means contamination peaks with high intensities can overshadow sample analytes (Hecht). MS2Optimisation
[0348] To build accurate MS2consensus spectra, the analysis workflow required modifications to increase the quantity of MS2spectra for each target metabolite (Figure 8). The recently acquired filtered analyte list was employed in a targeted inclusion list within the Orbitrap method (unique for each sample), where only analytes appearing on the list would be targeted for fragmentation. As a result, fragmentation data was gathered for all sample analytes, the quantity of MS2spectra greatly increased for these analytes and therefore the quality of the resulting consensus spectra also improved.
[0349] Building MS2Consensus Spectra
[0350] A script was written to generate fragmentation consensus spectra for individual analytes from their combined MS2spectra. The script initially separated thousands of raw spectra into sections of MS1or MS2spectra. Parent masses from the MS2data were grouped with their corresponding MS2spectra, before combining all the MS2spectra for analytes with the same parent mass. For example, there may be 10 MS2spectra for an analyte with a mass-to- charge ratio of 543.21. These 10 fragmentation spectra would be grouped to the single parent mass at 543.21 to create a single consensus spectrum. To remove noise, peaks in this consensus spectra were removed unless they appeared in at least 35% of the MS2spectra - in this example, peaks would need to appear in at least 4 out of the 10 MS2spectra. An intensity filter was applied, removing peaks of intensity < 10,000 au., to further reduce noise (Figure 9). The final stage of data cleaning was the removal of peaks that arose due to isotopic patterns. Often MS2spectra will feature isotope patterns containing fragments exactly 1 m / z apart, exhibited in Figure 9, so the script combined these peaks to not falsely assign them as separate fragmentation entities. The remaining peaks in the consensus spectra were then recorded, resulting in an output of a parent ion tied to its fragmentation consensus spectra.
[0351] Second Round MS Cleanup
[0352] After the second round of MS analysis in which target analytes have been selected for fragmentation, despite successfully recording fragmentation spectra for these analytes, the instrument also provided fragmentation data for additional analytes not specified in the targeted inclusion list. This was especially noticeable for the archaea samples which had a high number of target analytes, thus requiring multiple targeted analyses and so increasing the number of these additional fragmented analytes. These analytes do not appear on the target list and were subsequently removed from the archaea dataset using an in-house script that compared the target list with the parent ion of the MS2spectra. Thus, we have high quality fragmentation spectra for our full list of sample analytes, devoid of contamination or noise, ready to be implemented in the tree of life algorithm. This criterion was not used for the single-species lineage example (recursive culturing experiment of E. coll), in order to increase the coverage in the samples to accommodate the challenge of differentiating more homogenous samples. Initial Data Analysis
[0353] In our analysed dataset, every sample consists of detected analytes (after rigorous filtering) and their associated MS2fragments. To explore the data, we plotted the masses of all analytes in each sample and the amount of MS2peaks in their tandem spectra (Figure 10). Samples that are closely associated may present similar scatter distributions.
[0354] We expanded this analysis by looking at the scatterplots and their associated Kernel Density Estimate (KDE) maps of all samples for each predefined group in our dataset (Figure 11) to visualize distinct patterns. The KDE bandwidth estimation was conducted using Scott’s rule (Scott) with a scaling factor of 0.55.
[0355] Spectra Matching
[0356] In order to validate the sample preparation procedure and the analytical method employed in the protocol, the detected analytes for the E. coli sample were elucidated by matching them with reference and theoretically predicted spectra of known E. coli metabolites stored in the ECMDB (Guo et al.). Firstly, the spectra of the detected analytes were matched against the NIST database using MS Dial, an open-source software pipeline (Tsugawa et al.). The spectra matches attained in this way were then filtered to only those of known E. coli metabolites contained in the ECMDB. The E. coli metabolites detected using this method are Glutathione, Fumaric acid, 3,4-Dihydroxy-L-phenylalanine, Trehalose, Glutathione disulfide, Thymidine, Coenzyme A, Riboflavin, Adenosine, Thiamine monophosphate, L-Glutamine, NADH, S-Lactoylglutathione, and Pyridoxal 5’-phosphate.
[0357] Additionally, we acquired the SMILES of all known E. coli metabolites from the ECMDB, and performed an in-silico fragmentation procedure using the CFM-ID4 docker software (Wang et al.) and obtained ESI MS / MS predicted spectra. When compared to the experimentally attained spectra of E. coli analytes, matches were determined if a Jaccard similarity index between the lists of MS2peaks was above 0.1 , and the difference between the masses was below two Daltons. The confirmed annotations attained using this method are listed in Table 7. The diversity of elucidated metabolites in the sample qualify our analytical procedure, showing we are able to capture many metabolites of varied chemical structures and properties.
[0358] Table 7 provides a list of confirmed E. coli metabolites that were matched to analytes in the E. coli sample through comparison to theoretically predicted spectra generated by CFM-ID4. The matches reported are for analytes with a Jaccard similarity score of above 0.1 to confirmed E. coli metabolites, and that their mass is within 2 Da from the target metabolite mass. Table 7. A list of confirmed E. coli metabolites
[0359] Phylogenetic Modelling
[0360] With the generalisation of molecular to joint assembly space, a corresponding Joint MA index can be calculated for a group of molecules. Calculations of molecular assembly for single analytes is based on the Recursive MA algorithm presented in the inventors previous work (WO 2021 / 186193), the contents of which are incorporated by reference in their entirety.
[0361] Here we generalise the algorithm to estimate joint molecular assembly (Joint MA), i.e. the size of the joint assembly space of a collection of molecules, representing an entire sample. The goal of this extended definition is to use the estimated overlap between the joint assembly spaces as an indicator of causal relatedness between two samples. Calculating Joint MA using Recursive MA operates as shown in Figure 12, and follows the following steps:
[0362] 1. All observed ions are arranged into a tree structure, stratified by MS level.
[0363] 2. A pseudo-precursor representing the entire sample is added at the top of this hierarchy, with all MS1ions as children.
[0364] 3. For each child ion, its complement is calculated as MWparent - MWchiid and attached to the same tree layer. These inferred neutral fragments are not directly detected via mass spectrometry and thus need to be calculated. These neutral fragments may be deduced at every tandem MS level beside the MS1analytes level.
[0365] 4. Starting with the root node (A) all possible constructions of internal / branch nodes from child or sibling nodes ultimately terminating in a leaf node are enumerated.
[0366] 5. The MA of leaf nodes is estimated using their molecular weight as MA ± uncertainty.
[0367] 6. For each possible construction, the Joint MA is estimated as the sum of leaf MA estimates plus the number of internal nodes.
[0368] 7. The shortest path is selected among all possible enumerated construction operations detected in the data as argminpath(MA).
[0369] The same algorithm can be used to calculate the Joint MA of two samples, Joint MA(A, B) for the union of molecules and fragments in samples A and B, requiring only a second top level pseudo-precursor to be added, as shown in Figure 13. Although tandem MS used in this study was limited to MS2searching for sub-fragments among sibling ions allows molecular hierarchies more than one level deep to be discovered.
[0370] Joint MA index values lower than the sum of individual MAs indicate possible shared causal pathways among molecules in the group.
[0371] Spurred by the Joint MA index’s ability to capture the causal relationship between molecules, we proposed a further generalisation of this metric to tandem MS data of whole samples instead of individual molecules. We hypothesised that this new metric, Joint Assembly Overlap (JAO), could serve as an indicator of hereditary connection between samples. We formulated JAO of sample A and B to lie between 0 and 1 , 0 indicating no overlap and 1 full entailment, i.e. all molecules within one sample are contained within the joint assembly of the other. JAO is calculated thus: Equation 2. with JMA standing for Joint MA. Tree of life inference using this metric starts by calculating JAO for each pair of samples. Following this initial calculation, an approximate iterative process is used, in which a matrix of similarity scores amongst all samples is generated and the model clusters the samples that exhibit the highest similarity score. Using the Weighted Pair Group Method with Arithmetic Mean (WPGMA) algorithm, the matrix is then transformed to include the new clade containing the newly clustered samples by averaging their similarity scores to all the others groups in the matrix. This procedure is repeated until reaching a single root cladal group.
[0372] To generate a similarity matrix for phylogenetic inference, we calculated the JAO for each pair of samples in the dataset (Figure 14). This matrix represents the overlap of the joint assembly spaces of all pair samples, or how similar they are in term of assembly. At this stage, we considered developing an assembly-based phylogenetic algorithm that will simulate ancestors based on the assembly overlap of their descendants. Eventually, we opted to employ a more rudimentary approach, a Weighted Pair Group Method with Arithmetic Mean (WPGMA) algorithm (Sokal et al.). This is an iterative model, that at each step identifies the samples with the highest JAO score in the given matrix and cluster them. The JAO matrix is then transformed to include the newly-formed clade (containing the clustered samples) by averaging their JAO similarity scores between them and all other groups in the matrix. This procedure is repeated until arriving at a single cladal group. The recorded grouping sequence of samples were used to generate the phylogenetic trees.
[0373] Biogenicity Determination and Cladal Domain Classification
[0374] For the differentiation between biological and non-biological samples, and for the classification of biological samples into cladal domains, we compared our assemblygenerated JAO metric with the Otsuka-Ochiai coefficient (OOC) (Verma et al). For two samples A and B, their OOC is calculated thus:
[0375] Equation 3.
[0376] We performed the classifications using either a leave-one-out cross-validation method or a blind spectral clustering method. In the first method, we iterated over the samples in the dataset, each time one sample was kept aside, and the rest were grouped according to their predefined real groups. In the Assembly approach, we operated over the JAO matrix, where the prediction was made based on the maximal JAO value for the examined kept-aside sample with other samples, assigning it a paired sample group. For the OOC approach, we collected the MS1data (or MS2) of all the samples in each assigned group, and calculated the OOC of the examined kept-aside sample with all the inspected categories. We then gave a prediction based on the highest calculated OOC value between the sample and the categories.
[0377] The metric looks at the intersection of the groups A and B and divides them by the square root of the multiplication of their sizes. For MS1, all analytes of all samples belonging to each of the examined groups were organized in an exhaustive set, which was inspected for its overlap with the analytes of the sample. The model then gave a prediction for classification based on the highest OOC score among the presented categories. For the combined MS1and MS2model, separate cross validation models were generated for MS1and MS2, and their OOC values for their group annotations were averaged. The predictions the combined model produced are based on the maximal averaged OOC values among the examined categories.
[0378] The results for the biogenicity determination of samples using JAO and OOC metrics are featured in Table 8A and 8B, and the relevant OOC distribution for samples of biological and non-biological sources are presented in Figure 15. This shows KDE plots of the OOC values for the different predictions attributed by the life discrimination classifier to samples of biological and non-biological sources. The OOC values are a result of averaging the values of the MS1and MS2prediction models. A) For non-biological samples, the OOC values for their predictions. B) For biological samples, the OOC values for their predictions.
[0379] The OOC distributions for the classification of all cladal domains are presented in Figure 16. The Figure shows KDE plots of the OOC values attributed by the domain classifier to samples of each cladal domains. The distributions show the OOC values of each predicted label, and for each set of samples belonging to their “true” group in separate figures. The OOC values are a result of averaging the values of the MS1and MS2prediction models for A) Archaea samples. B) Fungi samples. C) Bacteria samples. D) Plant samples. E) Animal samples.
[0380] Table 8A shows the initial discrimination between biological and non-biological samples, based on preliminary data sets. Precision and recall values are tabulated for each Source type and for each of the two examined methods.
[0381] Table 8A. biogenicity determination of samples using JAO and OOC metrics
[0382] Table 8B shows the discrimination between biological and non-biological samples for the full data sets. Compared to the preliminary data set, the full data set includes additional data from tests with multiple tissues samples and single species lineages, which are more challenging to analyse. Precision and recall values are tabulated for each Source type and for each of the two examined methods.
[0383] Table 8B. biogenicity determination of samples using JAO and OOC metrics
[0384] For the second method of spectral clustering, we ran the algorithm on the JAO matrix, testing different numbers of clusters. We used the scikit python package. We reported the Silhouette score for each number of clusters in the algorithm, and the mean homogeneity of clusters, defined as the fraction of the dominant group within each cluster.
[0385] In more detail, the second classification method was a blind annotation using a spectral clustering algorithm (von Luxburg). Spectral clustering operates on an affinity matrix, conveying the similarity observed between discrete samples, and effectively cluster them by deriving the eigenvalues of the matrix and reducing its dimensionality. We conducted the clustering algorithm first for life detection, running the complete dataset of samples with specified two clusters (Figure 17A), showing good classification. We then removed the non- biological samples and performed the clustering algorithm again for cladal domain classification. We tested varying numbers of clusters with the algorithm, which presented cluster identities transformation reflecting known trajectories of evolution, with some noise (Figure 17B). Overall, both blind classification endeavours resulted in successful separation of predefined groups. We then performed the same spectral clustering algorithm on a matrix composed of OOC values of MS2data, to test whether the heuristic metric could perform to a similar standard. We have found comparably successful predictions for biogenicity discrimination (Figure 18A) and cladal domain classification (Figure 18B).
[0386] Structural Elucidation of E. coli Analytes
[0387] The MS2files were converted to abf files using AbfConverter (Reifycs). Samples were then checked and searched using MS Dial, an open-source software pipeline (Tsugawa et al.). Samples were searched against the NIST database in .msp format. Tolerances were set at 0.01 Da in the MS1and 0.05 Da in the MS2. Retention time tolerance was 100 mins, and was not used for scoring or filtering. The matches were then crossed with known E. coli metabolites from a public repository (Guo et al.). In addition, we extracted the SMILES of all metabolites in the repository and in-silico diffracted them using the CFM-IF4 docker software (Wang et al.) to generate ESI MS / MS predicted spectra and compared them to the experimental spectra. Matches were attained if a Jaccard similarity index between the lists of MS2peaks is above 0.1 , and the difference between the masses is below two Daltons.
[0388] Statistical Comparison Between Phylogenetic Trees
[0389] We used the R packages TreeDist and Quartet to run the Generalized Robinson-Foulds and Quartet similarity calculations, respectively. We defined four constraint states for the tree, and for each state generated 100,000 random trees whose randomness followed these constraints. T-tests were then conducted on each of the distributions, which were treated as the null-hypothesis for the assembly tree.
[0390] Relevant aspects of the general theory of object assembly are described below, along with the theory of molecular assembly and the model of molecular synthesis used to calculate the molecular assembly index. Concepts related to Assembly Spaces and the Assembly Index are formalised in Marshall (2019).
[0391] Assembly Theory
[0392] According to AT, every molecule has an inherent complexity (Molecular Assembly, MA) corresponding to the shortest construction pathway of its molecular graph (Figure 19A and 19B). Central to AT is the concept of assembly space, composed of a molecule’s precursors along its shortest construction pathway, where each fragment is recursively built up from simpler building blocks. The smallest possible cardinality of this space is indicated by the target molecule’s MA index and quantifies the minimal number of constraints within chemical space necessary for its construction. It has been demonstrated that this metric can be experimentally inferred using mass spectrometry and other analytical techniques (Jirasek et al.), and empirically verified that organic molecules with high MA values are indicative of life, making MA a genome-agnostic method for life detection (Marshal et al (2021 )). The concept of assembly space can be generalized to apply to a set of distinct molecules, where the shortest and simultaneous construction pathway for the entire set of molecules together is called a Joint Assembly Space (JAS) (Figure 19C). The pathways in a JAS represent the minimal set of constraints required to construct a given ensemble of molecules simultaneously (Liu et al.), consisting of the smallest number of construction steps required to generate all the target molecules. This number is termed Joint Molecular Assembly (Joint MA). Within the framework of a generalized AT, JAS represents the Assembly Observed (subset of Assembly Contingent) measured at a single point in time (Sharma et al.). This can be used as a powerful tool to reveal contingent relationships between individual molecules or entire chemical systems. This is particularly important when the exact chemical reaction pathways are too complex or the temporal formation histories of molecules are unknown. The JAS does not have explicit knowledge of geological time, but for a set of complex molecules it represents the underlying causal relationships that have a necessary temporal ordering for that set of molecules to be co-constructed through an evolutionary process.
[0393] Incorporating complexity considerations into molecular analyses could provide further insights into the informational architecture that governs biological life-forms (Sharma et al.). A biological cell is a constructor of its own underlying molecular network, which stays largely at compositional homeostasis (composome, a quasi-stationary state) (Lancet et al.',
[0394] Rados et al.), an evident result of selection, through intricate molecular interactions. Evolution facilitates the transformation of these networks over time, allowing exploration of new molecular compositions and assembly spaces that lead to molecular and functional novelty (Sharma et al.). It has been shown that idiosyncratic signatures in the compositional metabolome of species may facilitate bottom-up inference about higher-order informational levels in the cell such as the phenome and the genome (Johnson et al.', Rosato et al.) (Figure 20A). In this light, the metabolome may be regarded as an expansive informational arena describing the historical contingency stored in molecular identity (Borenstein et al.).
[0395] It is thought that genome-agnostic molecular information coupled with AT may allow the inference of evolutionary trajectories, manifesting in the transformation through time of assembly spaces, from shared ancestors to current day species (Figure 20B).
[0396] Fingerprinting Idiosyncratic MS-Derived Profiles of Samples
[0397] The analytical process described in the ‘experimental methods’ section allowed us to gather information about the analytes residing in the organismal samples. Even though the analytes were not structurally elucidated, we found that non-targeted MS1and MS2information could be sufficient for species classification. Data exploration showed that this is particularly true for differentiating between living and non-living samples - analysis of inorganic samples revealed distinctly less intense peaks (Figure 21 A) and fewer analytes altogether (Figure 21 B). We approached each set of analytes as an idiosyncratic fingerprint of a sample’s molecular network, which allows its classification within a larger clade based on shared analyte information among the relevant samples (Figure 21 C). To validate our experimental approach, we matched the detected analytes from the E. coll sample to molecules with reference or predicted spectra (see experimental methods section). Most detected molecules do not map to known or predicted spectra of E. coll metabolites, possibly due to inadequate databases, inaccurate parameterization, or algorithmic limitations.
[0398] These analytes may also belong to an underground metabolism, encompassing reactions that are not genomically encoded and thus not readily investigated and recorded (D’Ari et al.). Nevertheless, the results reveal several universal biomolecules with high certainty (such as nucleobases, amino acids, sugars, and organic cofactors) as well as more specific bacterial lipids with lesser detection certainty (see ‘experimental methods’ section). This reaffirms that our analytical procedure is general enough to facilitate detection of distinct analyte chemistries across a wide mass-range, allowing faithful metabolic mapping of the organismal samples.
[0399] Genome-Agnostic Mapping of Phylogenetic Relations Between Organisms
[0400] To uncover the underlying cladistic architecture of the sampled organisms, we developed a numeric model based on the Recursive MA algorithm (Jirasek et al.) that operates on hierarchical tandem MS data to constrain and estimate MA values for target molecules (Figure 22A). The sets of molecular fragments in organismal samples form MS-derived “construction spaces”, as they are the results of empirical deconstruction of their detected precursor analytes, and therefore encompass structural motifs that appear in the molecular networks. These spaces are comparable to multimolecular JAS (Liu et al.), which refers to the combined assembly steps and substructures included in the graphical construction process of the examined set of target molecules. Deeper levels of tandem MS represent construction operations between smaller fragments, which scale up in complexity across the MS data levels towards their original detected molecules, emulating the JAS hierarchical structure (Figure 19). This comparison is especially potent considering all metabolites and their fragments originate from the same metabolic synthetic system, an elaborate network of interconnected pathways expanding from a shared set of basic building blocks. We therefore termed the MS-derived construction space as a Joint Fragment Assembly Space and set to investigate its causal relationships using AT analytical tools.
[0401] By employing the Recursive MA algorithm, we generated a Joint Fragment Assembly Space for each individual sample in our dataset, recursively constructing their composite analytes from tandem MS data. We used these inferred analyte hierarchies to deduce the Joint MA values of discrete samples, describing the estimated level of complexity of their assembly pathways (see ‘experimental methods’ section). By comparing the Joint MA of individual and groups of samples, we were able to estimate the Joint Assembly Overlap (JAO) of the samples examined (see ‘experimental methods’ section). This metric describes the relatedness of species in the dataset based on the similarity of the Joint Fragment Assembly Space for individual species to those of multispecies composites (see methods). This analysis does not rely on any prior molecular information of the detected analytes, thus enabling agnostic phylogenetic inference. The interrelated evolutionary trajectories of species from shared ancestors explain why there is a degree of shared MS2fragments among different molecules within and across samples. Indeed, a larger number of mostly lower-mass fragments occur in multiple samples, with the probability of shared fragments among samples diminishing at higher fragment masses (Figure 22B). This illustrates the overlapping assembly spaces among samples, supporting the applicability of this feature for elucidating evolutionary trajectories.
[0402] Biogenicity Determination and Cladal Domain Classification With Assembly Spaces To investigate the capacity of AT to correctly discriminate between defined groups of samples, we employed it for two discrete tasks.
[0403] First, we attempted to distinguish between biological and non-biological samples. We used two classification methods: a supervised leave-one-out cross validation method, and an unsupervised spectral clustering method (see methods). As a control, we compared the assembly-based JAO metric with the heuristic Otsuka Ochiai coefficient (OOC), which is similar to the Jaccard similarity index that was originally employed in biological studies (Verma et al.) (see methods). Both metrics presented comparably successful prediction results, easily determining the biogenicity of samples in both targeted and non-targeted classification methods (see experimental methods’ section). Notably, blind biogenicity determination can be reliably achieved beyond relational clustering by looking at the MA of molecules detected in samples (Marshall etal. (2021 )), and this can be clearly seen in our database (Figure 22C).
[0404] Second, we tested whether we can correctly annotate biological samples as belonging to different predefined cladal domains, including Animalia, Archaea, Bacteria, Fungi and Plantae. The assembly-based JAO metric performed marginally better than the heuristic OOC metric in the unsupervised clustering method, better capturing a cohesive partition of samples into domains with up to 5 clusters (Figure 22D). The grouping of samples alongside the clustering operation depicted very logical decision-making that follows evolutionary trajectories, where taxonomical groups separate according to their rank and chronology (see SI section 8). Noticeably, the JAO method gave better classification results than OOC in supervised clustering (Table 1). It also allowed closely related taxa to be distinguished. Precision and recall values are tabulated for sample group assignment. For MS1+MS2, the OOC values were calculated for MS1and MS2separately and were averaged before giving predictions.
[0405] Table 9. Cladal domain classification, using Joint Assembly Overlap (JAO) and the Otsuka Ochiai coefficient (OOC)
[0406] Where the assembly-based metric presented flawless prediction accuracy, the heuristic method confused some plant and bacteria samples as fungi using only MS1data and made even more errors when implementing MS2information. It appears that even though heuristic methods can roughly identify clusters, considering the hierarchical nature of the data allowed the assembly algorithm to better discriminate between samples, especially when the groups are less discrete and more challenging to partition in the dataset. Furthermore, the JAO algorithm provides an explanatory ground truth with the OOC metric as a direct consequence as a set-based measure. Therefore, their comparable performance in our testing suggests that the heuristic approach approximates assembly-based inference, which captures the causal structures of the samples.
[0407] Agnostic Tree of Life Inference Using Assembly Theory
[0408] Contrary to the cladal domain classifier, phylogenetic inference does not rely on database comparison or predefined group annotation. It is often performed without a priori taxonomical information, and requires more complete datasets, as it aspires to infer evolutionary trajectories across time at higher resolution and their interconnectedness. Using our phylogenetic pipeline, we generated an agnostic tree of life based on MS information elucidated via AT (Figure 23A). The tree displays rational phylogenetic clustering and distinction, closely reflecting evolutionary trends that appear in genome-based phylogenetic studies, most importantly the cohesiveness and consolidation order of general cladal domains (Figure 23B). In comparison, a tree generated using the heuristic Otsuka-Ochiai coefficient (OOC) did not adequately capture known and deducible evolutionary trends. Clustering choices reveal limited success in capturing cohesive cladal domains, independent of the level of tandem MS employed in the tree construction. It appears that a ground truth assemblybased approach provides for a much more accurate cladistic configuration.
[0409] Closer examination of the resultant tree of life model reveals excellent distinction between closely related taxa - such as similar species (Figure 24A). Samples that belong to the same species are distinguished, but still closely positioned together (see experimental methods), and there is logic to the clustering at the genus and family levels, such as the cohesive clustering of all Staphylococci species, and the partitioning of halophiles and acidophiles in the archaea domain. Furthermore, there are some interesting cladistic choices that could be a result of our analytical procedure and particular choice of samples. For instance, nuts and seeds tend to associate closely, possibly reflecting their unique metabolic constitutions which may not be representative of other plant tissue. Similarly, in the animalia kingdom, land and sea animals cluster separately, though this is not far from the genomic model for this set of samples (see Figure 24C). Those phenotypic mappings may be overcome by more representative sampling, and by increasing the analyte coverage of the samples, further establishing the unique identities of the molecular fingerprints necessary for classification. This may explain some of the clustering choices for samples of fewer analytes, such in the bacteria domain. However, these phylogenetic “errors” are relatively minor, and reaffirm the reliability of our approach. Altogether, it appears that the agnostic tree of life produced by our assembly-based phylogenetic model using non-targeted small metabolite information, offers evolutionary insights that appear consistent with consensual biological knowledge.
[0410] In addition to the phylogenetic tree included discussed above, we generated three additional trees for comparison and validation.
[0411] First, we looked at a tree where multiple samples of individual species are separately included (Figure 25). Samples that belong to the same species tend to cluster together, further validating our assembly algorithms. Next, we generated a phylogenetic tree based on the OOC algorithm, with MS1information (Figure 26) or with MS2information (Figure 27), and a OOC tree based on MS2information with multiple samples of individual species collapsed to discrete leaves (Figure 28). This simple variation of the Jaccard similarity index worked well for general domain classification, however it was less successful when challenged with the more demanding task of producing an accurate phylogenetic tree, which requires higher-resolution clustering. The difference between the MS1and the MS2trees appear to be nonsignificant.
[0412] Validation of Assembly Phylogenetics
[0413] To substantiate our tree of life model, we first inspected the data-driven decisions of our assembly-based phylogenetic model. We observed that samples that are more closely clustered on the tree tend to share a higher degree of high-masses fragments (Figure 24B). This can be gleaned from the increasing contribution of low-mass fragments and the decreasing contribution of high-mass fragments to the set of overlapping fragments between samples the more they are distanced on the tree.
[0414] In more detail, for validation of the generated assembly tree, we first looked at the distribution of the overlapping tandem MS fragments between samples clustered to different degrees on the tree. We used geodesic distance, the number of nodes in the shortest path on the tree between individual samples, as the distance measure. The individual histograms and their Kernel Density Estimate (KDE) plots can be seen in Figure 29. The contribution of low-mass fragments is decreasing, and the contribution of high-mass fragments is increasing, the closest the samples are clustered on the tree.
[0415] This finding substantiates the model and connects it further to the fundamentals of AT.
[0416] Comparison with the Genomic Tree
[0417] We compared our assembly-based tree to a consensual phylogenetic tree (generated from the TimeTree database (Kumar et al. (2022))), to examine whether our agnostic model conforms with models constructed from genomic information.
[0418] We opted to use the TimeTree database (Gregor et al.), a public knowledge-base for information on the evolutionary timescale of life. It consolidates data from thousands of published studies based mostly on genomics, but also on palaeontologic and geologic investigations. When building a tree for our specified species, TimeTree made a few substitutions of close species for which it had evolutionary data. Substitutions include Aeromonas sobria (replacing Aeromonas salmonicida), the genus Grifola (replacing the species Grifola frondose), Rheum rhaponticum (replacing Rheum x hybridum), Clausia trichosepala (replacing Brassica oleracea) and Musa balbisiana (replacing Musa acuminata). In contrast to our tree, which we limit to clustering only two clades at each step, this tree has a few multi-clades junctions in their tree, in accordance with complex phylogeny or ambiguity in published data. Nonetheless, the genomic tree follows similar trends to the assembly tree (Figure 30).
[0419] In order to statistically test the similarity between the two trees, we generated distributions of random trees with different levels of constraints and compared them to the genomic tree using Generalized Robinson-Foulds similarity (Smith (2020) (Figure 24C) and Quartet similarity as robust phylogenetic analytical metrics (Smith (2022).
[0420] The GRF metric quantifies the dissimilarity between a pair of trees by summing the number of splits (bipartitions of the set of leaves, corresponding to clades) that are unique to either tree. The metric also includes a generalization apparatus which integrates probabilistic measures of split similarity and allow tree similarity to be measured in natural informational units. Quartet similarity approaches tree similarity from a different direction. Instead of observing similarity in tree partitions, it quantifies symmetric differences by counting the number of four- taxon groups (quartets) that differ between two trees in their phylogenetic clustering order. By comparing all possible quartets, an exhaustive measure of tree similarity is generated.
[0421] In order to use these metrics, we employed the R packages TreeDist and Quartet in order to run the Generalized Robinson-Foulds and Quartet similarity calculations, respectively. We defined four constraint states for the tree, corresponding of the distributions in Figure 24C and Figure 31. The random distributions had constraints spanning from no limitations to randomness only within each defined cladal domain. For each state we generated 100,000 random trees following their randomness constraints. T-tests were then conducted on each of the distributions, which were treated as the null-hypothesis for the assembly tree. Both the Robinson-Foulds and the quartet tree similarities show that the assembly tree highly agrees with the genomic one and is highly non-random.
[0422] The results reveal that the assembly-based tree is substantially non-random and aligns well with the genomic tree with high statistical significance. In contrast, the OOC tree model that is not assembly-based presented low Robinson-Foulds similarity with the genomic tree, positioned among the random tree distributions.
[0423] Additional Comparison with the Genomic Tree using Multiple Tissue Samples
[0424] It was noticed that multicellular organisms tended to be clustered according to the type of tissue sampled, such as the clustering of the seeds and nut, as well as the fruit-based samples. This is indicative that morphological clustering could pose a challenge for phylogenetic inference. To investigate this effect, a new set of samples were assembled representing seven species of which three organ types were sampled - Leaves, Branches and Flowers. The samples were processed quickly and effectively on two occasions: all samples from five species (S. wallisi, B. integrifolia, R. ponticum, C.japonica (v) and C. japonica (sdd)) were processed and analysed together on the mass spectrometer, and later the other two species (C. superba and C. splendens) were processed and analysed together. The AcquireX software was used. The 21 samples were then analysed to construct their assembly spaces and calculate their JAO values (Figure 33), which were used to generate the assembly-based phylogenetic tree (Figure 32) in comparison to trees generated by the control metrics (Figure 34 and Figure 35). All these models presented relatively similar trees that were clearly and dominantly tracing phylogeny, much more than morphology. In fact, morphological clustering only appeared at the strain level between the two variants of C.japonica. The separate processing and analysis of C. superba and C. splendens could have affected the results somewhat, as they were persistently clustered together in all trees, though this is also a feature of the genomic-based tree constructed according to the TimeTree database. The assembly-based tree was compared to a consensual phylogenetic tree (acquired from the TimeTree database), to examine whether the resultant model conforms with one constructed from genomic information. Distributions were generated of random trees with different levels of constraints and were compared to the genomic tree using the Generalized Robinson-Foulds (GRF) similarity (Figure 31) and the Quartet similarity as robust phylogenetic analytical metrics. The GRF similarity metric describes the common taxa between the trees, i.e. the shared sets of samples that are hierarchically clustered in both tree models. The Quartet similarity metric describes the fraction of quartets of samples that are clustered similarly in the tree models. The results of both analyses reveal that all phylogenetic methods produce trees that are substantially non-random and align relatively well with the genomic tree with high statistical significance, and with the assembly-based JAO tree performing better than the controls. The MS1&2tree presented marginally lower prediction accuracy (with the GRF metric) than the MS1tree, indicating that the addition of MS2information adds some interference into the model by treating both types of observations as detected analytes of equal importance. From the data, it appears MS2information could be used to improve the prediction if processed through the pathway-dependent assembly framework.
[0425] The resulting tree model presents phylogeny-governed clustering choices mostly aligned with the genomic-based tree (Figure 32B). The only contribution of morphological clustering to the tree model occurred at the strain level, indicating that phenotypical features influence the eventual structure of the tree model only when the phylogenetic contribution becomes sufficiently minimal. This example emphasizes the importance of sample preparation and quality on credible inferences, which could likely explain clustering errors in our assemblybased preliminary tree of life model.
[0426] Additional Phylogenetic Inference of Single-Species Lineages
[0427] To further substantiate the MS-based phylogenetic approach, it was tested on a case where both morphological and genetic inferences would be challenging and less effective in predicting accurate phylogenetic structures. A bacterial growth system in which E. coll was recursively cultured and plated was used, keeping genomic variation to a minimum in the absence of any exerted selection pressure. In each recursive cycle, an E. coll culture was plated and two colonies were picked from the plate and cultured separately (Figure 36A), generating an overall tree-like experimental design comprising four bacterial generations (Figure 36B). In this system, the phylogenetic inference was not meant to predict the evolutionary history of distinct species, but that of related colonies of the same species, based on their detected molecular compositions. Therefore this system tests whether assembly-based phylogenetics is able to detect molecular transformations attributed to evolutionary processes on a much shorter timescale than those investigated in a tree of life model, a result of phenotypic variation that is likely not genomically encoded. In genomics, different regions present diverse mutation rates, and we similarly argue that different metabolomic features vary at disparate rates, and thus an exhaustive analysis of the cellular composome would yield transformations of phylogenetic importance.
[0428] The analysed samples revealed a significant degree of homogeneity, as expected. 2186 unique MS1analytes and 1925 unique MS2fragments were detected across the eight leaf nodes in the tree, each individual sample comprising about 1 ,400 analytes and 1 ,300 unique fragments. The low variation across the dataset signifies a pronounced difficulty to elucidate the experimental tree structure. A simple MS1Overlap approach gave an almost completely random tree prediction at the 62ndpercentile of all possible trees (see Figure 36C). The assembly-based JAO method performed significantly better, with slightly even better performance for the MS1&2Overlap method, producing tree predictions at the 96thand 99thpercentiles respectively. This means that only a few percentiles of all possible tree structures (a few thousands out of 135, 135 possible trees) are closer to the true experimental tree than the assembly-based tree model. This result also indicates that MS2information could be more valuable to the detection of evolution on a short timescale than just MS1analytes. This is further confirmed by the observation that the JAO and MS1&2Overlap predictions appear to be relatively indifferent to the size of the dataset, whereas the accuracy of the MS1Overlap approach goes down the more samples are included in the tree (Figure 36D). Notably, the assembly-based metric seems to be more consistent and less influenced by outliers than the two control methods, showing lower variance of tree prediction accuracies. These results illustrate the adeptness and reliability of AT to correctly infer phylogenetic and causal relationships from molecular observations in highly homogenous systems that would prove challenging for previously established techniques.
[0429] The samples were shown to be highly similar to each other, a level of homogeneity expected for related samples of the same species (Figure 37). Nonetheless, peripheral differences in the molecular compositions of the samples could be indicative of cross-generational changes that could be used to trace the lineages of the strain. The phylogenetic metrices (MS1, MS1&2and JAO) were employed to generate phylogenetic models and compared them to the experimental tree structure, using the GRF (Figure 38) and Quartet (Figure 39) tree similarity heuristics. A jackknife cross validation analysis was also performed, using both the GRF and Quartet heuristics, to test the reliability of our phylogenetic metrices in generating accurate tree predictions (Figure 40A and B).
[0430] Comments
[0431] The above examples show how assembly theory may be used to devise a simple quantitative system for detecting, distinguishing between and classifying living species without reference to their genetic material, or indeed any prior knowledge of their biochemistry, using a robust general analytical workflow. The examples show how a non-targeted genome-agnostic direct examination of molecular compositional information of living species and can be modelled for accurate phylogenetic inference. It is thought that this agnostic assembly-based tree model can be used to map the evolution of life on earth. This methodology has the potential to provide an efficient diagnosis tool of species complementary to genome-based efforts, and should enable pushing our understanding of life, its origin and evolution, beyond current models and definitions. This approach is shown to work both over longer term evolutionary trajectories and short-term phenotypic trajectories, leading to divergence in populations where no detectable genomic evolution has yet sufficiently occurred.
[0432] Abbreviations
[0433] Assembly Pathway - the shortest sequence of join operations required to graphically construct a target object, such as a molecule.
[0434] Assembly Space - the set of precursors in the construction of target objects, each recursively built up from simpler building blocks.
[0435] Composome - a quasistationary state in a dynamic description of compositional molecular assemblies. Cells exist in compositional homeostasis, and therefore their idiosyncratic molecular network can be termed a cellular composome.
[0436] Da - Daltons
[0437] GD - Geodesic Distance, the length (or number of edges) of the shortest path between two vertices (or nodes) in a network graph.
[0438] GRF - Generalized Robinson-Foulds metric, wildly used for quantifying the similarity between phylogenetic trees.
[0439] JAO - Joint Assembly Overlap, a metric describing the similarity of several individual JFAS, and used here to compare pairs of discrete samples.
[0440] JAS - Joint Assembly Space, an Assembly Space that describes the shortest sequence of join operations to simultaneously construct multiple target objects.
[0441] JFAS - Joint Fragment Assembly Space, a multimolecular hierarchical construction space inferred from tandem MS information that is comparable to JAS.
[0442] Joint MA - Joint Molecular Assembly index, the overall number of steps in JAS to construct all targets molecules.
[0443] KDE - Kernel Density Estimate.
[0444] MA - Molecular Assembly index, the number of steps in the Assembly Pathway of a target molecule, describing its complexity.
[0445] MS - Mass spectrometry, an analytical instrument used to process the samples and measure the mass-to-charge ratio of ions. Tandem MS refer to the act of fragmenting analytes detected in the MS data into smaller fragment ions.
[0446] OOC - Otsuka-Ochiai coefficient, a heuristic metric describing the similarity between discrete datasets.
[0447] Recursive MA - an algorithm for estimating the MA of analytes from MS data by recursively assembling them from their tandem MS fragments.
[0448] WPGMA - Weighted Pair Group Method with Arithmetic Mean, a recursive grouping algorithm used for building a phylogenetic tree from a similarity matrix. References
[0449] A number of publications are cited above in order to more fully describe and disclose the invention and the state of the art to which the invention pertains. Full citations for these references are provided below. The entirety of each of these references is incorporated herein.
[0450] Azwanida, N. N. Med Aromat Plants 4, 2167-0412 (2015).
[0451] Borenstein, E. et al.Proc. Natl. Acad. Sei. 105, 14482-14487 (2008).
[0452] Broeckling, C. D. et al. Anal. Chem. 90, 8020-8027 (2018).
[0453] Chambers, M. C. et al. Nat. Biotechnol. 30, 918-920 (2012).
[0454] Chen, M. et al.. J. Zhejiang Univ. Sci. B 15, 333-342 (2014).
[0455] Ciccarelli, F. D. et al. Science 311 , 1283-1287 (2006).
[0456] D’Ari, R. & Casadesus, J. BioEssays 20, 181-186 (1998).
[0457] Darwin, C. & Bynum, W. F. The Origin of Species by Means of Natural Selection: Or, the Preservation of Favored Races in the Struggle for Life. (AL Burt New York, 2009).
[0458] Davies, V. et al. Anal. Chem. 93, 5676-5683 (2021).
[0459] Delsuc, F. et al. Nat. Rev. Genet. 6, 361-375 (2005).
[0460] Eight key rules for successful data-dependent acquisition in mass spectrometry-based metabolomics - Defossez - 2023 - Mass Spectrometry Reviews - Wiley Online Library. Fundamentals and Advances of Orbitrap Mass Spectrometry - Hecht - Major Reference Works - Wiley Online Library.
[0461] Gorbalenya, A. E. et al. Nat. Microbiol. 5, 668-674 (2020).
[0462] Gregor, R. et al. ISME J. 16, 1262-1274 (2022).
[0463] Guo, A. C. et al. Nucleic Acids Res. 41 , D625-D630 (2012).
[0464] Hug, L. A. et al. A new view of the tree of life. Nat. Microbiol. 1, 1-6 (2016).
[0465] Jirasek, M. et al. Determining Molecular Complexity using Assembly Theory and Spectroscopy. Preprint at https: / / doi.org / 10.48550 / arXiv.2302.13753 (2023).
[0466] Johnson, C. H. et al. Nat. Rev. Mol. Cell Biol. 17, 451-459 (2016).
[0467] Kapli, P. et al.Nat. Rev. Genet. 21 , 428-444 (2020).
[0468] Koelmel, J. P. et al. J. Am. Soc. Mass Spectrom. 28, 908-917 (2017).
[0469] Kumar, K. et al. Ultrason. Sonochem. 70, 105325 (2021).
[0470] Kumar, S. et al. Mol. Biol. Evol. 39, msac174 (2022).
[0471] Lancet, D. et al. J. R. Soc. Interface 15, 20180159 (2018).
[0472] Liu, Y. et al. Sci. Adv. 7, eabj2465 (2021 ).
[0473] Marshall, S. M. et al. Nat. Commun. 12, 3033 (2021).
[0474] Mayr, E. The Growth of Biological Thought: Diversity, Evolution, and Inheritance. (Harvard University Press, 1982).
[0475] Mills, C. L. et al. Com put. Struct. Biotechnol. J. 13, 182-191 (2015).
[0476] Mukherjee, S. et al. Nat. Biotechnol. 35, 676-683 (2017).
[0477] Oyinloye, T. M. & Yoon, W. B. Processes 8, 354 (2020).
[0478] Parks, D. H. et al. Nat. Biotechnol. 36, 996-1004 (2018).
[0479] Parks, D. H. et al. Nat. Microbiol. 2, 1533-1542 (2017).
[0480] Peterson, B. W., Sharma, P. K., van der Mei, H. C. & Busscher, H. J. Appl. Environ. Microbiol. 78, 120-125 (2012).
[0481] Rados, D., Donati, S., Lempp, M., Rapp, J. & Link, H. Iscience 25, (2022). Rosato, A. et al. Metabolomics 14, 1-20 (2018).
[0482] Scott, D. W. Multivariate Density Estimation: Theory, Practice, and Visualization. (Wiley, 1992). del: 10.1002 / 9780470316849.
[0483] Sharma, A. et al. Nature 1-8 (2023).
[0484] Smith, M. R. Bioinformatics 36, 5007-5013 (2020).
[0485] Smith, M. R. Syst. Biol. 71 , 1255-1270 (2022).
[0486] Sokal R. R„ Michener C. D. Univ Kans Sci Bull 38, 1409-1438 (1958).
[0487] Tsugawa, H. et al. Nat. Methods 12, 523-526 (2015).
[0488] Verma, V. & Aggarwal, R. K. Soc. Netw. Anal. Min. 10, 43 (2020).
[0489] Von Luxburg, U. Stat. Comput. 17, 395-416 (2007).
[0490] Wang, F. et al. Anal. Chem. 93, 1 1692-11700 (2021 ).
[0491] WO 2021 / 186193
Claims
Claims:
1. A method for estimating the molecular complexity of one sample or a plurality of samples, wherein each sample includes two or more components, the method comprising:(a) obtaining MS / MS data for the sample or the plurality of samples, the MS data comprising MS1data for ions of the components, and MSndata for the nthgeneration of ion fragments formed from the ions of the components, wherein n is 2 or more;(b) identifying the ion fragments in the MSndata for the ions of the components identified in the MS1data, and;(c) calculating a joint molecular assembly (JMA) for the one sample or the plurality of samples from the ions of the components identified in the MS1data and the ion fragments identified in the MSndata, wherein the JMA is a measure of the complexity of the sample or the plurality of samples, which represents the complexity of the components and the shared complexity amongst the components.
2. The method of claim 1, wherein the method is for estimating the molecular complexity of two samples.
3. The method of claim 1 or 2, wherein step (b) further comprises determining the neutral fragments formed during MSnfragmentation, and wherein in step (c) the JMA is calculated from the ions of the components identified in the MS1data, the ion fragments identified in the MSndata and the neutral fragments.
4. The method of any preceding claim, wherein in step (c) the JMA is calculated by:(I) determining a shortest construction pathway for the ions of the components within the sample or plurality of samples from the fragments, optionally together with neutral fragments;(II) calculating a joint molecular assembly (JMA) based on the location of the ions of the components and the fragments within the shortest construction pathway.
5. The method of claim 4, wherein determining a shortest construction pathway for the ions of the components within the sample or plurality of samples further comprises stratifying the ions of the components and fragments into groups based on the MSndata in which they were identified.
6. The method of claim 4 or 5, wherein in step (II) the JMA is calculated by determining if each of the ions of the components or the fragments are at a terminus of the shortest construction pathway, and determining if each of the ions of the components or the fragments are intermediates on the shortest construction pathway.
7. The method of any one of claims 4 to 6, wherein in step (ii) the JMA is calculated by: estimating a molecular assembly (MAT) for each of the ions of the component and each fragment, which is a terminus of the construction pathway; calculating a sum of the MAT(Z MAT); calculating the total number of ions of components and fragments which are intermediates on the construction pathway (Nj); calculating a joint molecular assembly (JMA) as the sum of Z MATand Nj.
8. The method of claim 7, wherein the MATis estimated based on the molecular weight of the component and / or fragment.
9. The method of any preceding claim, wherein in step (a) obtaining the MS / MS data comprises performing MS / MS on the sample or the plurality of samples.
10. The method of any preceding claim, wherein in step (a) obtaining the MS / MS data comprises performing MS / MS on one or more samples and obtaining MS / MS data from a database for one or more samples.11 . The method of any preceding claim, further comprising a step of preparing the sample or the plurality of samples, wherein the preparation step precedes step (a), preferably wherein the preparation step is the same for each of the plurality of samples.
12. The method of any preceding claim, wherein the preparation step comprises separating the two or more components by chromatography, such as liquid chromatography.
13. The method of any preceding claim, wherein each sample is an environmental sample, such as a biological sample, such as an extract from a biological source.
14. The method of any preceding claim, wherein each sample is a whole cell sample.
15. The method of any preceding claim, wherein the method is for estimating the phylogenetic clade of a sample or a plurality of samples.
16. A method for estimating a phylogenetic relationship of a first sample and a second sample, the method comprising: estimating the molecular complexity of the samples according to the method of claims1 to 15, to determine the JMA of the first sample and the second sample; and estimating the phylogenetic relationship of the first sample and second sample from the similarity of the JMA of the samples.
17. The method of claim 16, wherein the similarity of the JMA of the samples is calculated from the ratio of common elements to unique elements of the JMA of the first sample and second sample.
18. The method of claim 16 or 17, wherein the similarity of the JMA of the samples is a joint assembly overlap (JAO), wherein JAO is calculated based on equation 1 :]MA(A) + JMA (B) — JMA(A + B)Equation 4. JAO (A B) max (JMA(A),JMA(By) where JMA(A) is the JMA for a first sample, JMA(B) is the JMA for a second sample JMA(A+B) is the JMA for the first and second samples, and max(JMA(A),JMA(B)) is the maximum value for JMA(A) and JMA(B).
19. A method for estimating a phylogenetic character of a first sample, the method comprising:(a) estimating the molecular complexity of the first sample and a second sample according to the method of the first aspect, to determine the JMA of the first sample and second sample; and(b) estimating the phylogenetic character of the first sample from the similarity of the JMA of the first and second samples.
20. The method of claim 19, wherein the method is for estimating the phylogenetic clade of a first sample, wherein step (a) is repeated with one or more different second samples; and in step (b) the phylogenetic clade is estimated for the first sample from the similarity of the JMA of the first and second samples.21 . The method of claims 20, further comprising a step of generating a phylogenetic tree from the estimated phylogenetic clades of the samples.
22. The method of claim 19, wherein the method is for estimating a characteristic of a first sample, the method comprising: comparing the similarity of the JMA of the first sample and second sample to a threshold value, wherein the threshold value is indicative of the characteristic.
23. The method of claim 22, wherein the method is for identifying a candidate pharmaceutical or agrochemical; the method further comprising: identifying a first sample having a similarity of the JMA of the first sample and second sample which is greater than a threshold value as a candidate agrochemical or pharmaceutical.
24. The method of claim 22, wherein the method is for identifying diseased cells in a first sample, wherein the first sample contains a suspected diseased cell; the method further comprising: identifying a sample having a similarity of the JMA of the first sample and second sample greater than a threshold value as diseased cells.
25. The method of claim 22, wherein the method is for identifying treatment-resistance pathogens in a first sample, wherein the first sample is suspected treatment-resistance pathogens; the method further comprising: identifying a sample having a similarity of the JMA of the first sample and second sample greater that a threshold value as a treatment-resistance pathogen.
Citation Information
Patent Citations
Method for estimating molecular complexity
WO2021186193A1
Method for Estimating Molecular Complexity
US20230129380A1