Characterisation of taxonomically informative peptides
The method of protein analysis in hair samples through peptide comparison to reference sequences addresses the limitations of DNA-based and subjective methods, offering a reliable and objective approach for species identification in forensic science.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- CHEMCENTRE
- Filing Date
- 2026-01-14
- Publication Date
- 2026-07-23
AI Technical Summary
Current molecular techniques for analyzing non-human hair evidence often fail due to insufficient DNA availability and rely heavily on subjective morphological methods, lacking objective taxonomic resolution and statistical foundation, reducing the probative value of evidence in forensic investigations.
A method involving protein analysis of hair samples by digesting proteins to form peptides, comparing them to reference sequences (SEQ ID NOs: 15-310) for species identification, using mass spectrometry to match peptide sequences for accurate taxonomic classification.
Provides an objective and statistically grounded method for identifying the source of mammalian hair, enhancing the probative value of forensic evidence by ensuring accurate species identification.
Smart Images

Figure IMGF000026_0001_TABLE 
Figure IMGF000029_0001_TABLE 
Figure IMGF000033_0001_TABLE
Abstract
Description
Characterisation of Taxonomically Informative PeptidesTECHNICAL FIELD
[0001] The present invention relates to methods to identify mammalian species on the basis of protein analysis of hair samples.BACKGROUND ART
[0002] In a criminal investigation, the probative value of evidence is often determined by the objectivity of the analytical method used for examination. In recent years, molecular based analytical approaches to trace biological evidence analysis such as DNA have highlighted the limitations of some traditional forensic methods of analysis, particularly those reliant on pattern matching (such as fingerprinting, ballistics, and microscopic hair examination). Accordingly, the development of molecular approaches to trace analyses, which leave less room for subjective interpretation have become an important focus in the forensic science community. “Trace evidence” refers to particulate matter transferred anytime there is contact between objects and / or individuals or locations, which is collected during the investigation of a crime. Such trace evidence can include synthetic material, as well as biological material such as hair, bodily fluids, tissue, nail, or bone.
[0003] Hair is a prevalent and valuable form of trace evidence, due to the fact that it is chemically stable, persistent, and easily transferrable. Further, hairs are commonly shed by all mammals, making them ubiquitous to any physical crime scene, and easily collected by tape lift or forceps. As a result of the proximity that humans live and work with other species, particularly in the case of domestic pets, hair shafts collected in this way are likely to originate from a variety of sources, both human and non-human. The abundance and transferability of shed hairs makes them a particularly valuable form of animal trace, and the most common form of biological matter found in the homes of pet owners. Often, hairs from a non-human source can provide forensic intelligence that links an individual to another person, location or vehicle.
[0004] Current molecular techniques for analysing non-human biological evidence include genomic approaches, notably taxonomic classification by measuring differences in DNA sequences. Whilst such techniques have had considerable success analysing certain matrices, suitable DNA is not always extractable from hair shafts, particularly when analysing a portion from the proximal end. In contrast to its high protein content, hair may not contain sufficient DNA, which is the preferred evidence type for identification in forensic investigations owing to its specificity for individualization. Further, the analysis of non-human hair evidence often falls to traditional methods of analysis such as microscopy, that are inherently subjective and heavily reliant on the expertise of the examiner to interpret morphological features for taxonomic source determination.These methodologies lack an objective taxonomic resolution and a solid statistical foundation, reducing the probative value of evidence in court.
[0005] There is a need to develop new objective taxonomic and solid statistically founded methods of hair analysis; or at least the provision of methods of hair analysis to complement existing methods. The present invention seeks to provide an improved or alternative method for analysing the source of hair samples.
[0006] The previous discussion of the background art is intended to facilitate an understanding of the present invention only. The discussion is not an acknowledgement or admission that any of the material referred to is or was part of the common general knowledge as at the priority date of the application.SUMMARY OF INVENTION
[0007] The present disclosure provides a method of identifying the source of a mammalian hair, said method comprising the steps of:a) obtaining a sample of hair;b) digesting the proteins in the hair to form peptides;c) analysing the peptides;d) comparing the analysis from step (c) to any one or more of reference SEQ ID NOs: 15- 310,wherein a match between one or more peptides and one or more of reference SEQ ID NOs: 15-310 indicates the mammal the hair is derived from.
[0008] In one aspect, a match between one or more masses relating to reference SEQ ID NOs: 15-310 indicates the mammalian species the hair is derived from.
[0009] The method may include the further steps of washing the hair before digestion to remove contaminants, and / or solubilising the hair before digestion to increase the ability of the enzymes to digest the hair. The method may include the further steps of purifying the digestate and / or extracting the peptides from the digestate.
[0010] The present disclosure further provides a method of identifying a mammal, said method comprising the steps of:a) obtaining a sample of hair from the mammal;b) digesting the proteins in the hair to form peptides;c) analysing the peptides;d) comparing the analysis from step (c) to any one or more of reference SEQ ID NOs: 15- 310,wherein a match between one or more peptides and one or more of reference SEQ ID NOs: 15-310 indicates the mammal the hair is derived from.
[0011] The present disclosure further provides for the use of any one or more of reference SEQ ID NOs: 15-310 in a method of identifying the source of a mammalian hair, said method comprising the steps of:a) obtaining a sample of hair;b) digesting the proteins in the hair to form peptides;c) analysing the peptides;d) comparing the analysis from step (c) to any one or more of reference SEQ ID NOs: 15- 310,wherein a match between one or more peptides and one or more of reference SEQ ID NOs: 15-310 indicates the mammal the hair is derived from.
[0012] The present disclosure further provides a method of processing a mammalian hair, said method comprising the steps of:a) obtaining a sample of hair;b) digesting the proteins in the hair to form peptides;c) analysing the peptides;d) comparing the analysis from step (c) to any one or more of reference SEQ ID NOs: 15- 310,wherein a match between one or more peptides and one or more of reference SEQ ID NOs: 15-310 indicates the mammal the hair is derived from.
[0013] The present disclosure further provides a method of processing a mammalian hair, said method comprising the steps of:a) obtaining a sample of hair;b) digesting the proteins in the hair to form peptides;c) analysing the peptides;d) detecting a match between one or more peptides and one or more of reference SEQ ID NOs: 15-310.
[0014] Optionally, the methods comprise the further step of comparing the analysis from step (c) to the peptides of any of reference SEQ ID NOs: 12-14, wherein a match between one or more of the peptides and one or more of reference SEQ ID NOs: 12-14 indicates the hair is derived from a mammal.
[0015] Optionally, the methods may comprise the step of purifying the digestate during the preparation, before analysing the peptides.
[0016] Optionally, the methods comprise the step of extracting the peptides from the hair digestate before analysing the peptides.
[0017] The present invention further provides a kit comprising(i) a set of amino acid sequence reference standards, being one or more sequences of any of SEQ ID NOs: 12-310; and(ii) instructions for use of the amino acid sequence reference standards for the identification of the source of a mammalian hair or identification of a mammal.
[0018] The present invention further provides a kit for identifying a mammal from a hair, said kit comprising:a) one or more reagents in the form of:(i) one or more reagents for washing the hair;(ii) one or more reagents for partially or fully solubilising the hair;(iii) one or more reagents for digesting the proteins in the hair to form peptides;(iv) one or more reagents for extracting the peptides from the digestate;(v) one or more reagents for purifying the peptides in the digestate; and / or(vi) one or more standard reference peptides that correspond in sequence to any one or more of SEQ ID NOs: 12-310,b) instructions for use,wherein the instructions for use provide instructions to carry out the steps of:a) obtaining a sample of hair from the mammal;b) digesting the proteins in the hair to form peptides;c) analysing the peptides;d) comparing the analysis from step (c) to any one or more of reference SEQ ID NOs: 15-wherein a match between one or more peptides and one or more of reference SEQ ID NOs: 15-310 indicates the mammalian species the hair is derived from.
[0019] In one aspect of the above methods and kits, a match between one or more masses relating to reference SEQ ID NOs: 15-310 indicates the mammalian species the hair is derived from.BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Further features of the present invention are more fully described in the following description of several non-limiting embodiments thereof. This description is included solely for the purposes of exemplifying the present invention. It should not be understood as a restriction on the broad summary, disclosure or description of the invention as set out above. The description will be made with reference to the accompanying drawings in which:Figure 1 is a proteomics workflow for the identification of taxonomically diagnostic hair peptides A. Protein solubilisation and protease digestion of the hair. B. Shotgun proteomic data acquisition by high resolution LC / MS / MS C. MS data processed by database searching with protein database containing reference species. D. Protein composition of each species evaluated, and abundant proteins identified. E. Data driven bioinformatics approach is used for the amino acid sequence identification of tryptic peptides of abundant proteins, that contain taxonomically informative sequence variations, by aligning abundant proteins alongside analogous proteins from closely related species and de novo sequence determination. F. Combined keratin and KAP database for analysis of unknown hair G. Evaluation of taxonomic resolution achievable on a protein level by combined database searching. H. Taxonomically informative peptides evaluated for sensitivity and selectivity to compile a panel of diagnostic markers.Figure 2 is a heat map illustrating the percent coverage of protein sequences from the proteomes of each of the 15 species of interest, when processing is performed with a combined protein database.Figure 3 is a heat map illustrating the abundance and specificity of each peptide marker identified in this study.DESCRIPTION OF INVENTIONDetailed Description of the Invention
[0021] The field of proteomics has undergone significant advancements in recent decades, primarily due to the rise of mass spectrometry (MS)-based techniques. Researchers can now routinely perform in-depth protein sequence analysis with speed and sensitivity, allowing protein analysis to emerge as a viable alternative for the molecular identification of the taxonomic origin of biological material.
[0022] Proteins, composed of long chains of amino acids, are the fundamental molecular building blocks of all biological material. As proteins are derived from the genetic code, inheritable differences between and within species are reflected within the primary structure of proteins in the form of amino acid sequence variations. To this end, elucidation of protein sequences extracted from biological matter have been used to determine the taxonomic origin of the material in a similar way to DNA barcoding. Hair, being rich in protein content, is a suitable matrix for proteomics, containing over 300 proteins, with keratins and keratin-associated proteins (KAPs) being predominant.
[0023] This study explored the use of proteomics to achieve taxonomic classification of individual hair shafts. The present method allows genus or species level resolution for the taxonomic source identification of trace amounts of hair.
[0024] Highly abundant hair peptides, common to all hair proteomes of interest, may be validated as internal control markers. These peptides allowed the monitoring of the solubilisation / digestion and optional isolation efficiency and may serve as a reliable indicator that the fibre in question is a mammalian hair and not a synthetic fibre.
[0025] The terms “peptide”, “oligopeptide”, and “polypeptide” are interchangeable and these terms, along with the term “protein”, refer to chains of amino acids. Generally, peptides, oligopeptides and polypeptides are shorter, having a chain length of about 2-50 amino acids, and proteins have a chain length of more than about 50 amino acids. However, the amino acid chain length used for determination of the taxonomic origin of the material is not restricted to a certain size.Method of Identification
[0026] The present disclosure provides a method of identifying the source of a mammalian hair, said method comprising the steps of:a) obtaining a sample of hair;b) digesting the proteins in the hair to form peptides;c) analysing the peptides;d) comparing the analysis from step (c) to any one or more of reference SEQ ID NOs: 15- 310,:wherein a match between one or more peptides or ions with one or more of reference SEQ ID NOs: 15-310 indicates the mammal the hair is derived from.
[0027] The peptides generated in step (b) may be further converted into peptide ions, and / or the reference peptides of SEQ I D NOs: 15-310 may be in the form of ions. If the peptides to be tested and / or the reference peptides are in the form of ions, the analysis of step (c) may be in the form of mass spectrometry, which may be a high resolution or lower (unit mass) resolution mass spectrometer, for example triple-quadrupole mass spectrometer. Tables 22-35 provide information on the ions of SEQ ID NOs: 15-310. In one aspect, a match between one or more masses relating to reference SEQ ID NOs: 15-310 indicates the mammalian species the hair is derived from.
[0028] The present disclosure further provides a method of identifying a mammal, said method comprising the steps of:a) obtaining a sample of hair from the mammal;b) digesting the proteins in the hair to form peptides;c) analysing the peptides;d) comparing the analysis from step (c) to any one or more of reference SEQ ID NOs: 15- 310,wherein a match between one or more peptides and one or more of reference SEQ ID NOs: 15-310 indicates the mammal the hair is derived from.
[0029] The present disclosure further provides for the use of any one or more of reference SEQ ID NOs: 15-310 in a method of identifying the source of a mammalian hair, said method comprising the steps of:a) obtaining a sample of hair;b) digesting the proteins in the hair to form peptides;c) analysing the peptides;d) comparing the analysis from step (c) to any one or more of reference SEQ ID NOs: 15- 310,wherein a match between one or more peptides and one or more of reference SEQ ID NOs: 15-310 indicates the mammal the hair is derived from.
[0030] The present disclosure further provides a method of processing a mammalian hair, said method comprising the steps of:a) obtaining a sample of hair;b) digesting the proteins in the hair to form peptides;c) analysing the peptides;d) comparing the analysis from step (c) to any one or more of reference SEQ ID NOs: 15- 310,wherein a match between one or more peptides and one or more of reference SEQ ID NOs: 15-310 indicates the mammal the hair is derived from.
[0031] The present disclosure further provides a method of processing a mammalian hair, said method comprising the steps of:a) obtaining a sample of hair;b) digesting the proteins in the hair to form peptides;c) analysing the peptides;d) detecting a match between one or more peptides and one or more of reference SEQ ID NOs: 15-310.
[0032] The present disclosure further provides a method for determining the origin of an unknown hair, said method comprising the steps of:a) obtaining a sample of the hair;b) digesting the proteins in the hair to form peptides;c) analysing the peptides;e) detecting a match between one or more peptides and one or more of reference SEQ ID NOs: 15-310,wherein a match between one or more peptides and one or more of reference SEQ ID NOs: 15-310 indicates the mammal the hair is derived from.
[0033] The present disclosure further provides a method for associating a protein in a hair with an animal species, said method comprising the steps of:a) obtaining a sample of the hair;b) digesting the proteins in the hair to form peptides;c) analysing the peptides;d) detecting a match between one or more peptides and one or more of reference SEQ ID NOs: 15-310,wherein a match between one or more peptides and one or more of reference SEQ ID NOs: 15-310 indicates the mammal the hair is derived from.
[0034] The present disclosure further provides a method for determining the origin of a hair recovered at a crime scene, said method comprising the steps of:a) obtaining a sample of the hair;b) digesting the proteins in the hair to form peptides;c) analysing the peptides;d) detecting a match between one or more peptides and one or more of reference SEQ ID NOs: 15-310,wherein a match between one or more peptides and one or more of reference SEQ ID NOs: 15-310 indicates the mammal the hair is derived from.
[0035] The present disclosure further provides a method for associating a hair with an animal of origin, said method comprising the steps of:a) obtaining a sample of hair;b) digesting the proteins in the hair to form peptides;c) analysing the peptides;d) detecting a match between one or more peptides and one or more of reference SEQ ID NOs: 15-310,wherein a match between one or more peptides and one or more of reference SEQ ID NOs: 15-310 indicates the mammal the hair is derived from.
[0036] The present disclosure further provides a method for associating a non-human hair with an animal species, said method comprising the steps of:a) obtaining a sample of hair;b) digesting the proteins in the hair to form peptides;c) analysing the peptides;d) detecting a match between one or more peptides and one or more of reference SEQ ID NOs: 15-310,wherein a match between one or more peptides and one or more of reference SEQ ID NOs: 15-310 indicates the mammal the hair is derived from.
[0037] The present disclosure further provides a method of assigning a taxonomic classification to an unknown hair sample, said method comprising the steps of:a) obtaining a sample of hair;b) digesting the proteins in the hair to form peptides;c) analysing the peptides;d) detecting a match between one or more peptides and one or more of reference SEQ ID NOs: 15-310,wherein a match between one or more peptides and one or more of reference SEQ ID NOs: 15-310 indicates the mammal the hair is derived from.
[0038] The present disclosure further provides a molecular tool for the taxonomic classification of an unknown hair sample, said method comprising the steps of:a) obtaining a sample of hair;b) digesting the proteins in the hair to form peptides;c) analysing the peptides;d) detecting a match between one or more peptides and one or more of reference SEQ ID NOs: 15-310,wherein a match between one or more peptides and one or more of reference SEQ ID NOs: 15-310 indicates the mammal the hair is derived from.
[0039] The present disclosure further provides a molecular tool for identifying a hair as nonhuman, said method comprising the steps of:a) obtaining a sample of hair;b) digesting the proteins in the hair to form peptides;c) analysing the peptides;d) detecting a match between one or more peptides and one or more of reference SEQ ID NOs: 15-310,wherein a match between one or more peptides and one or more of reference SEQ ID NOs: 15-310 indicates the mammal the hair is derived from.
[0040] A method of analysis for the evidentiary triage of hair, for the preservation of human hair evidence, said method comprising the steps of:a) obtaining a sample of the hair;b) digesting the proteins in the hair to form peptides;c) analysing the peptides;d) detecting a match between one or more peptides and one or more of reference SEQ ID NOs: 15-310,wherein a match between one or more peptides and one or more of reference SEQ ID NOs: 15-310 indicates the mammal the hair is derived from, andwherein part of the hair is used as the sample for steps (a)-(d), and part of the hair is preserved for further analysis techniques.
[0041] A method of biological evidentiary analysis for the preservation of human hair evidence, said method comprising the steps of:a) obtaining a sample of the hair;b) digesting the proteins in the hair to form peptides;c) analysing the peptides;d) detecting a match between one or more peptides and one or more of reference SEQ ID NOs: 15-310wherein a match between one or more peptides and one or more of reference SEQ ID NOs: 15-310 indicates the mammal the hair is derived from, andwherein part of the hair is used as the sample for steps (a)-(d), and part of the hair is preserved for further analysis techniques.
[0042] The mammal may be a placental mammal. In another example, the mammal may be a marsupial or a monotreme.
[0043] The further analysis techniques may be, for example, DNA isolation, microscopic comparison etc.
[0044] Optionally, the methods comprise the step of comparing the analysis from step (c) to the peptides of any of reference SEQ ID NOs: 12-14, wherein a match between one or more of thepeptides and one or more of reference SEQ ID NOs: 12-14 indicates the hair is derived from a non-human mammal.
[0045] Optionally, the methods comprise the steps for comparing the analysis from step (c) to the peptides of any of reference SEQ ID NOs: 15-37, wherein a match between one or more of the peptides and one or more of reference SEQ ID NOs: 15-37 indicates the hair is derived from the Order comprising Homo sapiens, and optionally SEQ ID NOs: 31-37 identifies the hair as being from H. sapiens.
[0046] Optionally, the methods may comprise the step of washing the hair sample before solubilisation and digestion. Optionally, the methods may comprise the step of purifying the digestate before analysing the peptides. Optionally, the methods may comprise the step of extracting the peptides from the hair digestate before analysing the peptides. These steps may be used to reduce or remove contaminants that would interfere or otherwise change the results of the determination of the amino acid sequence or the nature of the peptides. The contaminants may be proteinaceous contaminants or non-proteinaceous contaminants.Peptide matching
[0047] The peptides obtained from the methods of the present invention will optionally match reference sequences SEQ ID NOs: 15-310. By “match” or “matching” or “matches”, it is meant that the amino acid sequences are substantially the same, with only 5, 4, 3, 2, 1 or optionally zero differences between the peptides from the hair and any one or more of reference SEQ ID NOs: 15-310. This includes situations where the genomic nucleic acid sequence coding for the amino acid sequence of the peptides is not identical between samples but, due to redundancy in the nucleic acid code, the discrepancies in the genomic nucleic acid sequence do not lead to changes in the amino acid sequence of the keratin or KAP.
[0048] If the peptides are sequenced, optionally the match comprises at least 75%, 80%, 85%, 90%, 95% or 100% identity between the peptide and the amino acid sequence of any one or more of the reference sequences of SEQ ID NOs: 15-310.
[0049] If the matching is via, for example mass to charge ratio (m / z) (eg using mass spectrometry), optionally without sequencing the peptide, then optionally the match is within at least 1 Dalton of a mass to charge ratio representing the mass of one or more of the reference sequences of SEQ ID NOs: 15-310. Information regarding the ions of SEQ ID NOs: 15-310 are provided in T ables 22-35. If the peptides or reference sequences are labelled, tagged or otherwise modified, including upon ionisation or by post-translational modification (PTM; which may be introduced endogenously or through the processing method), their m / z may differ from that provided in Tables 22-35. In such a case, the disclosure addresses the equivalence of the endogenous, unmodified mass (that is, the mass without the charge, label, tag or othermodification) of the peptide to the endogenous, unmodified mass of the reference sequences of SEQ ID NOs: 12-14 and SEQ ID NOs: 15-310.
[0050] The reference sequences SEQ ID NOs: 15-310 (and SEQ ID NOs: 12-14) may be labelled, tagged, or otherwise modified. This may change the mass from the endogenous mass listed in Table 15. However, the use of the sequences of any of SEQ ID NOs: 12-14 and SEQ ID NOs: 15-310 in any form as reference standards for the determination of the species from which a hair is derived falls within the scope of the present invention.
[0051] Optionally, the match is to one or more of SEQ ID NOs: 31, 32, 33, 34, 35, 36, 37, 53, 54, 55, 56, 57, 74, 75, 76, 77, 78, 79, 80, 87, 88, 102, 122, 134, 148, 149, 150, 151, 152, 153, 154, 172, 184, 185, 186, 197, 198, 199, 200, 201, 202, 203, 204, 206, 207, 208, 209, 210, 211, 212, 213, 214, 215, 231, 237, 238, 239, 240, or 241. As can be seen in Table 20, these sequences are specific to the placental mammal species being identified.
[0052] Optionally, the match is to one or more of SEQ ID NOs: 270, 271, 272, 273, 276, 277, 278, 279, 280, 283, 284, 285, 286, 288, 289, 290, 291, 292, 293, 294, 295, 296, 297, 298, 299, 300, 301, 302, 303, 304, 305, 306, 307, 308, 309, or 310. As can be seen in Table 21, these sequences are specific to the marsupial species being identified.
[0053] Optionally, the match is to two or more of SEQ ID NOs: 15-310. The match may be to three or more, four or more, or any relevant number of matches, taking into account the number of sequences provided for each mammal in the Tables.
[0054] The method may comprise the further step of matching one, two or three of SEQ ID NOs: 12-14 in the peptides of the hair. These sequences provide control sequences, allowing the identification of the hair as being from a mammal (not, for example, a polyester fibre), and that the hair protein digestion, peptide extraction, sequence determination and sequence comparison has occurred successfully.
[0055] The method may also be used to determine whether a hair is human or non-human. For such purposes, a match to any one or more of SEQ ID NOs: 15-37 would identify a hair as being from the Order comprising Homo sapiens, and optionally SEQ ID NOs: 31-37 would identify a hair as being from H. sapiens.Hair
[0056] The term “hair” as used herein refers to protein fibres, principally keratin, that are grown by mammals as a covering for thermoregulation, protection, sensory purposes, waterproofing, camouflage etc. The term “hair” as used herein includes wool, fur, whiskers, pubic hair, beard hair, etc. As keratin is also the key structural material making up nails, horns, claws, and hooves,sources of keratin other than hair (for example nails, horns, baleen, etc) may also be used in the methods of the present disclosure and the term “hair” encompasses these other keratin sources.
[0057] The growth of the hair follicle is cyclical. Stages of rapid growth and elongation of the hair shaft alternate with periods of quiescence and regression driven by apoptotic signals. The life stages of hair include: anagen (growth); catagen (regression); telogen (rest); and exogen (shedding). Hair used in the methods of the present invention may be obtained from any these hair life stages.
[0058] Previous identification methods associated with hair have focused on obtaining nucleic acids from the root or bulb of the hair. This root or bulb tissue contains cells that yield DNA and RNA. In contrast, the hair shaft, the preferred source of peptides for the present invention, may or may not contain DNA. Furthermore, telomeric hair or exogenic hair generally does not have a root or bulb attached as it is shed from the skin due to senescence, leaving the root or bulb behind in the skin. As a result, the hair shaft is seldom used in identification, except through visual assessment. The present invention may be used to determine peptide sequences from the root or bulb of a hair, as these proteins are similar in structure to the hair proteins. The method does not determine the sequence of the root or bulb DNA or RNA. However, the present invention may be used in addition to species identification based on any nucleic acid that is present.
[0059] Types of hair included:- definitive, which may be shed after reaching a certain length;- vibrissae, which are sensory hairs and are most commonly whiskers;- pelage or fur, which consists of guard hairs, under-fur, and awn hair;- spines, which are a type of stiff guard hair used for defence in, for example, porcupines; - bristles, which are long hairs usually used in visual signals, such as the mane of a lion;- velli, often called "down fur", which insulates newborn mammals; and- wool, which is long, soft, and often curly.
[0060] The source of the hair may be a pelt or other processed skin, a crime scene, etc. and the species identification of keratinous materials may be used to identify or eliminate the presence of a subject at a crime scene, in textile production to detect fraud in textiles made from luxury fibres such as cashmere, and in the determination of the source of materials from the illegal trade in endangered animals or animal parts (e.g. rhino horn powder or whale baleen).
[0061] In one aspect, the protein in the hairs analysed in the present invention are keratin proteins, optionally a-keratin proteins, for example keratin type I and type II; or keratin-associated proteins (KAPs). Mammals have a number of keratin genes; for example, humans have 54functional keratin genes - 28 in the keratin type I family and 26 in the keratin type II family. In one example, the keratin proteins analysed in mammals may be chosen from type I keratins 31, 33A and / or 34; and / or type II keratins 81 , 83 and / or 86.
[0062] As keratin is the key structural material making up hair, nails, horns, claws, hooves, and the outer layer of skin among mammals, sources of keratin other than hair (for example nails, horns, etc) may also be used in the methods of the present disclosure. For the purposes of this disclosure, the term “hair” encompasses these other keratin sources.
[0063] The hair may be from a subject chosen from the following: Homo sapiens (human), Canis lupus (domestic dog), Vulpes vulpes (red fox), Capra hircus (domestic goat), Ovis aries (domestic sheep), Bos taurus (cattle), Equus caballus (horse), Felis catus (domestic cat), Sus scrofa (wild boar), Rattus norvegicus (brown rat), Mus musculus (house mouse), Oryctolagus cuniculus (rabbit), Cavia porcellus (guinea pig), Camelus dromedarius (camel), or Vicugna pacos (alpaca).
[0064] The hair may be from a subject chosen from the following: monito del monte (Dromiciops gliroides), yellow footed antechinus (Antechinus flavipes), agile gracile opossum (Gracilinanus agilis), greater bilby (Macrotis lagotis), grey short tailed opossum (Monodelphis domestica), tammar wallaby (Notamacropus eugenii), sugar glider (Papuan subspecies; Petaurus breviceps papuanus), koala (Phascolarctos cinereus), Tasmanian devil (Sarcophilus harrisii), fat tailed dunnart (Sminthopsis crassicaudata), common brushtail possum (Trichosurus vulpecula), and common wombat (Vombatus ursinus).
[0065] The hair may be from a subject chosen from the following: platypus (Ornithorhynchus anatinus), and short beaked echidna (Tachyglossus aculeatus).
[0066] The hair may be from a subject chosen from the following: golden spiny mouse (Acomys russatus), nine-banded armadillo (Dasypus novemcinctus), Chinese hamster (Cricetulus griseus), Philippine tarsier (Carlito syrichta), or meerkat (Suricata suricatta). The hair may be from a subject chosen from the following: the Ursidae family (Ursus americanus, Urus arctos, Ursus maritimus, Ailuropoda melanoleuca lesser Egyptian jerboa (Jaculus jaculus gelada (Theropithecus gelada).Hair Washing
[0067] The hair sample may be washed before solubilisation and digestion. The optional washing step may be carried out using water, or other suitable solvent or reagents or procedures to remove contaminants. For example, surfactants may be used to remove oils from the hair before solubilisation and digestion. The contaminants may be proteinaceous contaminants or non-proteinaceous contaminants. Removal of non-proteinaceous components of the hair such as lipids may aid in the effectiveness of solubilisation / digestion. Washing of the hair may also beperformed to ensure that the subsequently recovered protein originates from the sample and not for example an environmental contaminant or other mammalian source.Hair Solubilisation
[0068] Hair protein solubilisation is difficult due to extensive cross-linking by cystine disulfide bonds and poor solubility of hair keratins and keratin-associated proteins (KAPs), which may prevent solubilization of approximately 15% of the constituent protein, even in the presence of strong denaturants.
[0069] The hair may be fully or partially solubilised before or during the digestion step. Solubilisation of the proteins of the hair may be carried out using chemical and / or physical means.
[0070] Suitable reagents for chemical solubilisation include, for example, surfactants and denaturants. Examples of reagents for solubilisation of hair include: sodium dodecanoate (SD), ammonium bicarbonate (ABC), dithiothreitol (DTT), trifluoroacetic acid, Triton X-100, urea, 2-mercaptoethanol, sodium dodecyl sulfate (SDS), thiourea, and iodoacetamide. The solubilisation may be carried out using a combination of two or more suitable reagents, for example a combination of two or more surfactants.
[0071] The solubilisation of hair may further involve a physical disruption step, for example the use of agitation, heating, or mechanical homogenization (e.g. a bead mill or freezing and grinding).
[0072] The solubilisation may be carried out in one step, or may comprise two or more solubilisation steps. One or more of the solubilisation steps may also be carried out simultaneously. For example, the solubilisation may comprise the steps of manual disruption, followed by one or more chemical solubilisation steps using appropriate reagents. Reduction and / or alkylation may assist in preparing the hair proteins for more efficient digestion and may be included before, during or after digestion.Hair Digestion
[0073] Digestion of the proteins of the hair may be carried out using a variety of suitable reagents, for example enzymes such as proteases. The digestion may be carried out using a combination of two or more suitable reagents, for example a combination of two or more proteases. Examples of suitable protease enzymes include trypsin, proteinase K, keratinases (such as subtilisin, aminopeptidase).
[0074] The digestion may be carried out in one step, or may comprise two or more digestion steps.
[0075] Solubilisation of the hair may occur before or at the same time as digestion of the hair.
[0076] The hair may be washed or otherwise decontaminated before solubilisation and / or digestion, to reduce contamination with other proteins or lipids. For example, the hair may be washed with an organic solvent such as ethanol, methanol, isopropanol, dichloromethane, ethyl acetate, chloroform, etc., or with an aqueous solvent such as water etc., or with a solution or suspension of an appropriate surfactant, including SD or SDS, before solubilisation and / or digestion.Digestate Purification and Peptide Extraction
[0077] The hair digestate may be purified of any contaminants that would interfere or otherwise change the results of the determination of the amino acid sequence or the nature of the peptides before analysis (for example determining the amino acid sequences or the nature of the peptides). Alternately or additionally, the peptides may be extracted from the hair digestate to eliminate any contaminants that would interfere or otherwise change the results of the determination of the amino acid sequence or the nature of the peptides before analysis (for example determining the amino acid sequences or the nature of the peptides). The contaminants may be proteinaceous contaminants or non-proteinaceous contaminants.
[0078] The purification and / or extraction may be carried out by the use of any suitable reagents or procedures to remove contaminants. The purification and / or extraction may comprise steps such as washing, and centrifugation. Further purification and / or extraction procedures including solid-phase extraction (SPE), suspension trapping and liquid-liquid extraction may also be used. Peptide Analysis
[0079] The peptides analysed using the method of the present invention may be from 5 to 40 amino acids in length, for example 7 to 36 amino acids in length. Sequences shorter than 5 amino acids may not be unique enough to allow determination of species, and sequences larger than 40 bases may be difficult to analyse and / or sequence.
[0080] The analysis method used must be able to separate and identify individual peptides in a mixture of peptides produced by digesting a hair sample. The method must be sensitive enough to detect a difference of one amino acid between sequences of up to 40 amino acids in length.
[0081] Analysis of the peptides may be carried out by any suitable mass based, or mass / charge based method, such as by mass spectrometry (MS) or tandem mass spectrometry (MS / MS). Preferably the peptides are ionized by electrospray ionization (ESI).
[0082] Alternatively, analysis of the peptides by mass spectrometry may use other ionization techniques including matrix-assisted laser desorption ionization (MALDI) or Desorptive electrospray ionisation (DESI).
[0083] The analysis may be untargeted shotgun analysis, targeted analysis to specifically measure only the claimed peptides of interest, or targeted processing of shotgun-acquired data.
[0084] Analysis of the peptides by mass spectrometry may be by direct entry of the sample into the MS by for example; infusion, liquid injection or liquid ejection, and may also include separation of the peptides or peptide ions using for example ion mobility or a liquid chromatographic technique (LC) such as high-performance liquid chromatography (HPLC), ultra-high-performance liquid chromatography (LIHPLC), capillary electrophoresis (CE), or paper chromatography, or combinations of such methods. In one example, the peptides may be analysed using LC-MS / MS. In the present method, optionally carbamidomethylation of cysteine may be considered a fixed modification of the peptides. Optionally, deamidation of glutamine and asparagine, and the oxidation of methionine, may be considered variable modifications of the peptides.
[0085] In the workflow developed, quality control measures included setting an abundance threshold, with only peptide-spectrum matches (PSMs) derived from precursor ions exceeding the threshold considered as possible valid identifications. This approach may help mitigate the risk of false PSM matches caused by noisy spectra that can mimic authentic peptide fragmentation patterns. Additional measures included subtracting 1-5% of signal from the previously run sample to limit the impact of carryover in the liquid chromatography (LC) system, a consequence of the sensitivity of the analytical platform and the resultant analyses. Sample carryover is particularly relevant for the most highly abundant peptides, especially those that have comparatively higher hydrophobicity.
[0086] Alternative strategies, such as incorporating additional solvent blanks and extending wash cycles between samples, may also be used. While these approaches can increase the overall runtime and reduce throughput, they offer an effective means of mitigating carryover if throughput is not a limiting factor.
[0087] The potential of inaccurate peptide spectral matches that are not eliminated by standard quality control measures may require additional interrogation of the precursor and / or product ion mass spectra, particularly when used in forensic proteomics. This is analogous to the process of interrogating spectral library matches based on product ion mass spectra of other biomolecules, a quality control measure regularly carried out in forensic mass spectrometry workflows. Additionally, it demonstrates the benefit of using a panel of markers, rather than relying on a single diagnostic peptide for taxonomic classification. In one aspect, identification may be made based on the presence of at least 10 diagnostic peptides, for example human identification may involve as many as 23 peptides, dependant on the evolutionary relatedness of species requiring identification. This approach significantly mitigates the risk of incorrect identifications and enhances the reliability of taxonomic classifications.Peptide comparison
[0088] The peptide sequences prepared by the methods of the present invention may be compared to the amino acid sequences of any of SEQ ID NOs: 15-310 by any suitable sequence analysis program that allows the alignment and comparison of two or more peptide amino acid sequences. For example, the sequences may be compared using BLAST (basic local alignment search tool).
[0089] Alternatively, the comparison may be made on the basis of a peptide’s mass to charge ratio, by comparing the mass to charge ratio of the peptide to that of one or more of those derived from SEQ IN NOs: 15-310. The identification may also include the measurement of the mass to charge ratio of one or more product ions (MS2), resulting from fragmentation of the precursor peptide ion. For example, MS / MS may measure the peptide as precursor ion and product ion mass pairs.
[0090] Alternatively, the comparison may be made on the basis of a peptide’s mass to charge ratio after the peptide’s mass has been altered, such as by post-translational modification of one or more of SEQ ID NOs: 15-310. The identification may also include the measurement of the mass to charge ratio of one or more product ions (MS2), resulting from fragmentation of the precursor peptide ion.Kits
[0091] The present invention further provides a kit comprising:(i) a set of amino acid sequence reference standards, being one or more sequences of any of SEQ ID NOs: 12-310; and(ii) instructions for use of the amino acid sequence reference standards for the identification of the source of a mammalian hair or identification of a mammal.
[0092] The present invention provides a kit for identifying a mammal from a hair, said kit comprising:a) one or more reagents in the form of:(i) one or more reagents for washing the hair;(ii) one or more reagents for partially or fully solubilising the hair;(iii) one or more reagents for digesting the proteins in the hair to form peptides;(iv) one or more reagents for extracting the peptides from the digestate;(v) one or more reagents for purifying the peptides in the digestate; and / or(vi) one or more standard reference peptides that correspond in sequence to any one or more of SEQ ID NOs: 12-310,b) instructions for use,wherein the instructions for use provide instructions to carry out the steps of:a) obtaining a sample of hair from the mammal;b) digesting the proteins in the hair to form peptides;c) analysing the peptides;d) comparing the analysis from step (c) to any one or more of reference SEQ ID NOs: 15- 310,wherein a match between one or more peptides and one or more of reference SEQ ID NOs: 15-310 indicates the mammalian species the hair is derived from. Alternatively, a match between one or more masses relating to reference SEQ ID NOs: 15-310 indicates the mammalian species the hair is derived from.General
[0093] Those skilled in the art will appreciate that the invention described herein is susceptible to variations and modifications other than those specifically described. The invention includes all such variation and modifications. The invention also includes all of the steps, features, formulations and compounds referred to or indicated in the specification, individually or collectively and any and all combinations or any two or more of the steps or features.
[0094] Each document, reference, patent application or patent cited in this text is expressly incorporated herein in their entirety by reference, which means that it should be read and considered by the reader as part of this text. That the document, reference, patent application or patent cited in this text is not repeated in this text is merely for reasons of conciseness.
[0095] Any manufacturer’s instructions, descriptions, product specifications, and product sheets for any products mentioned herein or in any document incorporated by reference herein, are hereby incorporated herein by reference, and may be employed in the practice of the invention.
[0096] The present invention is not to be limited in scope by any of the specific embodiments described herein. These embodiments are intended for the purpose of exemplification only. Functionally equivalent products, formulations and methods are clearly within the scope of the invention as described herein.
[0097] The invention described herein may include one or more range of values (eg. size, chargestate, mass accuracy, displacement and field strength etc). A range of values will be understoodto include all values within the range, including the values defining the range, and values adjacent to the range which lead to the same or substantially the same outcome as the values immediately adjacent to that value which defines the boundary to the range. Accordingly, unless indicated to the contrary, the numerical parameters set forth in the specification and claims are approximations that may vary depending upon the desired properties sought to be obtained by the present invention. Hence “about 80 %” means “about 80 %” and also “80 %”. At the very least, each numerical parameter should be construed in light of the number of significant digits and ordinary rounding approaches.
[0098] Throughout this specification, unless the context requires otherwise, the word “comprise” or variations such as “comprises” or “comprising”, will be understood to imply the inclusion of a stated integer or group of integers but not the exclusion of any other integer or group of integers. It is also noted that in this disclosure and particularly in the claims and / or paragraphs, terms such as “comprises”, “comprised”, “comprising” and the like can have the meaning attributed to it in U.S. Patent law; e.g., they can mean “includes”, “included”, “including”, and the like; and that terms such as “consisting essentially of’ and “consists essentially of’ have the meaning ascribed to them in U.S. Patent law, e.g., they allow for elements not explicitly recited, but exclude elements that are found in the prior art or that affect a basic or novel characteristic of the invention.
[0099] Other definitions for selected terms used herein may be found within the detailed description of the invention and apply throughout. Unless otherwise defined, all other scientific and technical terms used herein have the same meaning as commonly understood to one of ordinary skill in the art to which the invention belongs. The term “active agent” may mean one active agent, or may encompass two or more active agents.
[0100] The following examples serve to more fully describe the manner of using the above-described invention, as well as to set forth the best modes contemplated for carrying out various aspects of the invention. It is understood that these methods in no way serve to limit the true scope of this invention, but rather are presented for illustrative purposes.EXAMPLES
[0102] Further features of the present invention are more fully described in the following nonlimiting Examples. This description is included solely for the purposes of exemplifying the present invention. It should not be understood as a restriction on the broad description of the invention as set out above.Example 1Identification of members of the class MammaliaHair procurement and storageReference hair samples were collected from 15 species of interest:- human (Homo sapiens)’,- common domesticated animals: (Felis catus), dog (Canis lupus familiaris), guinea pig (Cav / a porcellus), horse (Equus caballus), cow (Bos taurus), goat (Capra hircus), sheep (Ovis aries), and alpaca (Viugna pacos)’,- wild animals: rat (Rattus rattus), red fox (Vulpes vulpes)',- animals that have both wild and domesticated populations: mouse (Mus musculus), dromedary (Camelus dromedarius), rabbit (Oryctolagus cuniculus), and pig (Sus scrofa).
[0103] Hair samples were obtained from local farmers and pet owners. Shed hair samples of approximately ten shafts were acquired from each individual sampled, to allow for multiple analyses. Samples were stored away from sunlight at room temperature within the NATA accredited Forensic Science Laboratory at ChemCentre, Western Australia.Protein extraction
[0104] The preparative workflow used for hair protein isolation was a modification of the methods by Goecker et. al. for protein extraction from human hair (Goecker, et al., Optimal processing for proteomic genotyping of single human hairs. Forensic Sci Int Genet, 2020.47:102314). The method was adapted for suitability to the increased shaft diameter of hair samples from some of the non-human species, including introduction of a longer duration of agitation during the solubilisation phase, as further described.
[0105] Protein was extracted from the reference hair samples of forensically relevant sample sizes (between 0.1-3 cm of a single hair shaft), with the hair length used roughly relative to the width of the hair so that approximately 100 pg of hair was prepared for analysis. Hair from each species was extracted in triplicate. Briefly, hair shafts were washed with 50:50 dichloromethane (DCM):ethyl acetate (EtOAc), then dissected into even segments of approximately 1 cm beforewashing with a solution of 2% (w / v) sodium dodecanoate (SD) and 50 mM ammonium bicarbonate (ABC).
[0106] The cleaned hair segments were added to a Lo-Bind tube along with 100 pL solution of 2% sodium dodecanoate (SD), 50 mM ammonium bicarbonate (ABC) and 10 pL of 1 M dithiothreitol (DTT). A single ceramic ruptor bead (OMNI International) was added to each tube and the samples were sonicated at approximately 50°C for 45 minutes. Tubes were then transferred to an Omni bead mill homogenizer and agitated at 0.8 metres per second (m / s), or approximately 2 Hz, for 100 min.
[0107] The solubilised hair was then alkylated by the addition of 40 pL of 0.5 M iodoacetamide (IAA) and agitated on an Eppendorf Thermomixer C at a speed of 300 rpm at room temperature in the dark. After 40 minutes of agitation, the solution was adjusted to pH 3 with the addition of 2 pL of trifluoroacetic acid (TFA). The detergent was then removed through biphasic extraction, requiring the addition and removal of three 175 pL additions of EtOAc, mixing by vortex and brief centrifugation to induce phase separation between each addition. The pH was then adjusted to approximately 8.5 by the addition of 6.3 pL 1M ABC and 0.6 pL of ammonium hydroxide (NH4OH) to prepare for tryptic digestion of the extracted proteins.
[0108] T ryptic digestion was achieved by the addition of 10 pg of solubilised trypsin (Promega sequencing grade modified), with samples then mixed at 600 rpm in an Eppendorf Thermomixer C for 16 hours at 37°C. The tryptic digests were purified by centrifugal filtration (Ultrafree-MC 0.45 pm PTFE (hydrophilic) membrane), producing a filtrate suitable for mass spectral analysis.
[0109] Hair samples from an additional two individuals of each species were later extracted in the same manner to assess reproducibility of the workflow.Shotgun proteomic data collection on nano-LC-Orbitrap-MS
[0110] Peptides from the extracted reference samples were resolved on a Thermo Scientific™ UltiMate™ 3000 RSLCnano LIHPLC system using a Thermo Scientific™ EASY-Spray™ HPLC C18 column (250 mm x 75 pm) with 2 pm particle size (100 A pores), preceded by a Thermo Scientific™ Acclaim™ PepMap™ 100 C18 HPLC Column (70 mm x 75 pm) with a 3 pm particle size. Resolution was achieved with the following gradient 2.5% B 0 - 2 min, 7.5 % B 3 min, 32 % B 63 min, 42.5 % B 66 min, 99% B 71 -80 min, zebra wash 80 - 98 min, where buffer A is 0.1% formic acid in water, buffer B 80:20 acetonitrile (ACN): 0.1% formic acid. Additional samples were resolved using a Thermo Scientific™ Vanquish™ Neo RSLCnano UHPLC system equipped with an lonoptics Aurora Ultimate 25 cm x 75 pm ID, 1.7 pm column.
[0111] Data was collected on a Thermo Scientific Orbitrap Exploris™ 240 mass spectrometer. The orbitrap mass resolution was 120,000, with a scan range of m / z 350-1200 full MS AGC target was 300%. Spray voltage was set to 2.0 kV and ion transfer tube temperature at275 °C. Dynamic exclusion was used with exclusion after n times and an exclusion window of 13 secs. Data dependent mode was set to Cycle Time. For fragmentation spectra, a resolution of 15,000, Isolation window was set at m / z 1.5, normalised collision energy was set at 30 eV. The AGC target value was set to standard, with max injection time on auto.Data analysis
[0112] The LC / MS / MS data acquired from the reference hair extracts was analysed by database search using PEAKS Studio11 (BioInformatics Solutions Inc., Canada) using the method of Zhang, J., et al. (PEAKS DB: de novo sequencing assisted database search for sensitive and accurate peptide identification. Mol Cell Proteomics, 2012. 11(4): M111.010587). The database for each reference species was downloaded from the National Centre of Biotechnology Information (NCBI; https: / / www.ncbi.nlm.nih.gov) and used to carry out the database searches on the reference samples.
[0113] The following search parameters were used: carbamidomethylation of cysteine was considered a fixed modification, whilst deamidation of glutamine and asparagine, and the oxidation of methionine were set as variable modifications. Max variable post translational modifications (PTMs) per peptide were capped at three. Error tolerance was set at 10 ppm for precursor masses, and 0.02 Da for fragment ions. The false discovery rate (FDR) was set to 1 %. BLAST alignments
[0114] Database search results from the shotgun data were interrogated to profile the proteomes of each reference sample and identify the proteins that extracted with greatest abundance. The most abundant keratin proteins for each of these species were aligned alongside the equivalent keratins of a closely related species and the orthologous keratin from Homo sapiens. These alignments were performed using the protein BLAST (BLASTp) alignment tool available from NCBI (Altschul et al 1997), with visual comparisons conducted manually to identify tryptic peptides containing amino acid variations between the protein sequence of the comparison species. Peptides that contained variations from both of the closely related species, and the human protein sequence, were subject to a unique peptide search through the Uniprot database using the peptide search function (https: / / www.uniprot.org / peptide-search), and through the NCBI protein database using the BLASTp tool, to confirm if the peptide was present in the proteome of another annotated species. Peptides confirmed as being present exclusively in the proteome of the species of interest or in a small selection of closely related species were listed as potential taxonomically informative markers for use in a forensic workflow for taxonomic identification. Combined database searching
[0115] The database search results from the reference samples were used to create an inhouse combined keratin database. The combined database contains the FASTA files of eachkeratin and KAP that appeared in the database search results of the reference species when searched against the source species NCBI database.
[0116] The LC / MS / MS data from each of the 45 reference samples was then reprocessed by database searching, again with PEAKs Studio 11, but this time using the inhouse combined keratin and KAP database. The database search was carried out with the same parameters as the initial source database search. The LC / MS / MS data from the second sample set, including hair from two additional individuals from each species extracted in duplicate, was processed with the inhouse keratin database. In this case the search space was reduced by searching without variable modifications, other search parameters remained the same.Positive peptide spectral matches (PSMs)
[0117] Due to the high standard of quality control required for a forensic application, several quality control measures were put in place during the data analysis process. These measures included setting a peptide abundancy threshold of 10,000 to prevent misidentifications resulting from noisy spectra. To prevent false positives due to a carryover effect, 5% of the previous samples total ion chromatogram (TIC) was subtracted from the next acquired data file in the analytical sequence. Further, each peptide spectrum match made by the database searching algorithm was manually interrogated. PSMs were only accepted if the match was made on the native peptide or with the only modifications being carbamidomethylation (fixed modification), the retention time matched the expected retention time based on reference hair samples, the MS2spectra included a good spread of y ions and two or more b ions (this was relaxed for peptides < 10 AA’s) and that the quality of the MS2spectra match the abundance of the precursor ion. ResultsProtein overview
[0118] Database search results from the 15 species of interest were reviewed to characterise the protein composition of the extracted hair proteome for each species. Good protein recovery was achieved for all 45 samples analysed, with several of the most abundant keratin and KAPs recovered with greater than 90% amino acid sequence coverage (Table 1). This demonstrates the utility of the modified preparative workflow as a universal method for efficient protein extraction from a variety of mammalian hairs of wide-ranging characteristics and morphologies. For Homo sapiens recovery averaged 4,945 unique peptides per sample and a combined total of 6,070 unique peptides were identified from the technical triplicates of one individual.
[0119] For each of the hairs analysed, there was an even spread of type I and type II keratins, with the database search results achieving greatest coverage on the same keratin proteins for many species. Type I keratins 31, 33A and 34 and type II keratins 81, 83 and 86 were the best recovered proteins overall amongst the species analysed, regularly achieving greater than 90%sequence coverage. Whilst many of the smaller KAPs were also identified with sequence coverages > 90%, the KAPs of greatest abundance varied more significantly from species to species.Table 1: Summary of protein amino acid sequence coverage from triplicate reference hair samples extracted from 15 species analysed by shotgun proteomics and processed by database search using PEAKS> > >Combined keratin database
[0120] The proteomics data from the reference samples of each species were processed with the inhouse combined keratin database. The results were analysed to compare the total protein coverage achieved by each sample for eight of the most abundant proteins from each of the species of interest. Examining the degree of sequence alignment between the sample-derived peptides and those present in the combined database allowed for evaluation of the specificity of protein sequence matching. This approach enabled the determination of whether a given hair sample preferentially matched to the proteome of its source species’, thereby assessing the accuracy and discriminatory power of the combined database at a protein level. The 8 proteins used in this comparison were type I proteins; KRT 31, 33A, 33B, 34 and type II proteins KRT 81, 83, 85 and 86. The only exceptions were for species that had demonstrated poor or no recovery of one of these proteins when LC / MS / MS data was processed with that species source database. In these cases, a substitute keratin of the same type was selected for that species. For example, KRT 35 was used in place of KRT 34 for Ovis aries.
[0121] As the heat maps in Figure 2 demonstrate, the best amino acid sequence match consistently occurred for the keratin proteins belonging to the species’ of origin. Further, thehuman hair samples did not match favourably with proteins from any other species in the panel, allowing unambiguous identification of the human hair sample from amongst this panel of species.
[0122] The same data analysis workflow was performed on the 10 most abundant KAPs extracted from each reference sample. Interestingly, these smaller proteins provided greater distinction between the species, with a clear match between each reference sample and the KAP sequences from the source database.Taxonomically informative peptide identification
[0123] The approach to biomarker identification involved BLAST alignments of the most abundant keratin proteins, determined experimentally, against analogous proteins of closely related species. Tryptic peptides in regions of sequence variation were assessed for their level of taxonomic discrimination by querying protein databases (NCBI and Uniprot). This process resulted in the identification of more than 250 potential biomarkers, each exhibiting varying degrees of specificity. These peptides offered taxonomic discrimination at the species, genus, or family level, or were present in the proteomes of a restricted number of species. Database search output from the processing of reference hair samples with the combined keratin database was interrogated to validate the potential diagnostic peptides.
[0124] Peptides that were identified by PSM in reference samples, had consistent recovery and met the quality control criteria for identification, were included on the panel. The final panel resulted in 227 peptide markers - 182 of these markers were shared between 5 or fewer species, 45 belonged only to the species of interest and one other species (2 species in total), and 62 were species-specific (unique to the species of interest). Species-specific markers (> 1) were identified for each of the species, other than Vulpes vulpes, which shared all peptides identified as taxonomically discriminating with closely related Vulpes Iagos (arctic fox). The Vulpes species required at least two marker peptides for discrimination.
[0125] In Tables 2-17 below, some sequences are given more than one SEQ ID NO, although the sequences are the same. In such cases, both identifiers are listed in the Table. Diagnostic peptide panel
[0126] The diagnostic panel was used to evaluate the database search results of hair extracts from an additional two individuals per species (extracted in duplicate). This provided a final dataset of seven hair extracts taken from three different individuals for each of the 15 species of interest (105 hair extracts in total).Homo sapiens
[0127] Bioinformatics analysis of the human hair proteome alongside the protein sequences of other closely related primates led to the identification of 28 peptides deemed to betaxonomically informative. Seven of these peptides were confirmed to be exclusive to the hair proteome of the genus Homo, of which Homo sapiens is the only extant species.
[0128] Each Homo-specific marker was recovered consistently in all human hair samples analysed in this study, providing confidence that an unambiguous assignment down to the species level could be made from a trace amount of human hair. A further four peptide markers provided family level discrimination, shared with only a handful of closely related Hominidae species. The remaining peptides are all primate specific, providing greater confidence in an identification without contributing to improving the taxonomic resolution already achieved.Canidae family
[0129] Two species from the Canidae family were of interest: the domestic dog (Canis lupus familiaris), a common domestic species, and the red fox (Vulpes vulpes), a prevalent and widespread pest-species in Australia. Through bioinformatics efforts, eight peptide markers were identified that are unique to the Canidae family and common to both V. vulpes and C. lupus familiaris. Seven of these markers were confirmed in all Canidae hair samples analysed.
[0130] Five peptides specific for the species C. lupus familiaris and subspecies C. lupus dingo were identified, with four of these peptides recovered in all domestic dog samples analysed, enabling species level identification of these extracts. Four peptides were identified as specific to the genus Vulpes, unique to the hair proteomes of two species, V. vulpes and V. lagopus (Arctic fox). However, only two of these markers were consistently extracted in all Vulpes samples. An additional peptide marker, present in both the genus Vulpes and the Canidae species Nyctereutes procyonoides (Raccoon dog), was included in the panel. Although this peptide represents lower taxonomic specificity, it was consistently extracted in high abundance in each of the fox samples, providing a reliable marker capable of differentiating between Vulpes and Canis species.Felis species
[0131] The genus Felis is one of fourteen genera that belong to the family Felidae. The recent evolutionary divergence of the Felidae family complicates the identification of a single peptide marker that can reliably achieve genus-level discrimination within the Felis hair proteome.
[0132] In this study, the combination of ACLPCLPAASCGPGAVR (SEQ ID NO: 133) in the presence of QWSSAEQLQSCQAEIIELR (SEQ ID NO: 132) or YSSQLGQVQCMITNVESQLAEIR (SEQ ID NO: 128), present in all feline hair samples analysed, provides genus level discrimination.
[0133] Table 2 details the Felidae species that share peptide markers identified as taxonomically informative for Felis catus, the domestic house cat. As many wild cat species do not occur in Australia outside of captivity, family-level discrimination may be sufficient forgenerating forensic intelligence in local contexts. A combination of markers is required for genus level discrimination of the Felis catus.Table 2: Peptides that provide taxonomic discrimination of Felidae species.(+) peptide is present in species proteome# - Genus: Acinonyx-,A- Genus: Neofelis-, % - Genus: Panthera-, * - Genus: LynxBovidae family
[0134] Three of the species included in this study belong to the family Bovida e : Ovis aries (sheep), Bos taurus (cattle), and Capra hircus (goat).
[0135] The more closely related species, Ovis aries and Capra hircus both belong to the subfamily Caprinae. The present study identified two Capra specific markers, with a further eight being specific to C. hircus, as well as an additional six Ovis markers, with two being specific to O. aries. Notably, one of the Ovis markers (LGCGSGFR) identified in this study appears by BLAST search to be unique to Ovis aries, whereas previously identified markers were also found in Ovis ammon polii. Species level identification was achieved for all Caprinae species analysed.
[0136] The genus Bos of the Bovidae family belongs to the subfamily bovinae, which contains both wild and domestic species of cattle. A total of 15 family limiting peptides were identified from the hair proteome of Bos taurus, of these four were genus specific markers, and two were speciesspecific for Bos taurus. Species-specific resolution was achieved for all but one of the Bos taurus samples analysed, with the remaining achieving genus level classification.Equus species
[0137] Twenty one peptides were identified in the hair proteome of Equus caballus, the domestic horse, each deemed to be taxonomically informative. Of these, only one was found to be species-specific. However, thirteen peptides were genus-specific, with five shared exclusively with the closely related species Equus przewalskii (Mongolian wild horse).
[0138] LEVAVSQAEQQGEVALTDAR (SEQ ID NO: 122) the peptide unique to Equus caballus was recovered in all horse-hair samples analysed, achieving species level discrimination. Two species belonging to the genus Equus are found in Australia, Equus caballus and Equus asinus (donkey). Eleven of the diagnostic markers identified in the hair proteome are shared between E. caballus and E. asinus. A hair sample obtained from E. asinus was extracted in duplicate alongside the other reference samples, the sample was found to have the expected 11 / 20 peptide markers that are shared with E. caballus.Sus family
[0139] In Australia, all wild pigs (Sus scrofa) are descendants of various domestic pig breeds (Sus scrofa domestica). Consequently, all Suidae found locally whether wild or domestic belong to the genus Sus. One of the S. scrofa samples analysed in this study was obtained from a domestic miniature pig, whilst the remaining two hair samples were collected from wild boar. Twenty taxonomically informative peptides were identified in the hair proteome of S. scrofa. Seven of these peptides are unique to only S. scrofa, five which were consistently extracted from all pig hair samples, enabling species-level discrimination. Seven peptides shared only with the Phacochoerus africanus (Common warthog) provided family level discrimination, and were included on the panel.Muridae family
[0140] The Muridae family is a large group of rodents that includes both the house mouse (Mus musculus) and the brown rat (Rattus norvegicus). Despite belonging to the same family, these species are readily distinguishable at the protein level. Six peptides were found to be specific for the genus Rattus, and one of the six was species-specific for Rattus norvegicus. Sixteen peptides were considered as diagnostic in the identification of hair from the Mus musculus, eleven of these family specific for the Muridae family, seven provided greater discrimination to the genus level, and three are species-specific for Mus musculus.Oryctolaqus species
[0141] Oryctolagus cuniculus (European rabbit), a member of the Leporidae family, is one of the most widespread mammal species both globally and across Australia. Although occasionally kept as a pet, the rabbit is also a prevalent pest. Twenty-one of the most sensitive and selective peptides for O. cuniculus were chosen, with ten which are species-specific. All twenty-onemarkers were consistently extracted in each animal sample, allowing confident species level identification.Cavia species
[0142] Cavia porcellus (Guinea pig) is a species of rodent belonging to the family Caviidae. This species is not found naturally in the wild but is a popular domestic mammal. Eleven diagnostic markers were identified for this species, 10 of these providing species level identification.Camelidae family
[0143] Two species from the family Camelidae were analysed - Camelus dromedarius (dromedary camel) and Vicugna paces (alpaca). Bioinformatics alignments of Camelidae proteins identified five diagnostic peptides shared by both species. Additionally, 10 peptides specific to the genus Camelus were found. Of these, one peptide was unique to C. dromedarius, while the remaining were shared with the closely related C. bactrianus and C. ferus. The species-specific peptide TVNALEVELQAQYSLR (SEQ ID NO: 231) was recovered in all dromedary samples analysed. Ten additional diagnostically informative peptides were found in the alpaca hair proteome, five of which are species-specific. A majority of these (four of five) were consistently recovered in all . paces, allowing unambiguous species classification of the hair samples. Assorted families
[0144] Further sequences to identify several assorted families such as pica - Ochotonidae; and leaf-nosed bats - Phyllostomidae, are also provided.Example 2Identification of members of the sub-class Marsupialia of the class Mammalia
[0145] Hair samples from various marsupials were collected and processed as discussed in Example 1.Results
[0146] Sequences were obtained that allowed identification of hair samples as being from the mammalian clade Metatheria (mammals more closely related to marsupials than placental mammals). Table 16 provides sequences that permit identification of marsupials such as the monito del monte (Dromiciops gliroides), yellow footed antechinus (Antechinus flavipes), agile gracile opossum (Gracilinanus agilis), greater bilby (Macrotis lagotis), grey short tailed opossum (Monodelphis domestica), tammar wallaby (Notamacropus eugenii), sugar glider (Papuan subspecies; Petaurus breviceps papuanus), koala (Phascolarctos cinereus), Tasmanian devil(Sarcophilus harrisii), fat tailed dunnart (Sminthopsis crassicaudata), common brushtail possum (Trichosurus vulpecula), and common wombat (Vombatus ursinus).
[0147] Some of the sequences of Table 16 are also indicative of the monotremes platypus (Ornithorhynchus anatinus), and short beaked echidna (Tachyglossus aculeatus). However, using two or more of the listed sequences for identification purposes will allow hair from the marsupials to be distinguished from the monotremes.
[0148] Some of the sequences of T able 16 are also indicative of placental mammals such as the golden spiny mouse (Acomys russatus), nine-banded armadillo (Dasypus novemcinctus), Chinese hamster (Cricetulus g rise us), Philippine tarsier (Carlito syrichta), or meerkat (Suricata suricatta). However, using two or more of the listed sequences for identification purposes will allow hair from the marsupials to be distinguished from the placental mammals.Phascolarctos genus
[0149] One of the sequences used to distinguish the koala (Phascolarctos cinereus), may also be used to distinguish the monito del monte (Dromiciops gliroides).
[0150] One of the sequences used to distinguish the koala (Phascolarctos cinereus), may also be used to distinguish members of the Ursidae family (Ursus americanus, Urus arctos, Ursus maritimus, Ailuropoda melanoleuca). However, using two or more of the listed sequences for identification purposes will allow hair from the koala to be distinguished from the Ursides.Vombatus genus
[0151] One of the sequences used to distinguish the wombat (Vombatus ursinus), may also be used to distinguish the lesser Egyptian jerboa (Jaculus jaculus). Another may also be used to distinguish the gelada (Theropithecus gelada). However, using two or more of the listed sequences for identification purposes will allow hair from the wombat to be distinguished from the placental mammals.Macrotis genus
[0152] One of the sequences used to distinguish the greater bilby (Macrotis lagotis), may also be used to distinguish the platypus (Ornithorhynchus anatinus). However, using two or more of the listed sequences for identification purposes will allow hair from the bilby to be distinguished from the platypus.Table 4: Control peptides to identify mammalian hairTable 5: Peptides to identify members of the Homo genusTable 6: Peptides to identify members of the Canidae familyTable 7: Peptides to identify members of the Bovidae familyTable 8: Peptides to identify members of the Equus genusTable 9: Peptides to identify members of the Felis genusTable 10: Peptides to identify members of the Sus genusTable 11: Peptides to identify members of the Muridae familyTable 12: Peptides to identify members of the Oryctolagus genusTable 13: Peptides to identify members of the Cavia genusTable 14: Peptides to identify members of the Camelidae familyTable 15: Assorted familiesTable 16: Peptides to identify members of the Marsupial sub-class of the MammalsTable 17: Peptides to identify members of the Phascolarctos genusTable 18: Peptides to identify members of the Vombatus genusTable 19: Peptides to identify members of the Macrotis genusTable 19: Peptides to identify members of the Trichosurus genusTable 20: Preferred peptide sequences for placental mammal species specific identificationTable 21: Preferred peptide sequences for marsupial species specific identification
[0153] The species markers above demonstrated high specificity and sensitivity for the correct identification of hair samples (Figure 3), with 97.5% of the expected markers successfully recovered.
[0154] Shotgun proteomics allows the elucidation of the sequence of thousands of peptides in a single acquisition from a trace amount of sample, thus providing many possible peptide markers to achieve the taxonomic resolution desired. This depth of analysis allowed the identification of a series of taxonomically informative peptides for each of the species of interest.
[0155] Genus level discrimination was achieved consistently for hair samples from the species analysed, using a fraction of the peptide markers identified during the discovery process. While the desired taxonomic resolution was often achieved with just a single marker, the occurrence of false negatives (absence of markers for the species of origin) suggest robustness is best achieved using a panel of biomarkers for confirmation of identification, particularly in forensic settings.
[0156] These data demonstrate the capability of a proteomics approach for the taxonomic classification of trace amounts of biological materialab le 22: Ion information for peptide sequences; base sequence ion mass and post-translation ion massable 23: Ion information for peptide sequences; base sequence ion mass and post-translation ion massable 24: Ion information for peptide sequences; base sequence ion mass and post-translation ion massab le 25: Ion information for peptide sequences; base sequence ion mass and post-translation ion massab le 26: Ion information for peptide sequences; base sequence ion mass and post-translation ion massab le 27: Ion information for peptide sequences; base sequence ion mass and post-translation ion massable 28: Ion information for peptide sequences; base sequence ion mass and post-translation ion massable 29: Ion information for peptide sequences; base sequence ion mass and post-translation ion massab le 30: Ion information for peptide sequences; base sequence ion mass and post-translation ion massable 31: Ion information for peptide sequences; base sequence ion mass and post-translation ion massable 32: Ion information for peptide sequences; base sequence ion mass and post-translation ion massable 33: Ion information for peptide sequences; base sequence ion mass and post-translation ion massable 34: Ion information for peptide sequences; base sequence ion mass and post-translation ion massable 35: Ion information for peptide sequences; base sequence ion mass and post-translation ion mass
Claims
1. CLAIMS1. A method of identifying the source of a mammalian hair, said method comprising the steps of:a) obtaining a sample of hair;b) digesting the proteins in the hair to form peptides;c) analysing the peptides;d) comparing the analysis from step (c) to any one or more of reference SEQ ID NOs: 15- 310,wherein a match between one or more peptides and one or more of reference SEQ ID NOs: 15-310 indicates the mammal the hair is derived from.
2. A method of identifying a mammal, said method comprising the steps of:a) obtaining a sample of hair from the mammal;b) digesting the proteins in the hair to form peptides;c) analysing the peptides;d) comparing the analysis from step (c) to any one or more of reference SEQ ID NOs: 15- 310,wherein a match between one or more peptides and one or more of reference SEQ ID NOs: 15-310 indicates the mammal the hair is derived from.
3. The method of claim 1 or 2, comprising the additional step of:- comparing the analysis from step (c) to the peptides of any of reference SEQ ID NOs:12-14,wherein a match to between one or more peptides and one or more of reference SEQ ID NOs: 12-14 indicates the hair is derived from a mammal.
4. The method of any of the above claims, wherein the mammal is a marsupial.
5. The method of any of the above claims, wherein the mammal is chosen from:Homo sapiens (human), Canis lupus (domestic dog), Vulpes vulpes (red fox), Capra hircus (domestic goat), Ovis aries (domestic sheep), Bos taurus (cattle), Equus caballus (horse), Felis catus (domestic cat), Sus scrofa (wild boar), Rattus norvegicus (brown rat), Mus musculus (house mouse), Oryctolagus cuniculus (rabbit), Cavia porcellus (guinea pig), Camelus dromedarius (camel), or Vicugna pacos (alpaca);monito del monte (Dromiciops gliroidesy yellow footed antechinus (Antechinus flavipes), agile gracile opossum (Gracilinanus agilis), greater bilby (Macrotis lagotis), grey short tailed opossum (Monodelphis domestica), tammar wallaby (Notamacropus eugenii), sugar glider (Papuan subspecies; Petaurus breviceps papuanus), koala (Phascolarctos cinereus), Tasmanian devil (Sarcophilus harrisii), fat tailed dunnart (Sminthopsis crassicaudata), common brushtail possum (Trichosurus vulpecula), and common wombat (Vombatus ursinusyplatypus (Ornithorhynchus anatinus), and short beaked echidna (Tachyglossus aculeatusygolden spiny mouse (Acomys russatus), nine-banded armadillo (Dasypus novemcinctus), Chinese hamster (Cricetulus griseus), Philippine tarsier (Carlito syrichta), or meerkat (Suricata suricattayUrsus americanus, Urus arctos, Ursus maritimus, Ailuropoda melanoleuca, lesser Egyptian jerboa (Jaculus jaculus), gelada (Theropithecus gelada).
6. The method of any of the above claims, wherein the hair sample is washed before step (b).
7. The method of any of the above claims, wherein the method comprises the step of (i) purifying the digestate of contaminants; and / or (ii) extracting the peptides from the hair digestate, before analysing the peptides.
8. A kit comprising:(i) a set of amino acid sequence reference standards, being one or more sequences of any of SEQ ID NOs: 12-310; and(ii) instructions for use of the amino acid sequence reference standards for the identification of the source of a mammalian hair or identification of a mammal.
9. A kit for of identifying a mammal from a hair, said kit comprising:a) reagents in the form of:(i) one or more reagents for washing the hair;(ii) one or more reagents for digesting the proteins in the hair to form peptides;(iii) one or more reagents for extracting the peptides from the digestate; and / or(iv) one or more standard reference peptides that correspond in sequence to any one or more of SEQ ID NOs: 12-310b) instructions for use,wherein the instructions for use provide instructions to carry out the steps of: a) obtaining a sample of hair from the mammal;b) digesting the proteins in the hair to form peptides;c) analysing the peptides;d) comparing the analysis from step (c) to any one or more of reference SEQ ID NOs:15-310,wherein a match between one or more peptides and one or more of reference SEQ ID NOs: 15-310 indicates the mammalian species the hair is derived from.
10. The method of any one of claims 1 to 7, or kit of any one of claims 8 to 9, wherein the peptides are analysed by sequencing.
11. The method or kit of claim 10, wherein the sequencing is used to compare the peptide, and the match comprises at least 75%, 80%, 85%, 90%, 95% or 100% identity between the peptide and the amino acid sequence of any one or more of the reference sequences of SEQ ID NOs: 15-310.
12. The method of any one of claims 1 to 7, or kit of any one of claims 8 to 9, wherein the peptides are analysed by mass to charge ratio (m / z).
13. The method or kit of claim 12, wherein the mass to charge ratio is used to compare the peptide, and the match is within at least 1 Dalton of the mass to charge ratio of one or more of the reference sequences of SEQ ID NOs: 15-310.