Method for determining contribution of dog variety to test dog genome
By comparing the DNA methylation map of test dogs with the reference DNA methylation map of different dog breeds, the problem of difficulty in determining the contribution of dog breeds to the test dog genome is solved in the prior art, and the accurate identification of contributions to dog genome and prevention of health risks is achieved.
Patent Information
- Application Number
- CN202380076634.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-11-08
- Filing Date
- 2023-11-06
- Publication Date
- 2025-06-13
AI Technical Summary
Existing methods for predicting dog breeds are based mainly on single nucleotide polymorphisms in genetics or short tandem repetitions, making it difficult to effectively determine the contribution of dog breeds to testing dog genomes.
The contribution of dog breeds to the test dog genome is determined by providing DNA methylation maps from test dogs and comparing them with reference DNA methylation maps from different dog breeds.
This method can accurately determine the contribution of dog breeds to the test dog genome, thereby selecting appropriate dietary, drug or lifestyle regimens for dogs to prevent or reduce the risk of developing diseases.
Smart Images

Figure CN120153098A_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This application claims priority to U.S. Provisional Application Serial No. 63 / 423,671, filed on November 8, 2022, the entire disclosure of which is incorporated herein by reference. Technical Field
[0003] The present invention relates to methods of using DNA methylation maps to determine the contribution of dog breeds to a test dog's genome. The present invention further relates to methods of selecting a diet, lifestyle, or pharmaceutical regimen for a dog based on the contribution of at least one dog breed to the test dog's genome determined by a DNA methylation map, for example to prevent or reduce the risk of the dog developing a disease. Background Art
[0004] The ability to determine information about a dog's breed is desirable to inform the dog's lineage and provide information about its general health and well - being, such as its disease susceptibility.
[0005] The domestic dog (Canis familiaris) is a single species that has been divided into over 400 phenotypically distinct genetic isolates, called breeds, 152 of which are recognized by the American Kennel Club in the United States. Different breeds of dogs are characterized by a range of morphologies, behaviors, and disease susceptibilities.
[0006] Due to inbreeding programs used to produce specific morphologies, different diseases are known to segregate in different purebred dog populations. Methods for identifying dog breeds can be used to demonstrate that a dog belongs to a particular breed. Additionally, in the case of mixed - breed dogs, the ability to identify the contribution of different breeds to the mixed dog's genome and the specific characteristics of those contributions can be used to determine the likely characteristics of the mixed - breed dog (e.g., disease susceptibility).
[0007] Existing methods for predicting breed are based on genetics, where each breed is defined by, for example, a set of single - nucleotide polymorphisms or short tandem repeats (STRs).
[0008] However, there is a need for additional methods for determining the contribution of dog breeds to a test dog's genome. Summary of the Invention
[0009] In a first aspect, the present invention provides a method for determining the contribution of a dog breed to a test dog's genome, the method comprising:
[0010] a) providing a DNA methylation map from a sample obtained from the test dog; and
[0011] b) Determining the contribution of dog breed to the genome of the test dog by comparing at least a portion of the DNA methylation profile of the test dog with reference DNA methylation profiles from different dog breeds.
[0012] The present invention also provides a method for determining the contribution of dog breed to the genome of a test dog, the method comprising:
[0013] a) Providing a DNA methylation profile from a sample obtained from the test dog; and
[0014] b) Determining the contribution of dog breed to the genome of the test dog by comparing at least a portion of the DNA methylation profile of the test dog with reference DNA methylation profiles from at least one reference dog breed.
[0015] Suitably, assessing the breed of a dog can allow for the selection of a diet, medication, or lifestyle that can help improve the health and well-being of the dog. For example, the assessment can be based on the size of the breed determined to contribute to the genome of the test dog to recommend a diet suitable for, for example, the size of the dog breed. Another example is based on breed body type, as some breeds can be classified as athletic or robust breeds. The breed defined as contributing to the genome of the test dog can be a pure breed (e.g., as defined by the American Kennel Club), or a clade or cluster (e.g., as defined by a phylogenetic breed wheel - e.g., Parker et al.; Cell Reports; 2017; 19,697 - 708).
[0016] Assessing the breed composition of a dog can be particularly useful for mixed-breed dogs. For example, the epigenetic portion of each breed in a mixed-breed dog can be used to specifically assess which pure-breed characteristics are passed on to the mixed-breed dog. This determination can then be used to specify a diet, medication, or lifestyle to improve the health and well-being of the mixed-breed dog.
[0017] For example, many purebreds have a susceptibility to specific diseases or medical conditions. For example, the Afghan hound is prone to glaucoma, hepatitis, and hypothyroidism; the Basenji is prone to Escherichia coli enteritis and pyruvate kinase deficiency; the Beagle is prone to bladder cancer and deafness; the Bernese Mountain dog is prone to cerebellar ataxia; the Border Terrier is prone to oligodendroglioma; and the Labrador Retriever is prone to food allergies. Of the genetic diseases found in dogs, 46% are thought to occur primarily or only in one or a few breeds (Patterson et al. (1988) J Am. Vet. Med. Assoc. 193:1131). Thus, for the purpose of proactively considering the health risks of an individual test animal, information about the genomic contribution of one or more breeds to the genome of a test animal is particularly valuable to the owner or caregiver of a mixed-breed canine. For example, the genetic diseases of a mixed-breed dog found to be a mixture of a Newfoundland and a Bernese Mountain dog can be proactively monitored, which occur at a rare frequency in the general dog population but at a significant frequency in these specific breeds; thus, this type of mixed-breed individual would benefit from screening for histiocytic sarcoma. Health-related information can also include potential treatments, special diets or products, diagnostic information, and insurance information.
[0018] In another aspect, the present invention provides a method for selecting a diet, medication, or lifestyle program for a test dog, the method comprising:
[0019] a) providing a DNA methylation profile from a sample obtained from the test dog;
[0020] b) determining the breed contribution to the genome of the test dog by comparing at least a portion of the DNA methylation profile of the test dog with reference DNA methylation profiles from different dog breeds; and
[0021] c) selecting a suitable diet, medication, or lifestyle program for the test dog based on the breed contribution to the genome of the test dog determined in step b).
[0022] In another aspect, the present invention provides a method for preventing or reducing the risk of a test dog developing a disease; the method comprising:
[0023] a) providing a DNA methylation profile from a sample obtained from the test dog;
[0024] b) determining the contribution of dog breeds to the genome of the test dog by comparing at least a portion of the DNA methylation profile of the test dog with reference DNA methylation profiles from different dog breeds; wherein at least one dog breed contributing to the genome of the test dog is associated with a predisposition to develop a disease; and
[0025] c) selecting a diet, medication, or lifestyle regimen for the test dog based on the contribution of at least one dog breed to the genome of the test dog determined in step b);
[0026] wherein the medication, lifestyle, or diet regimen prevents or reduces the risk of the test dog developing the disease.
[0027] Accordingly, the present invention enables the selection of an appropriate diet, medication, or lifestyle regimen for a dog based on the contribution of dog breeds to the genome of the test dog as determined by a DNA methylation profile.
[0028] As used herein, "selecting an appropriate diet, medication, or lifestyle for a dog" may also encompass "recommending a diet, medication, or lifestyle for a dog" or "providing a recommended diet, medication, or lifestyle for a dog".
[0029] The disease may be associated with the incidence or predicted incidence of: (i) a tissue; (ii) an organ; or (iii) a physiological system, such as the immune system, gastrointestinal system, urinary system, muscular system, cardiovascular system, and / or nervous system.
[0030] The disease may be osteoarthritis, dementia, cognitive dysfunction, pre-diabetic condition, diabetes, cancer, heart disease, obesity, gastrointestinal disorders, incontinence, kidney disease, sarcopenia, vision loss, hearing loss, osteoporosis, cataracts, cerebrovascular disease, and / or liver disease.
[0031] Suitably, the disease is a breed-related disease. For example, breed-related diseases may be osteoarthritis, dementia, cognitive dysfunction, pre-diabetic condition, diabetes, cancer, heart disease, obesity, gastrointestinal disorders, incontinence, kidney disease, sarcopenia, vision loss, hearing loss, osteoporosis, cataracts, cerebrovascular disease, and / or liver disease.
[0032] The method may optionally further comprise administering a diet, medication, or lifestyle regimen to the dog.
[0033] The lifestyle or diet regimen may be a dietary intervention. The dietary intervention may be a calorie-restricted diet, geriatric diet, or low-protein diet.
[0034] The present invention also provides a dietary intervention for preventing or treating a disease in a dog, wherein the dietary intervention is administered to a dog having a breed contribution determined by the method of the present invention.
[0035] The present invention further provides a computer-readable medium comprising instructions which, when executed, cause one or more processors to perform the methods of the present invention.
[0036] The present invention also provides a computer system for selecting a diet, drug, or lifestyle regimen for a test dog, the computer system being programmed to perform the following steps:
[0037] a) determining the contribution of dog breed to the test dog's genome by comparing at least a portion of the DNA methylation profile obtained from the test dog with reference DNA methylation profiles from different dog breeds; and
[0038] b) selecting a suitable drug, lifestyle, or diet regimen for the test dog based on the contribution of the dog breed to the test dog's genome determined in step a).
[0039] In another aspect, the present invention provides a computer program product comprising computer-executable instructions for causing a programmable computer to determine the contribution of dog breed to a test dog's genome by comparing at least a portion of the DNA methylation profile obtained from the test dog with reference DNA methylation profiles from different dog breeds.
[0040] In another aspect, the present invention provides a computer program product comprising computer-executable instructions for causing a programmable computer to select a diet, drug, or lifestyle regimen for a test dog by the following steps: a) determining the contribution of dog breed to the test dog's genome by comparing at least a portion of the DNA methylation profile obtained from the test dog with reference DNA methylation profiles from different dog breeds; and b) selecting a suitable drug, lifestyle, or diet regimen for the test dog based on the contribution of the dog breed to the test dog's genome determined in step a). BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 - UMAP of all samples using all fully observed methylation sites illustrates that methylation can accurately classify beagles and labrador retrievers. A binomial LASSO classifier trained on 2 / 3 of the data was able to accurately classify 2 breeds using 19 methylation sites.
[0042] Figure 2 - Examples of breeds classified as robust or athletic
[0043] Figure 3 - Exemplary breed clades, as classified by Parker et al. (Cell Reports; 2017; 19, 697-708).
[0044] Figure 4 - T-SNE of 200 selected methylation sites used in the classifier. Darker colors indicate misclassified dogs, while lighter colors indicate correctly classified dogs, and the shapes indicate the true breeds. All data except for one dog were fit using an SVM classifier to determine predictions, and the breed of the withheld dog was predicted using the model. Dogs that appeared as outliers in the above T-SNE were excluded from the training set. Detailed Description
[0045] The preferred features and embodiments of the present invention will now be described by way of non-limiting examples. Those skilled in the art will understand that they can combine all the features of the present invention disclosed herein without departing from the scope of the present invention disclosed.
[0046] It must be noted that, as used herein and in the appended claims, the singular forms "a", "an", and "the" include plural referents unless the context clearly dictates otherwise.
[0047] As used herein, the terms "comprising" and "consisting of" are synonymous with "including" or "containing", and are inclusive of end values or open-ended, and do not exclude additional unrecited members, elements or method steps. The terms "comprising" and "consisting of" also include the term "consisting of".
[0048] Numeric ranges include the numbers defining the range.
[0049] The publications discussed herein are provided only for their disclosure prior to the filing date of the present patent application. Nothing herein is to be construed as an admission that such publications constitute prior art to the appended claims herein.
[0050] The methods and systems disclosed herein can be used by veterinarians, healthcare professionals, laboratory technicians, pet care providers, etc.
[0051] subject
[0052] The present method is directed to canine subjects. Accordingly, the subjects of the present invention are dogs.
[0053] breed
[0054] The present invention relates to a method for determining the contribution of dog breeds to the genome of a test dog using DNA methylation maps. Specifically, the present invention provides a method for determining the contribution of dog breeds to the genome of a test dog, the method comprising: a) providing a DNA methylation map of a sample obtained from the test dog; and b) determining the contribution of dog breeds to the genome of the test dog by comparing at least a portion of the DNA methylation map of the test dog with reference DNA methylation maps from different dog breeds.
[0055] Step b) of the method may comprise determining the contribution of dog breeds to the genome of the test dog by comparing at least a portion of the DNA methylation map of the test dog with reference DNA methylation maps from at least one reference dog breed.
[0056] Accordingly, the present invention can be used to determine the breed of a dog or the probability that it belongs to a given breed.
[0057] Suitably, the method can be used to determine the contribution of one or more dog breeds to the genome of the test dog.
[0058] Suitably, the genome of the test dog is from a mixed breed dog, and the method of the present invention can be used to determine the contribution of one or more dog breeds to the genome of the mixed breed dog.
[0059] The method can be used to determine the contribution of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, at least 30, at least 50, at least 100, at least 150, at least 200, at least 300, or at least 400 dog breeds to the genome of the test dog.
[0060] The method can be used to determine the contribution of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, at least 30, at least 50, at least 100, at least 150, or at least 200 dog breeds to the genome of the test dog.
[0061] This method can be used to determine the contribution of about 1 to about 400, about 1 to about 300, about 1 to about 200, about 1 to about 100, about 1 to about 50, about 1 to about 20, or about 1 to about 10 dog breeds to the genome of a test dog. This method can be used to determine the contribution of about 3 to about 400, about 3 to about 300, about 3 to about 200, about 3 to about 100, about 3 to about 50, about 3 to about 20, or about 3 to about 10 dog breeds to the genome of a test dog. This method can be used to determine the contribution of about 5 to about 400, about 5 to about 300, about 5 to about 200, about 5 to about 100, about 5 to about 50, about 5 to about 20, or about 5 to about 10 dog breeds to the genome of a test dog.
[0062] This method can be used to determine the contribution of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 dog breeds to the genome of a test dog.
[0063] Suitably, this method can be used to determine the contribution of at least one, at least two, or at least three dog breeds to the genome of a test dog.
[0064] Suitably, this method can be used to determine the contribution of one, two, or three dog breeds to the genome of a test dog.
[0065] Suitably, this method can be used to determine the contribution of two dog breeds to the genome of a test dog.
[0066] Suitably, this method can be used to determine the contribution of one dog breed to the genome of a test dog.
[0067] The dog breed can be a pure breed, such as the breeds defined by the American Kennel Club (http: / / www.akc.org / ). Examples of pure breeds include, but are not limited to, Afghan Hound, Airedale Terrier, Akita, Alaskan Malamute, American Eskimo Dog, American Foxhound, American Hairless Rat Terrier, American Staffordshire Terrier, American Water Spaniel, Australian Cattle Dog, Australian Shepherd, Australian Terrier, Basenji, Basset Hound, Beagle, Bearded Collie, Bedlington Terrier, Belgian Laekenois, Belgian Malinois, Belgian Sheepdog, Belgian Tervuren, Bernese Mountain Dog, Bichon Frise, Bloodhound, Border Collie, Border Terrier, Borzoi, Boston Terrier, Bouvier des Flandres, Boykin Spaniel, Boxer, Briard, Brittany, Bulldog, Brussels Griffon, Bullmastiff, Bull Terrier, Cairn Terrier, Cardigan Welsh Corgi, Cavalier King Charles Spaniel, Chesapeake Bay Retriever, Chihuahua, Chinese Crested DogCrested), Chinese Shar-Pei, Chow Chow, Clumber Spaniel, Cocker Spaniel, Collie, Curly-Coated Retriever, Dachshund, Dalmatian, Dandie Dinmont Terrier, Doberman Pinscher, Dogo Canario, English Cocker Spaniel, English Foxhound, English Setter, English Springer Spaniel, Entlebucher Mountain Dog, Field Spaniel, Flat-Coated Retriever, French Bulldog, German Longhaired Pointer, German Shepherd Dog, German Shorthaired Pointer, German Wirehaired Pointer, Giant Schnauzer, Golden Retriever, Gordon Setter, Great Dane, Great Pyrenees, Greater Swiss Mountain Dog, Greyhound, Harrier, Havanese, Ibizan Hound, Irish Setter, Irish Terrier, Irish Water Spaniel, Irish Wolfhound, Italian Greyhound, Jack Russell Terrier, Keeshond, Kerry BlueTerrier), Komondor, Kuvasz, Labrador Retriever, Leonberger, Lhasa Apso, Lowchen, Maltese, Manchester Terrier - Standard, Manchester Terrier - Toy, Mastiff, Miniature Bull Terrier, Miniature Pinscher, Miniature Poodle, Miniature Schnauzer, Large Munsterlander, Neapolitan Mastiff, Newfoundland, New Guinea Singing Dog, Norwegian Elkhound, Norwich Terrier, Old English Sheepdog, Papillon, Pekingese, Pembroke Welsh Corgi, Petit Basset Griffon Vendeen, Pharaoh Hound, Pointer, Polish Lowland Sheepdog, Pomeranian, Portuguese Water Dog, Presa Canario, Pug, Puli, Pumi, Rhodesian Ridgeback, Rottweiler, Saint Bernard, Saluki, Samoyed, Schipperke, Scottish Deerhound, Scottish Terrier, Silky Terrier, Shetland Sheepdog, Shiba Inu, Shih Tzu, Siberian Husky, Smooth Fox Terrier, SoftCoated Wheaten Terrier, Spinone Italiano, Staffordshire Bull Terrier, Standard Poodle, Standard Schnauzer, Sussex Spaniel, Tibetan Spaniel, Tibetan Terrier, Toy Fox Terrier, Toy Poodle, Vizsla, Weimaraner, Welsh Springer Spaniel, Welsh Terrier, West Highland White Terrier, Wirehaired Pointing Griffon, Whippet, and Yorkshire Terrier.
[0068] Suitably, a dog breed can refer to a group or clade of purebreds classified based on one or more criteria.
[0069] Based on genetic distance, optionally combined with additional factors such as migration and whole-genome haplotype sharing analysis, breeds can be classified into clades. An example of breed clade classification is described in Parker et al. (Cell Reports; 2017; 19, 697 - 708), where 161 dog breed analyses were classified into 23 breed clades. The exemplary breed clades classified by Parker et al. are shown in Figure 3 in.
[0070] Suitably, the breed clade can be selected from wild dogs, Basenjis, Asian Spitz, Asian Toy, Nordic Spitz, Schnauzers, Small Spitz, Toy Spitz, Hungarian, Poodle, American Terrier, American Toy, Pinscher, Terrier, New World, Mediterranean, Scent Hound, Retriever, Pointer Setter, Continental Herder, UK Rural, Drover, Alpine, and European Mastiff. For example, the breed clade can include Figure 3 the pure breeds shown in
[0071] Breeds can be classified based on the size of the breed, such as based on the average size of the breed. For example, breeds can be classified as toy, small, medium, large, or giant breeds. Suitably, dog breeds can be classified based on the weight of the dog. Suitably, dog breeds can be classified based on the average weight of dogs of a given breed. A "miniature breed" can refer to a breed with an average weight of less than 5 kg. A "small breed" can refer to a breed with an average weight between 5 kg and 10 kg. A "medium breed" can refer to a breed with an average weight between 10 kg and 25 kg. A "large breed" can refer to a breed with an average weight between 25 kg and 40 kg. A "giant breed" can refer to a breed with an average weight exceeding 40 kg.
[0072] Suitably, breeds can be classified based on body type. For example, breeds can be classified as stocky or athletic. Certain breeds can be grouped as stocky or athletic using methods such as those described in EP1983842. Suitably, body type is affected by and depends on multiple factors, including body mass index, body composition, daily energy requirements, resting metabolic rate, dog breed, and genetic differentiation during the breeding history.
[0073] The body mass index can be calculated by the formula: body weight (kg) / [shoulder height (m)] 2 . Stocky dogs typically have a body mass index greater than 90 kg / m 2 and athletic dogs typically have a body mass index less than 90 kg / m2 Body mass index. Examples of typical BMI values for strong or athletic dogs are:
[0074] robust dog athletic dog <![CDATA[Saint Bernard: 158.2 kg / m 2 > <![CDATA[Greyhound: 54.8 kg / m 2 > <![CDATA[Bull dog: 211.5 kg / m 2 > <![CDATA[Irish Setter: 61.6 kg / m 2 > <![CDATA[Pekingese: 99.6 kg / m 2 > <![CDATA[Fox Terrier: 51.8 kg / m 2 >
[0075] Examples of breeds classified as strong or athletic are shown in Figure 2 .
[0076] Classifying a dog as strong or athletic may also be influenced by the dog's breeding history. For example, a dog may have a breeding history and genetic background different from the breed category in which it is primarily classified. Generally, dogs with some sporting blood in their breeding history tend to maintain a sporting conformation as the dominant phenotype and have higher energy requirements. For example, a Great Dane belonging to the working and guard dog group (and thus should be classified as a strong dog) may be classified as athletic due to its conformation and breeding history (sight hound blood). It has a distinct athletic body type, i.e., a deep chest and a thin abdomen, and a high daily energy requirement to maintain its proper weight.
[0077] Evaluating the breed composition of a dog can be particularly useful for mixed-breed dogs. Specifically, methylation maps can be used to determine the contribution of different breeds to the genome of the test dog. For example, methylation maps can be used to determine the percentage contribution of different dog breeds to the genome of the test dog. Methylation maps can also be used to identify regions of the test dog genome that are similar to or have been inherited from a given breed. This can be particularly advantageous when a given genomic region or locus is known to co-segregate with, for example, disease susceptibility or particularly behavioral traits.
[0078] sex
[0079] Suitably, the sex of the dog can be classified as male or female. Suitably, the sex of the dog can be included in the method (e.g., in the regression analysis as described herein).
[0080] chronological age
[0081] Chronological age can be defined as the amount of time elapsed from the birth of the subject to a given date. Chronological age can be expressed in years, months, days, etc.
[0082] Suitably, the method can be applied to dogs of any chronological age.
[0083] Suitably, the chronological age of the dog can be included in the method (e.g., in the regression analysis as described herein).
[0084] biological age
[0085] For example, depending on genetics, nutrition, and lifestyle, the rate of aging of an individual may be slower or faster than their chronological age. Thus, chronological age may not always reflect the biological rate of aging of an individual. Accordingly, the biological age of an individual (based on, for example, clinical biochemistry and cell biology metrics) may differ from that of other individuals of the same chronological age. Methods for determining biological age are known in the art and include, for example, methods that utilize methylation profiles, clinical chemistry profiles, or telomere length.
[0086] Suitably, the method can be applied to dogs of any biological age.
[0087] Suitably, the biological age of a dog can be included in the method (e.g., in a regression analysis as described herein).
[0088] sample
[0089] The invention includes the step of providing a DNA methylation profile from one or more samples obtained from a subject.
[0090] The invention includes the step of determining a DNA methylation profile from one or more samples obtained from a subject.
[0091] Suitably, the sample is a blood, hair follicle, oral swab, saliva, fecal, or tissue sample.
[0092] Suitably, the sample is derived from blood. The sample can comprise blood components or can be whole blood. The sample preferably comprises whole blood. The sample can include peripheral blood mononuclear cells (PBMCs) or a lymphocyte sample. Techniques for collecting a sample from a subject and extracting DNA (e.g., genomic DNA) from the sample are well known in the art.
[0093] Suitably, the sample is a hair follicle, oral swab, or saliva sample. Such sample types are particularly suitable if the sample is provided, for example, outside a veterinary setting—e.g., using a kit for home use.
[0094] DNA methylation
[0095] DNA methylation is the process of covalently adding a methyl group (CH 3 ) to a cytosine base that is part of a DNA molecule. In vivo, this process is catalyzed by the DNA methyltransferase (Dnmt) family, which produces the modified cytosine by transferring a methyl group from S-adenosylmethionine (SAM). The cytosine is modified at the 5th carbon atom, and the modified residue is called 5-methylcytosine (5mC). DNA methylation can also include 5-hydroxymethylcytosine (5hmc).
[0096] DNA methylation is an example of an epigenetic mechanism, i.e., it can modify gene expression without modifying the underlying DNA sequence. DNA methylation can inhibit gene expression, for example, by acting as a recruitment signal for repressors or by directly blocking the recruitment of transcription factors. DNA methylation mainly occurs at sites of adjacent cytosine and guanine that form dinucleotides (CpGs) in the genomes of mammalian somatic cells. Although non-CpG methylation is observed during embryonic development, these modifications are greatly reduced in most cell types in adults. CpG islands are DNA fragments with a high CpG density but are generally unmethylated. These regions are associated with promoter regions, especially those of housekeeping genes, and are thought to be maintained in a permissive state allowing gene expression.
[0097] The detection of specifically methylated DNA can be accomplished by a variety of methods (see, for example, Zuo et al., 2009; Epigenomics. 1(2):331 - 345) and Rauluseviciute et al.; Clinical Epigenetics; 2019; 11(193)). Many methods are available for detecting differentially methylated DNA at specific loci in samples such as blood, urine, feces, or saliva. These methods are able to distinguish 5-methylcytosine or methylated DNA from unmethylated DNA and subsequently quantify the proportion of methylated and unmethylated DNA at specific genomic loci.
[0098] The method can include using any suitable method to determine the DNA methylation profile of a dog. Suitable methods include, but are not limited to, those described below.
[0099] enzymatic methylation sequencing (EM-seq) )
[0100] Suitably, enzymatic methods are used to detect 5mC and 5hmC. For example, enzymatic methyl-sequencing (EM-seq) can be used.
[0101] Typically in EM-seq, in the first enzymatic step, 5mC is oxidized to 5hmC by the activity of Tet methylcytosine dioxygenase 2 (TET2), then to 5fC, and finally to 5caC. Additionally, the use of the T4-BGT enzyme glycosylates both pre-existing 5hmC and 5hmC generated by TET2 activity. In the second enzymatic step, after denaturation of double-stranded DNA, the enzyme apolipoprotein B mRNA editing enzyme catalytic polypeptide-like 3A (APOBEC3A) is used to deaminate cytosine, but not the oxidized or glycosylated forms of 5mC and 5hmC. Only unmethylated cytosine is deaminated to form a uracil base. Prior to the first enzymatic step, DNA fragments can be generated by mechanical shearing and end repair, A-tailing, and ligation to sequencing adapters, which can be performed using, for example DNA Ultra II reagent (NEB). After the second enzymatic step, the deaminated single-stranded DNA can be amplified by a PCR reaction using a polymerase that can amplify templates containing uracil (such as Q5U TM ), and the resulting library can be sequenced or analyzed in the same manner as DNA samples generated by bisulfite sequencing. The output of EM-seq is typically the same as whole-genome bisulfite sequencing, but uses fewer DNA-damaging reagents, which thus reduces sample loss and can be superior to samples prepared by bisulfite conversion in terms of coverage, sensitivity, and accuracy of cytosine methylation calling. An exemplary EM-seq method is described by Vaisvila et al. (Genome Research; 2021; 31:1-10).
[0102] bisulfite conversion-based method
[0103] When treated with sodium bisulfite, bisulfite conversion utilizes the selective conversion of unmethylated cytosine to uracil. Denatured DNA is treated with sodium bisulfite, which converts all unmodified cytosines to uracil, and subsequent PCR amplification converts these residues to thymine. Analysis of the resulting DNA sequence can be performed via many different methods, examples of which include, but are not limited to: denaturing gel electrophoresis, single-strand conformation polymorphism, melting curve, fluorescence real-time PCR (MethyLight), MALDI mass spectrometry, array hybridization, and sequencing (such as whole-genome bisulfite sequencing WGBS). Recently developed techniques (such as SeqCap Epi) enrich the sequences of interest prior to sequencing, which enables deeper coverage in a more focused region). Comparison of the sequence abundance in the bisulfite-converted sample to that of an untreated control allows analysis of methylation at the target site, where the proportion of the converted sequence indicates the methylation level at the target site.
[0104] Other variants of bisulfite conversion methods are available that are capable of distinguishing 5mC from the oxidized form 5-hydroxymethylcytosine (5hmC), which behaves the same as 5mC under standard bisulfite conversion, and are capable of detecting the further modification 5-formylcytosine (5fC). These methods, such as oxBS-Seq and redBS-Seq, utilize the oxidation and reduction of these markers to alter the sensitivity of each species to bisulfite conversion and quantify the amount of each modification at target loci by comparative analysis.
[0105] selective restriction endonuclease digestion method
[0106] Methods for analyzing DNA methylation patterns that exist can involve the use of restriction enzymes. These methods include, for example, restriction landmark genomic scanning (RLGS) (Costello et al., 2000; Nat Genet.; 24(2):132-8), methylation-sensitive representational difference analysis (MS-RDA) (Ushijima et al., Proc Natl Acad Sci U S A. March 18, 1997; 94(6):2284-9), and differential methylation hybridization (DMH) (Huang et al., Cancer Res. March 15, 1997; 57(6):1030-4). The digestion activity of restriction endonucleases can be methylation-dependent. This specificity can be used to distinguish methylated sequences from unmethylated sequences. Certain restriction enzymes (e.g., BstUI, HpaII, and NotI) are sensitive to methylated recognition sequences. Other restriction enzymes (such as McrBC) are specific for methylated sequences.
[0107] For example, differential methylation hybridization (DMH) (Huang et al., as described above]) requires initial fragmentation of the genome with a large number of genomic restriction enzymes (such as MseI), which fragments the genome into lengths less than 200 bp. After this step, digestion of the genomic fragments is carried out using a methylation-sensitive restriction endonuclease (MRE) or a mixture of MREs in some versions of the technique to improve coverage. Depending on the specificity of the one or more enzymes used, methylated sequences or unmethylated sequences will be degraded. The digested sequences will not be amplified in the subsequent PCR step. The resulting PCR products are suitable for further processing and analysis by a combination of sequencing or microarray hybridization with fluorescent dyes.
[0108] Suitably, the method utilizes a DNA methylation map generated by a method including the use of one or more MREs.
[0109] Suitable comparators can be used to study the methylation status between conditions. DNA from healthy subjects can be compared to that of elderly or diseased subjects to detect changes in methylation status (Huang et al., Hum Mol Genet. Mar 1999;8(3):459-70). Alternatively, a methylation-insensitive form of a secondary digestive enzyme, such as the HpaII isoschizomer MspI, can be used to generate control samples such that intra- or inter-genomic DNA methylation comparisons can be made (Khulan et al., Genome Res. Aug 2006;16(8):1046-55).
[0110] In some embodiments, methods for detecting methylation include randomly shearing or fragmenting genomic DNA, cutting the DNA with a methylation-dependent or methylation-sensitive restriction enzyme, and subsequently selectively identifying and / or analyzing the cut or uncut DNA. Selective identification can include, for example, separating the cut and uncut DNA (e.g., by size) and quantifying the cut or alternatively uncut sequence of interest. Alternatively, the method can encompass amplifying intact DNA after restriction enzyme digestion, such that only DNA that was not cut by the restriction enzyme in the amplified region is amplified. In some embodiments, gene-specific primers can be used for amplification. Alternatively, adapters can be added to the ends of randomly fragmented DNA, the DNA can be digested with a methylation-dependent or methylation-sensitive restriction enzyme, and primers that hybridize to the adapter sequence can be used to amplify the intact DNA. In this case, a second step can be performed to determine the presence, absence, or amount of a specific gene in the DNA amplification pool. In some embodiments, real-time quantitative PCR is used to amplify the DNA.
[0111] Suitably, digestion of a nucleic acid is detected by selective hybridization of a probe or primer to the undigested nucleic acid. Alternatively, the probe hybridizes selectively to both the digested and undigested nucleic acids, but for example electrophoresis facilitates discrimination between the two forms. Suitable detection methods for achieving selective hybridization with a hybridization probe include, for example, Southern or other nucleic acid hybridization.
[0112] Suitable hybridization conditions can be determined based on the melting temperature (Tm) of the nucleic acid duplex containing the probe. Those skilled in the art will appreciate that for each probe, the optimal hybridization reaction conditions should be determined empirically, but some general principles can be applied. Preferably, hybridization using short oligonucleotide probes is carried out at low to medium stringency. In the case of GC-rich probes or primers or longer probes or primers, high stringency hybridization and / or washing is preferred. High stringency is defined herein as hybridization and / or washing carried out in approximately 0.1×SSC buffer and / or approximately 0.1% (w / v) SDS or lower salt concentration and / or at a temperature of at least 65°C or equivalent conditions. The specific stringency levels mentioned herein encompass equivalent conditions using washing / hybridization solutions other than SSC known to those skilled in the art.
[0113] reduced representation bisulfite sequencing (RRBS)
[0114] Reduced Representation Bisulfite Sequencing (RRBS) uses the MspI restriction enzyme to enrich CpG-rich genomic regions - which cuts DNA at all CCGG sites regardless of their DNA methylation status at the CG site - and is capable of measuring DNA methylation levels at 5% to 10% of all CpG sites in the mammalian genome.
[0115] Thus, the method involves digesting DNA with methylation-insensitive MspI prior to bisulfite conversion and sequencing. Digesting genomic DNA with MspI produces fragments that always start with C (if the cytosine is methylated) or T (if the cytosine is not methylated and is converted to uracil in the bisulfite conversion reaction). This results in a non-random base pair composition. Additionally, due to the skewed frequency of C and T within the sample, the underlying composition is skewed. Various software for alignment and analysis are available, such as Maq, BS Seeker, Bismark, or BSMAP. Alignment with a reference genome allows the program to identify methylated base pairs within the genome.
[0116] affinity enrichment-based method
[0117] Differentiation between methylated DNA and unmethylated DNA can be achieved by using antibodies containing a methyl-CpG-binding domain (MBD), such as anti-5mC and / or methylated CpG-binding proteins. Antibodies to MBD-domain proteins are capable of specifically isolating methylated DNA relative to unmethylated DNA. The method using antibodies is generally referred to as MeDIP, while the method using methylated CpG-binding proteins is generally referred to as the MBD or MIRA method.
[0118] These methods require an initial fragmentation of the genome, which can be carried out by extensive genomic digestion with frequently cutting enzymes such as MseI, followed by affinity purification of the methylated fragments. The input DNA can be compared to the purified methylated DNA by microarray hybridization or sequencing to obtain a comparative analysis of the methylation levels at specific loci.
[0119] Other variants of affinity enrichment-based methods are available, such as MethylCap-Seq or MBD-Seq. These methods reduce sample complexity by using a salt gradient to elute methylated DNA fragments in a methyl-CpG-abundance-dependent manner, separating CpG islands and other highly methylated loci from loci with lower CpG density. These fractions can then be sequenced separately, thereby increasing sequence coverage.
[0120] single molecule sequencing and de novo methylation sequencing method
[0121] Contemporary sequencing methods are capable of directly sequencing individual molecules. Single molecule real-time (SMRT) DNA sequencing is available, for example, the Sequel system from Pacific Biosciences, and has been shown to be able to identify modified bases (such as methylated cytosine) based on polymerase kinetics. Nanopore sequencing devices (such as the MinION nanopore sequencer from Oxford Nanopore Technologies) that are capable of sequencing long stretches of DNA individually can also detect de novo base modifications, including methylation.
[0122] DNA methylation site
[0123] Suitably, a DNA methylation site can refer to the presence or absence of 5mC at a single cytosine, suitably a single CpG dinucleotide.
[0124] Suitably, a DNA methylation site can refer to the presence or absence of methylation across multiple CpG sites within a DNA region (i.e., the number of 5mC or the percentage of 5mC). Suitably, a DNA methylation site can refer to the methylation level across multiple CpG sites within a DNA region (i.e., the number of 5mC or the percentage of 5mC). A "DNA region" can refer to a specific portion of genomic DNA. These DNA regions can be specified by reference to a gene name or a set of chromosomal coordinates. Both gene names and chromosomal coordinates are well known and understood by those skilled in the art.
[0125] Suitably, gene names and / or coordinates can be based on the "Tasha" dog reference genome (https: / / www.ncbi.nlm.nih.gov / assembly / GCF_000002285.5; Jagannathan et al.; Genes (Basel); 2021; 12(6); 847).
[0126] For example, a DNA region can define a DNA portion near the promoter of a gene. The promoter region is known to be rich in CpG. By way of example, the DNA region can refer to about 3 kb upstream to about 3 kb downstream of the promoter; about 2 kb upstream to about 2 kb downstream; about 2 kb upstream to about 1 kb downstream; about 2 kb upstream to about 0.5 kb downstream; about 1 kb upstream to about 0.5 kb downstream; about 0.5 kb upstream to about 0.5 kb downstream. Suitably, the DNA region can refer to about 1 kb upstream to about 0.5 kb downstream of the promoter.
[0127] Suitably, the DNA region can contain or consist of CpG sites spaced less than about 5000, less than about 4000, less than about 3000, less than about 2000, less than about 1000, less than about 500, or less than about 200 bases apart.
[0128] Suitably, the DNA region can contain or consist of CpG sites spaced about 200 to about 5000, about 200 to about 4000, about 200 to about 3000, about 200 to about 2000, or about 200 to about 1000 bases apart.
[0129] Suitably, the DNA region can contain one or more CpG islands. Suitably, the DNA region can consist of CpG islands.
[0130] A "CpG island" can refer to a DNA region containing at least 200 bp, a GC percentage greater than 50%, and an observed to expected CpG ratio greater than 60%.
[0131] Suitably, the DNA methylation site does not contain X and / or Y chromosome CpG.
[0132] Suitably, the DNA methylation site does not contain CpG known to contain SNPs at CpG.
[0133] References to each gene / DNA region detailed above are to be understood as references to all forms of these molecules, as well as their fragments or variants. As will be understood by those skilled in the art, it is known that some genes exhibit allelic variation or single nucleotide polymorphisms between individuals. Variants include nucleic acid sequences from the same region sharing at least 90%, 95%, 98%, 99% sequence identity, i.e., having one or more deletions, additions, substitutions, reverse sequences, etc. relative to the DNA regions described herein. Thus, the present invention is to be understood as extending to such variants, which, for the purposes of this application, can achieve the same result despite minor genetic variations between the actual nucleic acid sequences of individuals. Accordingly, the present invention is to be understood as extending to all forms of DNA resulting from any other mutations, polymorphisms, or allelic variations.
[0134] For screening the methylation of these gene regions, it should be understood that the assay can be designed to screen for specific DNA. It is entirely within the skill of those in the art to select which strand to analyze and target based on chromosomal coordinates. In some cases, assays can be established to screen both strands.
[0135] “Methylation status” can be understood to refer to the presence, absence, and / or amount of methylation at one or more specific nucleotides within a DNA region. The methylation status of a specific DNA sequence (e.g., a DNA region as described herein) can indicate the methylation status of each base in the sequence or can indicate the methylation status of a subset of base pairs within the sequence (e.g., the methylation status of cytosine or the methylation status of one or more specific restriction enzyme recognition sequences), or can provide information regarding the regional methylation density within the sequence without providing exact information as to where methylation occurs within the sequence. The methylation status can optionally be represented or indicated by a “methylation value”.
[0136] Suitably, the EM-Seq strategy can be used to determine DNA methylation. In this method, the methylation level can be determined as the fraction of ‘C’ bases among the total ‘C’+‘U’ bases at the target CpG site “i” after enzyme and APOBEC3A conversion treatment. In other embodiments, the methylation level can be determined as the fraction of ‘C’ bases among the total ‘C’+‘T’ bases at site “i” after enzyme and APOBEC3A conversion treatment and subsequent nucleotide amplification. The average methylation level at each site can then be evaluated to determine whether one or more thresholds are met.
[0137] In some embodiments, particularly when using bisulfite conversion and sequencing methods, the methylation level can be determined as the fraction of 'C' bases in the total 'C' + 'U' bases at target CpG site "i" after bisulfite treatment. In other embodiments, the methylation level can be determined as the fraction of 'C' bases in the total 'C' + 'T' bases at site "i" after bisulfite treatment and subsequent nucleotide amplification. The average methylation level at each site can then be evaluated to determine whether one or more thresholds are met.
[0138] Alternatively, methylation values can be generated, for example, by quantifying the amount of intact DNA present after restriction digestion with a methylation-dependent restriction enzyme. In this example, if quantitative PCR is used to quantify a specific sequence in the DNA, an amount of template DNA approximately equal to that of a control treated mock indicates that the sequence is not highly methylated, while an amount substantially less than the amount of template present in the mock-treated sample indicates the presence of methylated DNA at that sequence. Thus, for example, the values (i.e., methylation values) from the above example represent the methylation status and can therefore be used as a quantitative measure of the methylation status. This is particularly useful when it is desired to compare the methylation status of a sequence in a sample to a threshold.
[0139] The present invention is not limited by the exact number of methylated residues that are considered indicative of a breed, as some variation will occur between samples. The present invention also need not be limited by the location of the methylated residues.
[0140] In one embodiment, a screening method can be employed that specifically targets the methylation status of one or more specific cytosine residues or the corresponding cytosine at position n+1 on the opposite DNA strand.
[0141] DNA methylation map
[0142] A "DNA methylation profile" or "methylation profile" can refer to the presence, absence, amount, or level of 5mC at one or more DNA methylation sites. Preferably, a "methylation profile" refers to the presence, absence, amount, or level of 5mC at multiple DNA methylation sites. Thus, the presence, absence, amount, or level of 5mC can be evaluated at each individual DNA methylation site within multiple sites and can contribute to determining the breed composition of a dog. The quality and / or efficacy of the method can therefore be improved by combining values from multiple DNA methylation markers.
[0143] Suitably, the breed profile of the present invention comprises methylation profiles from multiple methylation sites.
[0144] Suitably, the presence or absence of 5mC from at least 5, at least 10, at least 20, at least 50, at least 100, at least 200, at least 500, at least 1000, at least 2000, at least 5000, at least 10000, at least 50000, at least 100000, at least 250000 or at least 500000 DNA methylation sites can be used to determine the breed profile of a dog.
[0145] Suitably, a methylation profile can refer to the presence or absence of 5mC from at least 100, at least 200, at least 500, at least 1000 or at least 2000 DNA methylation sites.
[0146] Suitably, a methylation profile can refer to the presence or absence of 5mC from at least 5, at least 10, at least 20, at least 50, at least 100 or at least 200 methylation sites.
[0147] Suitably, a methylation profile can refer to the presence or absence of 5mC from about 5, about 10, about 20, about 50, about 100, about 200, about 500, about 1000 or about 2000 DNA methylation sites.
[0148] Suitably, a methylation profile can refer to the presence or absence of 5mC from about 100, about 200, about 500, about 1000 or about 2000 DNA methylation sites.
[0149] Suitably, a methylation profile can refer to the presence or absence of 5mC from about 5, about 10, about 20, about 50, about 100 or about 200 DNA methylation sites.
[0150] To generate a breed profile, an initial methylation profile can be processed or simplified to generate a restricted methylation profile, which can then be used to generate a breed profile.
[0151] For example, an initial methylation profile can be processed or simplified by, for example, using DNA regions instead of individual cytosines, by selecting a subset of methylation sites associated with a particular physiological or biochemical pathway, performing a correlation analysis and retaining one or more representative DNA methylation sites per cluster, or performing a differential analysis to pre-select DNA methylation sites or retain DNA methylation sites that vary more between breeds. Processing of DNA methylation sites can include calculating the mutual information for each site, ranking them according to this metric, and then selecting the top n to put into the model.
[0152] For example, a DNA region can be any DNA region as defined herein.
[0153] Suitably, a methylation map can refer to the DNA methylation sites of genes associated with a particular physiological or biochemical pathway. Thus, a methylation map can include methylation sites associated with a particular tissue, organ, or physiological system. Determining the status methylation sites associated with a particular tissue, organ, or physiological system can advantageously allow the method to be used in a manner focused on the pathologies and diseases of that tissue, organ, or physiological system. For example, if a particular breed of dog is known to be associated with muscle or cardiovascular disease, it may be advantageous to determine the methylation sites of that physiological system.
[0154] Suitably, the physiological system can be the immune system, gastrointestinal system, urinary system, muscular system, cardiovascular system, and / or nervous system.
[0155] The methylation map of a particular tissue, organ, or physiological system can be determined using a DNA methylation map that comprises or consists of methylation sites of genes preferentially or specifically expressed in that tissue, organ, or physiological system. The classification of genes by a particular tissue, organ, or physiological system is publicly available, for example, at Gene Ontology (http: / / geneontology.org / ), the KEGG pathway database (https: / / www.genome.jp / kegg / ), or MSIgDB (https: / / www.gsea-msigdb.org / gsea / msigdb / index.jsp).
[0156] In some embodiments, the threshold selects those sites with the highest ranked average methylation values for breed determination. For example, the threshold can be those sites with an average methylation level that is in the top 50%, top 40%, top 30%, top 20%, top 10%, top 5%, top 4%, top 3%, top 2%, or top 1% of the average methylation levels of all sites “i” tested for a predictor (e.g., breed identification).
[0157] Alternatively, the threshold can be those sites where the average methylation level is at a percentile rank greater than or equal to 50, 60, 70, 80, 90, 95, 96, 97, 98, or 99. In other embodiments, the threshold can be based on the absolute value of the average methylation level. For example, the threshold can be those sites where the average methylation level is greater than 99%, greater than 98%, greater than 97%, greater than 96%, greater than 95%, greater than 90%, greater than 80%, greater than 70%, greater than 60%, greater than 50%, greater than 40%, greater than 30%, greater than 20%, greater than 10%, greater than 9%, greater than 8%, greater than 7%, greater than 6%, greater than 5%, greater than 4%, greater than 3%, or greater than 2%. The relative threshold and the absolute threshold can be applied to the average methylation level at each site "i" either alone or in combination. As an illustration of the combined threshold application, a subset of sites can be selected that are in the top 3% of all sites tested by average methylation level and also have an absolute average methylation level greater than 6%. The result of this selection process is a DNA methylation profile of specific hypermethylated sites (e.g., CpG sites) that are considered to be the most informative for breed determination.
[0158] Suitably, the DNA methylation profile can include at least one methylation site as listed in Table 1.
[0159] Suitably, the methylation site can be defined as a methylation marker present in any one or more of SEQ ID NOs: 1 - 200. SEQ ID NOs: 1 - 200 show the sequences on either side of the methylation marker in the "Tasha" dog reference genome (https: / / www.ncbi.nlm.nih.gov / assembly / GCF_000002285.5; Jagannathan et al.; Genes (Bsael); 2021; 12(6); 847). The "CG" methylation marker is the 26th and 27th nucleotides in the sequence (i.e., there are 25 nucleotides before the methylation marker and 25 nucleotides after the methylation marker).
[0160] Suitably, the methylation site can be defined as the insertion position in the column labeled "Site" in Table 1. For example, for site chr10:10975030 - 10975032, the methylation marker is chr10:10975031.
[0161] Suitably, the DNA methylation profile can include at least 1, at least 2, at least 5, at least 10, at least 20, at least 50, at least 100, at least 150, at least 175 methylation sites as listed in Table 1 or preferably each methylation site therein.
[0162] Suitably, the DNA methylation profile can include each of at least 1, at least 2, at least 5, at least 10, at least 20, at least 50, at least 100, at least 125, or the methylation sites listed as "site number" 1 - 150 in Table 1.
[0163] Suitably, the DNA methylation profile can include each of at least 1, at least 2, at least 5, at least 10, at least 20, at least 50, at least 75, or the methylation sites listed as "site number" 1 - 100 in Table 1.
[0164] Suitably, the DNA methylation profile can include each of at least 1, at least 2, at least 5, at least 10, at least 20, at least 30, at least 40, or the methylation sites listed as "site number" 1 - 50 in Table 1.
[0165] Suitably, the DNA methylation profile can include each of at least 1, at least 2, at least 5, at least 10, at least 15, or the methylation sites listed as "site number" 1 - 20 in Table 1.
[0166] Suitably, the DNA methylation profile can include each of at least 1, at least 2, at least 3, at least 5, or the methylation sites listed as "site number" 1 - 10 in Table 1.
[0167] Suitably, the DNA methylation profile can include each of at least 1, at least 2, at least 3, or the methylation sites listed as "site number" 1 - 5 in Table 1.
[0168] Suitably, the DNA methylation profile can include at least one, at least two, at least five, at least ten, at least fifteen of the methylation sites listed in Table 5 or preferably each of them. The methylation profile is suitable for distinguishing, for example, beagle and Labrador retriever profiles.
[0169] determination of DNA methylation sites / methylation maps indicative of breed
[0170] The present invention includes using a DNA methylation profile to determine the contribution of dog breeds to the genome of a test dog. Thus, the present invention includes using a DNA methylation profile to determine the contribution of one or more dog breeds to the genome of a test dog. The contribution of one or more dog breeds to the genome of a test dog may also be referred to herein as a "breed profile".
[0171] For example, the provision of DNA methylation sites or DNA methylation profiles indicative of a breed can be achieved through a training dataset and machine learning methods. Suitably, the machine learning method can be a supervised machine learning method. Suitably, the model can be a multiple regression, support vector machine (SVM), or random forest model.
[0172] For example, the DNA methylation sites or DNA methylation profiles can be trained on a dataset including dogs with known breeds. Suitably, the DNA methylation sites or DNA methylation profiles can be trained on a dataset including dogs with a combination of known breed and known age and / or sex.
[0173] For example, a model of DNA methylation sites or DNA methylation profiles indicative of breed contribution can be provided by using a machine learning framework to train a methylation status dataset at multiple DNA methylation sites on a training dataset of dogs with known breeds and testing on a held-out cohort to validate the accuracy of the model.
[0174] The machine learning framework can include, for example, fitting a penalized regression to a training dataset of dogs with known breeds (and optionally age and / or sex) using the glmnet R package.
[0175] The machine learning framework can include fitting a multinomial model, random forest, SVM (support vector machine), penalized multinomial logistic regression, or other models for predicting multi-class outcomes.
[0176] Suitably, the machine learning framework can include fitting a penalized regression of breed explained by the DNA methylation profile (and optionally age and / or sex), such as elastic net regression.
[0177] Suitably, the machine learning framework can include fitting a penalized regression of breed explained by the DNA methylation profile, age, and sex, such as elastic net regression.
[0178] Suitably, the machine learning framework can be used to determine a model including a set of DNA methylation sites or DNA methylation profiles indicative of a breed.
[0179] The model can include the methylation status at multiple DNA methylation sites; wherein the methylation status at each site is considered by multiplying by a coefficient value in the model.
[0180] The coefficient value of each parameter generally depends on the measurement units of all variables in the model. As those skilled in the art will understand, the value of each coefficient value will thus depend on, for example, the number and nature of the different parameters used in the model and the nature of the training data provided. Thus, conventional statistical methods can be applied to the training dataset to obtain the coefficient values.
[0181] Suitably, sex can be encoded as a numerical value, where 0 represents female and 1 represents male.
[0182] Suitably, the machine learning platform can include one or more deep neural networks. A neural network is a collection of neurons (also called units) connected in an acyclic graph. Neural network models are typically organized into different layers of neurons. For most neural networks, the most common layer type is the fully connected layer, where neurons between two adjacent layers are fully pairwise connected, but neurons within a single layer do not share connections. One of the main characteristics of a deep neural network is that neurons are controlled by non-linear activation functions. This non-linearity combined with the deep architecture enables more complex combinations of input features, ultimately leading to a broader understanding of the relationships between them and thus to a more reliable final output. Deep neural networks have been applied to many types of data ranging from structured data to chemical descriptors or transcriptomic data.
[0183] Suitably, the machine learning platform includes one or more generative adversarial networks. Suitably, the machine learning platform includes an adversarial autoencoder architecture. Suitably, the machine learning platform includes feature importance analysis for ranking them by the importance of DNA methylation sites in breed determination.
[0184] comparison with a reference or control
[0185] The method may also include the step of comparing the DNA methylation differences at one or more sites in the test sample with one or more references or controls. The presence or absence of DNA methylation at one or more sites in the reference or control may be associated with the breed. In some embodiments, the reference value is a value previously obtained for a subject or group of subjects with a known breed. The reference value may be based on the known DNA methylation status at one or more sites from a group of subjects with a known breed, such as the average or median level.
[0186] The reference DNA methylation map may contain DNA methylation maps from at least 2, at least 4, at least 10, at least 20, at least 40, at least 80, at least 100, at least 150, at least 200, at least 300, or at least 400 dog breeds.
[0187] The reference DNA methylation map may contain DNA methylation maps from at least 2, at least 4, at least 10, at least 20, at least 40, at least 80, at least 100, at least 150, or at least 200 dog breeds.
[0188] enrichment and detection method
[0189] Determining a DNA methylation profile can include steps of enriching selected DNA regions in a DNA sample. For example, the method can include steps of enriching DNA regions in a DNA sample that contain DNA methylation sites that comprise the DNA methylation profile.
[0190] Suitable enrichment methods are known in the art and include, for example, amplification- or hybridization-based methods. Amplification enrichment generally refers to, for example, PCR-based enrichment using primers specific for the DNA regions to be enriched. Any suitable form of amplification can be used, such as polymerase chain reaction (PCR), rolling circle amplification (RCA), inverse polymerase chain reaction (iPCR), in situ PCR, strand displacement amplification, or cycling probe technology.
[0191] Hybridization enrichment or capture-based enrichment generally refers to using hybridization probes (or capture probes) that hybridize to the DNA regions to be enriched.
[0192] The hybridization probes can be directly attached to a solid support or can contain a moiety, such as biotin, to allow binding to a solid support (such as streptavidin-coated beads) suitable for capturing the biotin moiety. In either case, DNA containing a sequence complementary to the probe can be captured, thereby allowing separation of DNA containing the DNA regions of interest from DNA that does not contain the DNA regions of interest. Thus, such capture steps allow enrichment of the DNA regions of interest. For example, the DNA region can be a DNA region proximal to a gene promoter.
[0193] The arrays used herein can vary depending on the probe composition and the intended use of the array. For example, the nucleic acids (or CpG sites) detected in the array can be at least 10, 100, 1,000, 10,000, 100,000, 1 million, 10 million, 100 million, or more. Alternatively or additionally, the detected nucleic acids (or CpG sites) can be selected to be no more than 100 million, 10 million, 1 million, 100,000, 10,000, 1,000, 100, or fewer. Similar ranges can be obtained using nucleic acid sequencing methods, such as those known in the art; for example, next-generation or massively parallel sequencing.
[0194] Suitably, the enrichment step can be performed before or after the step of separating or differentiating between methylated and unmethylated DNA.
[0195] As used herein, the term "enrichment" or "enrichment of" "DNA" or "DNA region" means a process in which the (absolute) amount and / or proportion of DNA containing a desired sequence is increased compared to the amount and / or proportion of DNA containing the desired sequence in the starting material. In this regard, enrichment by amplification increases the amount and proportion of the desired sequence. Enrichment by capture-based enrichment increases the proportion of DNA containing the desired sequence.
[0196] After processing the DNA to distinguish methylated and unmethylated sites, the method may further include the step of identifying methylated or unmethylated sites (i.e., in the original sample).
[0197] The identifying step may include any suitable method known in the art, such as array detection or sequencing (e.g., next-generation sequencing).
[0198] The sequencing identification step preferably includes next-generation sequencing (massively parallel or high-throughput sequencing). Next-generation sequencing methods are well known in the art, and in principle, any method may be considered for the present invention. Next-generation sequencing techniques can be carried out according to the manufacturer's instructions (e.g., provided by Roche, Illumina, Applied Biosystems, PacBio, Oxford Nanopore, or MGI).
[0199] In a preferred embodiment, the sample is processed by using an enzymatic reaction to convert DNA methylation, performing whole-genome library preparation, and measuring the methylation profile by sequencing (EM-Seq).
[0200] In a particularly preferred embodiment, the sample is processed by using an enzymatic reaction to convert DNA methylation, performing whole-genome library preparation, hybridizing the whole-genome converted library preparation with capture probes (preferably capture probes capable of capturing DNA regions near gene promoters); and measuring the methylation profile by sequencing (EM-Seq).
[0201] method for selecting a dietary, pharmaceutical or lifestyle regimen for a dog
[0202] In another aspect, the present invention provides a method for selecting a diet, drug, or lifestyle regimen for a subject.
[0203] A diet, medication, or lifestyle regimen can be applied to a dog over any suitable period of time. For example, a diet, medication, or lifestyle regimen can be applied for at least 2 weeks, at least 4 weeks, at least 8 weeks, at least 16 weeks, at least 32 weeks, or at least 64 weeks. A diet, medication, or lifestyle regimen can be applied for at least 3 months, at least 6 months, at least 12 months, at least 24 months, at least 36 months, at least 48 months, or at least 60 months. Suitably, a diet, medication, or lifestyle regimen can be applied for at least 1 year, at least 2 years, at least 3 years, at least 4 years, at least 5 years, at least 6 years, at least 7 years, at least 8 years, at least 9 years, or at least 10 years. Suitably, a diet, medication, or lifestyle regimen can be applied throughout the life of the dog.
[0204] Suitably, the change is a dietary intervention as described herein. The term "dietary intervention" refers to an external factor applied to a subject that causes a change in the subject's diet. More preferably, the dietary intervention includes the administration of at least one dietary product or dietary regimen or nutritional supplement.
[0205] A dietary regimen can be a diet, a dietary regimen, a supplement or a supplement regimen, or a combination of a diet and a supplement, or a combination of a diet and multiple supplements.
[0206] The dietary intervention or dietary product described herein can be any suitable dietary regimen, such as a calorie-restricted diet, a geriatric diet, a low-protein diet, a phosphorus diet, a low-protein diet, a potassium-supplemented diet, a polyunsaturated fatty acid (PUFA)-supplemented diet, an antioxidant-supplemented diet, a vitamin B-supplemented diet, a liquid diet, a selenium-supplemented diet, an ω3-6 ratio diet, or a diet supplemented with carnitine, branched-chain amino acids or derivatives, nucleotides, niacinamide precursors (such as nicotinamide mononucleotide (MNM) or nicotinamide riboside (NR)), or any combination of the above.
[0207] Suitably, the dietary intervention or dietary product can be a calorie-restricted diet, a high-calorie diet, a geriatric diet, or a low-protein diet. Suitably, the dietary intervention or dietary product can be a calorie-restricted diet. Suitably, the dietary intervention or dietary product can be a low-protein diet.
[0208] The dietary intervention can be determined based on the baseline maintenance energy requirement (MER) of the dog or dog breed. Suitably, the MER can be the amount of food that stabilizes the dog's weight (change less than 5% within three weeks).
[0209] For example, some dog breeds are generally considered to benefit from a high-energy / high-protein diet; however, other breeds may have lower energy requirements, and thus the diet can be adjusted appropriately.
[0210] Suitably, the calorically restricted diet can account for about 50%, about 55%, about 60%, about 65%, about 75%, about 80%, about 85% or about 90% of the dog's MER. Suitably, the calorically restricted diet can account for about 60% or about 75% of the dog's MER.
[0211] Suitably, the low - protein diet can contain less than 20% protein (% dry matter). For example, the low - protein diet can contain less than 19% (% dry matter).
[0212] Dietary interventions can include foods, supplements, and / or beverages that contain nutrients and / or bioactive agents that mimic the benefits of caloric restriction (CR) without restricting daily caloric intake. For example, foods, supplements, and / or beverages can contain functional ingredients with similar benefits to CR. Suitably, the foods, supplements, and / or beverages can contain autophagy inducers. Suitably, the foods, supplements, and / or beverages can contain fruits and / or nuts (or their extracts). Suitable examples include, but are not limited to, pomegranates, strawberries, blackberries, camu camu, walnuts, chestnuts, pistachios, pecans. Suitably, the foods, supplements, and / or beverages can contain probiotics, with or without fruit extracts or nut extracts.
[0213] A dog food composition having a ratio of energy from protein to energy from fat below 0.80 can be advantageous for active dogs. High - protein and high - fat food compositions are particularly suitable for active dogs. Generally, the dog food composition for active dogs has about 20% to 30% protein and about 15% to 25% fat. In fact, an energy - dense food composition from fat will provide sufficient energy for the active dog's moderate - to - very - vigorous activities (i.e., brisk walking to running fast) that it spontaneously engages in. Additionally, a ratio of energy from protein to energy from fat has been found to be advantageous in such food compositions for maintaining the lean body mass of active dogs.
[0214] Similarly, a particularly well - adapted robust dog food composition can have a ratio of energy from protein to energy from fat in such food compositions greater than 0.80. More specifically, the protein content is about 20% to 30%, and the fat content is less than about 15%. Due to their low resting metabolic rate, such food compositions are ideally suited for robust dogs. The composition will have the effect of restricting the fat intake of robust dogs and thus restricting their tendency to become overweight.
[0215] A lifestyle change can be any of the changes described herein, such as a change in the exercise regimen.
[0216] Similar to dietary intervention, determining the breed contribution of breeds that typically benefit from a lot of exercise can allow for determining the appropriate exercise regimen to switch a test dog to.
[0217] Desirable activity levels and types may vary according to breed or breed classification. For example, a robust dog will spontaneously engage in mild (e.g., slow walking), moderate (e.g., brisk walking), or occasionally vigorous (e.g., running) types of activity. In contrast, a sporting dog will primarily be spontaneously engaged in moderate, vigorous, or very vigorous (e.g., fast running) activity. Among these different activity levels, dogs can be further classified as robust or sporting.
[0218] A drug regimen may refer to the administration of a treatment modality or protocol. The modality can be one for treating and / or preventing, for example, arthritis, dental disease, endocrine disorders, heart disease, diabetes, liver disease, kidney disease, prostate disorders, cancer, and behavioral or cognitive disorders. Suitably, prophylactic treatment can be administered to dogs identified as being at risk of such disorders due to the breed contribution of the breed associated with the disease. In other embodiments, dogs determined to be at risk of certain conditions due to breed contribution can be monitored more regularly so that diagnosis and treatment can be initiated as early as possible.
[0219] Accordingly, the present invention can advantageously achieve the identification of dogs that are expected to respond particularly well to a given intervention (e.g., a dietary, drug, or lifestyle protocol). Accordingly, the intervention can be applied in a more targeted manner to dogs expected to respond due to their breed profile.
[0220] use of dietary intervention
[0221] In one aspect, the present invention provides a dietary or drug intervention for treating and / or preventing diseases in dogs, wherein the dietary intervention is administered to dogs having a breed profile determined by the present method.
[0222] As described herein, the dietary intervention can be a dietary product or a dietary protocol or a nutritional supplement.
[0223] computer program product
[0224] The present method can be performed using a computer. Accordingly, the present method can be executed on a computer.
[0225] Suitably, the computer can prepare and share a report detailing the results of the present method.
[0226] The method described herein can be implemented as a computer program running on general-purpose hardware such as one or more computer processors. In some embodiments, the functions described herein can be implemented by a device such as a smartphone, a tablet terminal, or a personal computer.
[0227] In one aspect, the present invention provides a computer program product that includes computer-executable instructions for causing a programmable computer to determine the breed contribution of a dog breed to a test dog genome as described herein.
[0228] In another aspect, the present invention provides a computer program product that includes computer-executable instructions for causing a device to determine the contribution of dog breeds to a test dog genome; and, based on the contribution of dog breeds to the test dog genome determined using a DNA methylation map, optionally select a suitable diet, medication, or lifestyle regimen for the dog. Additional parameters or characteristics of the dog may also be provided to the computer program product. As described herein, the additional parameters or characteristics may include the age and gender of the dog.
[0229] In one embodiment, a user inputs the level of one or more DNA methylation markers as defined herein to the device, optionally together with age and gender. The device then processes the information and provides a determination of the breed map of the dog. Alternatively, the device then processes the information and provides a determination of a suitable diet, medication, or lifestyle regimen for the dog based on the breed map.
[0230] The device may generally be a server on a network. However, any device may be used as long as it can process biomarker data and / or additional parameters or characteristics using a processor, central processing unit (CPU), etc. For example, the device may be a smart phone, tablet terminal, or personal computer, and outputs information indicating the determined breed map of the dog or the determination of a suitable lifestyle or diet regimen for the dog based on the breed map.
[0231] Those skilled in the art will understand that they may freely combine all the features of the present invention described herein without departing from the scope of the present invention disclosed herein.
[0232] example
[0233] The present invention will now be further described by way of examples, which are intended to assist those skilled in the art in practicing the present invention and do not limit the scope of the present invention in any way.
[0234] Example 1 - Illustrative method for differentiating dog breeds using DNA methylation
[0235] identification of DNA methylation sites
[0236] Whole blood samples from a canine cohort were analyzed by performing DNA extraction, converting DNA methylation by using enzymatic reactions, performing whole-genome library preparation, hybridizing the whole-genome converted library preparation to capture probes for gene promoters, and measuring the methylation map by sequencing (EM-Seq).
[0237] The capture probes target approximately 40,000 targets (promoter regions - approximately 1 kb upstream to 0.5 kb downstream of the transcription start site). These target regions contain potential methylation sites of interest (single cytosine residues that can be methylated).
[0238] The following bioinformatics steps are performed after sequencing and before further analysis:
[0239] Perform fastq quality checks using fastQC - https: / / www.bioinformatics.babraham.ac.uk / projects / fastqc /
[0240] Perform adapter trimming using trimGalore - https: / / www.bioinformatics.babraham.ac.uk / projects / trim_galore /
[0241] Perform alignment to the dog genome using bwa-meth or Bismark - (https: / / github.com / brentp / bwa-meth or https: / / www.bioinformatics.babraham.ac.uk / projects / bismark / )
[0242] Mark duplicates using Picard - https: / / gatk.broadinstitute.org / hc / en-us / articles / 360037052812-MarkDuplicates-Picard-
[0243] Call methylation using Methyldackel - https: / / github.com / dpryan79 / MethylDackel
[0244] Methylation sites can be further filtered by: (i) removing sites that are (un)methylated in all samples and / or (ii) removing sites that do not have at least 5 counts in at least 90% of the samples.
[0245] distinguish different breeds
[0246] A subset of samples consisting of 66 samples from 2 breeds (Beagles and Labrador Retrievers) was analyzed. 5000 sites were randomly selected, and the methylation beta value (percentage of methylation per site) was measured for each dog.
[0247] Fit the binomial GLMnet (elastic net penalized binomial regression) to 2 / 3 of the data (training set), and estimate the parameters to maximize the AUC (area under the receiver operating characteristic curve). The LASSO component of the model acts as variable selection, and with α = 1 (LASSO regression) and λ = 0, we were able to identify 19 loci that could correctly classify the two varieties with an accuracy of 1 on the test set (see Figure 1 ).
[0248] The 19 loci and their coefficients are shown below (chromosome and position of the locus on the chromosome).
[0249] Table 5
[0250]
[0251]
[0252] Example 2 - Method for filtering sites to generate a restricted DNA methylation map for training breed identification
[0253] After determining the methylation status of the methylation sites in each sample (according to Example 1 - "Identification of DNA methylation sites"), the initial methylation map is filtered / processed to produce a restricted methylation map containing fewer discrete methylation sites. The purpose of this filtering is to provide a restricted methylation map containing, for example, 50,000 to 500,000 methylation sites that can be used to train the method.
[0254] The methods for filtering the initial methylation map include:
[0255] 1) Remove the sites that are (un)methylated in all samples (e.g., methylation% = 0 in all samples or methylation% = 1 in all samples)
[0256] 2) Remove the sites that do not have at least 5 counts in at least 90% of the samples
[0257] 3) Remove chromosome X and / or Y sites
[0258] Other potential filtering steps to reduce the number of discrete methylation sites include:
[0259] 1) Define methylation sites as DNA regions (CpG islands, sites separated by less than, for example, 1000 bp are considered the same)
[0260] 2) Restrict the targets to, for example, inflammatory sites, such as genes associated with inflammation from GO / KEGG pathways, and select the sites on these genes
[0261] 3) Correlation analysis and retention of representatives for each cluster (e.g., weighted gene co-expression network analysis (WGCNA) (Langfelder & Horvath; BMC Bioinformatics; 9(559);
[0262] 2008).
[0263] 4) Differential methylation analysis to pre-select DNA methylation sites (e.g., using logistic regression) or to identify sites that vary more between different breeds (e.g., using analysis of variance). If differential methylation analysis is performed (using the function preprocessQuantile in R), normalization (quantile) is performed before the filtering step.
[0264] 5) Reducing correlations and thus reducing the dimensionality of the feature space using, for example, EBmodule (Zollinger, A., Davison, A.C., & Goldstein, (2018).; Biostatistics, 19(2), 153 - 168. https: / / doi.org / 10.1093 / biostatistics / kxx032 ) ,
[0265] 6) Selection of top - site mutual information processing
[0266] identification of breed contributions based on DNA methylation
[0267] A dataset including information on breed, chronological age, and sex of the dog cohort is split into a training set and a test set (e.g., 2 / 3 of the data is used for training and 1 / 3 of the data is used for testing, thus ensuring a good partition of the metadata; e.g., similar proportions of each breed / each sex in the training set and the test set).
[0268] A prediction model is built using multinomial GLMnet (elastic net penalized binomial regression) to model breed as a function of the methylation profile. The parameters of the model as well as the penalty parameter are estimated using the glmnet package in R. The parameters of the model are tuned by using CV (10 - fold) and variables selected based on their importance in the model contribution.
[0269] The penalty parameter is selected to provide a reasonable number of DNA methylation sites (e.g., at most 100) that have a good model fit for the breed.
[0270] Then the breed identification model based on DNA methylation sites is evaluated on the test set. The model is evaluated based on κ, F1, precision, and recall.
[0271] Example 3 - Further dog breed classification using DNA methylation
[0272] Using a pet cohort consisting of 829 dogs and 20 breeds, we developed a breed classifier using only DNA methylation. Methylation β-values at sites near promoters were obtained by sequencing blood samples. First, the Boostme algorithm was used to estimate low coverage values (<15) and missing values (Zou, L.S., Erdos, M.R., Taylor, D. et al. BMC Genomics 19, 390 (2018).), which is a tree-based machine learning algorithm.
[0273] The X chromosome was removed. The EB module was used to select sites to reduce correlation and thus reduce the dimensionality of the feature space (Zollinger, A., Davison, A.C., & Goldstein, (2018).; Biostatistics, 19(2), 153 - 168. https: / / doi.org / 10.1093 / biostatistics / kxx032). By grouping sites for each target (1500bp around TSS) and each chromosome, we divided the dataset into chunks of approximately 5000 sites. Then, we calculated the correlation matrix on which we applied the EB module as described in Zollinger et al. above. Each module was then represented by the medoid site. Some sites did not belong to any module and were called dispersed sites. They were defined as described in Zollinger et al. The number of sites was reduced from 1.4mio to 471K.
[0274] The training cohort was divided into two subsets: a training set and a test set. The training set consisted of dogs that potentially shared family relationships, while the test set contained dogs unrelated to those in the training set. To ensure adequate representation of each breed in the training set, we excluded breeds with fewer than four samples. A total of 16 breeds out of the initial 20 breeds were included. Additionally, for breeds with fewer than ten dogs, we manually ensured that only one dog was included in the test set, while the remaining dogs were assigned to the training set.
[0275] To identify relevant methylation sites for breed prediction, we used mutual information (MI) to estimate the predictive power of sites. Sites were ranked according to the calculated values, and the top N sites were selected. N was determined empirically by fitting multiple classifiers with an N range. Since it performed best according to the average F1 across breeds, we selected the top 200 sites as the input to the classifier (see Table 1). Additionally, the F1 score across breeds was used as a performance metric to select the best model, namely a support vector machine (SVM) with a linear kernel.
[0276] To alleviate the severe imbalance in breed distribution, different weights were assigned to misclassifications of different breeds. The weights were defined as the reciprocals of the breed fractions in the training set. The model achieved an average F1 of 0.89. Note that the cost hyperparameter of the SVM was tuned to 0.1 using k-fold cross-validation (k = 10) (see Table 2 and Figure 4 ). Additionally, to obtain probabilities from the SVM outputs for each breed, Platt scaling was employed. Finally, to evaluate the generality of the model, we used a validation cohort consisting only of Labrador retrievers. All Labrador retrievers were correctly classified.
[0277] Using only the top 5, top 10, top 20, top 25, top 50, top 100, and top 150 loci from the complete list of loci shown in Table 1 also yielded further classification models; and each locus showed predictions for breed classification (see Table 4). These classifiers were generated by selecting the top n loci based on the mutual information from the most informative (top of the list) to the least informative. The average F1 scores for the classifiers using the top 5, top 10, top 20, top 25, top 50, top 100, and top 150 loci are shown in Table 4.
[0278] description of SVM algorithm output :
[0279] To generalize to multi-class classification, the SVM fit k(k - 1) / 2 binary classifiers, and a voting scheme was used to determine the class. This means that for 16 breeds, 120 binary classifiers were fit. Additionally, a total of 340 support vectors (SVs) were used to distinguish all classes. For each SV, coefficients were associated with each locus (340 × 200). Then, for each class except one, the coefficients of the SVs were given (340 × 15), and the decision thresholds for each trained binary classification problem were returned (120).
[0280] description of coefficients provided by MLR
[0281] Multinomial logistic regression (MLR) and SVM with a linear kernel achieved very similar performance. Therefore, the coefficients of the multinomial logistic regression are provided. MLR used the same 200 methylation loci as input. Similarly, differential weights were used to alleviate class imbalance. The weight of each data point was defined as the reciprocal of the breed fraction in the training data set. The hyperparameters of MLR were tuned for the training set using cross-validation: decay = 0.1. To concisely provide the coefficients of the model, we fit MLR with the same parameters but only using the top 10 loci. The coefficients for each breed for the 10 methylation loci are given in addition to the intercept (see Table 3).
[0282] To calculate the probability of belonging to a breed class,
[0283] should be calculated where β 0 is the intercept, β i is the coefficient, and x i is the methylation value at each locus i.
[0284] All publications mentioned in the above specification are incorporated herein by reference. Various modifications and variations of the methods, compositions, and uses disclosed herein will be apparent to those skilled in the art without departing from the scope and spirit of the invention. While the invention has been disclosed in connection with specific preferred embodiments, it should be understood that the invention as claimed should not be unduly limited to such specific embodiments. Indeed, various modifications to the modes for practicing the invention which are apparent to those skilled in the art are intended to fall within the scope of the following claims.
[0285] Table 1
[0286]
[0287]
[0288]
[0289]
[0290]
[0291]
[0292]
[0293]
[0294]
[0295] Table 2 - Summary metrics of SVM per-breed performance with a linear kernel classifier on an uncorrelated test set
[0296]
[0297] Table 3 - Coefficients of multinomial logistic regression for breed classifier using 10 methylation sites as features
[0298]
[0299]
[0300] Table 4 - F1 training cross-validation for different numbers of sites
[0301] number of sites F1_train_cross_validation 5 0.342923932 10 0.558679764 20 0.618192168 25 0.678316505 50 0.735451126 100 0.754503632 150 0.760859564 200 0.768644068
Claims
1. A method for determining the contribution of dog breeds to the genome of a test dog, the method comprising: a) providing a DNA methylation map from a sample obtained from the test dog; and b) determining the contribution of dog breeds to the genome of the test dog by comparing at least a portion of the DNA methylation map of the test dog with reference DNA methylation maps from different dog breeds.
2. A method for selecting a diet, drug, or lifestyle regimen for a test dog, the method comprising: a) providing a DNA methylation map from a sample obtained from the test dog; b) determining the contribution of dog breeds to the genome of the test dog by comparing at least a portion of the DNA methylation map of the test dog with reference DNA methylation maps from different dog breeds; and c) selecting a suitable diet, drug, or lifestyle regimen for the test dog based on the contribution of dog breeds to the genome of the test dog determined in step b).
3. A method for preventing or reducing the risk of a test dog developing a disease; the method comprising: a) providing a DNA methylation map from a sample obtained from the test dog; b) determining the contribution of dog breeds to the genome of the test dog by comparing at least a portion of the DNA methylation map of the test dog with reference DNA methylation maps from different dog breeds; wherein at least one dog breed contributing to the genome of the test dog is associated with a predisposition to developing the disease; and c) selecting a diet, drug, or lifestyle regimen for the test dog based on the contribution of the at least one dog breed to the genome of the test dog determined in step b); wherein the diet, drug, or lifestyle regimen prevents or reduces the risk of the test dog developing the disease.
4. The method according to any one of claims 1 to 3, wherein step a) comprises determining a DNA methylation map from a sample obtained from the test dog.
5. The method according to claim 4, wherein the DNA methylation is determined according to a method comprising one or more of the following aspects: (i) one or more of the following steps: (a) treating the sample DNA with APOBEC or using bisulfite conversion to deaminate unmethylated cytosine; (b) enrichment based on capture; and / or (c) high-throughput sequencing or arrays; or (ii) de novo sequencing.
6. The method according to any one of the preceding claims, wherein the contribution of at least two dog breeds to the genome of the test dog is determined.
7. The method according to any one of the preceding claims, wherein machine learning is used to compare the DNA methylation map of the test dog with reference DNA methylation maps from different dog breeds.
8. The method according to any one of the preceding claims, wherein the reference DNA methylation map comprises DNA methylation maps from at least 2, at least 4, at least 10, at least 20, at least 40, or at least 80 dog breeds.
9. The method according to any one of the preceding claims, wherein the DNA methylation map comprises at least one population-specific DNA methylation marker.
10. The method according to any one of the preceding claims, wherein the contribution of a dog breed to the genome of the test dog is used to distinguish two or more genetically related dog breeds.
11. The method according to any one of claims 1 to 9, wherein the contribution of a dog breed to the genome of the test dog is used to classify the test dog as: (i) an American Kennel Club registered breed; (ii) a genetic breed clade; (iii) breed size; and / or (iv) a robust or sporting breed.
12. The method according to claim 11, wherein the genetic breed clade is selected from wild dogs, Basenjis, Asian Spitz, Asian toy dogs, Nordic Spitz, Schnauzers, Miniature Spitz, Toy Spitz, Hungarian dogs, Poodles, American Terriers, American toy dogs, Dobermans, Terriers, New World dogs, Mediterranean dogs, scent hounds, retrievers, pointing setters, Continental herding dogs, British pastoral dogs, Zwergwachtel, Alpine dogs, and European Mastiffs.
13. The method according to any one of claims 2 to 12, wherein the diet, drug, or lifestyle regimen is a dietary intervention.
14. The method according to claim 13, wherein the dietary intervention is a calorie-restricted diet, a geriatric diet, or a low-protein diet.
15. The method according to any one of the preceding claims, wherein the sample is a blood sample.
16. The method according to any one of the preceding claims, wherein the DNA methylation profile comprises at least 10 methylation sites.
17. The method according to any one of the preceding claims, wherein the DNA methylation profile comprises at least one methylation site listed in Table 1.
18. The method according to claim 17, wherein the DNA methylation profile comprises at least 2, at least 5, at least 10, at least 20, at least 50, at least 100, at least 150, at least 175 methylation sites listed in Table 1 or each of the methylation sites therein.
19. The method according to any one of claims 3 to 18, wherein the disease is associated with the incidence or predicted incidence of: (i) a tissue; (ii) an organ; or (iii) a physiological system, such as the immune system, gastrointestinal system, urinary system, muscular system, cardiovascular system, and / or nervous system.
20. The method according to claim 19, the method further comprising applying to the dog a diet, drug, or lifestyle regimen suitable for improving the incidence or predicted incidence of the following identified in claim 19: a tissue; an organ; or a physiological system.
21. A computer-readable medium comprising instructions that, when executed, cause one or more processors to perform the method according to any one of claims 1 to 3 or 6 to 20.
22. A computer system for determining the contribution of a dog breed to the genome of a test dog, the computer system being programmed to compare at least a portion of a DNA methylation profile obtained from the test dog with reference DNA methylation profiles from different dog breeds.
23. A computer system for selecting a diet, medication, or lifestyle program for a test dog, the computer system being programmed to perform the following steps: a) determining the contribution of dog breed to the test dog's genome by comparing at least a portion of the DNA methylation profile obtained from the test dog with reference DNA methylation profiles from different dog breeds; and b) selecting a suitable diet, medication, or lifestyle program for the test dog based on the contribution of dog breed to the test dog's genome determined in step a).
24. A computer program product comprising computer-executable instructions for causing a programmable computer to determine the contribution of dog breed to the genome of a test dog by comparing at least a portion of the DNA methylation profile obtained from the test dog with reference DNA methylation profiles from different dog breeds.
25. A computer program product comprising computer-executable instructions for causing a programmable computer to select a diet, medication, or lifestyle program for a test dog by the following steps: a) determining the contribution of dog breed to the test dog's genome by comparing at least a portion of the DNA methylation profile obtained from the test dog with reference DNA methylation profiles from different dog breeds; and b) selecting a suitable diet, medication, or lifestyle program for the test dog based on the contribution of dog breed to the test dog's genome determined in step a).
Citation Information
Patent Citations
Method for improving dog food
EP1983842A1