System, method, and apparatus for predicting genetic ancestry

The system addresses the challenge of accurately predicting genetic ancestry and traits in mixed-breed animals by using a reference panel and machine learning to classify DNA segments, achieving efficient and precise breed classification and trait prediction.

JP2025176713APending Publication Date: 2025-12-04MARS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025136841
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-07-07
Filing Date
2025-08-20
Publication Date
2025-12-04

AI Technical Summary

Technical Problem

Current methods for animal genetic mapping struggle to accurately and efficiently evaluate mixed-breed genomic samples, leading to inaccurate results and wasted computational power due to the complexity of pet genomes resulting from interbreeding and increasing population genomic datasets.

Method used

A system utilizing computational and statistical methods to predict genetic ancestry and physical traits from raw DNA sequences, employing a large reference panel of animals with known genetic ancestry, and machine learning algorithms to assign genetic ancestry to small genome segments, followed by aggregation for breed classification and trait prediction.

Benefits of technology

Accurately predicts genetic ancestry and physical traits in companion animals, such as adult weight, with improved accuracy and efficiency, even for mixed-breed samples, by leveraging a reference panel and machine learning algorithms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025176713000001_ABST
    Figure 2025176713000001_ABST
Patent Text Reader

Abstract

To accurately and efficiently identify ancestral contributions.SOLUTION: A method comprises accessing a sample of genetic material containing raw genotypes, generating phased haplotypes based on the raw genotypes, generating local assignments for genetic populations for the phased haplotypes by machine learning algorithms based on comparisons between the phased haplotypes and a reference panel comprising multiple reference haplotypes associated with multiple reference populations, and sending instructions to a user device for presenting an output associated with a first animal to a user, wherein the output is generated based on the local assignments for the genetic populations.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Related Applications

[0001] This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 219,349, filed July 7, 2021, the entire contents of which are incorporated herein by reference and claim priority. [Technical Field]

[0002] SUMMARY OF THE INVENTION The embodiments described in this disclosure relate to systems and methods for predicting the genetic ancestry of animals based on input DNA sequences. [Background technology]

[0003] Current methods for animal genetic mapping suffer from an inability to accurately and efficiently evaluate mixed-breed genomic samples. Existing methods are unable to efficiently process large volumes of query sequences, nor can they accurately provide the origin of a given sample. As a result, current genomic analysis of pets (and other livestock) does not achieve satisfactory levels of accuracy for both single-origin and mixed-breed samples, resulting in wasted computational power and inaccurate results. The complexities associated with pet genomes are further compounded by the potential for complex genetic profiles resulting from interbreeding. Given the increasing size and complexity of population genomic datasets, as well as the increasing complexity of downstream genetic profiles, there is a need for systems and methods that can efficiently predict the local and global genetic ancestry of a given genomic sample with great accuracy and without significant computational overhead.

[0004] Information regarding genetic risk factors for disease development and clinical and veterinary recommendations can aid in optimal management, monitoring, and treatment of animals. Identifying ancestry contributions can be useful in determining these risk factors. Therefore, there is a need for methods and systems for accurately and efficiently identifying ancestry contributions. Summary of the Invention

[0005] The objects and advantages of the disclosed subject matter will be set forth in and become apparent from the following description, as well as learned by practice of the disclosed subject matter. Additional advantages of the disclosed subject matter will be realized and attained by the methods and systems particularly pointed out in the specification and claims, as well as from the appended drawings.

[0006] To achieve these and other advantages, and in accordance with the purposes of the disclosed subject matter as embodied and broadly described, the disclosed subject matter provides systems, methods, and apparatus that can be used to collect, receive, and / or analyze data. For example, certain non-limiting embodiments can be used to predict the genetic ancestry of animals.

[0007] In certain non-limiting embodiments, the present disclosure describes a system of computational and statistical methods for generating predictions of genetic ancestry and physical traits in companion animals from only their raw DNA sequences. The prediction system utilizes information from a large reference panel of animals with known genetic ancestry and traits to accurately assign genetic ancestry to small segments within the genome. The resulting segment classifications are then aggregated for each animal and used to predict whether an individual animal belongs to one of hundreds of predefined pure-breed or mixed-breed classes. Furthermore, the aggregated genetic ancestry classifications can be used to accurately predict physical traits, such as an animal's adult weight.

[0008] In certain non-limiting embodiments, one or more computing systems can access a sample of genetic material associated with the first animal. The sample of genetic material can include one or more raw genotypes. The computing system can then generate one or more phased haplotypes based on the one or more raw genotypes. The computing system can then generate, for the one or more phased haplotypes, one or more local assignments to one or more genetic populations based on a comparison between the one or more phased haplotypes and a reference panel including a plurality of reference haplotypes associated with a plurality of reference populations using one or more machine learning algorithms. The computing system can further transmit instructions to a user device to present an output associated with the first animal to a user. In some embodiments, the output can be generated based on the one or more local assignments to one or more genetic populations.

[0009] In certain non-limiting embodiments, one or more computer-readable non-transitory storage media comprising software are operable, when executed, to access a sample of genetic material associated with a first animal. The sample of genetic material may include one or more raw genotypes. The computer-readable non-transitory storage medium is further operable, when executed, to generate one or more phased haplotypes based on the one or more raw genotypes. The computer-readable non-transitory storage medium is further operable, when executed, to generate, for the one or more phased haplotypes, one or more local assignments to one or more genetic populations based on a comparison between the one or more phased haplotypes and a reference panel including a plurality of reference haplotypes associated with a plurality of reference populations, by one or more machine learning algorithms. The computer-readable non-transitory storage medium is further operable, when executed, to transmit instructions to a user device for presenting an output associated with the first animal to a user. In some embodiments, the output may be generated based on the one or more local assignments to one or more genetic populations.

[0010] In certain non-limiting embodiments, the system may include one or more processors and a non-transitory memory coupled to the processor, the non-transitory memory including instructions executable by the processor. The processor, when executing the instructions, is operable to access a sample of genetic material associated with the first animal. The sample of genetic material may include one or more raw genotypes. The processor, when executing the instructions, is further operable to generate one or more phased haplotypes based on the one or more raw genotypes. The processor, when executing the instructions, is further operable to generate, for the one or more phased haplotypes, one or more local assignments to one or more genetic populations based on a comparison between the one or more phased haplotypes and a reference panel including a plurality of reference haplotypes associated with a plurality of reference populations by one or more machine learning algorithms. The processor, when executing the instructions, is further operable to transmit instructions to a user device for presenting an output related to the first animal to a user. In some embodiments, the output may be generated based on the one or more local assignments to one or more genetic populations.

[0011] Furthermore, the disclosed embodiments of the method, computer-readable non-transitory storage medium, and system may have additional non-limiting features, as described below.

[0012] In certain non-limiting embodiments, the computing system can further generate one or more consensus genotypes based on one or more raw genotypes.The computing system can then generate one or more phased haplotypes based on one or more raw genotypes and one or more consensus genotypes.In some embodiments, generating can include phasing one or more raw genotypes and one or more consensus genotypes into maternal chromosomes and paternal chromosomes.In one feature, the one or more machine learning algorithms can include a positional Burrows-Wheeler transformation algorithm.

[0013] In certain non-limiting embodiments, the computing system can remove one or more errors associated with one or more local assignments to one or more genetic populations based on one or more machine learning algorithms. In one feature, the one or more machine learning algorithms can include hidden Markov models.

[0014] In certain non-limiting embodiments, the computing system can further determine one or more source populations associated with the first animal based on one or more local assignments to the one or more genetic populations. In some embodiments, determining the one or more source populations can include aggregating the one or more local assignments to the one or more genetic populations across both maternal and paternal chromosomes, calculating proportions associated with the one or more source populations based on the aggregating, and determining the one or more source populations based on the calculated proportions.

[0015] In certain non-limiting embodiments, the computing system may further partition the one or more local assignments for the one or more genetic populations into one or more maternally or paternally genetic groups. The partitioning may be based on one or more clustering algorithms.

[0016] In certain non-limiting embodiments, the computing system can further determine one or more genetic traits associated with the first animal based on one or more local assignments to one or more genetic populations and one or more source populations. In some embodiments, determining the one or more genetic traits can be further based on one or more of influential variant genotypes, genome-wide statistics, genomic principal component analysis (PCA) predictions, DNA methylation profiles, or polygenic risk scores. In some embodiments, the one or more genetic traits include one or more of adult weight range, genetic disease risk prediction or predisposition, nutritional recommendations, behavioral and temperament class prediction, lifespan estimate, all-cause mortality prediction in years, predicted pharmacological response, or recovery time range in hours for an injectable anesthetic.

[0017] In certain non-limiting embodiments, the computing system may further update one or more machine learning algorithms based on one or more new reference samples added to the reference panel. In some embodiments, the updating may include applying cross-validation across all samples in the reference panel, identifying one or more outliers based on results associated with the cross-validation by the detection algorithm, and removing the identified outliers from the reference panel. In some embodiments, the updating may further include generating one or more labels for one or more unlabeled samples in the reference panel, where the updating is based on the generated labels. The updating may be repeated until a predetermined accuracy level of the one or more machine learning algorithms is reached.

[0018] In certain non-limiting embodiments, the present disclosure provides kits for determining the local and global ancestry of an animal using any of the methods disclosed herein. In certain embodiments, the kit includes a sample collection device. In certain embodiments, the sample collection device includes a carrier and a reservoir. In certain embodiments, the carrier includes an absorbent member and the reservoir includes a shield. In certain embodiments, the kit further includes instructions for using the sample collection device and / or for collecting a sample.

[0019] In certain non-limiting embodiments, one or more computing systems can access a sample of genetic material associated with the first animal. The sample of genetic material can include one or more raw genotypes. The computing system can then generate one or more phased haplotypes based on the one or more raw genotypes. The computing system can then generate, for the one or more phased haplotypes, one or more local assignments to one or more genetic populations based on a comparison between the one or more phased haplotypes and a reference panel including a plurality of reference haplotypes associated with a plurality of reference populations using one or more machine learning algorithms. The computing system can then determine one or more source populations associated with the first animal based on the one or more local assignments to the one or more genetic populations. The computing system can then partition the one or more local assignments to the one or more genetic populations into one or more maternal genetic groups or paternal genetic groups. The computing system can then determine one or more genetic traits associated with the first animal based on the one or more local assignments to the one or more genetic populations and the one or more source populations. The computing system can further transmit instructions to a user device to present an output associated with the first animal to a user. In some embodiments, the output may be generated based on one or more genetic populations, one or more source populations, results associated with partitioning, or one or more local assignments to one or more genetic traits.

[0020] In certain non-limiting embodiments, one or more computer-readable non-transitory storage media comprising software are operable, when executed, to access a sample of genetic material associated with a first animal. The sample of genetic material may include one or more raw genotypes. The computer-readable non-transitory storage medium is further operable, when executed, to generate one or more phased haplotypes based on the one or more raw genotypes. The computer-readable non-transitory storage medium is further operable, when executed, to generate, for the one or more phased haplotypes, one or more local assignments to one or more genetic populations based on a comparison between the one or more phased haplotypes and a reference panel including a plurality of reference haplotypes associated with a plurality of reference populations, by one or more machine learning algorithms. The computer-readable non-transitory storage medium is further operable, when executed, to determine one or more source populations associated with the first animal based on the one or more local assignments to the one or more genetic populations. The computer-readable non-transitory storage medium comprising the software is further operable, when executed, to partition the one or more local assignments for the one or more genetic populations into one or more maternally or paternally genetic groups. The computer-readable non-transitory storage medium comprising the software is further operable, when executed, to determine one or more genetic traits associated with the first animal based on the one or more local assignments for the one or more genetic populations and the one or more source populations. The computer-readable non-transitory storage medium comprising the software is further operable, when executed, to transmit instructions to a user device for presenting an output associated with the first animal to a user. In some embodiments, the output may be generated based on the one or more local assignments for the one or more genetic populations, the one or more source populations, results associated with the partitioning, or the one or more genetic traits.

[0021] In certain non-limiting embodiments, the system may include one or more processors and a non-transitory memory coupled to the processor, the non-transitory memory including instructions executable by the processor. The processor, when executing the instructions, is operable to access a sample of genetic material associated with a first animal. The sample of genetic material may include one or more raw genotypes. The processor, when executing the instructions, is further operable to generate one or more phased haplotypes based on the one or more raw genotypes. The processor, when executing the instructions, is further operable to generate, for the one or more phased haplotypes, one or more local assignments to one or more genetic populations based on a comparison between the one or more phased haplotypes and a reference panel including a plurality of reference haplotypes associated with a plurality of reference populations by one or more machine learning algorithms. The processor, when executing the instructions, is further operable to determine one or more source populations associated with the first animal based on the one or more local assignments to the one or more genetic populations. The processor, when executing the instructions, is further operable to partition the one or more local assignments to the one or more genetic populations into one or more maternally or paternally genetic groups. The processor, when executing the instructions, is further operable to determine one or more genetic traits associated with the first animal based on one or more local assignments to the one or more genetic populations and one or more source populations. The processor, when executing the instructions, is further operable to send instructions to a user device for presenting to a user an output associated with the first animal. In some embodiments, the output may be generated based on one or more genetic populations, one or more source populations, results associated with the partitioning, or one or more local assignments to the one or more genetic traits.

[0022] Furthermore, the disclosed embodiments of the method, computer-readable non-transitory storage medium, and system may have additional non-limiting features, as described below.

[0023] In certain non-limiting embodiments, the computing system can further generate one or more consensus genotypes based on one or more raw genotypes.The computing system can then generate one or more phased haplotypes based on one or more raw genotypes and one or more consensus genotypes.In some embodiments, generating can include phasing one or more raw genotypes and one or more consensus genotypes into maternal chromosomes and paternal chromosomes.In one feature, the one or more machine learning algorithms can include a positional Burrows-Wheeler transformation algorithm.

[0024] In certain non-limiting embodiments, the computing system can remove one or more errors associated with one or more local assignments to one or more genetic populations based on one or more machine learning algorithms. In one feature, the one or more machine learning algorithms can include hidden Markov models.

[0025] In certain non-limiting embodiments, determining the one or more source populations may include aggregating one or more local assignments to one or more genetic populations across both maternal and paternal chromosomes, calculating proportions associated with the one or more source populations based on the aggregation, and determining the one or more source populations based on the calculated proportions. In some embodiments, the partitioning may be based on one or more clustering algorithms.

[0026] In certain non-limiting embodiments, determining the one or more genetic traits can be further based on one or more of influential variant genotypes, genome-wide statistics, genomic principal component analysis (PCA) predictions, DNA methylation profiles, or polygenic risk scores. In some embodiments, the one or more genetic traits include one or more of adult weight range, genetic disease risk prediction or predisposition, nutritional recommendations, behavioral and temperament class prediction, lifespan estimate, all-cause mortality prediction in years, predicted pharmacological response, or recovery time range in hours for injectable anesthetics.

[0027] In certain non-limiting embodiments, the computing system may further update one or more machine learning algorithms based on one or more new reference samples added to the reference panel. In some embodiments, the updating may include applying cross-validation across all samples in the reference panel, identifying one or more outliers based on results associated with the cross-validation by the detection algorithm, and removing the identified outliers from the reference panel. In some embodiments, the updating may further include generating one or more labels for one or more unlabeled samples in the reference panel, where the updating is based on the generated labels. The updating may be repeated until a predetermined accuracy level of the one or more machine learning algorithms is reached.

[0028] It is understood that both the foregoing general description and the following detailed description are exemplary and intended to provide further explanation of the disclosed subject matter as claimed. These and other features, aspects, and advantages of the present disclosure will become apparent from a reading of the following detailed description in conjunction with the accompanying drawings, which are briefly described below. The present disclosure includes any combination of two, three, four, or more of the above-described embodiments, as well as any combination of two, three, four, or more features or elements described in the present disclosure, whether or not such features or elements are explicitly combined in the description of a specific embodiment herein. The present disclosure, in any of its various aspects and embodiments, is intended to be read as a whole such that any separable features or elements of the disclosed invention should be considered as intended to be combinable unless the context clearly dictates otherwise. [Brief explanation of the drawings]

[0029] The foregoing and other objects, features, and advantages of the present disclosure will become apparent from the following description of the embodiments, as illustrated in the accompanying drawings, in which reference characters refer to the same parts throughout the various views. The drawings are not necessarily to scale, emphasis instead being placed upon illustrating the principles of the present disclosure. [Figure 1] FIG. 1 illustrates an exemplary workflow of a system in accordance with the subject matter of this disclosure. [Figure 2] FIG. 1 illustrates an exemplary workflow of a local ancestry classifier. [Figure 3-1] FIG. 10 shows several models showing the results of varying the length of a given subregion from 6 centimorgans to 48 centimorgans. [Figure 3-2] This is a continuation of Figure 3-1. [Figure 4] FIG. 10 illustrates an example of a perimeter match length in accordance with the subject matter of this disclosure. [Figure 5]FIG. 1 shows an exemplary comparison of the "chromosome painting" model (A) with the PBWT-based models (B and C) described in this disclosure. [Figure 6] FIG. 1 illustrates an exemplary smoothing process. [Figure 7A] FIG. 1 illustrates a confusion matrix associated with multiple animal species and / or animal breeds. [Figure 7B-1] FIG. 1 shows animal breed on the y-axis. [Figure 7B-2] This is a continuation of Figure 7B-1. [Figure 7B-3] This is a continuation of Figure 7B-2. [Figure 7C-1] FIG. 1 shows animal breeds on the x-axis. [Figure 7C-2] This is a continuation of Figure 7C-1. [Figure 7C-3] This is a continuation of Figure 7C-2. [Figure 8] FIG. 1 illustrates an exemplary sorting of chromosome pairs into maternal and paternal copies using k-means clustering. [Figure 9] FIG. 1 shows examples of principal components from global ancestry proportions of a chromosome set. [Figure 10] FIG. 10 illustrates exemplary results of an accuracy benchmark of the disclosed system against the prior art classifier RFMix. [Figure 11] FIG. 10 illustrates an exemplary receiver operating characteristic (ROC) curve for the global ancestry classifier. [Figure 12] FIG. 1 shows an exemplary regression of predicted adult weight versus true observed adult weight. [Figure 13] FIG. 1 shows an exemplary iterative refinement of a local ancestry reference panel using isolation forest techniques for anomaly detection. [Figure 14] FIG. 1 illustrates an exemplary method for ancestry prediction. DETAILED DESCRIPTION OF THE INVENTION

[0030] Mapping local and global genetic ancestry traits within pet populations is an ongoing aspect of population genetics research. In this context, the term "ancestor" refers to the source population from which a segment of DNA originates. Furthermore, the modifier "local ancestor" refers to the source population of the small fragments of DNA that make up a chromosome. Alternatively, the modifier "global ancestor" refers to one or more source populations that contribute to the entirety of all chromosomes. Local ancestry can assign a single source population to a localized segment of DNA, while global ancestry can represent the aggregation of local ancestry across all DNA segments in a genome. Global ancestry can be reported as the proportion of an organism's genome that originates from a specific source population. Importantly, both local and global ancestry classifications may rely on reference panels that typify DNA segments from all source populations. As sample sizes for population genomic data increase, the computational complexity of assigning new sequences to predefined population groups can become significant. In particular, for many pets, such as cats and dogs, and other livestock, genomic sequences may become intermingled as subsequent generations interbreed, creating more complex genomes.

[0031] There remains a need in the art for scalable systems and methods that can accurately and efficiently predict the genetic ancestry of a query sample, whether the sample is of single origin or mixed species. The subject matter of the present disclosure addresses this need through the following methods and systems.

[0032] Certain systems and methods according to the present embodiments use computational and statistical methods to generate predictions of genetic ancestry and physical traits in companion animals from only their raw DNA sequences. The present systems and methods can import batches of DNA sequences from a sample set with unknown genetic ancestry and then efficiently match this "query" set to a curated reference database of DNA sequences with known genetic ancestry and traits. In certain embodiments, information from a large reference panel of animals with known genetic ancestry and traits can be used to accurately assign genetic ancestry to small segments within the genome. The resulting segment classifications can then be aggregated for each animal and used to predict whether an individual animal belongs to one of hundreds of predefined pure-breed or mixed-breed classes. Furthermore, the aggregated genetic ancestry classifications can be used to accurately predict physical traits, such as an animal's adult weight. Details of the present embodiments are provided below. For clarity, and not by way of limitation, the detailed description of the present disclosure is divided into the following subsections: 1. Definition; 2. System overview; 3. Sequencing, kits, and treatment methods; and 4. Working Example

[0033] 1. definition The terms used herein generally have their ordinary meaning in the art, within the context of this disclosure and in the specific context in which each term is used. Certain terms are discussed below or elsewhere herein to provide additional guidance in describing the compositions and methods of the present disclosure and how to make and use them.

[0034] Unless otherwise defined, all technical and scientific terms used herein have the meanings that are commonly understood by those skilled in the art to which the present invention belongs.The following references provide those skilled in the art with the general definitions of many terms used in this disclosure: King, Mulligan, and Stansfield. A Dictionary of Genetics, Oxford University Press, 2013; Glossary of Bioinformatics Terms, Current Protocols in Bioinformatics, 35, 1934-3396, 2011; and Whole-Transcriptome Amplification of Single Cells for Next-Generation Sequencing, Current Protocols in Molecular Biology, 111, 1934-3639, 2015.As used herein, the following terms have the following meanings unless otherwise specified.

[0035] As used herein, the words "a" or "an," when used in conjunction with the term "comprising" in the claims and / or specification, can mean "one," but are also consistent with the meanings of "one or more," "at least one," and / or "one or more." Furthermore, the terms "having," "including," "containing," and "comprising" are interchangeable, and one of ordinary skill in the art will recognize that these terms are open-ended.

[0036] The term "about" or "approximately" means within an acceptable error range for a particular value as determined by one skilled in the art, which error range depends in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, "about" can mean within 3 standard deviations or more than 3 standard deviations, according to the practice in the art. Alternatively, "about" can mean a range of up to 20%, preferably up to 10%, more preferably up to 5%, and even more preferably up to 1% of a given value. Alternatively, particularly with respect to a system or process, the term can mean within a single order of magnitude, preferably within 5 times, and more preferably within 2 times of a value.

[0037] As used herein, "comprises," "comprising," or any other variation thereof, is intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements does not include only those elements, but may include other elements not expressly listed or elements inherent to such process, method, article, or apparatus.

[0038] As used herein, the term "local ancestry" refers to the ancestral origin of different chromosomal segments within an individual's genome. In certain embodiments, local ancestry is the call of a particular segment of a chromosome in an animal, such as a dog breed. In certain exemplary embodiments, local ancestry refers to an individual's genetic ancestors at a particular chromosomal location, and an individual may have 0, 1, or 2 copies of alleles from each ancestral population.

[0039] As used herein, the term "global ancestry" refers to the proportion of ancestry averaged across the genome of a subject. In certain embodiments, global ancestry is the proportion of calls across the entire genome of an animal, e.g., a dog breed.

[0040] As used herein, the term "haplotype" refers to a set of linked genes or other genetic markers that are inherited together as a unit. During meiosis, there is little or no recombination with corresponding regions on homologous chromosomes, so alleles are rarely shuffled between homologous regions. In certain embodiments, a stretch of DNA containing a haplotype is referred to as a "haplotype block." For example, and not by way of limitation, certain genes of the major histocompatibility complex in canines are closely linked at the DLA locus on chromosome 12 and behave as a haplotype, with alleles on the maternal and paternal chromosomes generally being transmitted to offspring in the same combination. In certain embodiments, the term "haplotype" refers to a single chromosome or a haploloid set of chromosomes. As used herein, the terms "haplotype estimation" or "haplotype phasing" refer to the process of statistically inferring haplotypes from genotype data.

[0041] As used herein, the term "centimorgan" or "cM" refers to a unit of measurement for the frequency of genetic recombination. One centimorgan corresponds to a 1% probability that a recombination event during meiosis (occurring during the formation of egg and sperm cells) will result in the separation of two markers on a chromosome from one another. On average, one centimorgan corresponds to approximately one million base pairs of the human genome.

[0042] As used herein, the term "phasing" refers to the process of assigning alleles (e.g., A, C, T, and G) to paternal and maternal chromosomes. The term typically applies to the type of DNA undergoing recombination (e.g., autosomal DNA or the X chromosome). In certain embodiments, phasing can help determine whether a match is on the paternal or maternal side, on both sides, or on neither side. In certain embodiments, phasing can also aid in the process of chromosome mapping (e.g., assigning a segment to a particular ancestor). Traditionally, the use of phased data reduces the number of false positive matches.

[0043] As used herein, the term "genotype" refers to the genetic makeup of an organism. For example, a genotype describes the complete set of genes in an organism, such as a dog. In certain embodiments, the term "genotype" refers to the alleles or variant types of genes carried by an organism. A particular genotype is described as homozygous if it is characterized by two identical alleles, and as heterozygous if the two alleles are different. As used herein, the process of determining a genotype is referred to as "genotyping." As used herein, "genotyping call" and variations thereof refer to estimating a genotype value from raw data or processed data.

[0044] The terms "nucleic acid molecule," "nucleotide sequence," and "polynucleotide," as used herein, refer to a single- or double-stranded covalently linked sequence of nucleotides in which the 3' and 5' ends of each nucleotide are linked by a phosphodiester bond. Nucleic acid molecules can contain deoxyribonucleotide or ribonucleotide bases and can be produced synthetically in vitro or isolated from natural sources.

[0045] The terms "polypeptide," "peptide," "amino acid sequence," and "protein," used interchangeably herein, refer to a molecule formed from the linkage of at least two amino acids. The linkage between one amino acid residue and the next is an amide bond, sometimes referred to as a peptide bond. Polypeptides can be obtained by any suitable method known in the art, including isolation from natural sources, expression in a recombinant expression system, chemical synthesis, or enzymatic synthesis. The terms apply to amino acid polymers in which one or more amino acid residues are artificial chemical mimetics of the corresponding naturally occurring amino acid, as well as to naturally occurring and non-naturally occurring amino acid polymers.

[0046] The terms "pet food" or "pet food composition" or "pet food product" or "finished pet food product" refer to a product or composition that is intended for consumption by a companion animal, such as a cat, dog, guinea pig, rabbit, bird, or horse, and that provides certain nutritional benefits. For example, without limitation, the companion animal may be a "domestic" dog, such as Canis lupus familiaris. In certain embodiments, the companion animal may be a "domestic" cat, such as Felis domesticus. "Pet food" or "pet food composition" or "pet food product" or "finished pet food product" includes any food, feed, snack, food supplement, liquid, beverage, treat, toy (chewable and / or consumable toy), meal replacer, or meal replacement.

[0047] For purposes of this disclosure, the terms "user," "subscriber," "consumer," or "customer" should be understood to refer to a user of the applications described herein and / or a consumer of data provided by a data provider. By way of example and not limitation, the terms "user" or "subscriber" may refer to a person receiving data over the Internet in a browser session or data provided by a service provider, or may refer to an automated software application that receives the data and stores or processes the data.

[0048] 2. System Overview Figure 1 shows an exemplary workflow 100 of a system according to the subject matter of this disclosure. The system can ingest a batch of DNA sequences from a sample set with unknown genetic ancestry and then efficiently match this "query" set to a curated reference database of DNA sequences with known genetic ancestry and traits. In certain embodiments, the DNA sequences encompassed by this disclosure include gene sequences and / or genetic markers. For example, but not limited to, genetic markers include single nucleotide polymorphisms (SNPs), short tandem repeats (STRs), base insertions and deletions (indels), and copy number variations (CNVs).

[0049] In certain embodiments, the system may include multiple individual component subsystems. These subsystems may include one or more of a local ancestry classifier, a global ancestry classifier, a genealogical ancestry predictor, a trait suite (e.g., physical, behavioral, and metabolic) predictor, or an automated system for improving the accuracy of the classifiers. Each of these subsystems may have its own functionality. When combined, these subsystems enable the overall system to generate predictions of genetic ancestry and physical traits in companion animals from only their raw DNA sequences.

[0050] In certain non-limiting embodiments, the local ancestry classifier can be associated with raw input genotypes 102, consensus genotypes 104, phased haplotypes 106, training panels 108a-108c, PBWT matching 110, raw local ancestry 112, HMM 114, and smoothed local ancestry 116. The local ancestry classifier can receive the raw input genotypes 102 and generate a consensus genotype 104 accordingly. In some embodiments, the raw input genotypes 102 can serve as a query genotype, and the consensus genotype 104 can serve as a reference genotype. The consensus genotypes 104 can then be processed into phased haplotypes 106 that can distinguish between maternal and paternal chromosomes. A matching process 110 (e.g., a positional Burrows-Wheeler transformation) can then partition the phased haplotypes 106 into multiple windows, which can be compared to the reference or training panel 108. The density of matches between the phased haplotypes 106 and the reference or training panel 108 can be calculated to generate a raw local ancestry 112, which can be defined as the reference population with the highest relative density of matches (or other criteria). The raw local ancestry 112 can then be used as input to a hidden Markov model (HMM) 114, which can remove or replace certain errors in the raw local ancestry 112 to generate a smoothed local ancestry 116. This smoothed local ancestry 116 can be output to an end user to indicate the relative origin of one or more chromosomes. By way of example, and not limitation, the output may include a detailed description of an animal's chromosomes, indicating exactly where the animal obtained each piece of DNA (e.g., Great Pyrenees, German Shepherd Dog, Beauceron, White Swiss Shepherd, Maremma Sheepdog, Chow Chow, Siberian Husky, Parson Russell Terrier, Border Terrier, and Hovawart).

[0051] In certain non-limiting embodiments, a global ancestry classifier can then use the smoothed local ancestry 116 to generate a global ancestry 118. This global ancestry 118 can be output to an end user, providing the relative contributions of different source populations in the animal's genome. By way of example and not limitation, the output may include the different breeds detected in the animal's DNA.

[0052] In certain non-limiting embodiments, a genealogical ancestry predictor can predict genealogical ancestry using smoothed local ancestry 116. By way of example, and not limitation, predicting genealogical ancestry can enable workflow 100 to provide pedigree 120. In some embodiments, k-means 122 (discussed in more detail below) can be applied to smoothed local ancestry 116 to generate pedigree 120 (or other genealogical information) for an animal.

[0053] In certain non-limiting embodiments, the trait suite predictor can use the global ancestry 118 to generate trait predictions or estimates for an animal based on specific genetic probabilities. The global ancestry 118 can be used as input for a meta-classifier 124, which can provide entire sample subpopulation labels. This meta-classifier 124 can identify one or more predicted classes / groups and confidence levels 126 for the input global ancestry 118. These classes / groups with confidence levels 126 can be further used (alone or in combination with additional genotypes) in various downstream applications 128, which can include predicting subjects' life spans, genetic predispositions, and other traits specific to their genomes. In some embodiments, the downstream applications 128 can take in additional genotypes 130 as inputs. These downstream applications 128 can also be used to improve the consumer experience 132, enabling the creation of applications or other services that provide predictions to end users.

[0054] In certain non-limiting embodiments, the automated system for improving accuracy can be associated with new reference samples 134, isolation forest outlier detection 136, and cross-validation 138. The automated system can evaluate new reference samples 134 added to the reference / training panel 108. This evaluation can first include performing cross-validation 138 across all samples of the candidate reference panel. The results of the cross-validation can then be used as input to a detection algorithm, such as the isolation forest outlier detection algorithm 136. By way of example and not limitation, based on the new reference samples 134a, cross-validation 138a, isolation forest outlier detection 136a, and training panel 108b, the automated system can improve the accuracy of the PBWT matching 110. By way of another example and not limitation, based on the new reference samples 134b, cross-validation 138b, isolation forest outlier detection 136b, and training panel 108d, the automated system can improve the accuracy of the meta-classifier 124.

[0055] Local Ancestor Classifier Conventional local ancestry classifiers can have significant limitations. By way of example and not limitation, they cannot be easily scaled to accommodate large reference panels, and they can require significant computational resources to generate predictions. By controlling these, local ancestry classifiers as disclosed herein can achieve improved accuracy over conventional methods and can easily accommodate much larger reference panels. In certain embodiments, local ancestry classifiers as disclosed herein can use a positional Burrows-Wheeler transform (PBWT) algorithm in conjunction with a mathematical approximation of a standard local ancestry model. In certain non-limiting embodiments, the standard local ancestry model can include "chromosome painting." As used herein, chromosome painting describes various techniques for characterizing chromosomal rearrangements, including, but not limited to, the employment of fluorescently labeled DNA probes. Furthermore, local ancestry classifiers as disclosed herein can leverage reference panels to learn common misclassifications and smooth the resulting assignments to improve overall accuracy. In some embodiments, the local ancestry classifier can reference a list or matrix containing common misclassification results to smooth the resulting classifications. Smoothing can remove commonly mistaken sequences and replace them with more likely substitutions. The degree to which local ancestry assignments are smoothed can be adjusted to accommodate both single-origin chromosomes and highly mixed chromosomes. By way of example and not limitation, smoothing can be adjusted to accommodate single-origin chromosomes, or alternatively, highly mixed chromosomes containing DNA from multiple origins.

[0056] FIG. 2 illustrates an exemplary workflow 200 of a local ancestry classifier. In a specific, non-limiting embodiment, a cloud data monitoring service 205 can periodically probe a cloud storage environment 210 for the presence of new query DNA sequences. The cloud storage environment 210 can be a scalable storage infrastructure. By way of example and not limitation, the query DNA sequence can include multiple genotype data organized into multiple haplotypes 215. The cloud data monitoring service 205 can retrieve the sequences upon detection of a positive signal and deposit the query batch into a high-performance computing environment. A computational configuration service 220 can then characterize the batch of ingested DNA sequences and configure a custom bioinformatic workflow. The computational configuration service 220 can compare the query haplotypes 215 to a reference panel of haplotypes 225 to generate a local ancestry profile 230. An emission / transition 235 can be generated based on the reference panel of haplotypes 225. In some embodiments, the local ancestry profile 230 and outputs / transitions 235 may then be smoothed based on HMM smoothing 240 to remove common errors. In some embodiments, the reference panel 225 may be used as part of a purebred training set 245, upon which the purebred classifier 250 may be trained. Once smoothing is complete, the smoothed local ancestry profile may be processed by the purebred classifier 250 to generate purebred meta-classifier labels. Finally, the labeled local ancestry profile may be output in a report 255. By way of example and not limitation, the report 255 may be in JavaScript Object Notation (JSON) format.

[0057] In certain non-limiting embodiments, the local ancestry classifier can predict a local ancestry label for a subject. The local ancestry classifier can select two samples: a first sample corresponding to a query nucleotide sequence and a second sample corresponding to a reference nucleotide sequence. The query nucleotide sequence can include one or more unknown ancestry labels, which can be selected from an ordered set of subpopulation labels. The reference nucleotide sequence can include one or more known genetic subpopulations corresponding to known nucleotide sequences. Each of the first and second samples can be further partitioned into subregions, also known as windows, for use in comparing the two samples. At least one subregion of the first sample can then be compared with at least one subregion of the second sample to identify nucleotide matches between the two samples. In this manner, the degree of similarity between the samples can be determined by counting the number of nucleotide matches between the first and second samples. Genetic subpopulations corresponding to and including one or more of the nucleotide matches can be selected from a known list of genetic subpopulation information. Based on the selected genetic subpopulations, local ancestry labels can be applied and optionally applied to one or more query nucleotide sequences.

[0058] In any of the exemplary methods, the identification of a nucleotide match may involve a variety of different factors and is not intended to limit the nucleotide match to an exact match between all elements of the two subregions being compared. For example, in certain non-limiting embodiments, one or more nucleotide matches may include at least one nucleotide sequence in a first sample that is identical to at least one nucleotide sequence in a second sample. In alternative embodiments, a nucleotide match may be determined when at least one nucleotide sequence in a first sample is identical to at least one nucleotide sequence in a second sample by a predetermined percentage. Furthermore, in non-limiting embodiments, each of the one or more nucleotide matches may include multiple nucleotides. In such embodiments, each of the multiple nucleotides may be identical between the first and second samples, or alternatively, each may meet a predetermined percentage of identity between the first and second samples. In further non-limiting embodiments, the nucleotide match may include adjacent nucleotides.

[0059] The number of nucleotide matches can be determined according to various methods.For example, this method can use the length of the number of adjacent nucleotides in the first sample (or a subregion of the first sample) that matches the number of adjacent nucleotides in the second sample (or a subregion of the second sample) to calculate the number of nucleotide matches.According to this non-limiting embodiment, the length of the number of adjacent nucleotides in at least one subregion of the first sample and / or the length of the number of adjacent nucleotides in at least one subregion of the second sample can be approximate or exact.

[0060] In non-limiting embodiments, at least one genetic subpopulation can be determined by examining nucleotide matches. For example, the genetic subpopulation can be selected based on a maximum number of nucleotide matches, a specified number of nucleotide matches, and / or a preselected number of nucleotide matches. In further embodiments, the subpopulation can be selected if the number of nucleotide matches exceeds a certain value or falls within a certain range. The local ancestry classifier can further identify and / or remove certain outliers from the population.

[0061] In certain non-limiting embodiments, the local ancestry classifier can assume the existence of a curated reference panel containing several haplotype sequences, each of which can be labeled according to its membership in several population groups. The goal of the local ancestry classifier can include classifying any query haplotype into one of the populations of the reference panel. The local ancestry classifier can begin by phasing both the query genotype and the reference genotype into maternal and paternal chromosomes. In certain non-limiting embodiments, phasing the maternal genome and the paternal genome can be performed using a phasing reference panel. Alternatively, for example, the phasing reference panel can be obtained by first performing cohort phasing using the local ancestry reference panel. This set of phased haplotypes can then be used as a panel for reference-based phasing. The phased data can then be partitioned into 5-centimorgan (cM) windows. By way of example and not limitation, a 5 cM window size can be selected to balance linkage disequilibrium in canids with restoring sufficient haplotype diversity to be informative. While a 5 cM window can be used in certain embodiments, other window lengths are also contemplated. Figure 3 illustrates several models showing the results of varying the length of a given subregion from approximately 6 centimorgans to approximately 48 centimorgans. As shown in Figure 3, windows of lengths 6 cM, 12 cM, 18 cM, 20 cM, 24 cM, 30 cM, 36 cM, and 48 cM can also be used. Additionally, windows of lengths less than 5 cM, for example, can be used to view a subregion of a target chromosome in more detail. For purposes of the presently disclosed subject matter, the window length can correspond to the length of any subregion in the first or second sample.

[0062] In certain non-limiting embodiments, population assignment for each window can be achieved by recovering all pairwise set-maximal matches between query haplotypes and reference haplotypes using a positional Burrows-Wheeler transformation algorithm. The density of set-maximal matches between a given query haplotype and all reference haplotypes can be calculated, and the reference population with the highest relative density can be selected as the "raw" assignment. A hidden Markov model (HMM) can then be run on the raw calls on windows grouped by chromosome to "smooth" the local ancestry assignments. Finally, global ancestry fractions can be aggregated from the local assignments and used in a global ancestry classifier to create population assignments for the entire diploid genome.

[0063] In certain non-limiting embodiments, the local ancestry classifier can recover matching short DNA segments. A set of algorithms specific to PBWT can efficiently recover matches between pairs of haplotype sequences in a collection. Some PBWT-based algorithms can iterate through a collection of haplotype sequences and recover a set of maximum matches, which can be defined as the set of other sequences that show the greatest local uninterrupted matches to the current sequence. In the disclosure presented herein, the collection of sequences can include both query haplotypes and reference haplotypes.

[0064] As mentioned above, PBWT can include a collection of related algorithms for fast sorting of binary matrices. This algorithm can operate on a binary matrix with N rows representing haplotypes and M columns representing biallelic DNA sites. The rows can be sorted starting from the leftmost column. As the algorithm progresses, two vectors can be updated for each site: the first is the rank of the haplotype (the position prefix array), and the second is a measure of the number of differences from the previous haplotype (the divergence array). The elements of the divergence array can be added between ordered haplotypes, resulting in the Hamming distance between the haplotypes. Matching haplotype sequences can be searched for by tracking which sequences are adjacent in the position prefix array and whether the divergence array is zero. Matching can be broken if the haplotypes are no longer adjacent or if the corresponding element in the divergence array is no longer zero. In certain non-limiting embodiments, the set maximum match may be the locally maximum match to a given sequence (over the interval ending at the current position) and may include one or more adjacent haplotypes with the longest match over that interval.

[0065] 4 is a diagram illustrating examples of neighboring match lengths according to the subject matter of this disclosure. A query sequence 410 is depicted at the bottom. Matches to a reference panel sequence 420 are shown at the top. Each of matches 420a-420c may correspond to its corresponding reference population label. The sum of the neighboring match lengths for each reference population can be considered proportional to the likelihood of the query sequence coming from that reference population.

[0066] Figure 5 shows an exemplary comparison between the "chromosome painting" model (A) and the PBWT-based model (B and C) described in this disclosure. In panel A, a query sequence is compared to all reference panel sequences, and the most likely path through the reference panel sequences can be used to label or "paint" the query chromosome. PBWT-based methods can sort sequences stepwise so that locally matching sequences are adjacent in the list. For example, in panel B, the PBWT algorithm sorts up to position 6, and in panel C, the algorithm sorts up to the final position. The query chromosome can be "painted" by evaluating which sequences are adjacent in the PBWT data structure. This simplification allows PBWT-based methods to easily scale to very large reference panel sizes. As shown in Figure 4, PBWT can be selected to sort at specific positions within the selected query genotype sequence, such as position 6 within the query genotype, and alternatively, at the end position of the query genotype. Alternative methods of achieving population assignment, such as chromosome painting, can also be used. The density of the set maximum matches between a given query haplotype and all reference haplotypes can be calculated, and the reference subpopulation with the highest relative density can be selected as the raw assignment. This density of matches can also correspond to the number of nucleotide matches between the selected samples.

[0067] In certain non-limiting embodiments, the local ancestry classifier can assume that a curated population reference panel containing a total of N phased haplotypes is available. Each reference haplotype can be assigned a single label from k, an ordered set of K source population labels. Furthermore,

[0068]

number

[0069] There can be a corresponding ordered set n of subpopulation sample sizes such that

[0070] In addition to the haplotypes of the reference panel, the local ancestry classifier can consider a single query haplotype that is assigned a label from k. After running the above PBWT-based algorithm, all set-maximal matches between the query haplotype and the haplotypes of the reference panel can be recovered. Each set-maximal match can be labeled by the reference population label of the matching haplotype (see Figure 4). To filter out small haplotype segments with high homozygosity across source populations (and therefore unlikely to arise from a recent common ancestor), set-maximal matches longer than 0.5 cM can be considered in the analysis. Each of the recovered set-maximal match lengths can be labeled with the label of the corresponding source population from k. The marginal sum of match lengths with label i is l. i It is written as follows.

[0071] In certain non-limiting embodiments, the local ancestry classifier can determine the probability that a query haplotype is sampled from source population i. These probabilities can form an ordered set of source population sampling probabilities p, which can be expressed as a probability mass function P(Q=i|p)=p i where Q is the source population label of the query haplotype. The query haplotype has a criterion max i p i According to the label Q=k i can be assigned.

[0072] In certain non-limiting embodiments, the above marginal match lengths can be formulated as statistics for estimating p. The marginal match lengths can be drawn from a sample space encompassing the total length of haplotypes in all source populations, and L i =L int ×n i where L int is the recombination distance of the genomic interval under consideration. The first-order statistic can be defined as: g i =li / L i .

[0073] This statistic allows us to estimate the proportion of all source population haplotypes that match the query, and in effect, g i It is expected to be <<1.

[0074] The parameters of a categorical distribution can be approximated by standardizing the statistics:

[0075]

number

[0076] In contrast to other local ancestry classification algorithms, the local ancestry classifier disclosed herein can be a simple moment-based estimator, which minimizes the reliance on complex underlying population genetic models that are often required when a Dirichlet distribution is used as a conjugate prior for the categorical distribution. Bayesian inference based on Dirichlet priors can require assumptions inherent in simulating high-probability population structure models whose parameters are uncharacterized, at the expense of increased scalability and computational time, and often with unknown improvements in accuracy. Rather than leveraging traditional simulation-based Bayesian inference, the local ancestry classifier disclosed herein focuses on improving assignment accuracy through the application of machine learning models trained on reference panel samples.

[0077] In certain non-limiting embodiments, the method of using the surrounding reference source population match length can also increase robustness against haplotype phasing errors present in the reference panel haplotypes. The rationale is that a long match interrupted by a phase switch can still be recovered by the algorithm as a separate match and contribute equally to the surrounding sum of match lengths. A scenario in which this may not be the case is when a phase switch interrupts a long match and one (or both) of the resulting match segments is too short (i.e., <0.5 cM) to be recorded by the method. In some embodiments, the estimate of the surrounding population match length can be reduced by up to 1 cM. One approach to address this case is to reduce the match length threshold.

[0078] In certain non-limiting embodiments, the local ancestry predictions can be smoothed. FIG. 6 illustrates an exemplary smoothing process. As shown in FIG. 6, the raw assignment data can be further smoothed to remove common errors or improve accuracy. For example, in certain non-limiting embodiments, a machine learning model can be run on the raw calls across all windows in the raw assignment dataset to smooth the local ancestry estimates. Various machine learning models can be used, including, but not limited to, hidden Markov models. Smoothing can also obtain global subpopulation proportions (i.e., global ancestry) from the local estimates, and these global estimates can be used in conjunction with multiple meta-classifiers to create overall sample subpopulation labels, and the local ancestry calls can be partitioned as either maternally or paternally inherited.

[0079] Hidden Markov models (HMMs) are widely used in population genomics to model the linearity of features along chromosomes. In the disclosure presented herein, an ordered sequence of local ancestry labels can be treated as an observation sequence in the HMM. In this framework, each reference population can be viewed as a latent variable, or "hidden state," of the query haplotype. The goal of employing an HMM in this manner is to eliminate spurious transitions between local ancestry assignments and correct common erroneous assignments. In certain non-limiting embodiments, the local ancestry classifier can prioritize HMM parameters that encourage linkage mixing, such as adding pseudocounts to transition probabilities to prevent zero probabilities, so that the local ancestry classifier can perform well on highly mixed samples. In some embodiments, the HMM can be trained on a reference panel where local ancestry assignments are assumed to be a reliable source.

[0080] HMM output probabilities can be estimated by a leave-one-out procedure applied to all reference panel haplotypes. Each reference haplotype is used as a query sequence and assigned a subpopulation label from the estimated parameters of the categorical distribution p. These estimates are aggregated across all N query haplotypes to form a K × K matrix binned by the haplotype's "true" subpopulation label. The elements of the resulting population confusion matrix can be used as the HMM's output probabilities. A transition matrix can also be learned from the estimated sequence of population labels for the reference panel haplotypes. Finally, a vector of probabilities starting from a given hidden state can be estimated from the global ancestry estimates obtained from PBWT-based calling. A separate HMM can be run for each chromosome using a backward-forward algorithm, and the Viterbi algorithm can be used to decipher the most likely path through the hidden states.

[0081] In certain non-limiting embodiments, the smoothing method may include multiple steps. By way of example and not limitation, the method may identify a first portion of at least one of two or more subregions of a first sample of genetic material. The method may then identify a second portion of at least one of the two or more subregions of the first sample of genetic material. The method may then replace the second portion with the first portion. The smoothing method may be performed when the second portion is commonly confused with the first portion, e.g., when identifying the second portion of the subregion as a particular breed is a common error and the first portion represents the correct breed. The smoothing method may help improve the accuracy of the overall workflow, resulting in more accurate breed identification. In some embodiments, the identification of commonly confused breeds and / or species may be facilitated by a confusion matrix. Figure 7A shows a confusion matrix associated with multiple animal species and / or animal breeds. Figure 7B shows animal breeds on the y-axis. Figure 7C shows animal species on the x-axis. 7A-7C show that the confusion matrix is ​​useful for variety and / or species identification.

[0082] The techniques utilized by the local ancestry classifier as described above allow it to improve accuracy over previous work and to be easily adapted to much larger reference panels than previous work. These advantages are described in the "Examples" section later in this disclosure, specifically in the "Benchmarking the Accuracy of Ancestry Classifiers" and "Benchmarking the Scalability of Classification Systems" sections.

[0083] Global Ancestry Classifier In certain non-limiting embodiments, the global ancestry classifier can predict the source population of an entire organism by considering the totality of local ancestry classifications. This can include organisms derived from a single source population, but can also include commonly encountered combinations (or mixtures) of source populations. For companion animals, a good example is predicting the "goldendoodle," a cross between a golden retriever and a poodle. Furthermore, the global ancestry classifier can weight specific DNA variants known to affect specific traits to refine predictions of source populations that otherwise cannot be distinguished at the genome-wide level. For example, variants in the fibroblast growth factor gene FGF5 are known to affect coat length in domestic dogs. For canine breeds with different coat lengths that otherwise cannot be distinguished at the genome-wide level, weighting FGF5 gene variants can accurately distinguish between long-haired and short-haired breeds.

[0084] In certain non-limiting embodiments, local ancestry assignments from the Viterbi pathway can be aggregated across both the maternal and paternal chromosome sets and used to calculate a global ancestry proportion for a given diploid sample. The global ancestry proportion can be used as a feature to predict the population label for the entire diploid sample using a random forest classifier. The prediction can be associated with a confidence score recalibrated by one or more algorithms. In some embodiments, the random forest classifier can be trained based on the leave-one-out results of the reference panel described above (after being run through an HMM).

[0085] The techniques utilized by the global ancestry classifier as described above enable it to have advantageous features and performance over previous work. Such advantages are described in the "Examples" section later in this disclosure, specifically in the "Benchmarking the Accuracy of Ancestor Classifiers" and "Benchmarking the Scalability of Classification Systems" and "Evaluating the Accuracy of Global Ancestor Classifiers" sections.

[0086] Genealogical Ancestry Predictor Because the local ancestry method predicts the source population for a single phased copy of a chromosome, the local ancestry prediction can be further partitioned into maternal and paternal inheritance. In certain non-limiting embodiments, the genealogical ancestry predictor for partitioning parent chromosomes can assume that the proportion of local ancestry that constitutes a single haploid copy of the genome is similar between different chromosomes. The genealogical ancestry predictor can then find the most likely partitioning of maternal and paternal chromosomes by minimizing the Euclidean distance between the complete complement of haploid chromosomes. Figure 8 shows an example of sorting chromosome pairs into maternal and paternal copies using k-means. Figure 8 shows an example of partitioning 38 canid chromosome pairs into maternal and paternal copies using k-means based on chromosome-specific local ancestry proportions.

[0087] In certain non-limiting embodiments, the genealogical ancestry predictor can use an eigendecomposition of a matrix of global ancestry proportions per chromosome. The rows of the matrix can be haploid chromosomes, and the columns can be source population labels. The resulting two components can be subjected to k-means with k=2 (maternal and paternal groupings are optional). The goal is to group chromosomes with similar ancestry configurations and use this criterion to partition each chromosome into a maternal set and a paternal set. This procedure can serve as a basis for reconstructing the pedigree of an individual companion animal. Figure 9 shows an example of principal components from the global ancestry proportions of a chromosome set. Figure 9 shows a plot of the first two principal components from the global ancestry proportions for each of the 38 pairs of canid chromosomes. The parent chromosome sets are arbitrarily labeled as maternally or paternally inherited.

[0088] The techniques utilized by the genealogical ancestry predictor as described above also allow the genealogical ancestry predictor to partition local ancestry predictions into maternal and paternal inheritance, which may be a unique feature.

[0089] Trait Suite Predictors In certain non-limiting embodiments, the output from the local ancestry classifier and / or the global ancestry classifier can be used as input for a trait suite predictor, which includes a series of trait prediction modules. These prediction modules can incorporate a variety of auxiliary inputs, including influential variant genotypes, genome-wide statistics (e.g., average homozygosity), genomic principal component analysis (PCA) predictions, DNA methylation profiles, and / or polygenic risk scores. By way of example and not limitation, the trait suite predictor can predict one or more of expected healthy adult weight, along with a genetic disease range prediction, risk prediction or predisposition, nutritional recommendations based on ancestry classification, behavioral and temperament class prediction, lifespan and all-cause mortality prediction in years, or predicted pharmacological response, recovery time range in hours for injectable anesthetics. In some embodiments, the nutritional recommendations can include recommendations for one or more pet food products, including commercial and / or personalized pet food products.

[0090] In certain non-limiting embodiments, the trait suite predictor can use the local ancestry classification to determine specific predictions or estimates of various characteristics of a subject, for example, by using the local ancestry labels to identify known genetic sequences that contribute to a particular trait. By way of example and not limitation, the trait suite predictor can use the local ancestry labels to identify one or more ranges of adult weight for a subject, identify one or more predispositions for one or more genetic diseases, provide one or more nutritional product recommendations and / or one or more nutritional regimen recommendations, estimate the subject's lifespan and / or life span, and / or predict one or more pharmacological responses for a subject.

[0091] As described above, by utilizing inputs from the local and global ancestry classifiers, as well as a variety of auxiliary inputs, the trait suite predictor is able to predict many more traits than previous work. Such advantages are illustrated in the "Examples" section later in this disclosure, specifically the "Trait Prediction Performance" section.

[0092] Automated system for improved accuracy The accuracy of the generated classifier may depend on the individual samples in the source population reference panel. By way of example and not limitation, if the source population reference panel contains inaccurate population labels, the accuracy of the entire workflow 100 of the system may be reduced. In certain non-limiting embodiments, the automated system for improving accuracy can evaluate new samples to be added to the reference panel. This evaluation may first include performing leave-one-out cross-validation across all samples in the candidate reference panel. The results of the cross-validation can then be used as input to a detection algorithm, such as an isolation forest outlier detection algorithm. The algorithm can identify certain samples as outliers by comparing them with the population labels and remove them from the reference panel. The automated system can be run iteratively as needed until a predetermined level of accuracy is reached, for example, until the precision and recall of the panel no longer significantly improves. In an alternative non-limiting embodiment, a machine learning algorithm can be used to generate labels for unlabeled samples. By way of example and not limitation, a semi-supervised machine learning label propagation algorithm can be used to automate the assignment of estimated labels to unlabeled samples.

[0093] As described above, an automated system for improving accuracy can utilize cross-validation of a reference panel with a leave-one-out procedure. In this scenario, each sample included in the reference panel can be iteratively removed from the panel and then run as a query sequence. The remaining query sequence can then be assigned a local ancestry label. This procedure can be repeated for all samples included in the reference panel. The samples can then be grouped by their estimated source population label. Using the local ancestry call as a feature, an isolation forest technique can be performed on each set of samples grouped by source population label. The number of tree partitions induced to isolate a given sample can be used as a decision function for identifying anomalies. If a forest of random trees generates a path length shorter than expected for a particular sample, the sample can be labeled an anomaly and removed from the reference panel. This procedure can be repeated until the improvement in weighted recall and precision falls below a pre-specified threshold.

[0094] The techniques utilized by the automated system for precision enhancement as described above enable the automated system to further improve the performance of the systems and subsystems as disclosed herein, and such advantages are described later in this disclosure in the "Examples" section, specifically in the "Performance of Automated Precision Enhancement" section.

[0095] 3. Sequencing, kits, and methods of treatment The present disclosure includes a method for sequencing the genome of an animal or pet. The term "animal" or "pet" as used in accordance with the present disclosure refers to domestic animals, including but not limited to domestic dogs, domestic cats, horses, cows, ferrets, rabbits, pigs, rats, mice, gerbils, hamsters, goats, etc. Domestic dogs and domestic cats are particularly non-limiting examples of pets. The term "animal" or "pet" as used in accordance with the present disclosure can also refer to wild animals, including but not limited to bison, elk, deer, venison, ducks, poultry, fish, etc.

[0096] As used herein, the terms "dog" or "canid" are used interchangeably and refer to any member of the Canidae family, including, but not limited to, Canis lupus, Canis familiaris, Canis latrans, Canis dingo, Lycaon pictus, Chrysocyon brachyurus, Atelocynus microtis, Cuon alpinus, Speothos venaticus, Nyctereutes procyonoides, Vulpes vulpes, and Alopex lagopus. In certain embodiments, the dog or canid is Canis familiaris.

[0097] In certain embodiments, the method includes obtaining a sample from an animal. In certain embodiments, the sample can be a bodily fluid obtained from the animal. In certain non-limiting embodiments, the sample can be saliva, sputum, blood, perspiration (e.g., sweat), pus, tears, mucosal excretion, vomit, urine, feces, semen, vaginal fluid, or other types of bodily fluid. In certain embodiments, the sample can be a non-fluid sample. In certain embodiments, the sample can be a cell-free sample. For example, without limitation, the sample is a cell-free nucleic acid sample. In certain embodiments, the sample can include cell-free deoxyribonucleic acid (DNA), cell-free ribonucleic acid (RNA), and / or cell-free protein. In certain embodiments, the sample can include one or more cells.

[0098] In certain embodiments, the sample may be a solid sample or a tissue sample. In certain embodiments, the sample may be a skin sample. In certain embodiments, the sample may be a cheek swab sample or a swab sample from a different body part. In certain embodiments, the sample may be a homogenous sample or a heterogenous sample. In certain embodiments, the sample may be a tumor sample. In certain embodiments, the sample may include one or more different biological samples. For example, but not by way of limitation, the sample may include saliva and skin tissue. In certain embodiments, the sample may be a plasma or serum sample.

[0099] In certain embodiments, the sample is a sputum sample. In certain embodiments, the sample is a saliva sample. In certain embodiments, the sample is a buccal swab sample.

[0100] In certain embodiments, a sample may be collected from an animal and preserved and / or stabilized until further processing and / or analysis. For example, without limitation, the sample may be preserved and / or stabilized for such use by incubation with a reagent. In certain embodiments, the reagent for preserving and / or stabilizing the sample may be any substance that acts on the collected sample to achieve a desired effect. In certain embodiments, the reagent may be in any suitable form, such as a fluid (e.g., a liquid, a gas, a solution, etc.) or a non-fluid (e.g., a solid powder, etc.). In certain embodiments, the reagent may preserve deoxyribonucleic acid (DNA), ribonucleic acid (RNA), protein, or other components of protein in the sample. In certain embodiments, the reagent may prevent changes to the cellular epigenome of one or more cells. In certain embodiments, the reagent may enable extraction of a desired molecule (e.g., a nucleic acid molecule) from cells from the collected sample. In certain embodiments, the reagent may be configured to process the collected sample and / or one or more components thereof in a separate process.

[0101] In another non-limiting example, collected samples may be stored intact until further processing and / or analysis. In certain embodiments, collected samples may be preserved and / or stabilized to prevent bacterial or fungal growth. In certain embodiments, collected samples may be stored for at least about 1 hour, about 2 hours, about 3 hours, about 4 hours, about 5 hours, about 6 hours, about 12 hours, about 1 day, about 2 days, about 3 days, about 4 days, about 5 days, about 6 days, about 7 days, about 1 week, about 2 weeks, about 3 weeks, about 4 weeks, about 1 month, about 2 months, about 3 months, about 4 months, about 5 months, about 6 months, about 1 year, about 2 years, about 3 years, or longer. In certain embodiments, collected samples may be preserved and stored for long periods at or below room temperature. In certain embodiments, collected samples may be preserved and stored for long periods at or below ambient temperature. In certain embodiments, collected samples may be stored at temperatures up to about 60°C.

[0102] In certain embodiments, stabilized and / or preserved samples may be further processed and analyzed at an external facility (e.g., a remote facility) For example, but not by way of limitation, nucleic acid molecules (e.g., DNA or RNA) may be isolated and extracted from the sample for amplification and / or sequencing applications.

[0103] After a sample is collected, it can be treated to extract nucleic acid molecules (e.g., DNA or RNA). In certain embodiments, DNA extraction methods include organic extraction methods (e.g., phenol-chloroform), non-organic methods (e.g., salting out and proteinase K treatment), and adsorption methods (e.g., silica gel membrane). Additional non-limiting examples of techniques for isolating nucleic acids include the Qiagen DNeasy kit™, Qiagen QIAamp Cador Pathogen Mini kit™, Nucleospin 96 Tissue kit (Macherey-Nagel), QIAzol Lysis Reagent, Qiagen RNeasy kit, Qiagen TurboCapture mRNA kit, and Isopropanol DNA Extraction.

[0104] In certain embodiments, the methods disclosed herein include detecting and quantifying the genome of an animal or pet. In certain embodiments, detecting and quantifying the genome includes isolating DNA from a sample and sequencing the DNA. In certain embodiments, detecting and quantifying the genome includes isolating DNA from a sample and quantifying the DNA (e.g., quantitative PCR).

[0105] Any suitable technique can be employed for detecting and quantifying the genome of an animal or pet. Examples of techniques for detecting and quantifying the genome of an animal or pet include, but are not limited to, 454 pyrosequencing, polymerase chain reaction (PCR), quantitative PCR (qPCR), shotgun sequencing, metagenomic sequencing, Illumina sequencing, PacBio sequencing, nanopore sequencing, and microarray genotyping. In certain non-limiting embodiments, the genome of an animal or pet can be determined by qPCR amplification and sequencing of specific loci. In certain embodiments, the sequencing method is 454-pyrosequencing. In certain embodiments, the sequencing method is Illumina sequencing. In certain embodiments, the sequencing method is whole genome sequencing. In certain embodiments, the method for detecting and quantifying the genome of an animal or pet is microarray genotyping. In certain embodiments, the microarray genotyping is Illumina Infinium BeadChip microarray genotyping.

[0106] The genome of the animal or pet can be further analyzed using any of the methods disclosed herein.

[0107] In certain embodiments, the present disclosure includes systems, devices, and methods that enable convenient and easy collection of samples at home, in the field, or remotely. For example, any user can collect a sample without direct supervision. In certain embodiments, the sample can be collected in a sample collection device. In certain embodiments, the sample collection device can include a reservoir preloaded with chemical reagents for preserving and / or storing the sample (e.g., nucleic acid molecules). In certain embodiments, the reservoir of the sample collection device can advantageously be shielded to prevent direct user exposure. In certain embodiments, the user can be provided with easy-to-understand instructions. In certain embodiments, the instructions can instruct how to use the device, how to collect a sample using the device, how to dispose of the device after use (e.g., by shipping to a remote location), how to access the sample analysis results, or other instructions. In certain embodiments, the collected sample can be transported, such as by shipping (e.g., by mail or via a courier), to a remote laboratory for further processing and / or analysis.

[0108] In certain embodiments, the sample collection device may include a carrier on which the biological sample is collected. In certain embodiments, the carrier may be an absorbent member. For example, but not by way of limitation, the carrier may be a swab, cotton, pad, sponge, foam, or other material or device capable of carrying a biological sample by absorption.

[0109] In certain embodiments, the present disclosure provides a kit. In certain embodiments, the kit includes a sample collection device. In certain embodiments, the sample collection device includes a reservoir and a carrier. In certain embodiments, the reservoir includes a reagent for stabilizing and / or preserving the sample. In certain embodiments, the reservoir includes a shield to protect the user from direct exposure to the reagent. In certain embodiments, the carrier includes an absorbent member. In certain embodiments, the carrier is a cotton swab. In certain embodiments, the reservoir and carrier are configured and arranged to limit or prevent spillage of the reagent or sample. In certain embodiments, the kit includes instructions. The instructions can be provided in a pamphlet or using an internet connection (e.g., using a QR code). For example, and without limitation, the instructions can include information on how to use the sample collection device, collect the sample, dispose of it, and access the analysis results of the sample.

[0110] In certain embodiments, the kit includes a container for shipping the sample collection device to a remote processing location. In certain embodiments, the kit includes a box, envelope, or other packaging material (e.g., insulation, self-sealing or other sealing mechanism, mailer, etc.). In certain embodiments, the kit includes a return label and / or a prepaid label.

[0111] In certain embodiments, the kit includes instructions on how to access the analysis results of the sample. In certain embodiments, the instructions may include a hyperlink or quick response code (e.g., a QR code) to enable access to a website or download of an application on a personal device (e.g., a smartphone). In certain embodiments, the results are provided in a report. In certain embodiments, the report is delivered to the user or healthcare provider (e.g., a veterinarian) by email or electronically. In certain embodiments, the report can be visualized on a personal device (e.g., a smartphone). In certain embodiments, the report may include customized recommendations.

[0112] In certain embodiments, the customized recommendation includes administering to the animal an individualized, nutritionally complete diet. For example, and without limitation, the customized recommendation may be one of the dietary regimens described in WO 2021 / 061743, the entire contents of which are incorporated by reference.

[0113] In certain embodiments, the customized recommendation includes administering a weight-gaining or weight-loss diet. In certain embodiments, the dietary regimen (e.g., weight-loss or weight-gaining diet) is adjusted based on the animal's current body weight and the animal's genome. In certain non-limiting embodiments, the diet is adjusted to about 4100 kcal (about 17154.4 kJ) / kg, about 4000 kcal (about 16736.0 kJ) / kg, about 3900 kcal (about 16317.6 kJ) / kg, about 3800 kcal (about 15899.2 kJ) / kg, about 3700 kcal (about 15480.8 kJ) / kg, about 3600 kcal (about 15062.4 kJ) / kg, about 3500 kcal (about 15062.4 kJ) / kg, ...3500 kcal (about 15062.4 kJ) / kg, about 3500 kcal (about 15062.4 kJ) / kg, about 3500 kcal (about 15062.4 kJ) / kg, about 3500 kcal (about 15062.4 kJ) / kg, 00kcal (approximately 14644.0kJ) / kg, about 3000kcal (approximately 12552.0kJ) / kg, about 2500kcal (approximately 10460.0kJ) / kg, about 2000kcal (approximately 8368.0kJ) / kg, about 1500kcal (approximately 6276.0kJ) / kg, about 1000kcal (approximately 4184.0kJ) / kg or less, or any intermediate value or range thereof. In certain non-limiting embodiments, the diet comprises a fat amount of about 20% w / w, 19% w / w, 18% w / w, 17% w / w, 16% w / w, 15% w / w, 14% w / w, 13% w / w, 12% w / w, 11% w / w, 10% w / w, 9% w / w, 8% w / w, 7% w / w, 6% w / w, 5% w / w, 4% w / w, 3% w / w, 2% w / w, 1% w / w or less, or any intermediate value or range thereof. In certain non-limiting embodiments, the diet comprises a carbohydrate amount of about 25% w / w, 20% w / w, 15% w / w, 10% w / w, 5% w / w, 1% w / w or less, or any intermediate value or range thereof. In certain non-limiting embodiments, the diet comprises an amount of protein of about 20% w / w, 25% w / w, 30% w / w, 35% w / w, 40% w / w, 45% w / w or more, or any intermediate value or range thereof. In certain non-limiting embodiments, the diet comprises an amount of dietary fiber of about 5% w / w, 10% w / w, 15% w / w, 20% w / w, 25% w / w, 30% w / w, 35% w / w, 40% w / w, 45% w / w or more, or any intermediate value or range thereof. Additional information regarding weight loss and weight gain diets is provided in WO 2018 / 129518, the entire contents of which are incorporated by reference.

[0114] In certain embodiments, the customized recommendation includes administering a diet to the animal to improve skin condition (e.g., hydration, texture, elasticity, integrity, barrier, etc.). In certain embodiments, the diet includes linoleic acid. In certain embodiments, the diet includes linoleic acid in an amount of about 7 g / Mcal (about 7 g / 4.184 MJ) to about 9 g / Mcal (about 9 g / 4.184 MJ). In certain embodiments, the diet includes linoleic acid in an amount of about 8 g / Mcal (about 8 g / 4.184 MJ). As used herein, the expression "x g / Mcal (x g / 4.184 MJ)" for a given substance in the diet means that the substance is included in an amount of x grams per Mcal (4.184 MJ) included in the diet. In certain embodiments, the diet includes linoleic acid and zinc. In certain embodiments, the diet comprises zinc in an amount of about 40 mg / Mcal (about 40 mg / 4.184 MJ) to about 60 mg / Mcal (about 60 mg / 4.184 MJ). In certain embodiments, the diet comprises zinc in an amount of about 50 mg / Mcal (about 50 mg / 4.184 MJ). Additional information regarding diets for improving skin conditions is described in WO 2020 / 055856, the entire contents of which are incorporated by reference.

[0115] Additional exemplary diets encompassed by this disclosure can be found in International Publication Nos. WO 2019 / 183557, WO 2019 / 144081, and U.S. Patent Application Publication No. 2022 / 0096537, the entire contents of each of which are incorporated by reference. [Example]

[0116] 4. Example The subject matter of the present disclosure provides improved accuracy of each subsystem in ancestry classification and trait classification, including, but not limited to, local ancestry classification and global ancestry classification, examples of which are described below.

[0117] Example 1: Benchmarking the accuracy of ancestry classifiers A publicly available dataset of 84,414 genetic variants genotyped in 4,368 dog samples from a group of 87 breeds was partitioned into a reference panel (n=4,168) and 200 single-origin query samples. The 200 single-origin query samples were then used to create 200 highly mixed synthetic samples. Both the single-origin query samples and the highly mixed query samples were subjected to local and global ancestry prediction using the system disclosed herein and RFMix. Because the true labels of the 200 query samples were known, the accuracy of the system disclosed herein could be compared with that of RFMix. The accuracy of the classifier was measured as the mean squared error (MSE) between the predicted ancestry proportions and the true proportions.

[0118] Figure 10 shows exemplary results of accuracy benchmarks of our system against the prior art classifier RFMix. Figure 10 and Table 1 show the distribution of MSE for 200 samples in each query set for RFMix and the disclosed system. For single-origin query samples, both the disclosed system and RFMix showed similarly high accuracy (Figure 10). A paired samples t-test showed no significant difference between the MSE of our system compared to RFMix for single-origin samples (t=-1.0749; P=0.2831). Conversely, the average MSE for single-origin and highly mixed samples was significantly different between the disclosed system and RFMix (t=14.1269; P<0.01).

[0119] [Table 1]

[0120] Example 2. Benchmarking the scalability of a classification system The scalability of the disclosed system was compared with that of RFMix. Dramatic differences were observed between the computational resources utilized by the disclosed system and RFMix. To generate the results reported here, the disclosed system required up to 2 Gb of RAM, and the entire workflow took an average of 6 minutes to complete for all chromosomes. However, RFMix required up to 60 Gb of RAM and took an average of 3 hours to complete a single chromosome dataset. To run both workflows in a commodity cloud environment, RFMix required an r5a.4xlarge instance type, currently priced at $0.904 per hour, with an average runtime of 3 hours to run a single chromosome dataset of 200 samples. These requirements translate to a cost of $0.515 per sample. The requirements of the disclosed system imply a cost of approximately $0.000384 per sample when running 200 samples for all chromosomes for an average of 6 minutes on an m5.4xlarge instance type, currently priced at $0.768 per hour. It should be noted that RFMix was not adapted for reference panels of more than 6,000 individuals, but the disclosed system performed efficiently with samples of more than 20,000 individuals.

[0121] Example 3: Assessing the accuracy of global ancestry classification As previously mentioned, previous work has not been able to predict the ancestry of an entire organism from local ancestry assignments. The embodiments disclosed herein characterized the accuracy of a global ancestry classifier using a stratified k-fold cross-validation procedure. Figure 11 shows an exemplary receiver operating characteristic (ROC) curve for the global ancestry classifier. The macro-recall using a publicly available reference panel was 0.9939, and the ROC curve in Figure 11 shows an area under the curve (AUC) of 0.9192.

[0122] In addition to predicting biological labels, specific genetic variants can be used in conjunction with global ancestry to refine predictions. In a proof-of-concept experiment, we used 10 genetic markers known to have a significant effect on phenotype to further classify otherwise indistinguishable subtypes of poodles (toy and miniature), collies (roughhaired and smoothhaired), and dachshunds (longhaired and shorthaired). Table 2 shows the accuracy of using these additional markers in the context of a random forest machine learning model.

[0123] [Table 2]

[0124] Example 4: Performance of trait prediction A decision-tree-based machine learning approach is employed in the trait suite predictor to build a model capable of predicting healthy adult body weight in companion animals. Inputs to the machine learning algorithm are global ancestry data plus genotype data for 39 size- and weight-related genetic markers, a sample training set of 16,168 canids with sex, neuter status, and weight data obtained from ongoing veterinary examinations.

[0125] 12 shows an exemplary regression of predicted adult weights versus true observed adult weights. FIG. 12 shows an exemplary regression analysis of predicted adult weights versus true observed adult weights using the exemplary adult weight prediction module as discussed above according to the present embodiments. Evaluation of the weight prediction model on a test set of samples yielded a mean absolute percentage error (MAPE) of 21.8%.

[0126] Example 5: Performance of automated refinement 13 shows an exemplary iterative improvement of a local ancestry reference panel using isolation forest technology for anomaly detection. FIG. 13 shows that there is an improvement in precision and recall of the reference panel upon application of further iterations of the cross-validation method, including isolation forest iterations, which remove mislabeled reference samples. In certain non-limiting embodiments, the cross-validation method can be supervised or semi-supervised.

[0127] Table 3 shows the accuracy of distinguishing subtypes as described above using supervised and semi-supervised label propagation, where semi-supervised label propagation was used to assign 50% of the subtype labels.

[0128] [Table 3]

[0129] FIG. 14 illustrates an exemplary method 1400 for ancestry prediction. The method may begin at step 1410, in which a computing system may access a sample of genetic material associated with a first animal, the sample of genetic material including one or more raw genotypes. At step 1420, the computing system may generate one or more phased haplotypes based on the one or more raw genotypes. At step 1430, the computing system may generate, for the one or more phased haplotypes, one or more local assignments to one or more genetic populations based on a comparison between the one or more phased haplotypes and a reference panel including multiple reference haplotypes associated with multiple reference populations, using one or more machine learning algorithms. At step 1440, the computing system may send instructions to a user device to present an output associated with the first animal to a user, the output being generated based on the one or more local assignments to one or more genetic populations. Certain embodiments may repeat one or more steps of the method of FIG. 14, as appropriate. 14 as occurring in a particular order, the present disclosure contemplates any suitable steps of the method of FIG. 14 occurring in any suitable order. Further, while the present disclosure describes and illustrates exemplary methods for ancestry prediction that include certain steps of the method of FIG. 14, the present disclosure contemplates any suitable method for ancestry prediction that includes any suitable steps, which may include all, some, or none of the steps of the method of FIG. 14, as appropriate. Further, while the present disclosure describes and illustrates particular components, devices, or systems that perform certain steps of the method of FIG. 14, the present disclosure contemplates any suitable combination of any suitable components, devices, or systems that perform any suitable steps of the method of FIG. 14.

[0130] Those skilled in the art will recognize that the disclosed methods and systems can be implemented in many ways and are therefore not limited by the exemplary embodiments and examples described above. In other words, functional elements performed by single or multiple components, and individual functions can be distributed among software applications at either the client level or the server level, or both, in various combinations of hardware and software or firmware. In this regard, any number of features of the different embodiments described herein can be combined in a single or multiple embodiments, and alternative embodiments having fewer or more than all of the features described herein are possible.

[0131] Also, all or part of the functionality may be distributed among multiple components in ways now known or that will become known in the future. Thus, countless software / hardware / firmware combinations are possible for implementing the functions, features, interfaces, and configurations described herein. Furthermore, the scope of the present disclosure covers conventionally known methods for implementing the described features, functions, and interfaces, as well as variations and modifications to the hardware or software or firmware components described herein that are now or will become understood by those skilled in the art.

[0132] Furthermore, method embodiments presented and described as flowcharts in this disclosure are provided as examples to provide a more complete understanding of the present technology. The disclosed methods are not limited to the operations and logical flow presented herein. Alternative embodiments are contemplated in which the order of various operations is changed and sub-operations described as part of a larger operation are performed independently.

[0133] While various embodiments have been described for purposes of this disclosure, such embodiments should not be construed as limiting the teachings of this disclosure to those embodiments. Various changes and modifications can be made to the elements and operations described above to obtain results that remain within the scope of the systems and processes described in this disclosure.

[0134] While the disclosed subject matter has been described herein in terms of certain preferred embodiments, those skilled in the art will recognize that various modifications and improvements can be made to the disclosed subject matter without departing from its scope. Furthermore, although individual features of one non-limiting embodiment of the disclosed subject matter may be discussed herein or shown in the drawings of one non-limiting embodiment and not shown in other embodiments, it will be apparent that individual features of one non-limiting embodiment may be combined with one or more features of another embodiment or features from multiple embodiments.

Claims

1. 1. A kit comprising a sample collection device, configured by one or more computing systems to: accessing a sample of genetic material associated with a first animal, the sample comprising one or more raw genotypes; generating one or more phased haplotypes based on the one or more raw genotypes; generating, for the one or more phased haplotypes, by one or more machine learning algorithms, one or more local assignments to one or more genetic populations based on a comparison between the one or more phased haplotypes and a reference panel comprising a plurality of reference haplotypes associated with a plurality of reference populations; transmitting instructions to a user device to present to a user an output associated with the first animal and generated based on the one or more local assignments to the one or more genetic populations; 10. A kit for determining the local and global ancestry of an animal by a method comprising:

2. The kit of claim 1 , wherein the sample collection device comprises a carrier and a reservoir.

3. The kit of claim 3 , wherein the carrier comprises an absorbent member and the reservoir comprises a shield.

4. The kit of claim 1 , further comprising instructions on how to use the sample collection device and / or how to collect a sample.