Methods and systems for cross-species transfer learning

Cross-species transfer learning using a generative machine learning model addresses the challenge of predicting animal responses by analyzing biological molecule amounts and features, improving clinical outcome predictions across species without relying on structural similarities.

WO2025151452A1PCT designated stage expired Publication Date: 2025-07-17BIOSTATE AI INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/010628
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-08
Filing Date
2025-01-07
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

Predicting an animal's response to a stimulus, such as a drug, across different biological species is challenging due to structural biological differences, leading to failures in traditional methods that rely on structural similarities like DNA sequences.

Method used

Employing cross-species transfer learning using a generative machine learning model trained on signal strength values, molecular feature codes, and experiment feature codes to predict clinical outcomes without relying on explicit structural similarities, allowing for improved prediction of an animal's response to a stimulus.

Benefits of technology

The method provides accurate clinical outcomes for subjects of a second species by leveraging machine learning to analyze biological molecule amounts and features, enhancing treatment prediction and administration strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025010628_17072025_PF_FP_ABST
    Figure US2025010628_17072025_PF_FP_ABST
Patent Text Reader

Abstract

Methods for training and using one or more machine learning models, e.g., a generative machine learning model, are described. The methods may comprise, for example, conducting experiments on subjects from a first species; measuring amounts of a biological smolecule from samples from the experiments; determining signal strength values corresponding to the measured amounts; determining feature codes corresponding to aspects of the experiments; creating vectors based on the signal strength values and the feature codes; and training a generative machine learning model to provide generated data for a subject from the second species, using the created vectors. The methods may also comprise generating generated data by inputting experiment data from a first species, into a trained generative machine learning model. The generated data can be inputted into another trained machine learning model to predict a clinical outcome for a subject from the second species. The second species can be human.
Need to check novelty before this filing date? Find Prior Art

Description

METHODS AND SYSTEMS FOR CROSS-SPECIES TRANSFER LEARNINGCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 618,879, filed January 8, 2024, which is incorporated herein by reference in its entirety.FIELD

[0002] The present disclosure relates generally to methods and systems of transfer learning in artificial intelligence. More specifically, the present disclosure relates to training and deploying a machine learning model based on training data comprising data from a first biological species, to predict a clinical outcome for a subject of a second biological species.BACKGROUND

[0003] Despite structural biological similarities (e.g., DNA sequences) across biological species, the response of an animal of a first species to a stimulus (e.g., drug) is challenging to predict based on known stimulus responses of a second species. This challenge persists even when the stimulus is identical across multiple biological species. For example, a response to a drug in a human, may not necessarily be similar to a response to a drug in a non-human, such as a rat. Accordingly, the majority of drugs selected based on experimental results from nonhumans fail to provide the expected outcome in clinical studies in humans. Improved methods are needed for determining an animal’s (e.g., a human’s) response to a stimulus, based on data from other biological species. The present disclosure addresses these needs.BRIEF SUMMARY

[0004] Disclosed herein are methods and systems for cross-species transfer learning in artificial intelligence. Existing methods for predicting an animal’s response to a stimulus (e.g., drug) from animals of other biological species are often based on structural biological similarities (e.g., DNA sequences) between the two or more species. Such predictions, however, often fail. The failure may result from the numerous structural biological differences that exist across species, such as differences in DNA sequences. The methods and systems describedherein include machine learning, such as the training of a generative machine learning model, to predict an animal’s response to a stimulus, based on the responses of animals of other species, to the same and / or different stimuli. That is, the methods and systems described herein use transfer learning. The prediction from the trained machine learning model can be used to provide a clinical outcome to an animal, e.g., subject. In doing so, the methods and systems described herein do not necessarily rely explicitly on the structural similarities of biological features across species, and provide a technical advantage over traditional methods.

[0005] In some aspects, disclosed herein is a method for training a generative machine learning model for providing generated data for a second species based on training data from a first species, comprising: conducting one or more experiments on one or more subjects from a first species; measuring one or more amounts of a biological molecule from one or more samples from the one or more experiments; determining one or more signal strength values corresponding to the one or more amounts; determining one or more molecular feature codes corresponding to one or more molecular features of the biological molecule; determining one or more experiment feature codes corresponding to one or more experiment features of an experiment from the one or more experiments; creating one or more vectors based on the one or more signal strength values, the one or more molecular feature codes, and the one or more experiment feature codes; and training the generative machine learning model to provide the generated data for a subject from the second species, using the created vectors.

[0006] In some aspects, disclosed herein is a method for generating generated data for a subject of a second species, comprising: inputting experiment data for one or more subjects from a first species into a generative machine learning model trained on a second species, the generative machine learning model having been trained by a method comprising: conducting one or more experiments on one or more subjects from the first species; measuring one or more amounts of a biological molecule from one or more samples from the one or more experiments; determining one or more signal strength values corresponding to the one or more amounts; determining one or more molecular feature codes corresponding to one or more molecular features of the biological molecule; determining one or experiment feature codes corresponding to one or more experiment features of an experiment from the one or more experiments; creatingone or more vectors based on the one or more signal strength values, the one or more molecular feature codes, and the one or more experiment feature codes; training the generative machine learning model to provide the generated data for the subject from the second species, using the created vectors; and receiving from the trained generative machine learning model, the generated data.

[0007] In some embodiments, the methods disclosed herein can further comprise: inputting the generated data into a second trained machine learning model configured to predict a clinical outcome for the subject from the second species; and receiving from the second trained machine learning model, the predicted clinical outcome for the subject from the second species. In some embodiments, the second species can be a human. In any of the embodiments herein, the second species can be identical to the first species.

[0008] In any of the embodiments herein, the predicted clinical outcome can comprise determining a disease signal in the subject from the second species. In any of the embodiments herein, the predicted clinical outcome can comprise determining a diagnosis for the subject from the second species. In any of the embodiments herein, the predicted clinical outcome can comprise determining a therapy for the subject from the second species. In some embodiments, the determining the therapy for the subject from the second species can comprise determining a dosage of the therapy. In some embodiments, the predicted clinical outcome can comprise administering the therapy to the subject from the second species. In any of the embodiments herein, the predicted clinical outcome can comprise determining a prognosis for the subject from the second species. In some embodiments, the prognosis can comprise determining a response to a predetermined therapy from the subject from the second species. In some embodiments, the response comprises an adverse response to the predetermined therapy, for the subject from the second species. In some embodiments, the response can comprise a recovery from the predetermined therapy, for the subject from the second species. In any of the embodiments herein, the predicted clinical outcome can comprise determining a life expectancy of the subject from the second species.

[0009] In any of the embodiments herein, the measured one or more amounts can be relative to a predetermined baseline level. In some embodiments, the predetermined baseline level can bebased on observing the biological molecule in one or more healthy subjects of the first species. In any of the embodiments herein, the measured one or more amounts can be above a predetermined threshold amount. In some embodiments, the one or more signal strength values can be determined when the one or more amounts are above the predetermined threshold amount. In any of the embodiments herein, the measured one or more amounts can be greater than a first consensus amount of the biological molecule based on the observing the one or more healthy subjects of the first species. In some embodiments, the one or more signal strength values can be determined when the measured one or more amounts are greater than the first consensus amount of the biological molecule based on the observing the one or more healthy subjects of the first species. In any of the embodiments herein, the measured one or more amounts can be greater than the first consensus amount by more than a first multiple. In some embodiments, the first multiple can be two. In any of the embodiments herein, the first consensus amount can be a median amount of the biological molecule based on the observing the one or more healthy subjects of the first species.

[0010] In any of the embodiments herein, a signal strength value from the one or more signal strength values can correspond to an amount from the one or more measured values. In some embodiments, the experiment can comprise providing a stimulus to the one or more subjects from the first species. In some embodiments, the stimulus can be a drug. In some embodiments, the drug can be isoniazid, tetracycline, carbon tetrachloride, valproate, or a combination thereof. In any of the embodiments herein, the providing the stimulus can comprise administering a dosage of the drug that is at or above a predetermined dosage level. In any of the embodiments herein, the one or more signal strength values can be determined based on the one or more measured amounts deviating from a second consensus amount of the biological molecule in the first species, by more than a second multiple. In some embodiments, the second multiple is four. In any of the embodiments herein, the second consensus amount can be a median amount of the biological molecule of the first species. In any of the embodiments herein, the one or more signal strength values can be determined based on the one or more measured amounts deviating from the second consensus amount of the biological molecule in the first species, for more than a predetermined duration. In some embodiments, the predetermined duration can be eight days. In any of the embodiments herein, the biological molecule can be an RNA molecule. In someembodiments, the RNA molecule can be an mRNA transcript. In any of the embodiments herein, the determining the one or more measured amounts can be based on RNA sequencing. In any of the embodiments herein, the biological molecule can be a protein molecule. In any of the embodiments herein, the determining the one or more measured amounts can be based on mass spectrometry. In any of the embodiments herein, the biological molecule can be an epigenetic marker. In any of the embodiments herein, the biological molecule can be a methyl group on a nucleotide base. In any of the embodiments herein, the determining the one or more measured amounts can be based on methyl- sequencing. In any of the embodiments herein, the biological molecule can be a DNA molecule. In some embodiments, the DNA molecule can be a cell-free DNA (cfDNA). In any of the embodiments herein, the determining the one or more measured amounts can be based on DNA sequencing. In any of the embodiments herein, the biological molecule can be a metabolite. In any of the embodiments herein, the one or more molecular features can comprise an identifier of the biological molecule.

[0011] In any of the embodiments herein, the one or more experiment features comprise an identifier of the first species. In any of the embodiments herein, the one or more experiment features can comprise a description of the one or more experiments. In some embodiments, the description of the one or more experiments can comprise a description of the one or more samples. In any of the embodiments herein, the one or more samples can comprise liquid biopsy samples. In some embodiments, the liquid biopsy samples can comprise blood, plasma, cerebrospinal fluid, sputum, stool, urine, sweat, or saliva. In any of the embodiments herein, the one or more samples can comprise a tissue biopsy sample. In any of the embodiments herein, the one or more samples can be obtained from a necropsy of the one or more subjects. In any of the embodiments herein, the one or more molecular feature codes can correspond to the one or more molecular features according to a numeric index. In any of the embodiments herein, the one or more experiment feature codes can correspond to the one or more experiment features according to a numeric index. In any of the embodiments herein, the combining can comprise ordering the one or more signal strength values, the one or more molecular feature codes, or the one or more experiment feature codes, according to a method of ordering. In some embodiments, the method of ordering can be based on an unsupervised clustering of the one or more signal strength values. In any of the embodiments herein, the method of ordering can comprise ordering the one or moreexperiment feature codes before the one or more signal strength values or the one or more molecular feature codes. In any of the embodiments herein, the combining can comprise randomly ordering the one or more signal strength values, the one or more molecular feature codes, or the one or more experiment feature codes. In any of the embodiments herein, a percentage of the training data can comprise data describing the second species. In some embodiments, the percentage can be 0% or greater and 5% or less. In any of the embodiments herein, the percentage can be 1%. In any of the embodiments herein, the first species can be a mouse, a rat, a roundworm, a non-human primate, a fruitfly, a yeast, a rabbit, or a pig.

[0012] In any of the embodiments herein, the generative machine learning model can comprise a transformer model. In any of the embodiments herein, the generative machine learning model can represent a probability distribution. In any of the embodiments herein, the generative machine learning model can represent the probability distribution by learning the density or mass of the probability distribution. In some embodiments, the learning can be based on maximum likelihood estimation. In any of the embodiments herein, the generative machine learning model can be an autoregressive model, a normalizing flow model, an energy-based model, or a variational auto-encoder. In any of the embodiments herein, the generative machine learning model can represent the probability distribution by learning a model of the sampling process for the probability distribution. In any of the embodiments herein, the generative machine learning model can be a generative adversarial network. In any of the embodiments herein, the generated data can comprise generated RNA sequencing data. In any of the embodiments herein, the generated data can comprise generated mass spectrometry data. In any of the embodiments herein, the generated data can comprise generated DNA sequencing data. In any of the embodiments herein, the generated data can comprise generated methyl-sequencing data. In any of the embodiments herein, the inputted experiment data can comprise an input amount of an input biological molecule from an input sample from an input biological experiment. In any of the embodiments herein, the inputted experiment data can comprise an input molecular feature code corresponding to an input molecular feature of an input biological molecule. In any of the embodiments herein, the inputted experiment data can comprise an input experiment feature code corresponding to an input experiment feature of an input experiment. In some embodiments, the input experiment feature can comprise an input description of the subjectfrom the second species. In some embodiments, the input description of the subject from the second species can comprise an input label for the second species. In any of the embodiments herein, the input experiment feature can comprise an input label for the input sample. In any of the embodiments herein, the second trained machine learning model can be a classifier for predicting the clinical outcome for the second species. In any of the embodiments herein, the second trained machine learning model can use regression to predict the clinical outcome for the second species.

[0013] In some aspects, disclosed herein is a system comprising: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, can cause the system to: determine one or more signal strength values corresponding to one or more amounts of a biological molecule measured from one or more samples from one or more experiments; determine one or more molecular feature codes corresponding to one or more molecular features of the biological molecule; determine one or more experiment feature codes corresponding to one or more experiment features of an experiment from the one or more experiments; create one or more vectors based on the one or more signal strength values, the one or more molecular feature codes, and the one or more experiment feature codes; and train the generative machine learning model to provide generated data for a subject from a second species, using the created vectors.

[0014] In some aspects, disclosed herein is a system comprising: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: input experiment data for one or more subjects from a first species into a generative machine learning model trained on a second species, the generative machine learning model having been trained by a method comprising: determine one or more signal strength values corresponding to one or more amounts of a biological molecule measured from one or more samples from one or more experiments; determine one or more molecular feature codes corresponding to one or more molecular features of the biological molecule; determine one or more experiment feature codes corresponding to one or more experiment features of an experiment from the one or moreexperiments; create one or more vectors based on the one or more signal strength values, the one or more molecular feature codes, and the one or more experiment feature codes; train the generative machine learning model to provide generated data for a subject from a second species, using the created vectors; and receive from the trained generative machine learning model, the generated data.

[0015] In some embodiments, the system disclosed herein can comprise further instructions that, when executed by the one or more processors, cause the system to: input the generated data into a second trained machine learning model configured to predict a clinical outcome for the subject from the second species; and receive from the second trained machine learning model, the predicted clinical outcome for the subject from the second species. In any of the embodiments herein, the second species can be a human. In any of the embodiments herein, the second species can be identical to the first species. In any of the embodiments herein, the biological molecule can be an RNA molecule. In some embodiments, the RNA molecule can be an mRNA molecule. In any of the embodiments herein, determining the one or more amounts can be based on RNA sequencing. In any of the embodiments herein, a percentage of the training data can comprise data describing the second species. In some embodiments, the percentage can be 0% or greater and 5% or less. In any of the embodiments herein, the percentage can be 1%. In any of the embodiments herein, the one or more generative machine learning models can comprise a transformer model.

[0016] In some aspects, disclosed herein is a non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of a system, can cause the system to: determine one or more signal strength values corresponding to one or more amounts of a biological molecule measured from one or more samples from one or more experiments; determine one or more molecular feature codes corresponding to one or more molecular features of the biological molecule; determine one or more experiment feature codes corresponding to one or more experiment features of an experiment from the one or more experiments; create one or more vectors based on the one or more signal strength values, the one or more molecular feature codes,and the one or more experiment feature codes; and train the generative machine learning model to provide generated data for a subject from a second species, using the created vectors.

[0017] In some aspects, disclosed herein is a non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of a system, can cause the system to: input experiment data for one or more subjects from a first species into a generative machine learning model trained on a second species, the generative machine learning model having been trained by a method comprising: determine one or more signal strength values corresponding to one or more amounts of a biological molecule measured from one or more samples from one or more experiments; determine one or more molecular feature codes corresponding to one or more molecular features of the biological molecule; determine one or more experiment feature codes corresponding to one or more experiment features of an experiment from the one or more experiments; create one or more vectors based on the one or more signal strength values, the one or more molecular feature codes, and the one or more experiment feature codes; train the generative machine learning model to provide generated data for a subject from a second species, using the created vectors; and receive from the trained generative machine learning model, the generated data. In some embodiments, the non-transitory computer-readable storage medium described herein can further comprise instructions that, when executed by the one or more processors, cause the system to: input the generated data into a second trained machine learning model configured to predict a clinical outcome for the subject from the second species; and receive from the second trained machine learning model, the predicted clinical outcome for the subject from the second species. In any of the embodiments herein, the second species can be a human. In any of the embodiments herein, the second species can be identical to the first species. In any of the embodiments herein, the biological molecule can be an RNA molecule. In some embodiments, the RNA molecule can be an mRNA molecule. In any of the embodiments herein, determining the one or more amounts can be based on RNA sequencing. In any of the embodiments herein, a percentage of the training data can comprise data describing the second biological species. In some embodiments, the percentage can be 0% or greater and 5% or less. In any of the embodiments herein, the percentage can be 1%. In any of the embodiments herein, the one or more generative machine learning models can comprise a transformer model.BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Various aspects of the disclosed methods, devices, and systems are set forth with particularity in the appended claims. A better understanding of the features and advantages of the disclosed methods, devices, and systems will be obtained by reference to the following detailed description of illustrative embodiments and the accompanying drawings, of which:

[0019] FIG. 1 provides a non-limiting exemplary method for training a generative machine learning model.

[0020] FIG. 2 provides a non-limiting exemplary method for generating generated data from a trained generative machine learning model.

[0021] FIG. 3 provides a non-limiting exemplary method for predicting a clinical outcome from the generated data, using a second trained machine learning model.

[0022] FIG. 4 provides an example of one or more signal strength values, e.g., functional labels, and their corresponding experimental conditions.

[0023] FIG. 5 provides an example depicting mappings of the number of genes between two different biological species.

[0024] FIG. 6 provides an example of the one or more signal strength values, e.g., functional labels, for genes, and their mappings across two different biological species.

[0025] FIG. 7 provides a non-limiting exemplary schematic of one use case of cross-species transfer learning, based on signal strength values, e.g., functional labels, for genes, to predict a clinical outcome for a subject of a second species.

[0026] FIG. 8 depicts an exemplary computing device or system in accordance with one embodiment of the present disclosure.

[0027] FIG. 9 depicts an exemplary computer system or computer network, in accordance with some instances of the systems described herein.

[0028] FIG. 10 provides an example of longitudinal gene expression data in rats (Rattus norvegicus), in response to drugs.

[0029] FIG. 11 provides an example of data of the number of genes responding to a drug, or a combination of drugs, by a change in gene expression, in rats (Rattus norvegicus).

[0030] FIG. 12 provides an example of data (e.g., functional labels and descriptions of the corresponding experiments) corresponding to the one or more signal strength values, one or more molecular feature codes, and one or more experiment feature codes.

[0031] FIG. 13A provides an example of data for the 4930507D05Rik gene in mice, including corresponding one or more signal strength values, e.g., functional labels, and their corresponding gene expression values.

[0032] FIG. 13B provides an example of data for the Yipf7 gene in mice, including corresponding one or more signal strength values, e.g., functional labels, and their corresponding gene expression values.

[0033] FIG. 13C provides an example of data for the 170001 lL22Rik gene in mice, including corresponding one or more signal strength values, e.g., functional labels, and their corresponding gene expression values.

[0034] FIG. 13D provides an example of data for the WntlOa gene in mice, including corresponding one or more signal strength values, e.g., functional labels, and their corresponding gene expression values.

[0035] FIG. 13E provides an example of data for the Capnl 1 gene in mice, including corresponding one or more signal strength values, e.g., functional labels, and their corresponding gene expression values.

[0036] FIG. 13F provides an example of data for the C730014E05Rik gene in mice, including corresponding one or more signal strength values, e.g., functional labels, and their corresponding gene expression values.

[0037] FIG. 14A provides an example of data of the number of genes responding to a drug, or a combination of drugs, by a change in gene expression, in mice (Mus musculus).

[0038] FIG. 14B provides an example of data of the number of genes in an organ responding to, at least a drug, by a change in gene expression, in mice (Mus musculus).

[0039] FIG. 15A provides an example of data for the Bexl gene in rats, including corresponding one or more signal strength values, e.g., functional labels, and their corresponding gene expression values.

[0040] FIG. 15B provides an example of data for the Texl 1 gene in rats, including corresponding one or more signal strength values, e.g., functional labels, and their corresponding gene expression values.

[0041] FIG. 15C provides an example of data for the Add2 gene in rats, including corresponding one or more signal strength values, e.g., functional labels, and their corresponding gene expression values.

[0042] FIG. 15D provides an example of data for the Vill gene in rats, including corresponding one or more signal strength values, e.g., functional labels, and their corresponding gene expression values.

[0043] FIG. 15E provides an example of data for the Hba-a3 gene in rats, including corresponding one or more signal strength values, e.g., functional labels, and their corresponding gene expression values.

[0044] FIG. 15F provides an example of data for the RGD1565566 gene in rats, including corresponding one or more signal strength values, e.g., functional labels, and their corresponding gene expression values.

[0045] FIG. 15G provides an example of data for the Mybpc3 gene in rats, including corresponding one or more signal strength values, e.g., functional labels, and their corresponding gene expression values.

[0046] FIG. 15H provides an example of data for the LOC120100962 gene in rats, including corresponding one or more signal strength values, e.g., functional labels, and their corresponding gene expression values.

[0047] FIG. 16A provides an example of data of the number of genes responding to a drug, or a combination of drugs, by a change in gene expression, in rats.

[0048] FIG. 16B provides an example of data of the number of genes in an organ responding to, at least a drug, by a change in gene expression, in rats.

[0049] FIG. 17A provides an example of data of signal strength values with unique perfect matches between mouse Apol7a and rat Aacs genes.

[0050] FIG. 17B provides an example of data of signal strength values with unique perfect matches between mouse Mucl6 and rat Dusp29 genes.

[0051] FIG. 17C provides an example of data of signal strength values with unique perfect matches between mouse Disp2 and rat LOCI 20100015 genes.DETAILED DESCRIPTION

[0052] Predicting the function of an animal, such as the animal’s response to a stimulus, e.g., drug, is challenging. Providing such predictions, however, can be of critical usefulness, especially if the function of the animal, such as the animal’s response to a particular drug, is unknown. Disclosed herein are methods and systems for predicting a clinical outcome, e.g., a response to a drug, for a first subject. The disclosed methods can include conducting experiments on subjects from a first species. From the one or more experiments, amounts of a biological molecule can be measured. Based on the amounts, signal strength values can be determined. In addition, molecular feature codes relating to features of the biological molecule can be determined, as well as experiment feature codes relating to features of the experiments. Vectors can then be created, based on the signal strength values, the molecular feature codes, and the experiment feature codes. Based on the created vectors, generative machine learning model can be trained to provide generated data for a subject from a second species. The second species canbe human. The trained generative machine learning model can be used to produce generated data, upon receiving input experiment data from subjects from a first species. The generated data can be inputted into another trained machine learning model that predicts a clinical outcome for the second species. The predicted clinical outcome can be used to guide clinical procedures for subjects from the second species, e.g., human subjects.

[0053] Existing methods for predicting an animal’s (e.g., human’s) response to a stimulus (e.g., drug) from animals of other biological species (e.g., non-humans) are often based on structural biological similarities (e.g., DNA sequences) between the two or more species. For example, if the p53 gene shows increased expression in the lungs of rats, in response to tetracycline, then traditional methods may predict that the p53 gene homolog in humans may also show increased expression in the lungs of rats, in response to tetracycline, and may recommend a therapeutic strategy based on the increased p53 expression. Such predictions, however, often fail. The failure may result from the numerous structural biological differences that exist across species. For instances, the number of genes between biological species can differ. Accordingly, a one-to-one correspondence of a gene, such as p53, does not exist between biological species, such as rats and humans. One biological species may have multiple genes that are structurally similar (e.g., have similar sequences) to those of another sequence, but of those structurally similar genes, the functions of those genes may, in fact, differ between the biological species (e.g., due to evolutionary phenomena, such as subfunctionalization or neofunctionalization of genes). The overall differences in structural biological similarities across species can result in traditional methods relying on such structural similarities to struggle, when predicting the functions, e.g., responses, of a biological species, such as an animal’s response to a drug.

[0054] The methods and systems described herein use machine learning, such as the training of a generative machine learning model, to predict an animal’s response to a stimulus, based on the responses of animals of other species, to the same and / or different stimuli. That is, the methods and systems described herein use transfer learning. The prediction from the trained machine learning model can be used to provide a clinical outcome to an animal, e.g., subject. In doing so, the methods and systems described herein do not necessarily rely explicitly on thestructural similarities of biological features across species. The training data used when training the machine learning model, as well as the inputs to the trained machine learning model, need not provide structural biological similarities across species, such as DNA sequences and their differences across species, to predict a function of an animal, such as the animal’s response to a stimulus, e.g., drug. By eliminating the need to rely on structural biological similarities when predicting animal function, the methods and systems described herein provide a technical advantage over traditional methods.

[0055] In some aspects, described herein is a method for training a generative machine learning model for providing generated data for a second species based on training data from a first species, comprising: conducting one or more experiments on one or more subjects from a first species; measuring one or more amounts of a biological molecule from one or more samples from the one or more experiments; determining one or more signal strength values corresponding to the one or more amounts; determining one or more molecular feature codes corresponding to one or more molecular features of the biological molecule; determining one or more experiment feature codes corresponding to one or more experiment features of an experiment from the one or more experiments; creating one or more vectors based on the one or more signal strength values, the one or more molecular feature codes, and the one or more experiment feature codes; and training the generative machine learning model to provide the generated data for a subject from the second species, using the created vectors.

[0056] In some aspects, described herein is a method for generating generated data, comprising: inputting experiment data for one or more subjects from a first species into a generative machine learning model trained on a second species, the generative machine learning model having been trained by a method comprising: conducting one or more experiments on one or more subjects from the first species; measuring one or more amounts of a biological molecule from one or more samples from the one or more experiments; determining one or more signal strength values corresponding to the one or more amounts; determining one or more molecular feature codes corresponding to one or more molecular features of the biological molecule; determining one or experiment feature codes corresponding to one or more experiment features of an experiment from the one or more experiments; creating one or more vectors based on theone or more signal strength values, the one or more molecular feature codes, and the one or more experiment feature codes; training the generative machine learning model to provide the generated data for the subject from the second species, using the created vectors; and receiving from the trained generative machine learning model, the generated data. In some aspects, the methods described herein further comprise: inputting the generated data into a second trained machine learning model configured to predict a clinical outcome for the second species; and receiving from the second trained machine learning model, the predicted clinical outcome for the second species.Definitions

[0057] Unless otherwise defined, all of the technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art in the field to which this disclosure belongs.

[0058] As used in this specification and the appended claims, the singular forms “a”, “an”, and “the” include plural references unless the context clearly dictates otherwise. Any reference to “or” herein is intended to encompass “and / or” unless otherwise stated.

[0059] ‘About” and “approximately” shall generally mean an acceptable degree of error for the quantity measured given the nature or precision of the measurements. Exemplary degrees of error are within 20 percent (%), typically, within 10%, and more typically, within 5% of a given value or range of values.

[0060] As used herein, the terms "comprising" (and any form or variant of comprising, such as "comprise" and "comprises"), "having" (and any form or variant of having, such as "have" and "has"), "including" (and any form or variant of including, such as "includes" and "include"), or "containing" (and any form or variant of containing, such as "contains" and "contain"), are inclusive or open-ended and do not exclude additional, un-recited additives, components, integers, elements, or method steps.

[0061] As used herein, the terms “individual,” “patient,”, “animal”, or “subject” are used interchangeably and refer to any single animal, e.g., a mammal (including such non-humananimals as, for example, dogs, cats, horses, rabbits, zoo animals, cows, pigs, sheep, and nonhuman primates) for which treatment is desired. In particular embodiments, the individual, patient, or subject herein is a human.

[0062] It is understood that aspects and variations of the invention described herein include “consisting” and / or “consisting essentially of’ aspects and variations.

[0063] When a range of values is provided, it is to be understood that each intervening value between the upper and lower limit of that range, and any other stated or intervening value in that states range, is encompassed within the scope of the present disclosure. Where the stated range includes upper or lower limits, ranges excluding either of those included limits are also included in the present disclosure.

[0064] Some of the analytical methods described herein include mapping sequences to a reference sequence, determining sequence information, and / or analyzing sequence information. It is well understood in the art that complementary sequences can be readily determined and / or analyzed, and that the description provided herein encompasses analytical methods performed in reference to a complementary sequence.

[0065] The section headings used herein are for organization purposes only and are not to be construed as limiting the subject matter described. The description is presented to enable one of ordinary skill in the art to make and use the invention and is provided in the context of a patent application and its requirements. Various modifications to the described embodiments will be readily apparent to those persons skilled in the art and the generic principles herein may be applied to other embodiments. Thus, the present invention is not intended to be limited to the embodiment shown but is to be accorded the widest scope consistent with the principles and features described herein.

[0066] The figures illustrate processes according to various embodiments. In the exemplary processes, some blocks are, optionally, combined, the order of some blocks is, optionally, changed, and some blocks are, optionally, omitted. In some examples, additional steps may be performed in combination with the exemplary processes. Accordingly, the operations asillustrated (and described in greater detail below) are exemplary by nature and, as such, should not be viewed as limiting.

[0067] The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.Methods for training a generating machine learning model to provide a clinical outcome

[0068] The methods disclosed herein comprise a method for training a generative machine learning model for providing generated data for a second species based on training data from a first species. The resulting trained generative machine learning model can output generated data for a subject of a second species. The generated data can be inputted into another trained machine learning model configured to predict a clinical outcome for a subject from the second species. The second species can be human. The methods described herein allow for improvements in predicting the clinical outcome for a subject, such as treatments that can be administered to a subject.

[0069] FIG. 1 shows an exemplary schematic showing a general process 100 for training a generative machine learning model for providing generated data for a second species based on training data from a first species. The method can include: conducting one or more experiments on one or more subjects from a first species; (102); measuring one or more amounts of a biological molecule from one or more samples from the one or more experiments (104); determining one or more signal strength values corresponding to the one or more amounts (106); determining one or more molecular feature codes corresponding to one or more molecular features of the biological molecule (108); determining one or more experiment feature codes corresponding to one or more experiment features of an experiment from the one or more experiments (110); creating one or more vectors based on the one or more signal strength values, the one or more molecular feature codes, and the one or more experiment feature codes (112); and training the generative machine learning model to provide the generated data for a subject from the second species, using the created vectors (114).

[0070] FIG. 2 shows an exemplary schematic showing a general process 200 for generating generated data for a subject of a second species, comprising: inputting experiment data for one or more subjects from a first species into a generative machine learning model trained on a second species, the generative machine learning model having been trained by a method (202); and receiving from the trained generative machine learning model, generated data (204).

[0071] FIG. 3 shows an exemplary schematic showing a general process 300, comprising: inputting experiment data for one or more subjects from a first species into a generative machine learning model trained on a second species, the generative machine learning model having been trained by a method (302); receiving from the trained generative machine learning model, generated data (304); inputting the generated data into a second trained machine learning model configured to predict a clinical outcome for the second species (306); and receiving from the second trained machine learning model, the predicted clinical outcome for the second species (308).

[0072] Process 100, 200 or 300 can be performed, for example, using one or more electronic devices implementing a software platform. In some examples, process 100, 200, or 300 is performed using a client-server system, and the blocks of process 100, 200, or 300 are divided up in any manner between the server and a client device. In other examples, the blocks of process 100, 200, or 300 are divided up between the server and multiple client devices. Thus, while portions of process 100, 200, or 300 are described herein as being performed by particular devices of a client-server system, it will be appreciated that process 100, 200, or 300 is not so limited. In other examples, process 100, 200, or 300 is performed using only a client device or only multiple client devices. In process 100, 200, or 300 some blocks are, optionally, combined, the order of some blocks is, optionally, changed, and some blocks are, optionally, omitted. In some examples, additional steps may be performed in combination with the process 100, 200, or 300. Accordingly, the operations as illustrated (and described in greater detail below) are exemplary by nature and, as such, should not be viewed as limiting.

[0073] At 102 in FIG. 1, one or more experiments on one or more subjects from a first species are conducted.

[0074] At 104 in FIG. 1, one or more amounts of a biological molecule from one or more samples from the one or more experiments are measured.

[0075] At 106 in FIG. 1, one or more signal strength values corresponding to the one or more amounts are determined. The one or more signal strength values can be an integer or a float, and can be a positive, or a negative number. The one or more signal strength values can be zero. The signal strength value can, for example, be values in a functional label. A functional label can be a quantitative description of an experiment comprising, e.g., a vector of values, and can describe a functional response of a subject to a stimulus. The signal strength value can represent the strength of a signal from an experiment, such as a signal representing a measured amount of a biological molecule.

[0076] At 108 in FIG. 1, one or more molecular feature codes corresponding to one or more molecular features of the biological molecule are determined. The one or more molecular features can include a gene name, as well as the raw expression value of the gene. The one or more molecular features can include a log2-transformed value of the raw expression value of the gene. The molecular feature codes can be expressed as a numeric value, such as an integer or a float, and can be positive, zero or negative. The molecular feature code can be assigned based on a numeric index that maps a non-numeric descriptor, such a string, against a numeric value.

[0077] At 110 in FIG. 1, one or more experiment feature codes corresponding to one or more experiment features of an experiment from the one or more experiments are determined. The one or more experiment features can be metadata regarding the experiment. The one or more experiment features can include the species, e.g., biological species, from which the data derives, the sample type from which the data derives (e.g., whole blood), the relative date at which the sample was collected, or the treatment in the past duration, e.g., the treatment in the last 24 hours, including whether or not a treatment was administered at all to the subject. The experiment feature codes can be expressed as a numeric value, such as an integer or a float, and can be positive, zero or negative. The experiment feature code can be assigned based on a numeric index that maps a non-numeric descriptor, such a string or a sequence of strings, against a numeric value. The experiment being described by the experiment feature codes can comprise providing a stimulus to the one or more subjects from the first species. The stimulus can be adrug. The drug can be isoniazid, tetracyline, carbon tetrachloride, valproate, or a combination thereof. The stimulus can comprise administering a dosage of the drug that is at or above a predetermined dosage level, where the predetermined dosage level can be higher than the FDA (Food and Drug Administration)-approved drug dosage or equivalent, in animal models.

[0078] At 112 in FIG. 1, one or more vectors based on the one or more signal strength values, the one or more molecular feature codes, and the one or more experiment feature codes are created. The creating of the one or more vectors can comprise ordering the one or more signal strength values, the one or more molecular feature codes, and / or the one or more experiment feature codes, according to a method of ordering. The method of ordering can be based on unsupervised clustering. The unsupervised clustering can refer to methods of organizing data based on statistical similarities and / or dissimilarities within the data.

[0079] At 114 in FIG. 1, a generative machine learning model to provide generated data for a subject from a second species, is trained, using the created vectors. The percentage of the training data, including the created vectors, can comprise data describing the second species, e.g., humans. The percentage can be 0% or greater, and 5% or less. More specifically, the percentage can be 0%, 0.01%, 0.05%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, or 5%.

[0080] At 202 in FIG. 2, experiment data for one or more subjects from a first species are inputted into a generative machine learning model trained on a second species, the generative machine learning model having been trained by a method.

[0081] At 204 in FIG. 2, generated data is received from the trained generative machine learning model. The generated data can be a matrix or a vector that can represent biological data. The generated data need not be directly derived from a biological experiment, but can instead be derived from a statistical distribution that is generated based on biological data. The generated data can emulate or approximate data from an actual biological experiment, e.g., the generated data can be simulated data. The generated data can be simulated data approximating RNA sequencing data.

[0082] At 302 in FIG. 3, experiment data for one or more subjects from a first species are inputted into a generative machine learning model trained on a second species, the generative machine learning model having been trained by a method.

[0083] At 304 in FIG. 3, generated data is received from the trained generative machine learning model.

[0084] At 306 in FIG. 3, the generated data is inputted into a second trained machine learning model configured to predict a clinical outcome for the second species. A clinical outcome can comprise determining a response to a predetermined therapy from the subject from the second species. The predetermined therapy can be a therapy that is known in the art to be effective or efficacious, given a clinical condition of a subject. For example, the predetermined therapy can comprise administering an effective dosage of a drug that is known to be alleviate disease progression or disease status in a subject with the disease. The response to the predetermined therapy can comprise an adverse response, which can include a subject relapsing after being prescribed a therapy, such as the predetermined therapy. An adverse response can also include an allergic reaction to the therapy, such as the predetermined therapy. The response to the predetermined therapy can comprise a recovery from the predetermined therapy, such as a recovery that is the result of the predetermined therapy. For example, prescribing the predetermined therapy may result in the alleviating of a disease status or disease progression, and such alleviating may comprise the recovery. The clinical outcome can comprise determining a life expectancy of the subject from the second species. The life expectancy can be a predicted time of death for a subject. The life expectancy can be expressed as an age at which the subject dies, or the amount of time remaining before a subject dies.

[0085] At 308 in FIG. 3, the predicted clinical outcome for the second species is received from the second trained machine learning model.

[0086] FIG. 4 depicts signal strength values, e.g., functional labels, for genes that are applicable cross-species. Some of the signal strength values refer to the expression of a gene in an organ for a healthy subject, relative to other genes in the healthy subject. Signal strength values for a specific gene, for a specific organ, e.g., the heart, may be zero, indicating that theexpression in the heart for that gene is within 4-fold of median expression across all tissues. Signal strength values for a specific gene, for a specific organ, e.g., the lung, may be below zero, indicating that the gene expression in the lung for that gene is below 4-fold of median expression across all tissues. Signal strength values for a specific gene, for a specific organ, e.g., the liver, may be above zero, indicating that the gene expression in the liver for that gene is above 4-fold of median expression across all tissues. Signal strength values can correspond to the relative expression of the gene at a fixed time after the administration of the drug (e.g., 3 days).

[0087] The signal strength values can correspond to the maximum relative expression change of the gene within a number of days after the administration of the drug (e.g., strongest effect within the next 14 days). A signal strength value can be a value of zero for a specific gene, if its expression does not change by more than a factor of four, following the dosing of a drug. A signal strength value can possess a value less than zero for a specific gene, if its expression decreases by at least a factor of four, following the dosing of a drug. A signal strength value can possess a value more than zero for a specific gene, if its expression increases by at least a factor of four, following the dosing of a drug. The signal strength value can also reflect the duration of an effect on the measured amounts of a biological level, e.g., on gene expression. For example, a signal strength value for a gene can be -2, if its expression decreases by at least a factor of four for at least 8 out of 10 days, following the dosing of a drug (a sustained down-regulation response), and a signal strength value for the gene can be -1, if its expression decreases by at least a factor of four for 4 out of 10 days or fewer, following the dosing of the drug.

[0088] FIG. 5 illustrates gene mapping across two different species using signal strength values, e.g., functional labels, for genes. In FIG. 5, the training species, e.g., the first species, has N total genes and the target species, e.g., the second species, has M total genes, and in general, N M. In the absence of information, any training species gene could correspond to any target species gene. Signal strength values, e.g., functional labels, can be used to map the M genes to the N genes. In the right-most panel of FIG. 5, the training species gene M correspond to target species gene 1. In addition to the simplest case of 1:1 correspondence of genes across the two species, there can also be more complex situations where a) multiple training species genes map to a single target species gene, or 2) one training species gene maps to multiple targetspecies genes. Furthermore, the signal strength values for genes given a set of experiments, such as a set of stimuli, e.g., drugs, may be insufficient to fully resolve the mapping for some genes. For example, a lack of expression or response for different sets of genes in the training species and target species may result in insufficient mapping.

[0089] FIG. 6 provides an additional illustration of gene mapping between a training species, e.g., first species, and a target species, e.g., second species. The vector of numbers next to each gene shows the signal strength values, e.g., the functional labels, listed in a fixed pre-determined order. Four different types of gene mapping are shown: Gene T1 in the training species and gene XI in the target species are a unique perfect match, meaning that T1 and XI are the only genes across both species that exhibit the exact same pattern of signal strength values. Gene T2 and X2 are a unique closest match, in that gene T2’s signal strength values have the smallest edit distance (e.g., Levenshtein Distance) to gene X2’s signal strength values, out of all target species genes, and gene X2’s signal strength values have the smallest edit distance to T2’s signal strength values out of all training species genes. Target species gene X3 is an imperfect match to genes T3, T5, and TN in the training species, because genes T3, T5, and TN’s signal strength values all have similar or the same edit distances to X3’s signal strength values. Target species genes X4 and XM exhibit imperfect separation, because X4 and XM’s signal strength values have similar or same edit distances to gene T4’s signal strength values. It is expected that as the dimensionality of the signal strength values increase — for example, through more experiments, such as the administration of more different types of drugs, or the analysis of gene expression in more tissue types — the number of unique closest matches may increase, and all other types of matches may decrease.

[0090] FIG.7 provides one example use case, e.g., application, of cross-species transfer learning, based on signal strength values for genes. FIG. 7 depicts how a particular disease human subject (a member of the second species) would react to a particular drug treatment, Treatment A. A large amount of outcomes data on the rat training species (the first species) dosed with the same Treatment A are available, from which the following observations can be grouped: group 1) some rats’ diseases progressed with no improvement; group 2) some rats fully recovered from their disease but also suffered a severe adverse event due to the drug; and group3) some rats fully recovered from their disease with no negative impact. A generative machine learning model pre-trained on a large amount of rat biostate data (e.g., 95% or more) and a moderate amount of human biostate data (e.g., 5% or less), then fine-tuned on Treatment A may determined that the human subject best resembles rats in group 2, based on the similarities in gene expression patterns and signal strength values (e.g., functional labels) between the human subject and the rat group 2. From the above information, the generative machine learning model may then predict that the human subject would likely recover from their disease, but also suffer an adverse event.Systems

[0091] Also disclosed herein are systems designed to implement any of the disclosed methods for training one or more machine learning models to provide a clinical outcome to a subject, e.g., a second subject. The systems may comprise, e.g., one or more processors, and a memory unit communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: determine one or more signal strength values corresponding to one or more amounts of a biological molecule measured from one or more samples from one or more experiments; determine one or more molecular feature codes corresponding to one or more molecular features of the biological molecule; determine one or more experiment feature codes corresponding to one or more experiment features of an experiment from the one or more experiments; create one or more vectors based on the one or more signal strength values, the one or more molecular feature codes, and the one or more experiment feature codes; and train the generative machine learning model to provide generated data for a subject from a second species, using the created vectors. The second species can be human. The biological molecule can be an RNA molecule, and the RNA molecule can be an mRNA molecule. For the described systems, the one or more amounts can be determined based on RNA sequencing.

[0092] Also disclosed herein are systems designed to implement any of the disclosed methods, where the systems may comprise, e.g., one or more processors, and a memory unit communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: input experiment data for oneor more subjects from a first species into a generative machine learning model trained on a second species, the generative machine learning model having been trained by a method comprising: determine one or more signal strength values corresponding to one or more amounts of a biological molecule measured from one or more samples from one or more experiments; determine one or more molecular feature codes corresponding to one or more molecular features of the biological molecule; determine one or more experiment feature codes corresponding to one or more experiment features of an experiment from the one or more experiments; create one or more vectors based on the one or more signal strength values, the one or more molecular feature codes, and the one or more experiment feature codes; train the generative machine learning model to provide generated data for a subject from a second species, using the created vectors; and receive from the trained generative machine learning model, the generated data.

[0093] Also disclosed herein are systems comprising further instructions that, when executed by the one or more processors, cause the system to: input the generated data into a second trained machine learning model configured to predict a clinical outcome for the subject from the second species; and receive from the second trained machine learning model, the predicted clinical outcome for the subject from the second species.Computer systems and networks

[0094] FIG. 8 illustrates an example of a computing device or system in accordance with one embodiment. Device 800 can be a host computer connected to a network. Device 800 can be a client computer or a server. As shown in FIG. 8, device 800 can be any suitable type of microprocessor-based device, such as a personal computer, workstation, server or handheld computing device (portable electronic device) such as a phone or tablet. The device can include, for example, one or more processor(s) 810, input devices 820, output devices 830, memory or storage devices 840, communication devices 860, and nucleic acid sequencers 870. Software 850 residing in memory or storage device 840 may comprise, e.g., an operating system as well as software for executing the methods described herein. Input device 820 and output device 830 can generally correspond to those described herein, and can either be connectable or integrated with the computer.

[0095] Input device 820 can be any suitable device that provides input, such as a touch screen, keyboard or keypad, mouse, or voice-recognition device. Output device 830 can be any suitable device that provides output, such as a touch screen, haptics device, or speaker.

[0096] Storage 840 can be any suitable device that provides storage (e.g., an electrical, magnetic or optical memory including a RAM (volatile and non-volatile), cache, hard drive, or removable storage disk). Communication device 860 can include any suitable device capable of transmitting and receiving signals over a network, such as a network interface chip or device. The components of the computer can be connected in any suitable manner, such as via a wired media (e.g., a physical system bus 880, Ethernet connection, or any other wire transfer technology) or wirelessly (e.g., Bluetooth®, Wi-Fi®, or any other wireless technology).

[0097] Software module 850, which can be stored as executable instructions in storage 840 and executed by processor(s) 810, can include, for example, an operating system and / or the processes that embody the functionality of the methods of the present disclosure (e.g., as embodied in the devices as described herein).

[0098] Software module 850 can also be stored and / or transported within any non-transitory computer-readable storage medium for use by or in connection with an instruction execution system, apparatus, or device, such as those described herein, that can fetch instructions associated with the software from the instruction execution system, apparatus, or device and execute the instructions. In the context of this disclosure, a computer-readable storage medium can be any medium, such as storage 840, that can contain or store processes for use by or in connection with an instruction execution system, apparatus, or device. Examples of computer- readable storage media may include memory units like hard drives, flash drives and distribute modules that operate as a single functional unit. Also, various processes described herein may be embodied as modules configured to operate in accordance with the embodiments and techniques described above. Further, while processes may be shown and / or described separately, those skilled in the art will appreciate that the above processes may be routines or modules within other processes. 1

[0099] Software module 850 can also be propagated within any transport medium for use by or in connection with an instruction execution system, apparatus, or device, such as those described above, that can fetch instructions associated with the software from the instruction execution system, apparatus, or device and execute the instructions. In the context of this disclosure, a transport medium can be any medium that can communicate, propagate or transport programming for use by or in connection with an instruction execution system, apparatus, or device. The transport readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic or infrared wired or wireless propagation medium.

[0100] Device 800 may be connected to a network (e.g., network 904, as shown in FIG. 9 and / or described below), which can be any suitable type of interconnected communication system. The network can implement any suitable communications protocol and can be secured by any suitable security protocol. The network can comprise network links of any suitable arrangement that can implement the transmission and reception of network signals, such as wireless network connections, T1 or T3 lines, cable networks, DSL, or telephone lines.

[0101] Device 800 can be implemented using any operating system, e.g., an operating system suitable for operating on the network. Software module 850 can be written in any suitable programming language, such as C, C++, Java or Python. In various embodiments, application software embodying the functionality of the present disclosure can be deployed in different configurations, such as in a client / server arrangement or through a Web browser as a Web-based application or Web service, for example. In some embodiments, the operating system is executed by one or more processors, e.g., processor(s) 810.

[0102] Devices 800 and 906 may communicate, e.g., using suitable communication interfaces via network 904, such as a Local Area Network (LAN), Virtual Private Network (VPN), or the Internet. In some embodiments, network 904 can be, for example, the Internet, an intranet, a virtual private network, a cloud network, a wired network, or a wireless network. Devices 800 and 906 may communicate, in part or in whole, via wireless or hardwired communications, such as Ethernet, IEEE 802.11b wireless, or the like. Additionally, devices 800 and 906 may communicate, e.g., using suitable communication interfaces, via a second network, such as a mobile / cellular network. Communication between devices 800 and 906 may furtherinclude or communicate with various servers such as a mail server, mobile server, media server, telephone server, and the like. In some embodiments, Devices 800 and 906 can communicate directly (instead of, or in addition to, communicating via network 904), e.g., via wireless or hardwired communications, such as Ethernet, IEEE 802.11b wireless, or the like. In some embodiments, devices 800 and 906 communicate via communications 908, which can be a direct connection or can occur via a network (e.g., network 804).

[0103] One or all of devices 800 and 906 generally include logic (e.g., http web server logic) or are programmed to format data, accessed from local or remote databases or other sources of data and content, for providing and / or receiving information via network 904 according to various examples described herein.ENUMERATED EMBODIMENTS

[0104] The following enumerated embodiments are representative of some aspects of the invention.1. A method for training a generative machine learning model for providing generated data for a second species based on training data from a first species, comprising: conducting one or more experiments on one or more subjects from a first species; measuring one or more amounts of a biological molecule from one or more samples from the one or more experiments; determining one or more signal strength values corresponding to the one or more amounts; determining one or more molecular feature codes corresponding to one or more molecular features of the biological molecule; determining one or more experiment feature codes corresponding to one or more experiment features of an experiment from the one or more experiments; creating one or more vectors based on the one or more signal strength values, the one or more molecular feature codes, and the one or more experiment feature codes; andtraining the generative machine learning model to provide the generated data for a subject from the second species, using the created vectors. A method for generating generated data for a subject of a second species, comprising: inputting experiment data for one or more subjects from a first species into a generative machine learning model trained on a second species, the generative machine learning model having been trained by a method comprising: conducting one or more experiments on one or more subjects from the first species; measuring one or more amounts of a biological molecule from one or more samples from the one or more experiments; determining one or more signal strength values corresponding to the one or more amounts; determining one or more molecular feature codes corresponding to one or more molecular features of the biological molecule; determining one or experiment feature codes corresponding to one or more experiment features of an experiment from the one or more experiments; creating one or more vectors based on the one or more signal strength values, the one or more molecular feature codes, and the one or more experiment feature codes; training the generative machine learning model to provide the generated data for the subject from the second species, using the created vectors; and receiving from the trained generative machine learning model, the generated data. The method of embodiment 2, further comprising: inputting the generated data into a second trained machine learning model configured to predict a clinical outcome for the subject from the second species; and receiving from the second trained machine learning model, the predicted clinical outcome for the subject from the second species. The method of embodiment 2, wherein the second species is a human.The method of any of embodiments 1-4, wherein the second species is identical to the first species. The method of any of embodiments 3-5, wherein the predicted clinical outcome comprises determining a disease signal in the subject from the second species. The method of any of embodiments 3-6, wherein the predicted clinical outcome comprises determining a diagnosis for the subject from the second species. The method of any of embodiments 3-7, wherein the predicted clinical outcome comprises determining a therapy for the subject from the second species. The method of embodiment 8, wherein the determining the therapy for the subject from the second species comprises determining a dosage of the therapy. The method of embodiment 8, wherein the predicted clinical outcome comprises administering the therapy to the subject from the second species. The method any of embodiments 3-10, wherein the predicted clinical outcome comprises determining a prognosis for the subject from the second species. The method of embodiment 11, wherein the prognosis comprises determining a response to a predetermined therapy from the subject from the second species. The method of embodiment 12, wherein the response comprises an adverse response to the predetermined therapy, for the subject from the second species. The method of embodiment 12, wherein the response comprises a recovery from the predetermined therapy, for the subject from the second species. The method of any of embodiments 1-14, wherein the predicted clinical outcome comprises determining a life expectancy of the subject from the second species. The method of any of embodiments 1-15, wherein the measured one or more amounts are relative to a predetermined baseline level.The method of embodiment 16, wherein the predetermined baseline level is based on observing the biological molecule in one or more healthy subjects of the first species. The method of any of embodiments 1-17, wherein the measured one or more amounts are above a predetermined threshold amount. The method of embodiment 18, wherein the one or more signal strength values are determined when the one or more amounts are above the predetermined threshold amount. The method of any of embodiments 17-19, wherein the measured one or more amounts are greater than a first consensus amount of the biological molecule based on the observing the one or more healthy subjects of the first species. The method of embodiment 20, wherein the one or more signal strength values are determined when the measured one or more amounts are greater than the first consensus amount of the biological molecule based on the observing the one or more healthy subjects of the first species. The method of embodiment 20 or 21, wherein the measured one or more amounts are greater than the first consensus amount by more than a first multiple. The method of embodiment 22, wherein the first multiple is two. The method of any of embodiments 20-23, wherein the first consensus amount is a median amount of the biological molecule based on the observing the one or more healthy subjects of the first species. The method of any of embodiments 1-24, wherein a signal strength value from the one or more signal strength values corresponds to an amount from the one or more measured values. The method of embodiment 25, wherein the experiment comprises providing a stimulus to the one or more subjects from the first species.The method of embodiment 26, wherein the stimulus is a drug. The method of embodiment 27, wherein the drug is isoniazid, tetracycline, carbon tetrachloride, valproate, or a combination thereof. The method of any of embodiments 26-28, wherein the providing the stimulus comprises administering a dosage of the drug that is at or above a predetermined dosage level. The method of any of embodiments 1-29, wherein the one or more signal strength values are determined based on the one or more measured amounts deviating from a second consensus amount of the biological molecule in the first species, by more than a second multiple. The method of embodiment 30, wherein the second multiple is four. The method of embodiment 30 or 31, wherein the second consensus amount is a median amount of the biological molecule of the first species. The method of any of embodiments 1-32, wherein the one or more signal strength values are determined based on the one or more measured amounts deviating from the second consensus amount of the biological molecule in the first species, for more than a predetermined duration. The method of embodiment 33, wherein the predetermined duration is eight days. The method of any of embodiments 1-34, wherein the biological molecule is an RNA molecule. The method of embodiment 35, wherein the RNA molecule is an mRNA transcript. The method of any of embodiments 1-36, wherein the determining the one or more measured amounts is based on RNA sequencing. The method of any of embodiments 1-37, wherein the biological molecule is a protein molecule.The method of any of embodiments 1-38, wherein the determining the one or more measured amounts is based on mass spectrometry. The method of any of embodiments 1-39, wherein the biological molecule is an epigenetic marker. The method of any of embodiments 1-40, wherein the biological molecule is a methyl group on a nucleotide base. The method of any of embodiments 1-41, wherein the determining the one or more measured amounts is based on methyl- sequencing. The method of any of embodiments 1-42, wherein the biological molecule is a DNA molecule. The method of embodiment 43, wherein the DNA molecule is a cell-free DNA (cfDNA). The method of any of embodiments 1-44, wherein the determining the one or more measured amounts is based on DNA sequencing. The method of any of embodiments 1-45, wherein the biological molecule is a metabolite. The method of any of embodiments 1-46, wherein the one or more molecular features comprise an identifier of the biological molecule. The method of any of embodiments 1-47, wherein the one or more experiment features comprise an identifier of the first species. The method of any of embodiments 1-48, wherein the one or more experiment features comprise a description of the one or more experiments. The method of embodiment 49, wherein the description of the one or more experiments comprise a description of the one or more samples.The method of any of embodiments 1-50, wherein the one or more samples comprise liquid biopsy samples. The method of embodiment 51, wherein the liquid biopsy samples comprise blood, plasma, cerebrospinal fluid, sputum, stool, urine, sweat, or saliva. The method of any of embodiments 1-52, wherein the one or more samples comprise a tissue biopsy sample. The method of any of embodiments 1-53, wherein the one or more samples are obtained from a necropsy of the one or more subjects. The method of any of embodiments 1-54, wherein the one or more molecular feature codes correspond to the one or more molecular features according to a numeric index. The method of any of embodiments 1-55, wherein the one or more experiment feature codes correspond to the one or more experiment features according to a numeric index. The method of any of embodiments 1-56, wherein the combining comprises ordering the one or more signal strength values, the one or more molecular feature codes, or the one or more experiment feature codes, according to a method of ordering. The method of embodiment 57, wherein the method of ordering is based on an unsupervised clustering of the one or more signal strength values. The method of embodiment 57 or 58, wherein the method of ordering comprises ordering the one or more experiment feature codes before the one or more signal strength values or the one or more molecular feature codes. The method of any of embodiments 57-59, wherein the combining comprises randomly ordering the one or more signal strength values, the one or more molecular feature codes, or the one or more experiment feature codes. The method of any of embodiments 1-60, wherein a percentage of the training data comprise data describing the second species.The method of embodiment 61, wherein the percentage is 0% or greater and 5% or less. The method of embodiment 61 or 62, wherein the percentage is 1%. The method of any of embodiments 1-63, wherein the first species is a mouse, a rat, a roundworm, a non-human primate, a fruitfly, a yeast, a rabbit, or a pig. The method of any of embodiments 1-64, wherein the generative machine learning model comprises a transformer model. The method of embodiment 1-65, wherein the generative machine learning model represents a probability distribution. The method of embodiment 1-66, wherein the generative machine learning model represents the probability distribution by learning the density or mass of the probability distribution. The method of embodiment 67, wherein the learning is based on maximum likelihood estimation. The method of any of embodiments 1-68, wherein the generative machine learning model is an autoregressive model, a normalizing flow model, an energy-based model, or a variational auto-encoder. The method of any of embodiments 1-69, wherein the generative machine learning model represents the probability distribution by learning a model of the sampling process for the probability distribution. The method of any of embodiments 1-70, wherein the generative machine learning model is a generative adversarial network. The method of any of embodiment 1-71, wherein the generated data comprise generated RNA sequencing data.The method of embodiment 1-72, wherein the generated data comprise generated mass spectrometry data. The method of any of embodiments 1-73, wherein the generated data comprise generated DNA sequencing data. The method of any of embodiments 1-74, wherein the generated data comprise generated methyl-sequencing data. The method of any of embodiment 2-75, wherein the inputted experiment data comprise an input amount of an input biological molecule from an input sample from an input biological experiment. The method of any of embodiments 2-76, wherein the inputted experiment data comprise an input molecular feature code corresponding to an input molecular feature of an input biological molecule. The method of any of embodiments 75-77, wherein the inputted experiment data comprise an input experiment feature code corresponding to an input experiment feature of an input experiment. The method of embodiment 78, wherein the input experiment feature comprises an input description of the subject from the second species. The method of embodiment 79, wherein the input description of the subject from the second species comprises an input label for the second species. The method of any of embodiments 75-80, wherein the input experiment feature comprises an input label for the input sample. The method of any of embodiments 3-81, wherein the second trained machine learning model is a classifier for predicting the clinical outcome for the second species. The method of any of embodiments 3-82, wherein the second trained machine learning model uses regression to predict the clinical outcome for the second species.A system comprising: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: determine one or more signal strength values corresponding to one or more amounts of a biological molecule measured from one or more samples from one or more experiments; determine one or more molecular feature codes corresponding to one or more molecular features of the biological molecule; determine one or more experiment feature codes corresponding to one or more experiment features of an experiment from the one or more experiments; create one or more vectors based on the one or more signal strength values, the one or more molecular feature codes, and the one or more experiment feature codes; and train the generative machine learning model to provide generated data for a subject from a second species, using the created vectors. A system comprising: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: input experiment data for one or more subjects from a first species into a generative machine learning model trained on a second species, the generative machine learning model having been trained by a method comprising: determine one or more signal strength values corresponding to one or more amounts of a biological molecule measured from one or more samples from one or more experiments; determine one or more molecular feature codes corresponding to one or more molecular features of the biological molecule;determine one or more experiment feature codes corresponding to one or more experiment features of an experiment from the one or more experiments; create one or more vectors based on the one or more signal strength values, the one or more molecular feature codes, and the one or more experiment feature codes; train the generative machine learning model to provide generated data for a subject from a second species, using the created vectors; and receive from the trained generative machine learning model, the generated data. The system of embodiment 85, comprising further instructions that, when executed by the one or more processors, cause the system to: input the generated data into a second trained machine learning model configured to predict a clinical outcome for the subject from the second species; and receive from the second trained machine learning model, the predicted clinical outcome for the subject from the second species. The system of any of embodiments 84-86, wherein the second species is a human. The system of any of embodiments 84-87, wherein the second species is identical to the first species. The system of any of embodiments 84-88, wherein the biological molecule is an RNA molecule. The system of embodiment 89, wherein the RNA molecule is an mRNA molecule. The system of any of embodiments 84-90, wherein determining the one or more amounts is based on RNA sequencing. The system of any of embodiments 84-91, wherein a percentage of the training data comprise data describing the second species. The system of embodiment 92, wherein the percentage is 0% or greater and 5% or less.The system of embodiment 92 or 93, wherein the percentage is 1%. The system of any of embodiments 84-94, wherein the one or more generative machine learning models comprises a transformer model. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of a system, cause the system to: determine one or more signal strength values corresponding to one or more amounts of a biological molecule measured from one or more samples from one or more experiments; determine one or more molecular feature codes corresponding to one or more molecular features of the biological molecule; determine one or more experiment feature codes corresponding to one or more experiment features of an experiment from the one or more experiments; create one or more vectors based on the one or more signal strength values, the one or more molecular feature codes, and the one or more experiment feature codes; and train the generative machine learning model to provide generated data for a subject from a second species, using the created vectors. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of a system, cause the system to: input experiment data for one or more subjects from a first species into a generative machine learning model trained on a second species, the generative machine learning model having been trained by a method comprising: determine one or more signal strength values corresponding to one or more amounts of a biological molecule measured from one or more samples from one or more experiments; determine one or more molecular feature codes corresponding to one or more molecular features of the biological molecule;determine one or more experiment feature codes corresponding to one or more experiment features of an experiment from the one or more experiments; create one or more vectors based on the one or more signal strength values, the one or more molecular feature codes, and the one or more experiment feature codes; train the generative machine learning model to provide generated data for a subject from a second species, using the created vectors; and receive from the trained generative machine learning model, the generated data. The non-transitory computer-readable storage medium of embodiment 97, further comprising instructions that, when executed by the one or more processors, cause the system to: input the generated data into a second trained machine learning model configured to predict a clinical outcome for the subject from the second species; and receive from the second trained machine learning model, the predicted clinical outcome for the subject from the second species. The non-transitory computer-readable storage medium of any of embodiments 96-98, wherein the second species is a human. . The non-transitory computer-readable storage medium of any of embodiments 96-99, wherein the second species is identical to the first species. . The non-transitory computer-readable storage medium of any of embodiments 96-100, wherein the biological molecule is an RNA molecule. . The non-transitory computer-readable storage medium of embodiment 101, wherein the RNA molecule is an mRNA molecule. . The non-transitory computer-readable storage medium of any of embodiments 96- 102, wherein determining the one or more amounts is based on RNA sequencing.104. The non-transitory computer-readable storage medium of any of embodiments 96- 103, wherein a percentage of the training data comprise data describing the second biological species.105. The non-transitory computer-readable storage medium of embodiment 104, wherein the percentage is 0% or greater and 5% or less.106. The non-transitory computer-readable storage medium of embodiment 104 or105, wherein the percentage is 1%.107. The non-transitory computer-readable storage medium of any of embodiments 96-106, wherein the one or more generative machine learning models comprises a transformer model.EXAMPLES

[0105] The following examples further demonstrate to one skilled in the art how to make and use the methods and systems described herein, and are not intended to limit the scope of the claimed invention.

[0106] The data for the following examples were derived according to the following:1. Animal study design

[0107] Rats were acclimatized prior to an experiment and were housed in enriched and ventilated housing cages throughout the experimental phase. Cage litter was changed at least once a week. The rats were housed under specific-pathogen-free (SPF) conditions, and were subjected to a normal 12 hour light and 12 hour night cycle, at 22 ± 2°C and 50 ± 10 % relative humidity. Chow and water were available ad libitum. The health conditions were examined and recorded daily by a veterinarian. 15 male SD rats and 15 female SD rats were housed for 5 days with daily 200 pL blood extraction. 15 male C57BL / 6 mice and 15 female kc57BL / 6 mice were housed for 5 days with daily 50 pL blood extraction. Rats and mice in group 1 (the control group) were intraperitoneally injected with 0.9% sodium chloride on day 3. Rats and mice in group 2 (an experimental group) were intraperitoneally injected with tetracyline dissolved insterile saline on day 3. Rats and mice in group 3 (an experimental group) were intraperitoneally injected with valproic acid diluted in sterile saline on day 3. Rats and mice in group 4 (an experimental group) were administered a 1:1 mixture of com or olive oil and carbon tetrachloride (CC14) orally on day 3. Rats and mice in group 5 (an experimental group) were intraperitoneally injected with isoniazid dissolved in sterile saline on day 3. All rats and mice were sacrificed on day 5. A detailed summary of the animal study design is shown in Table 1 below.Table 1. Animal study design2. Bio-samples collection and storage

[0108] 200 pL of whole blood was collected daily from the same jugular vein site of each rat. 50 pL of whole blood was collected daily from the same jugular vein site of each mouse.

[0109] Whole blood samples were immediately processed for RNA extraction. On the first day of the study, rats were fasted overnight and then euthanized by cervical dislocation. Wholeblood was immediately collected via heart puncture under RNase-free conditions. 1 mL of whole blood underwent the standard peripheral blood mononuclear cell (PBMC) separation process, and the rest was preserved at -80 °C for RNA extraction.

[0110] On the end day of the study, liver, left and right kidneys, lung, heart, back skin, duodenum, brain, pancreas, and reproductive organs (testis for males and ovaries for females) were collected, and underwent a standard tissue RNA extraction process. The rest were preserved at -80 °C.3. RNA extraction from whole blood samples and liver samples

[0111] For RNA extraction from whole blood, 100 pL of whole blood was mixed with 700 pL of TRIzol reagent. The resulting mixture then underwent a standard TRIzol and chloroform RNA extraction method. Extracted RNA was further purified using an Automatic Nucleotide Isolation Machine. For RNA extraction from liver samples, around 30 mg of chopped liver tissue was mixed with 700 pL of TRIzol reagent and Lysing MatrixD. The resulting mixture was ground for 2 minutes and then mixed with 140 pL of chloroform. After 2 minutes of incubation at room temperature and 10 minutes of centriguation at 12000 rpm at 4 °C, the supernatant was transferred to a tube and was further purified with MagaBio Plus Total RNA Purification kit. The purified RNA was quantified with a Nanodrop 2000 (ThermoFisher, USA). The quality of RNA was measured by using RNA ScreenTape Assay (Agilent, USA) and 4200 Tapestation System (Agilent, USA).4. mRNA Capture and Release

[0112] VAHTS mRNA Capture Beads was used for the enrichment of mRNA from total RNA extracted from blood and tissue samples. During the mRNA capture and release process, mRNA was fragmented.5. cDNA Synthesis and Library Preparation

[0113] Library preparation of the mRNA included three steps: 1) reverse transcription, 2) adaptor ligation, and 3) index PCR amplification. Reverse transcription and adaptor ligation wasperformed with a standard protocol. Index PCR amplification was performed with standard protocol. One round of 1.6 x beads purification was then performed after the adaptor ligation. One round of 1.6 x beads purification was then performed after the index PCR amplification. Library DNA was quantified, and then analyzed.6. Sequencing

[0114] 2 x 150 paired-end sequencing was performed.7. Alignment

[0115] For rat data, the FASTQ file was initially trimmed to remove Illumina adaptor sequences, and subsequently aligned with the rat (Rattus norvegicu.s) reference genome (NCBI GCF_015227675.2, genome assembly mRatBN7.2) using HISAT2 (default parameters). This process generated a raw gene expression file.

[0116] For mouse data, the FASTQ file was initially trimmed to remove Illumina adaptor sequences, and then subsequently aligned with the C57BL / 6J mouse reference genome (NCBI GCF_000001635.27, genome assembly GRCm39) using HISAT2 (default parameters). This process generated a raw gene expression file.8. Normalizing Reads

[0117] To minimize differences in sequence depth across various samples, normalization was performed as follows:

[0118] 1. Calculate the total reads for each sample.

[0119] 2. Adjust the total reads to 1 million reads for each sample.

[0120] 3. Back-calculate the reads per gene in each sample.

[0121] Normalized reads are referred to as reads per million total reads (RPM).Example 1

[0122] FIG. 10 illustrates the experimental setup in Sprague-Dawley rats for measuring the functional labels, e.g., the signal strength values and their represented experiment features, of genes. The functional labels are based on gene expression changes in the rats. 200 pL of blood was collected from each rat every day. A drug of interest was administered to the rat just prior to the Day 3 blood collection. Blood continued to be collected from the rat on a daily basis for a number of days before the experiment ends. Sample experimental results for two drugs (tetracycline and isoniazid) were provided. The three traces in each figure represent three separate rat subjects. The Ahil gene's signal strength value, e.g., functional label, from responding to drugs is 0 (no change) for isoniazid and +1 (up-regulated) for tetracycline. The Arfl gene's signal strength value, e.g., functional label, from responding to drugs is 0 (no change) for isoniazid and -1 (down-regulated) for tetracycline. The Chd8 gene's signal strength value, e.g., function label, from responding to drugs is -1 (down-regulated) for isoniazid and 0 (no change) for tetracycline.

[0123] FIG.10 summarizes in a Venn diagram, the initial experimental findings for the rat models’ gene expression responses to the four drugs: tetracycline, isoniazid, carbon tetrachloride, and valproate. FIG. 11 indicates the number of genes that respond to the drugs, including both up-regulation and down-regulation. For example, a total of 3,206 genes were responsive to tetracycline and have signal strength value other than zero. Of these, 1,323 genes were responsive only to tetracycline but not to isoniazid, carbon tetrachloride, or valproate. These genes would have signal strength values, e.g., functional labels, of (+1, 0, 0, 0) or (-1, 0, 0, 0). For the four drug molecules tested, sustained up-regulation or down-regulation responses were not observed for any genes. From the result that every single sector of the Venn diagram in FIG.10 has a non-zero number of genes, the testing of additional drugs would likely improve the separation of genes by function, and result in improved gene mapping between a training species and a target species.

[0124] FIG. 12 illustrates an example tokenization process for a biostate as input to a generative machine learning model, such as a transformer-based deep neural network. A biostate can be considered to be a quantitative high-dimensional biological state of an organism at aparticular point in time. The biostate can comprise gene expression data, such as the signal strength values, e.g., functional labels, for genes in an experiment. Each biostate is inputted into the generative machine learning model by first starting with a global token, e.g., experiment feature codes, that describes the species, the specimen type, the date of the one or more experiments for the biostate, and any recent treatments, e.g., drug administrations, since the previous biostate. Additional subject information such as sex, age, body weight, genetic variants of interest, and disease states can also be incorporated into the global token. In some implementations, such subject information may not be incorporated into the global token, e.g., experiment feature codes, because they can be inferred from the biostate information — for example, female- specific gene expression may indicate female sex.

[0125] A number of gene tokens, e.g., molecular feature codes and their corresponding signal strength values, can follow the global token, e.g., the experiment feature codes, such that one gene token corresponds to a gene. Each gene token can minimally comprise the gene’s molecular feature codes and signal strength values, including the gene’s expression in a sample. The gene expression can be expressed in log2 units normalized to a total RNA expression or reads. The molecular feature code can additionally include the conventional name of the gene. The molecular feature code can include the genetic coordinates of the gene, including the chromosome and nucleotide start and end positions.

[0126] In transformer-based neural networks, attention drops off quadratically, based on the distance between two tokens. Consequently, selecting which genes to include as gene tokens, as well as deciding the order when inputting the gene tokens after the global token, into the generative machine learning model, can have a significant impact on the amount of data needed to train the machine learning model to a particular level of performance. Naive approaches of ordering, such as including all genes as gene token in alphabetical order by gene name can be used, but may be highly suboptimal. Other neural network models based on state-space models, e.g., Mamba, can exhibit linear drop-offs of attention based on the distance between tokens, which can significantly improve the length of the model’s context windows. Regardless, a method of ordering tokens, e.g., the signal strength values, the molecular feature codes, and the experiment feature codes, is important for the efficient and / or performative training of thegenerative machine learning model. The ordering of the gene tokens, e.g., molecular feature codes and signal strength values for the genes, can be based on an unsupervised clustering analysis of the signal strength values, e.g., functional labels, of all the genes. This approach would more closely group genes that are more functionally similar to each other, which can result in easier machine learning training using smaller datasets. The ordering of the gene tokens can be determined empirically by assessing the performance of the generative machine learning model with a fixed sized training dataset.

[0127] Only genes with expression above a threshold can be tokenized as gene tokens, and all genes with expression below the threshold can be omitted. Based on the experiments described herein, only about 10000 of the 25 000 genes in the rat genome are expressed at significant levels above the RNA sequencing noise floor, without incurring dramatic increases in sequencing costs. Consequently, roughly 60% of the genes need be omitted, leading to biostates with a significantly smaller number of gene tokens.

[0128] Gene tokens can be created for only genes with expression different by more than a factor of two from the median expression level of the gene in healthy animals from the same species and sample type. The vast majority of healthy biostates can have a small number of gene tokens, including their molecular feature codes, such as possibly zero gene tokens. The advantage of such an approach is that biostates are represented in a highly parsimonious manner, resulting in a few tokens per biostate. The disadvantage of such an approach is that information from many weaker signals (e.g., 1.3x or 0.8x median expression) could be lost.Example 2

[0129] FIGS. 13A-13F provide examples of data for various genes in mice (Mus musculus), including their gene expression values, and their corresponding signal strength values, e.g., functional labels. For each FIGS. 13A-F, the left side shows a table of the gene expression levels for the indicated gene, for three drugs — tetracyline, valproate, or carbon tetrachloride (CC14) — for four organs — liver, lung, kidney, or heart. More specifically, the indicated gene expression is the mean gene expression change in log2 units after for n=6 samples. For each figure in FIGS. 13A-F, the right side shows a table of the signal strength values corresponding to the geneexpression values depicted on the left side of the figure, for each of the same three drugs and four organs (tetracyline, valprorate, and CC14, for the drugs, and liver, lung, kidney, and heart, for the organs). The signal strength values and their represented experimental conditions, e.g., the administered drugs, and the organs from which the gene expression values were measured, comprise functional labels for the indicated gene.

[0130] The indicated gene in FIG. 13A is the 4930507D05Rik gene in mice. The expression of the 4930507D05Rik gene in the liver was affected by all three of the tested drugs. The indicated gene in FIG. 13B is the Yipf7 gene in mice. The expression of the Yipf7 gene in the lung was affected by all three of the tested drugs. The indicated gene in FIG. 13C is the 170001 lL22Rik gene in mice. The expression of the 170001 lL22Rik gene in the kidney was affected by all three of the tested drugs. The indicated gene in FIG. 13D is the WntlOa gene in mice. The expression of the WntlOa gene in the heart was affected by all three of the tested drugs. The indicated gene in FIG. 13E is the Capnl 1 gene in mice. The expression of the Capnl 1 gene is affected by tetracyline for all four organs. The indicated gene in FIG. 13F is the C730014E05Rik gene in mice. The expression of the C730014E05Rik gene is affected by CC14 for all four organs.

[0131] FIG. 14 summarizes the functional labels, e.g., the gene expression responses, for all measured genes, for the three drugs — tetracyline, valproate, CC14 — for the four organs — liver, lung, kidney, heart — in mice. FIG. 14A shows the number of genes that showed significant changes in gene expression in any organ, e.g., an increase or a decrease in gene expression, relative to the median expression of that gene when no drugs were provided, in response to the three drugs, and thus have non-zero values for their signal strength values. According to FIG. 14A, 3572 genes responded to tetracyline, of which 1216 genes responded to only tetracyline, 606 genes responded to valproate and tetracyline, 869 genes responded to tetracyline and CC14, and 881 genes responded to all three of the drugs. Furthermore, according to FIG. 14A, 2972 genes responded to valproate, of which 899 genes responded to only valproate, and 586 genes responded to both valproate and CC14. Furthermore, according to FIG. 14A, 4494 genes responded to CC14, of which 2158 genes responded to only CC14. FIG. 14B shows the number of genes that showed significant changes in gene expression, e.g., an increase or decrease in geneexpression, relative to the median expression of that gene when no drugs were provided, in the four organs, in response to any of the three drugs, and thus have non-zero values for their signal strength values.Example 3

[0132] FIGS. 15A-15H provides examples of data for various genes in rats (Rattus norvegicus), including their gene expression values, and their corresponding signal strength values, e.g., functional labels. For each figure in FIGS. 15A-H, the left side shows a table of the gene expression levels for the indicated gene, for four drugs — tetracyline, valproate, carbon tetrachloride (CC14), or iosniazid — for four organs — liver, lung, kidney, or heart. More specifically, the indicated gene expression is the mean gene expression change in log2 units after for n=6 samples. For each figure in FIGS. 15A-H, the right side shows a table of the signal strength values corresponding to the gene expression values depicted on the left side of the figure, for each of the same three drugs and four organs (tetracyline, valprorate, CC14, and isoniazid, for the drugs, and liver, lung, kidney, and heart, for the organs). The signal strength values and their represented experimental conditions, e.g., the administered drugs, and the organs from which the gene expression values were measured, comprise functional labels for the indicated gene.

[0133] The indicated gene in FIG. 15A is the Bexl gene in rats. The expression of the Bexl gene in the liver was affected by all four of the tested drugs. The indicated gene in FIG. 15B is the Texl 1 gene in rats. The expression of the Texl 1 gene in the lung was affected by all four of the tested drugs. The indicated gene in FIG. 15C is the Add2 gene in rats. The expression of the Add2 gene in the kidney was affected by all four of the tested drugs. The indicated gene in FIG. 15D is the Vill gene in rats. The expression of the Vill gene in the heart was affected by all four of the tested drugs. The indicated gene in FIG. 15E is the Hba-a3 gene in rats. The expression of the Hba-a3 gene was affected by tetracyline for all four organs. The indicated gene in FIG. 15F is the RGD1565566 gene in rats. The expression of the RGD1565566 gene was affected by tetracycline and valproate for all four organs. The indicated gene in FIG. 15G is the Mybpc3 gene in rats. The expression of the Mybpc3 gene was affected by CC14 for all four organs. Theindicated gene in FIG. 15H is the LGC120100962 gene in rats. The expression of the LOCI 20100962 gene was affected by isoniazid for all four organs.

[0134] FIG. 16 summarizes the functional labels, e.g., the gene expression responses, for all measured genes, for the four drugs — tetracyline, valproate, CC14, isoniazid — for the four organs — liver, lung, kidney, heart — in mice. FIG. 16A shows the number of genes that showed significant changes in gene expression in any organ, e.g., an increase or a decrease in gene expression, relative to the median expression of that gene when no drugs were provided, in response to the four drugs, and thus have non-zero values for their signal strength values. FIG. 16B shows the number of genes that showed significant changes in gene expression, e.g., an increase or decrease in gene expression, relative to the median expression of that gene when no drugs were provided, in the four organs, in response to any of the three drugs, and thus have nonzero values for their signal strength values.Example 4

[0135] FIGS. 17A-17C provide an example of data where two different genes between two different species exhibit a unique perfect match. A unique perfect match occurs when only a single pair of genes across species — e.g., a gene for each species — exhibits the exact same functional labels. That is, the genes have the exact same signal strength values, for the same represented experiment conditions. FIG. 17A provides an example of data where the functional labels for the Apol7a gene in mice are identical to the functional labels for the Aacs gene in rats. FIG. 17B provides an example of data where the functional labels for the Mucl6 gene in mice are identical to the functional labels for the Dusp29 genes in rats. FIG. 17C provides an example of data where the functional labels for the Disp2 gene in mice are identical to the functional labels for the LOCI 20100015 gene in rats.

[0136] It should be understood from the foregoing that, while particular implementations of the disclosed methods and systems have been illustrated and described, various modifications can be made thereto and are contemplated herein. It is also not intended that the invention be limited by the specific examples provided within the specification. While the invention has been described with reference to the aforementioned specification, the descriptions and illustrations ofthe preferable embodiments herein are not meant to be construed in a limiting sense.Furthermore, it shall be understood that all aspects of the invention are not limited to the specific depictions, configurations or relative proportions set forth herein which depend upon a variety of conditions and variables. Various modifications in form and detail of the embodiments of the invention will be apparent to a person skilled in the art. It is therefore contemplated that the invention shall also cover any such modifications, variations and equivalents.

Claims

CLAIMSWhat is claimed is:

1. A method for training a generative machine learning model for providing generated data for a second species based on training data from a first species, comprising: conducting one or more experiments on one or more subjects from a first species; measuring one or more amounts of a biological molecule from one or more samples from the one or more experiments; determining one or more signal strength values corresponding to the one or more amounts; determining one or more molecular feature codes corresponding to one or more molecular features of the biological molecule; determining one or more experiment feature codes corresponding to one or more experiment features of an experiment from the one or more experiments; creating one or more vectors based on the one or more signal strength values, the one or more molecular feature codes, and the one or more experiment feature codes; and training the generative machine learning model to provide the generated data for a subject from the second species, using the created vectors.

2. A method for generating generated data for a subject of a second species, comprising: inputting experiment data for one or more subjects from a first species into a generative machine learning model trained on a second species, the generative machine learning model having been trained by a method comprising: conducting one or more experiments on one or more subjects from the first species; measuring one or more amounts of a biological molecule from one or more samples from the one or more experiments; determining one or more signal strength values corresponding to the one or more amounts; determining one or more molecular feature codes corresponding to one or more molecular features of the biological molecule;determining one or experiment feature codes corresponding to one or more experiment features of an experiment from the one or more experiments; creating one or more vectors based on the one or more signal strength values, the one or more molecular feature codes, and the one or more experiment feature codes; training the generative machine learning model to provide the generated data for the subject from the second species, using the created vectors; and receiving from the trained generative machine learning model, the generated data.

3. The method of claim 2, further comprising: inputting the generated data into a second trained machine learning model configured to predict a clinical outcome for the subject from the second species; and receiving from the second trained machine learning model, the predicted clinical outcome for the subject from the second species.

4. The method of claim 2, wherein the second species is a human.

5. The method of any of claims 1-4, wherein the second species is identical to the first species.

6. The method of any of claims 3-5, wherein the predicted clinical outcome comprises one or more of the following: determining a disease signal in the subject from the second species; determining a diagnosis for the subject from the second species; determining a therapy for the subject from the second species; determining a prognosis for the subject from the second species; and determining a life expectancy of the subject from the second species.

7. The method of any of claims 1-6, wherein the measured one or more amounts are:relative to a predetermined baseline level, optionally wherein the predetermined baseline level is based on observing the biological molecule in one or more healthy subjects of the first species; above a predetermined threshold amount, optionally wherein the one or more signal strength values are determined when the one or more amounts are above the predetermined threshold amount; greater than a first consensus amount of the biological molecule based on the observing the one or more healthy subjects of the first species, optionally wherein the one or more signal strength values are determined when the measured one or more amounts are greater than the first consensus amount of the biological molecule based on the observing the one or more healthy subjects of the first species; greater than the first consensus amount by more than a first multiple, optionally wherein the first multiple is two8. The method of claim 7, wherein the first consensus amount is a median amount of the biological molecule based on the observing the one or more healthy subjects of the first species.

9. The method of any of claims 1-8, wherein a signal strength value from the one or more signal strength values corresponds to an amount from the one or more measured values.

10. The method of claim 9, wherein the experiment comprises providing a stimulus to the one or more subjects from the first species.

11. The method of claim 10, wherein the stimulus is a drug.

12. The method of claim 11, wherein the drug is isoniazid, tetracycline, carbon tetrachloride, valproate, or a combination thereof.

13. The method of claim 11 or 12, wherein the providing the stimulus comprises administering a dosage of the drug that is at or above a predetermined dosage level.

14. The method of any of claims 1-13, wherein the one or more signal strength values are determined based on the one or more measured amounts deviating from a second consensus amount of the biological molecule in the first species, by more than a second multiple, optionally wherein the second multiple is four, and optionally wherein the second consensus amount is a median amount of the biological molecule of the first species; or determined based on the one or more measured amounts deviating from the second consensus amount of the biological molecule in the first species, for more than a predetermined duration, optionally wherein the predetermined duration is eight days.

15. The method of any of claims 1-14, wherein the biological molecule is an RNA molecule.

16. The method of claim 15, wherein the RNA molecule is an mRNA transcript.

17. The method of any of claims 1-16, wherein the determining the one or more measured amounts is based on RNA sequencing.

18. The method of any of claims 1-17, wherein the biological molecule is: a protein molecule; an epigenetic marker; a methyl group on a nucleotide base; a DNA molecule, optionally wherein the DNA molecule is a cell-free DNA (cfDNA); or a metabolite.

19. The method of any of claims 1-18, wherein the determining the one or more measured amounts is based on: mass spectrometry; methyl-sequencing; orDNA sequencing.

20. The method of any of claims 1-19, wherein the one or more molecular features comprise an identifier of the biological molecule.

21. The method of any of claims 1-20, wherein the one or more experiment features comprise: an identifier of the first species; a description of the one or more experiments; or a description of the one or more samples.

22. The method of any of claims 1-21, wherein the one or more samples comprise: liquid biopsy samples, optionally wherein the liquid biopsy samples comprise blood, plasma, cerebrospinal fluid, sputum, stool, urine, sweat, or saliva; or a tissue biopsy sample.

23. The method of any of claims 1-22, wherein the one or more molecular feature codes correspond to the one or more molecular features according to a numeric index.

24. The method of any of claims 1-23, wherein the one or more experiment feature codes correspond to the one or more experiment features according to a numeric index.

25. The method of any of claims 1-24, wherein the combining comprises ordering the one or more signal strength values, the one or more molecular feature codes, or the one or more experiment feature codes, according to a method of ordering.

26. The method of claim 25, wherein the method of ordering: is based on an unsupervised clustering of the one or more signal strength values; or comprises ordering the one or more experiment feature codes before the one or more signal strength values or the one or more molecular feature codes.

27. The method of claim 25 or 26, wherein the combining comprises randomly ordering the one or more signal strength values, the one or more molecular feature codes, or the one or more experiment feature codes.

28. The method of any of claims 1-27, wherein a percentage of the training data comprise data describing the second species. optionally, wherein the percentage is 0% or greater and 5% or less. optionally, wherein the percentage is 1%.

29. The method of any of claims 1-28, wherein the first species is a mouse, a rat, a roundworm, a non-human primate, a fruitfly, a yeast, a rabbit, or a pig.

30. The method of any of claims 1-29, wherein the generative machine learning model: comprises a transformer model; represents a probability distribution; represents the probability distribution by learning the density or mass of the probability distribution, optionally wherein the learning is based on maximum likelihood estimation;is an autoregressive model, a normalizing flow model, an energy-based model, or a variational auto-encoder; represents the probability distribution by learning a model of the sampling process for the probability distribution; and / or is a generative adversarial network.

31. The method of any of claim 1-30, wherein the generated data comprise: generated RNA sequencing data; generated mass spectrometry data; generated DNA sequencing data; and / or generated methyl- sequencing data.

32. The method of any of claim 2-31, wherein the inputted experiment data comprise: an input amount of an input biological molecule from an input sample from an input biological experiment; an input molecular feature code corresponding to an input molecular feature of an input biological molecule; and / or an input experiment feature code corresponding to an input experiment feature of an input experiment, optionally wherein the input experiment feature comprises an input description of the subject from the second species, optionally wherein the input description of the subject from the second species comprises an input label for the second species; or optionally wherein the input experiment feature comprises an input label for the input sample.

33. The method of any of claims 3-32, wherein the second trained machine learning model: is a classifier for predicting the clinical outcome for the second species; and / or the second trained machine learning model uses regression to predict the clinical outcome for the second species.

34. A system comprising: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: determine one or more signal strength values corresponding to one or more amounts of a biological molecule measured from one or more samples from one or more experiments; determine one or more molecular feature codes corresponding to one or more molecular features of the biological molecule; determine one or more experiment feature codes corresponding to one or more experiment features of an experiment from the one or more experiments; create one or more vectors based on the one or more signal strength values, the one or more molecular feature codes, and the one or more experiment feature codes; and train the generative machine learning model to provide generated data for a subject from a second species, using the created vectors.

35. A system comprising: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to:input experiment data for one or more subjects from a first species into a generative machine learning model trained on a second species, the generative machine learning model having been trained by a method comprising: determine one or more signal strength values corresponding to one or more amounts of a biological molecule measured from one or more samples from one or more experiments; determine one or more molecular feature codes corresponding to one or more molecular features of the biological molecule; determine one or more experiment feature codes corresponding to one or more experiment features of an experiment from the one or more experiments; create one or more vectors based on the one or more signal strength values, the one or more molecular feature codes, and the one or more experiment feature codes; train the generative machine learning model to provide generated data for a subject from a second species, using the created vectors; and receive from the trained generative machine learning model, the generated data.

36. The system of claim 35, comprising further instructions that, when executed by the one or more processors, cause the system to: input the generated data into a second trained machine learning model configured to predict a clinical outcome for the subject from the second species; and receive from the second trained machine learning model, the predicted clinical outcome for the subject from the second species.

37. The system of any of claims 34-36, wherein the second species is: a human; and / or identical to the first species.

38. The system of any of claims 34-37, wherein the biological molecule is an RNA molecule,optionally, wherein the RNA molecule is an mRNA molecule.

39. The system of any of claims 34-38, wherein determining the one or more amounts is based on RNA sequencing.

40. The system of any of claims 34-39, wherein a percentage of the training data comprise data describing the second species, optionally wherein the percentage is 0% or greater and 5% or less; or optionally wherein the percentage is 1%.

41. The system of any of claims 34-40, wherein the one or more generative machine learning models comprises a transformer model.

42. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of a system, cause the system to: determine one or more signal strength values corresponding to one or more amounts of a biological molecule measured from one or more samples from one or more experiments; determine one or more molecular feature codes corresponding to one or more molecular features of the biological molecule; determine one or more experiment feature codes corresponding to one or more experiment features of an experiment from the one or more experiments; create one or more vectors based on the one or more signal strength values, the one or more molecular feature codes, and the one or more experiment feature codes; and train the generative machine learning model to provide generated data for a subject from a second species, using the created vectors.

43. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of a system, cause the system to:input experiment data for one or more subjects from a first species into a generative machine learning model trained on a second species, the generative machine learning model having been trained by a method comprising: determine one or more signal strength values corresponding to one or more amounts of a biological molecule measured from one or more samples from one or more experiments; determine one or more molecular feature codes corresponding to one or more molecular features of the biological molecule; determine one or more experiment feature codes corresponding to one or more experiment features of an experiment from the one or more experiments; create one or more vectors based on the one or more signal strength values, the one or more molecular feature codes, and the one or more experiment feature codes; train the generative machine learning model to provide generated data for a subject from a second species, using the created vectors; and receive from the trained generative machine learning model, the generated data.

44. The non-transitory computer-readable storage medium of claim 43, further comprising instructions that, when executed by the one or more processors, cause the system to: input the generated data into a second trained machine learning model configured to predict a clinical outcome for the subject from the second species; and receive from the second trained machine learning model, the predicted clinical outcome for the subject from the second species.

45. The non-transitory computer-readable storage medium of any of claims 42-44, wherein the biological molecule is an RNA molecule, optionally wherein the RNA molecule is an mRNA molecule.

46. The non-transitory computer-readable storage medium of any of claims 42-45, wherein determining the one or more amounts is based on RNA sequencing.

47. The non-transitory computer-readable storage medium of any of claims 42-46, wherein the one or more generative machine learning models comprises a transformer model.

Citation Information

Patent Citations

  • Workflow for generating compounds with biological activity against a specific biological target

    US20210057050A1

  • Inter-model prediction score recalibration during training

    US20230207064A1

  • Integrated host-microbe metagenomics of cell-free nucleic acid for sepsis diagnosis

    WO2023224913A1