Methods and systems of generative artificial intelligence for predicting biomolecule amounts

A generative AI model trained on functional biomolecule amounts addresses the limitations of structural data-dependent methods by improving the prediction of biomolecule levels and clinical outcomes across species and anatomical regions.

WO2025179196A1PCT designated stage Publication Date: 2025-08-28BIOSTATE AI INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/016872
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-21
Filing Date
2025-02-21
Publication Date
2025-08-28

AI Technical Summary

Technical Problem

Existing methods for predicting biomolecule amounts across biological contexts, such as gene expression levels, often fail due to structural biological differences across individual subjects and species, limiting their accuracy and foresight in predicting organismal functions.

Method used

A generative artificial intelligence (AI) model is trained using functional biological data, such as measured biomolecule amounts, excluding structural data like protein structures, to predict biomolecule amounts personalized to a subject's physiological system or anatomical region, which can then be used to inform clinical outcomes.

Benefits of technology

Improves the accuracy of predicting biomolecule amounts and clinical outcomes by leveraging functional biological data, reducing the need for invasive sampling and enhancing the predictive power of machine learning models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025016872_28082025_PF_FP_ABST
    Figure US2025016872_28082025_PF_FP_ABST
Patent Text Reader

Abstract

Methods for training and deploying a generative artificial intelligence (Al) model for predicting biomolecule amounts is described. The methods may comprise, for example, a method of training a generative Al model based on training data from a training biological species, and a method for deploying the trained generative Al model to generate generated data for a target biological species. The methods may also comprise, for example, a method of inferring target amounts of a first biomolecule in a target anatomical region.
Need to check novelty before this filing date? Find Prior Art

Description

METHODS AND SYSTEMS OF GENERATIVE ARTIFICIAL INTELLIGENCE FORPREDICTING BIOMOLECULE AMOUNTSCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 556,265, filed February 21, 2024, which is incorporated herein by reference in its entirety.FIELD

[0002] The present disclosure relates generally to methods and systems of generative artificial intelligence for predicting biomolecular states. More specifically, the present disclosure relates to training and deploying a machine learning model for predicting biomolecule amounts for various biological contexts, including across biological species, across anatomical regions within a biological species, and / or based on training data exclusionary of protein structure data.BACKGROUND

[0003] Biologists have long appealed to biological structures, such as sequences, to predict biological functions. For example, nucleic acid sequences are often used to predict protein functions, and even cellular or organ-level functions, among other higher order features. Such predictions, however, are often limited in their accuracy and foresight. Biological structures often lack obvious correlations with the functional aspects of the organism, especially as those functional aspects relate more broadly to organismal function, such as physiology-specific biomolecule expression programs and their associated amounts of biomolecules, e.g., gene expression programs and their associated transcripts. Improved methods are needed for predicting biological function in a subject. The present disclosure addresses these needs.BRIEF SUMMARY OF THE INVENTION

[0004] Disclosed herein are methods and systems for predicting biomolecule amounts across biological contexts. Existing methods for predicting biomolecule amounts across biological contexts often leverage structural biological data. Such methods, however, often fail. The failure may result from the numerous structural biological differences that exist across individual subjects and / or species, such as differences in DNA sequences and / or differences in proteinstructures. The methods and systems described herein use a generative artificial intelligence (Al) model to predict amounts of biomolecules from a subject. For example, the methods and system described herein can be used to predict gene expression amounts personalized to a physiological system or an anatomical region of the subject. In some instances, predicting the biomolecule amounts can be for a tissue for which procuring a biological sample is invasive. The generative Al model can be trained on training data relating to measured amounts of biomolecules, such as gene expression data. The training data can derive from physiologies for which samples are readily accessible. Training the generative Al model can comprise determining a summary value based on the measured amounts of biomolecules. The training data can be exclusionary of structural data, such as protein structure data. The generated data generated from the generative Al model can be used as inputs to other machine learning models, e.g., classifiers, to provide clinical information regarding a subject.

[0005] In some aspects, disclosed herein is a method for training a generative artificial intelligence (Al) model for providing generated data for a target biological species based on training data from a training biological species, comprising: conducting an experiment on a subject from the training biological species; measuring one or more amounts of a biomolecule from a sample from the experiment; creating a vector based on the measured one or more amounts of the biomolecule; determining a summary value based on the created vector; and training the generative Al model to provide the generated data for a subject from the target biological species, using the determined summary value.

[0006] In some aspects, disclosed herein is a method for generating generated data for a subject of a target biological species, comprising: inputting experiment data for a subject from a target biological species into a generative Al model trained on a training biological species, the generative Al model having been trained by a method comprising: conducting an experiment on a subject from the training biological species; measuring one or more amounts of a biomolecule from a sample from the experiment; creating a vector based on the measured one or more amounts of the biomolecule; determining a summary value based on the created vector; and training the generative Al model to provide the generated data for a subject from the target biological species, using the determined summary value; and receiving the generated data.

[0007] In some embodiments, the method can further comprise: inputting the generated data into an additional trained machine learning model configured to predict a clinical outcome for the subject from the target biological species; and receiving from the additional trained machine learning model, the predicted clinical outcome for the subject from the target biological species.

[0008] In some embodiments, the predicted clinical outcome can comprise determining a disease signal in the subject. In any of the embodiments herein, the predicted clinical outcome can comprise determining a diagnosis for the subject. In any of the embodiments herein, the predicted clinical outcome can comprise determining a therapy for the subject. In some embodiments, the determining the therapy for the subject can comprise determining a dosage of the therapy. In some embodiments, the predicted clinical outcome can comprise administering the therapy to the subject.

[0009] In any of the embodiments herein, the predicted clinical outcome can comprise determining a prognosis for the subject. In some embodiments, the prognosis can comprise determining a response from the subject to a predetermined therapy. In some embodiments, the response can comprise an adverse response to the predetermined therapy for the subject. In some embodiments, the response can comprise a recovery from the predetermined therapy for the subject.

[0010] In any of the embodiments herein, the target biological species can be a human. In any of the embodiments herein, the determined summary value can be based on one or more relative amounts of the biomolecule. In some embodiments, the one or more relative amounts of the biomolecule can comprise comparing the measured one or more amounts to measured one or more amounts from a control group. In any of the embodiments herein, the summary value can comprise determining a norm. In some embodiments, the norm can be an Lp norm. In any of the embodiments herein, the norm can be an LI norm, an L2 norm, or an L max norm.

[0011] In some aspects, disclosed herein is a method for training a generative Al model for inferring one or more target amounts in a target anatomical region, comprising: conducting an experiment on a subject; measuring one or more amounts of a biomolecule from a sample from the experiment; and training the generative Al model to infer the one or more target amounts in the target anatomical region, based on the one or more measured amounts.

[0012] In some aspects, disclosed herein is a method of inferring one or more target amounts of a first biomolecule in a target anatomical region, comprising: conducting an experiment to obtain a sample from an input anatomical region of a subject; measuring, from the sample, one or more amounts of the first biomolecule or a second biomolecule; inputting the one or more measured amounts into a trained generative Al model; and inferring, from the one or more measured amounts, the one or more target amounts of the first biomolecule in the target anatomical region.

[0013] In any of the embodiments herein, the obtaining the sample comprises non-invasively obtaining the sample. In any of the embodiments herein, the obtaining the sample comprises drawing blood from the subject. In any of the embodiments herein, the sample is obtained from a necropsy of the subject. In any of the embodiments herein, the input anatomical region comprises peripheral blood mononuclear cells, red blood cells, or a combination thereof. In any of the embodiments herein, the target anatomical region comprises pancreas.

[0014] In some aspects, disclosed herein is a method for training a generative Al model for providing generated data comprising: conducting an experiment on a subject; measuring one or more amounts of a biomolecule from a sample from the experiment; including, in a training dataset exclusionary of pre-existing structural data, a value based on the one or more measured amounts; and training the generative Al model to provide the generated data, based on the training dataset.

[0015] In some aspects, disclosed herein is a method for generating generated data, comprising: inputting experiment data into a generative Al model, the generative Al model having been trained by a method comprising: conducting an experiment on a subject; measuring one or more amounts of a biomolecule from a sample from the experiment; including, in a training dataset exclusionary of pre-existing structural data, a value based on the one or more measured amounts; and training the generative Al model to provide the generated data, based on the training dataset; and receiving the generated data.

[0016] In some aspects, disclosed herein is a method for training a generative Al model for providing generated data comprising: conducting an experiment on a subject; measuring one or more amounts of a biomolecule from a sample from the experiment; including, in a training dataset, a value based on the one or more measured amounts; excluding from the training dataset,pre-existing structural data corresponding to a biomolecule; and training the generative Al model to provide the generated data, based on the training dataset.

[0017] In some aspects, disclosed herein is a method for generating generated data, comprising: inputting experiment data into a generative Al model, the generative Al model having been trained by a method comprising: conducting an experiment on a subject; measuring one or more amounts of a biomolecule from a sample from the experiment; including, in a training dataset, a value based on the one or more measured amounts; excluding from the training dataset, preexisting structural data corresponding to a biomolecule; and training the generative Al model to provide the generated data, based on the training dataset; and receiving the generated data.

[0018] In any of the embodiments herein, the experiment can comprise providing a stimulus to the subject. In some embodiments, the stimulus can be a drug. In some embodiments, the drug can be isoniazid, tetracycline, carbon tetrachloride, valproate, or a combination thereof. In any of the embodiments herein, the providing the stimulus can comprise administering a dosage of the drug that is at or above a predetermined dosage level. In any of the embodiments herein, the one or more measured amounts can be based on RNA sequencing. In any of the embodiments herein, the one or more measured amounts can be based on mass spectrometry. In any of the embodiments herein, the one or more measured amounts can be based on methyl-sequencing. In any of the embodiments herein, the one or more measured amounts can be based on DNA sequencing. In any of the embodiments herein, the sample can comprise liquid biopsy samples. In some embodiments, the liquid biopsy sample can comprise blood, plasma, cerebrospinal fluid, sputum, stool, urine, sweat, or saliva. In any of the embodiments herein, the sample can comprise a tissue biopsy sample.

[0019] In any of the embodiments herein, the generated data can comprise generated RNA sequencing data, generated mass spectrometry data, generated DNA sequencing data, generated methyl-sequencing data or a combination thereof.

[0020] In some aspects, disclosed herein is a method of tokenizing data exclusionary of preexisting structural data, to generate tokens for a generative Al model, comprising: determining a signal strength value from the data; determining an entropy value based on the signal strength value; and outputting tokens based on the determined entropy value.

[0021] In some aspects, disclosed herein is a method of tokenizing data to generate tokens for a generative Al model, comprising: excluding from the data, pre-existing structural data corresponding to a biomolecule; determining a signal strength value from the data; determining an entropy value based on the signal strength value; and outputting tokens based on the determined entropy value. In some embodiments, the entropy value can be a Shannon entropy value. In any of the embodiments herein, the entropy value can be based on determining the probability of a correspondence between an amount of a first biomolecule of a first biological species and the amount of a second biomolecule of a second biological species.

[0022] In some embodiments, the determining the probability of the correspondence can be based on a predetermined probability sharpness value, a predetermined control expression weighting, or a combination thereof. In any of the embodiments herein, the pre-existing structural data comprises sequence data or molecular structure data. In some embodiments, the sequence data can comprise DNA sequences, RNA sequences, or amino acid sequences. In any of the embodiments herein, the sequence data comprises sequence similarity values. In some embodiments, the molecular structure data can comprise protein structure data. In any of the embodiments herein, the excluding the pre-existing structural data is based on a correspondence value indicating the correspondence between the pre-existing structural data and the entropy value. In any of the embodiments herein, the biomolecule can be an RNA molecule. In some embodiments, the RNA molecule can be an mRNA transcript. In any of the embodiments herein, the biomolecule can be a protein molecule. In any of the embodiments herein, the biomolecule can be an epigenetic marker. In any of the embodiments herein, the biomolecule can be a methyl group on a nucleotide base. In any of the embodiments herein, the biomolecule can be a DNA molecule. In some embodiments, the DNA molecule can be a cell-free DNA (cfDNA). In any of the embodiments herein, the biomolecule can be a metabolite.

[0023] In some aspects, disclosed herein is a system comprising: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: conduct an experiment on a subject from the training biological species; measure one or more amounts of a biomolecule from a sample from the experiment; create a vector based on the measured one or more amounts of the biomolecule; determine a summary value based on the created vector; andtrain the generative Al model to provide the generated data for a subject from the target biological species, using the determined summary value.

[0024] In some aspects, disclosed herein is a system comprising: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: input experiment data for a subject from a target biological species into a generative Al model trained on a training biological species, the generative Al model having been trained by a method comprising: conducting an experiment on a subject from the training biological species; measuring one or more amounts of a biomolecule from a sample from the experiment; creating a vector based on the measured one or more amounts of the biomolecule; determining a summary value based on the created vector; and training the generative Al model to provide the generated data for a subject from the target biological species, using the determined summary value; and receive the generated data.

[0025] In some embodiments, the system can comprise further instructions that, when executed by the one or more processors, cause the system to: input the generated data into an additional trained machine learning model configured to predict a clinical outcome for the subject from the target biological species; and receive from the additional trained machine learning model, the predicted clinical outcome for the subject from the target biological species.

[0026] In some aspects, disclosed herein is a system comprising: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: conduct an experiment on a subject; measure one or more amounts of a biomolecule from a sample from the experiment; and train the generative Al model to infer the one or more target amounts in the target anatomical region, based on the one or more measured amounts.

[0027] In some aspects, disclosed herein is a system comprising: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: conduct an experiment to obtain a sample from an input anatomical region of a subject; measure, from the sample, one or more amounts of the first biomolecule or a second biomolecule; input the one ormore measured amounts into a trained generative Al model; and infer, from the one or more measured amounts, the one or more target amounts of the first biomolecule in the target anatomical region.

[0028] In some aspects, disclosed herein is a system comprising: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: conduct an experiment on a subject; measure one or more amounts of a biomolecule from a sample from the experiment; include, in a training dataset exclusionary of pre-existing structural data, a value based on the one or more measured amounts; and train the generative Al model to provide the generated data, based on the training dataset.

[0029] In some aspects, disclosed herein is a system comprising: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: input experiment data into a generative Al model, the generative Al model having been trained by a method comprising: conducting an experiment on a subject; measuring one or more amounts of a biomolecule from a sample from the experiment; including, in a training dataset exclusionary of pre-existing structural data, a value based on the one or more measured amounts; and training the generative Al model to provide the generated data, based on the training dataset; and receive the generated data.

[0030] In some aspects, disclosed herein is a system comprising: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: conduct an experiment on a subject; measure one or more amounts of a biomolecule from a sample from the experiment; include, in a training dataset, a value based on the one or more measured amounts; exclude from the training dataset, pre-existing structural data corresponding to a biomolecule; and train the generative Al model to provide the generated data, based on the training dataset.

[0031] In some aspects, disclosed herein is a system comprising: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: inputexperiment data into a generative Al model, the generative Al model having been trained by a method comprising: conducting an experiment on a subject; measuring one or more amounts of a biomolecule from a sample from the experiment; including, in a training dataset, a value based on the one or more measured amounts; excluding from the training dataset, pre-existing structural data corresponding to a biomolecule; and training the generative Al model to provide the generated data, based on the training dataset; and receive the generated data.

[0032] In some aspects, disclosed herein is a system comprising: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: determine a signal strength value from the data; determine an entropy value based on the signal strength value; and output tokens based on the determined entropy value.

[0033] In some aspects, disclosed herein is a system comprising: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: exclude from the data, pre-existing structural data corresponding to a biomolecule; determine a signal strength value from the data; determine an entropy value based on the signal strength value; and output tokens based on the determined entropy value.

[0034] In some aspects, disclosed herein is a non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of a system, cause the system to: conduct an experiment on a subject from the training biological species; measure one or more amounts of a biomolecule from a sample from the experiment; create a vector based on the measured one or more amounts of the biomolecule; determine a summary value based on the created vector; and train the generative Al model to provide the generated data for a subject from the target biological species, using the determined summary value.

[0035] In some aspects, disclosed herein is a non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of a system, cause the system to: input experiment data for a subject from a target biological species into a generative Al model trained on a trainingbiological species, the generative Al model having been trained by a method comprising: conducting an experiment on a subject from the training biological species; measuring one or more amounts of a biomolecule from a sample from the experiment; creating a vector based on the measured one or more amounts of the biomolecule; determining a summary value based on the created vector; and training the generative Al model to provide the generated data for a subject from the target biological species, using the determined summary value; and receive the generated data. In some embodiments, the non-transitory computer-readable storage medium further comprising instructions that, when executed by the one or more processors, cause the system to: input the generated data into an additional trained machine learning model configured to predict a clinical outcome for the subject from the target biological species; and receive from the additional trained machine learning model, the predicted clinical outcome for the subject from the target biological species.

[0036] In some aspects, disclosed herein is a non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of a system, cause the system to: conduct an experiment on a subject; measure one or more amounts of a biomolecule from a sample from the experiment; and train the generative Al model to infer the one or more target amounts in the target anatomical region, based on the one or more measured amounts.

[0037] In some aspects, disclosed herein is a non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of a system, cause the system to: conduct an experiment to obtain a sample from an input anatomical region of a subject; measure, from the sample, one or more amounts of the first biomolecule or a second biomolecule; input the one or more measured amounts into a trained generative Al model; and infer, from the one or more measured amounts, the one or more target amounts of the first biomolecule in the target anatomical region.

[0038] In some aspects, disclosed herein is a non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of a system, cause the system to: conduct an experiment on a subject; measure one or more amounts of a biomolecule from a sample from the experiment; include, in a training dataset exclusionary of pre-existing structural data, a value based on theone or more measured amounts; and train the generative Al model to provide the generated data, based on the training dataset.

[0039] In some aspects, disclosed herein is a non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of a system, cause the system to: input experiment data into a generative Al model, the generative Al model having been trained by a method comprising: conducting an experiment on a subject; measuring one or more amounts of a biomolecule from a sample from the experiment; including, in a training dataset exclusionary of pre-existing structural data, a value based on the one or more measured amounts; and training the generative Al model to provide the generated data, based on the training dataset; and receive the generated data.

[0040] In some aspects, disclosed herein is a non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of a system, cause the system to: conduct an experiment on a subject; measure one or more amounts of a biomolecule from a sample from the experiment; include, in a training dataset, a value based on the one or more measured amounts; exclude from the training dataset, pre-existing structural data corresponding to a biomolecule; and train the generative Al model to provide the generated data, based on the training dataset.

[0041] In some aspects, disclosed herein is a non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of a system, cause the system to: input experiment data into a generative Al model, the generative Al model having been trained by a method comprising: conducting an experiment on a subject; measuring one or more amounts of a biomolecule from a sample from the experiment; including, in a training dataset, a value based on the one or more measured amounts; excluding from the training dataset, pre-existing structural data corresponding to a biomolecule; and training the generative Al model to provide the generated data, based on the training dataset; and receive the generated data.

[0042] In some aspects, disclosed herein is a non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which whenexecuted by one or more processors of a system, cause the system to: determine a signal strength value from the data; determine an entropy value based on the signal strength value; and output tokens based on the determined entropy value.

[0043] In some aspects, disclosed herein is a non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of a system, cause the system to: exclude from the data, preexisting structural data corresponding to a biomolecule; determine a signal strength value from the data; determine an entropy value based on the signal strength value; and output tokens based on the determined entropy value.BRIEF DESCRIPTION OF THE DRAWINGS

[0044] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawings will be provided by the Office upon request and payment of the necessary fee.

[0045] Various aspects of the disclosed methods, devices, and systems are set forth with particularity in the appended claims. A better understanding of the features and advantages of the disclosed methods, devices, and systems will be obtained by reference to the following detailed description of illustrative embodiments and the accompanying drawings, of which:

[0046] FIG. 1 provides a non-limiting exemplary method for training a generative artificial intelligence (Al) model.

[0047] FIG. 2 provides a non-limiting exemplary method for generating generated data from a trained generative Al model.

[0048] FIG. 3 provides a non-limiting exemplary method for training a generative Al model.

[0049] FIG. 4 provides a non-limiting exemplary method for inferring target amounts of a biomolecule in a target region, using a generative Al model.

[0050] FIG. 5 provides a non-limiting exemplary method for training a generative Al model.

[0051] FIG. 6 provides a non-limiting exemplary method for generating generated data from a trained generative Al model.

[0052] FIG. 7 provides a non-limiting exemplary method for training a generative Al model.

[0053] FIG. 8 provides a non-limiting exemplary method for generating generated data from a trained generative Al model.

[0054] FIG. 9 provides a non-limiting exemplary method for tokenizing data for a generative Al model.

[0055] FIG. 10 provides a non-limiting exemplary method for tokenizing data for a generative Al model.

[0056] FIG. 11 depicts example schematics for performing and analyzing experiments relating to cross-species gene-mapping.

[0057] FIG. 12 provides example data indicating gene expression profiles across organs, in response to drugs.

[0058] FIG. 13 provides example data indicating that protein structure data similarities do not correlate with gene expression profiles in response to drugs.

[0059] FIG. 14 provides example data depicting cross-species functional gene mapping based on gene expression profiles.

[0060] FIG. 15 provides example data depicting longitudinal whole blood gene expression profiles in response to drugs versus organ endpoint gene expression profiles in response to drugs.

[0061] FIG. 16 depicts an exemplary computing device or system in accordance with one embodiment of the present disclosure.

[0062] FIG. 17 depicts an exemplary computer system or computer network, in accordance with some instances of the systems described herein.DETAILED DESCRIPTION

[0063] Methods and systems for training or deploying generative artificial intelligence (Al) for predicting biomolecule amounts are described. The methods and system described herein can be used to predict gene expression amounts personalized to a physiological system or an anatomical region of the subject. In some instances, predicting the biomolecule amounts can be for a tissue for which procuring a biological sample is invasive. The generative Al model can be trained on training data relating to measured amounts of biomolecules, such as gene expression data. The training data can derive from physiologies for which samples are readily accessible. Training the generative Al model can comprise determining a summary value based on the measured amounts of biomolecules. The training data can be exclusionary of structural data, such as protein structure data. The generated data generated from the generative Al model can be used as inputs to other machine learning models, e.g., classifiers, to provide clinical information regarding a subject.

[0064] Drug testing in animal models is a critical and mandatory component of preclinical studies for drug development. Drug testing in animals can lead to investigational new drug applications for regulatory agencies, such as the Food and Drug Administration or the European Medicines Agency. Drug testing in animals is primarily focused on toxicology and safety, and in doing so, aims to capture the systems biology effects of novel compounds on various organ systems. However, roughly 40% of new drug candidates tested in animal models — including non-human primate models — fail Phase I human clinical trials, often because of severe adverse events in humans. This failure is despite the more than 98% shared genetic homology between non-human primates and humans.

[0065] To address such failures in translational medicine, the methods and systems described herein leverage functional biological data, such as data based on measured amounts of biomolecules, including gene expression data, e.g., gene expression levels in response to drug treatments. The methods and systems described herein use the functional biological data, such as the gene expression data, to train or deploy generative Al models that predict biomolecule amounts. The leveraging of functional biological data to predict biomolecule amounts is in stark contrast to existing methods that require or rely on structural biological data, such as nucleic acid sequences or protein structures. The training or deploying of generative Al models based on functional biological data allows for improved accuracies in predicting a subject’s response to atreatment, such as a drug, given a different subject’s response to the treatment, or given a different aspect of the subject’s response to the treatment (e.g., given the response of a particular physiological system or anatomical region of the subject). The methods and systems described herein can exclude the use of biological structural data, such as protein structure data, during the training or deploying of the generative Al model, and / or during the tokenizing of the data. The methods and systems described herein can comprise the determining of a summary value from data based on measured amounts of biomolecules.

[0066] Described herein are methods and systems used to predict gene expression amounts personalized to a physiological system or an anatomical region of the subject. In some instances, predicting the biomolecule amounts can be for a tissue for which procuring a biological sample is invasive. The generative Al model can be trained on training data relating to measured amounts of biomolecules, such as gene expression data. The training data can derive from physiologies for which samples are readily accessible. Training the generative Al model can comprise determining a summary value based on the measured amounts of biomolecules. The training data can be exclusionary of structural data, such as protein structure data. The generated data generated from the generative Al model can be used as inputs to other machine learning models, e.g., classifiers, to provide clinical information regarding a subject.Definitions

[0067] Unless otherwise defined, all of the technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art in the field to which this disclosure belongs.

[0068] As used in this specification and the appended claims, the singular forms “a”, “an”, and “the” include plural references unless the context clearly dictates otherwise. Any reference to “or” herein is intended to encompass “and / or” unless otherwise stated.

[0069] “About” and “approximately” shall generally mean an acceptable degree of error for the quantity measured given the nature or precision of the measurements. Exemplary degrees of error are within 20 percent (%), typically, within 10%, and more typically, within 5% of a given value or range of values.

[0070] As used herein, the terms "comprising" (and any form or variant of comprising, such as "comprise" and "comprises"), "having" (and any form or variant of having, such as "have" and "has"), "including" (and any form or variant of including, such as "includes" and "include"), or "containing" (and any form or variant of containing, such as "contains" and "contain"), are inclusive or open-ended and do not exclude additional, un-recited additives, components, integers, elements, or method steps.

[0071] As used herein, the terms “individual,” “patient,” or “subject” are used interchangeably and refer to any single animal, e.g., a mammal (including such non-human animals as, for example, dogs, cats, horses, rabbits, zoo animals, cows, pigs, sheep, and non-human primates) for which treatment is desired. In particular embodiments, the individual, patient, or subject herein is a human.

[0072] The terms “cancer” and “tumor” are used interchangeably herein. These terms refer to the presence of cells possessing characteristics typical of cancer-causing cells, such as uncontrolled proliferation, immortality, metastatic potential, rapid growth and proliferation rate, and certain characteristic morphological features. Cancer cells are often in the form of a tumor, but such cells can exist alone within an animal, or can be a non-tumorigenic cancer cell, such as a leukemia cell. These terms include a solid tumor, a soft tissue tumor, or a metastatic lesion. As used herein, the term “cancer” includes premalignant, as well as malignant cancers.

[0073] As used herein, “treatment” (and grammatical variations thereof such as “treat” or “treating”) refers to clinical intervention (e.g., administration of an anti-cancer agent or anticancer therapy) in an attempt to alter the natural course of the individual being treated, and can be performed either for prophylaxis or during the course of clinical pathology. Desirable effects of treatment include, but are not limited to, preventing occurrence or recurrence of disease, alleviation of symptoms, diminishment of any direct or indirect pathological consequences of the disease, preventing metastasis, decreasing the rate of disease progression, amelioration or palliation of the disease state, and remission or improved prognosis.

[0074] As used herein, the term “subgenomic interval” (or “subgenomic sequence interval”) refers to a portion of a genomic sequence.

[0075] As used herein, the term "subject interval" refers to a subgenomic interval or an expressed subgenomic interval (e.g., the transcribed sequence of a subgenomic interval).

[0076] As used herein, the terms “variant sequence” or “variant” are used interchangeably and refer to a modified nucleic acid sequence relative to a corresponding “normal” or “wild-type” sequence. In some instances, a variant sequence may be a “short variant sequence” (or “short variant”), i.e., a variant sequence of less than about 50 base pairs in length.

[0077] The terms “allele frequency” and “allele fraction” are used interchangeably herein and refer to the fraction of sequence reads corresponding to a particular allele relative to the total number of sequence reads for a genomic locus.

[0078] The terms “variant allele frequency” and “variant allele fraction” are used interchangeably herein and refer to the fraction of sequence reads corresponding to a particular variant allele relative to the total number of sequence reads for a genomic locus.

[0079] It is understood that aspects and variations of the invention described herein include “consisting” and / or “consisting essentially of’ aspects and variations.

[0080] When a range of values is provided, it is to be understood that each intervening value between the upper and lower limit of that range, and any other stated or intervening value in that states range, is encompassed within the scope of the present disclosure. Where the stated range includes upper or lower limits, ranges excluding either of those included limits are also included in the present disclosure.

[0081] Some of the analytical methods described herein include mapping sequences to a reference sequence, determining sequence information, and / or analyzing sequence information. It is well understood in the art that complementary sequences can be readily determined and / or analyzed, and that the description provided herein encompasses analytical methods performed in reference to a complementary sequence.

[0082] The section headings used herein are for organization purposes only and are not to be construed as limiting the subject matter described. The description is presented to enable one of ordinary skill in the art to make and use the invention and is provided in the context of a patentapplication and its requirements. Various modifications to the described embodiments will be readily apparent to those persons skilled in the art and the generic principles herein may be applied to other embodiments. Thus, the present invention is not intended to be limited to the embodiment shown but is to be accorded the widest scope consistent with the principles and features described herein.

[0083] The figures illustrate processes according to various embodiments. In the exemplary processes, some blocks are, optionally, combined, the order of some blocks is, optionally, changed, and some blocks are, optionally, omitted. In some examples, additional steps may be performed in combination with the exemplary processes. Accordingly, the operations as illustrated (and described in greater detail below) are exemplary by nature and, as such, should not be viewed as limiting.

[0084] The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.Methods for training or deploying a generative artificial intelligence (Al) model for predicting biomolecule amounts

[0085] The disclosed methods relate to training or deploying a generative Al model for predicting biomolecule amounts. The methods and system described herein can be used to predict gene expression amounts personalized to a physiological system or an anatomical region of the subject. In some instances, predicting the biomolecule amounts can be for a tissue for which procuring a biological sample is invasive. The generative Al model can be trained on training data relating to measured amounts of biomolecules, such as gene expression data. The training data can derive from physiologies for which samples are readily accessible. Training the generative Al model can comprise determining a summary value based on the measured amounts of biomolecules. The training data can be exclusionary of structural data, such as protein structure data. The generated data generated from the generative Al model can be used as inputs to other machine learning models, e.g., classifiers, to provide clinical information regarding a subject.

[0086] FIG. 1 shows an exemplary schematic showing general process 100 for training a generative Al model to provide generated data for a target species based on training data from a training species. The method can include: conducting an experiment on a subject from the training species (102); measuring one or more amounts of a biomolecule from a sample from the experiment (104); creating a vector based on the one or more amounts of the biomolecule (106); determining a summary value based on the created vector (108); and training the generative Al model to provide the generated data for a subject from the target species, using the determined summary value (110).

[0087] At 102 in FIG. 1, an experiment on a subject from the training biological species is conducted.

[0088] At 104 in FIG. 1, one or more amounts of a biomolecule from a sample from the experiment are measured.

[0089] At 106 in FIG. 1, a vector based on the measured one or more amounts of the biomolecule is created.

[0090] At 108 in FIG. 1, a summary value based on the created vector is determined.

[0091] At 110 in FIG. 1, the generative Al model is trained to provide the generated data for a subject from the target biological species, using the determined summary value.

[0092] FIG. 2 shows an exemplary schematic showing general process 200 for generating generated data for a subject of a target species. The method can include: inputting experiment data for a subject from a target species into a generative Al model trained on a training species, the generative Al model having been trained by a method comprising: conducting an experiment on a subject from the training species; measuring one or more amounts of a biomolecule from a sample from the experiment; determining a summary value based on the one or more measured amounts of the biomolecule; and training the generative Al model to provide generated data for a subject from the target species, using the determined summary value (202); and receiving the generated data (204).

[0093] At 202 in FIG. 2, experiment data for a subject from a target biological species is inputted into a generative Al model trained on a training biological species. The generative Al model can be trained according to at least a portion of the elements described in process 100.

[0094] At 204 in FIG. 2, the generated data is received.

[0095] FIG. 3 shows an exemplary schematic showing general process 300 for training a generative Al model for inferring one or more target amounts in a target anatomical region. The method can include: conducting an experiment on a subject (302); measuring one or more amounts of a biomolecule from a sample from the experiment (304); and training the generative Al model to infer the one or more target amounts in the target anatomical region, based on the one or more measured amounts (306).

[0096] At 302 in FIG. 3, an experiment is conducted on a subject.

[0097] At 304 in FIG. 3, one or more amounts of a biomolecule from a sample from the experiment is measured.

[0098] At 306 in FIG. 3, the generative Al model is trained to infer the one or more target amounts in the target anatomical region, based on the one or more measured amounts. Training the generative Al model can comprise providing training data, which can include publicly available data, e.g., publicly available databases or the scientific literature. The training data can also comprise data obtained by an experimenter and / or not shared publicly.

[0099] FIG. 4 shows an exemplary schematic showing general process 400 for inferring one or more target amounts of a first biomolecule in a target anatomical region, comprising: conducting an experiment to obtain a sample from an input anatomical region of a subject (402); measuring, from the sample, one or more amounts of the first biomolecule or a second biomolecule (404); inputting the one or more measured amounts into a trained generative Al model (406); and inferring, from the one or more measured amounts, the one or more target amounts of the first biomolecule in the target anatomical region (408).

[0100] At 402 in FIG. 4, an experiment is conducted to obtain a sample from an input anatomical region of subject.

[0101] At 404 in FIG. 4, from the sample, one or more amounts of the first biomolecule or a second biomolecule are measured.

[0102] At 406 in FIG. 4, the one or more measured amounts are inputted into a trained generative Al model. The generative Al model can be trained according to at least a portion of the elements described in process 300.

[0103] At 408 in FIG. 4, the one or more target amounts of the first biomolecule in the target anatomical region are inferred from the one or more measured amounts.

[0104] FIG. 5 shows an exemplary schematic showing general process 500 for training a generative Al model for providing generated data. The method can include: conducting an experiment on a subject (502); measuring one or more amounts of a biomolecule from a sample from the experiment (504); including, in a training dataset exclusionary of pre-existing protein structure data, a value based on the one or more measured amounts (506); and training the generative Al model to provide the generated data, based on the training dataset (508).

[0105] At 502 in FIG. 5, an experiment on a subject is conducted.

[0106] At 504 in FIG. 5, one or more amounts of a biomolecule are measured from a sample from the experiment.

[0107] At 506 in FIG. 5, a value based on the one or more measured amounts is included in a training dataset exclusionary of pre-existing structural data. The pre-existing structural data can relate to data regarding the three-dimensional conformations of biomolecules, such as x-ray crystallography data, NMR spectroscopy data, or cryo-electron microscopy data. The preexisting structural data need not comprise the entire structure of the biomolecule, and can comprise, instead, a partial structure of the biomolecule, such as various secondary structure motifs. The pre-existing structural data is not limited to data regarding the three-dimensional conformations of biomolecules, but relates to structural data, in general, such as sequence data, e.g., nucleic acid sequence data or polypeptide sequence data.

[0108] At 508 in FIG. 5, the generative Al model is trained to provide the generated data, based on the training dataset.

[0109] FIG. 6 shows an exemplary schematic showing general process 600 for generating generated data. The method can include: inputting experiment data into a generative Al model, the generative Al model having been trained by a method comprising: conducting an experiment on a subject; measuring one or more amounts of a biomolecule from a sample from the experiment; including, in a training dataset exclusionary of pre-existing protein structure data, a value based on the one or more measured amounts; and training the generative Al model to provide the generated data, based on the training dataset (602); and receiving the generated data (604).

[0110] At 602 in FIG. 6, the experiment data is inputted into a generative Al model. The generative Al model can be trained according to at least a portion of the elements described in process 500.

[0111] At 604 in FIG. 6, the generated data is received.

[0112] FIG. 7 shows an exemplary schematic showing general process 700 for training a generative Al model for predicting generated data. The method can include: conducting an experiment on a subject (702); measuring one or more amounts of a biomolecule from a sample from the experiment (704); including, in a training dataset, a value based on the one or more measured amounts (706); excluding from the training dataset, pre-existing protein structure data corresponding to the biomolecule (708); and training the generative Al model to provide the generated data, based on the training dataset (710).

[0113] At 702 in FIG. 7, an experiment is conducted on a subject.

[0114] At 704 in FIG. 7, one or more amounts of a biomolecule from a sample from the experiment are measured.

[0115] At 706 in FIG. 7, a value based on the one or more measured amounts is included in a training dataset.

[0116] At 708 in FIG. 7, pre-existing structural data corresponding to a biomolecule are excluded from the training dataset. The pre-existing structural data can relate to data regarding the three-dimensional conformations of biomolecules, such as x-ray crystallography data, NMRspectroscopy data, or cryo-electron microscopy data. The pre-existing structural data need not comprise the entire structure of the biomolecule, and can comprise, instead, a partial structure of the biomolecule, such as various secondary structure motifs. The pre-existing structural data is not limited to data regarding the three-dimensional conformations of biomolecules, but relates to structural data, in general, such as sequence data, e.g., nucleic acid sequence data or polypeptide sequence data.

[0117] At 710 in FIG. 7, the generative Al model is trained to provide the generated data, based on the training dataset.

[0118] FIG. 8 shows an exemplary schematic showing general process 800 for generating generated data. The method can include: inputting experiment data into a generative artificial intelligence model, the generative artificial intelligence model having been trained by a method comprising: conducting an experiment on a subject; measuring one or more amounts of a biomolecule from a sample from the experiment; including, in a training dataset, a value based on the one or more measured amounts; excluding from the training dataset, pre-existing protein structure data corresponding to the biomolecule; and training the generative artificial intelligence model to provide the generated data, based on the training dataset (802); and receiving the generated data (804).

[0119] At 802 in FIG. 8, experiment data is inputted into a generative Al model. The generative Al model can be trained according to at least a portion of the elements described in process 700.

[0120] At 804 in FIG. 8, the generated data is received.

[0121] FIG. 9 shows an exemplary schematic showing general process 900 for tokenizing data exclusionary of pre-existing structural data, to generate tokens for a generative Al model. The method can include: determining a signal strength value from the data (902); determining an entropy value based on the signal strength value (904); and outputting tokens based on the determined entropy value (906).

[0122] At 902 in FIG. 9, a signal strength value is determined from the data.

[0123] At 904 in FIG. 9, an entropy value based on the signal strength value is determined.

[0124] At 906 in FIG. 9, tokens are outputted based on the determined entropy value.

[0125] FIG. 10 shows an exemplary schematic showing general process 1000 for tokenizing data to generate tokens for a generative Al model. The method can include: excluding from the data, pre-existing structural data corresponding to a biomolecule (1002); determining a signal strength value from the data (1004); determining an entropy value based on the signal strength value (1006); and outputting tokens based on the determined entropy value (1008).

[0126] Processes 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 can be performed, for example, using one or more electronic devices implementing a software platform. In some examples, processes 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 are performed using a client-server system, and the blocks of processes 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 are divided up in any manner between the server and a client device. In other examples, the blocks of processes 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 are divided up between the server and multiple client devices. Thus, while portions of processes 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 are described herein as being performed by particular devices of a client-server system, it will be appreciated that processes 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 are not so limited. In other examples, processes 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 are performed using only a client device or only multiple client devices. In processes 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 some blocks are, optionally, combined, the order of some blocks is, optionally, changed, and some blocks are, optionally, omitted. In some examples, additional steps may be performed in combination with the process 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000. Accordingly, the operations as illustrated (and described in greater detail below) are exemplary by nature and, as such, should not be viewed as limiting.Samples

[0127] The disclosed methods and systems may be used with any of a variety of samples (also referred to herein as specimens) comprising nucleic acids (e.g., DNA or RNA) that are collected from a subject (e.g., a patient). Examples of a sample include, but are not limited to, a tumor sample, a tissue sample, a biopsy sample e.g., a tissue biopsy, a liquid biopsy, or both), a blood sample (e.g., a peripheral whole blood sample), a blood plasma sample, a blood serum sample, alymph sample, a saliva sample, a sputum sample, a urine sample, a gynecological fluid sample, a circulating tumor cell (CTC) sample, a cerebral spinal fluid (CSF) sample, a pericardial fluid sample, a pleural fluid sample, an ascites (peritoneal fluid) sample, a feces (or stool) sample, or other body fluid, secretion, and / or excretion sample (or cell sample derived therefrom). In certain instances, the sample may be frozen sample or a formalin-fixed paraffin-embedded (FFPE) sample.

[0128] In some instances, the sample may be collected by tissue resection (e.g., surgical resection), needle biopsy, bone marrow biopsy, bone marrow aspiration, skin biopsy, endoscopic biopsy, fine needle aspiration, oral swab, nasal swab, vaginal swab or a cytology smear, scrapings, washings or lavages (such as a ductal lavage or bronchoalveolar lavage), etc..

[0129] In some instances, the sample is a liquid biopsy sample, and may comprise, e.g., whole blood, blood plasma, blood serum, urine, stool, sputum, saliva, or cerebrospinal fluid. In some instances, the sample may be a liquid biopsy sample and may comprise circulating tumor cells (CTCs). In some instances, the sample may be a liquid biopsy sample and may comprise cell- free DNA (cfDNA), circulating tumor DNA (ctDNA), or any combination thereof.

[0130] In some instances, the sample may comprise one or more premalignant or malignant cells. Premalignant, as used herein, refers to a cell or tissue that is not yet malignant but is poised to become malignant. In certain instances, the sample may be acquired from a solid tumor, a soft tissue tumor, or a metastatic lesion. In certain instances, the sample may be acquired from a hematologic malignancy or pre-malignancy. In other instances, the sample may comprise a tissue or cells from a surgical margin. In certain instances, the sample may comprise tumor-infiltrating lymphocytes. In some instances, the sample may comprise one or more non- malignant cells. In some instances, the sample may be, or is part of, a primary tumor or a metastasis (e.g., a metastasis biopsy sample). In some instances, the sample may be obtained from a site (e.g., a tumor site) with the highest percentage of tumor (e.g., tumor cells) as compared to adjacent sites (e.g., sites adjacent to the tumor). In some instances, the sample may be obtained from a site (e.g., a tumor site) with the largest tumor focus (e.g., the largest number of tumor cells as visualized under a microscope) as compared to adjacent sites (e.g., sites adjacent to the tumor).

[0131] In some instances, the disclosed methods may further comprise analyzing a primary control (e.g., a normal tissue sample). In some instances, the disclosed methods may further comprise determining if a primary control is available and, if so, isolating a control nucleic acid (e.g., DNA) from said primary control. In some instances, the sample may comprise any normal control (e.g., a normal adjacent tissue (NAT)) if no primary control is available. In some instances, the sample may be or may comprise histologically normal tissue. In some instances, the method includes evaluating a sample, e.g., a histologically normal sample (e.g., from a surgical tissue margin) using the methods described herein. In some instances, the disclosed methods may further comprise acquiring a sub-sample enriched for non-tumor cells, e.g., by macro-dissecting non-tumor tissue from said NAT in a sample not accompanied by a primary control. In some instances, the disclosed methods may further comprise determining that no primary control and no NAT is available, and marking said sample for analysis without a matched control.

[0132] In some instances, samples obtained from histologically normal tissues (e.g., otherwise histologically normal surgical tissue margins) may still comprise a genetic alteration such as a variant sequence as described herein. The methods may thus further comprise re-classifying a sample based on the presence of the detected genetic alteration. In some instances, multiple samples (e.g., from different subjects) are processed simultaneously.

[0133] The disclosed methods and systems may be applied to the analysis of nucleic acids extracted from any of variety of tissue samples (or disease states thereof), e.g., solid tissue samples, soft tissue samples, metastatic lesions, or liquid biopsy samples. Examples of tissues include, but are not limited to, connective tissue, muscle tissue, nervous tissue, epithelial tissue, and blood. Tissue samples may be collected from any of the organs within an animal or human body. Examples of human organs include, but are not limited to, the brain, heart, lungs, liver, kidneys, pancreas, spleen, thyroid, mammary glands, uterus, prostate, large intestine, small intestine, bladder, bone, skin, etc.

[0134] In some instances, the nucleic acids extracted from the sample may comprise deoxyribonucleic acid (DNA) molecules. Examples of DNA that may be suitable for analysis by the disclosed methods include, but are not limited to, genomic DNA or fragments thereof, mitochondrial DNA or fragments thereof, cell-free DNA (cfDNA), and circulating tumor DNA(ctDNA). Cell-free DNA (cfDNA) is comprised of fragments of DNA that are released from normal and / or cancerous cells during apoptosis and necrosis, and circulate in the blood stream and / or accumulate in other bodily fluids. Circulating tumor DNA (ctDNA) is comprised of fragments of DNA that are released from cancerous cells and tumors that circulate in the blood stream and / or accumulate in other bodily fluids.

[0135] In some instances, DNA is extracted from nucleated cells from the sample. In some instances, a sample may have a low nucleated cellularity, e.g., when the sample is comprised mainly of erythrocytes, lesional cells that contain excessive cytoplasm, or tissue with fibrosis. In some instances, a sample with low nucleated cellularity may require more, e.g., greater, tissue volume for DNA extraction.

[0136] In some instances, the nucleic acids extracted from the sample may comprise ribonucleic acid (RNA) molecules. Examples of RNA that may be suitable for analysis by the disclosed methods include, but are not limited to, total cellular RNA, total cellular RNA after depletion of certain abundant RNA sequences (e.g., ribosomal RNAs), cell-free RNA (cfRNA), messenger RNA (mRNA) or fragments thereof, the poly(A)-tailed mRNA fraction of the total RNA, ribosomal RNA (rRNA) or fragments thereof, transfer RNA (tRNA) or fragments thereof, and mitochondrial RNA or fragments thereof. In some instances, RNA may be extracted from the sample and converted to complementary DNA (cDNA) using, e.g., a reverse transcription reaction. In some instances, the cDNA is produced by random-primed cDNA synthesis methods. In other instances, the cDNA synthesis is initiated at the poly (A) tail of mature mRNAs by priming with oligo(dT)-containing oligonucleotides. Methods for depletion, poly(A) enrichment, and cDNA synthesis are well known to those of skill in the art.

[0137] In some instances, the sample may comprise a tumor content (e.g., comprising tumor cells or tumor cell nuclei), or a non-tumor content (e.g., immune cells, fibroblasts, and other nontumor cells). In some instances, the tumor content of the sample may constitute a sample metric. In some instances, the sample may comprise a tumor content of at least 5-50%, 10-40%, 15-25%, or 20-30% tumor cell nuclei. In some instances, the sample may comprise a tumor content of at least 5%, at least 10%, at least 20%, at least 30%, at least 40%, or at least 50% tumor cell nuclei. In some instances, the percent tumor cell nuclei (e.g., sample fraction) is determined (e.g., calculated) by dividing the number of tumor cells in the sample by the total number of all cellswithin the sample that have nuclei. In some instances, for example when the sample is a liver sample comprising hepatocytes, a different tumor content calculation may be required due to the presence of hepatocytes having nuclei with twice, or more than twice, the DNA content of other, e.g., non-hepatocyte, somatic cell nuclei. In some instances, the sensitivity of detection of a genetic alteration, e.g., a variant sequence, or a determination of, e.g., micro satellite instability, may depend on the tumor content of the sample. For example, a sample having a lower tumor content can result in lower sensitivity of detection for a given size sample.

[0138] In some instances, as noted above, the sample comprises nucleic acid e.g., DNA, RNA (or a cDNA derived from the RNA), or both), e.g., from a tumor or from normal tissue. In certain instances, the sample may further comprise a non-nucleic acid component, e.g., cells, protein, carbohydrate, or lipid, e.g., from the tumor or normal tissue.Subjects

[0139] In some instances, the sample is obtained (e.g., collected) from a subject (e.g., patient) with a condition or disease (e.g., a hyperproliferative disease or a non-cancer indication) or suspected of having the condition or disease. In some instances, the hyperproliferative disease is a cancer. In some instances, the cancer is a solid tumor or a metastatic form thereof. In some instances, the cancer is a hematological cancer, e.g., a leukemia or lymphoma.

[0140] In some instances, the subject has a cancer or is at risk of having a cancer. For example, in some instances, the subject has a genetic predisposition to a cancer (e.g., having a genetic mutation that increases his or her baseline risk for developing a cancer). In some instances, the subject has been exposed to an environmental perturbation (e.g., radiation or a chemical) that increases his or her risk for developing a cancer. In some instances, the subject is in need of being monitored for development of a cancer. In some instances, the subject is in need of being monitored for cancer progression or regression, e.g., after being treated with an anti-cancer therapy (or anti-cancer treatment). In some instances, the subject is in need of being monitored for relapse of cancer. In some instances, the subject is in need of being monitored for minimum residual disease (MRD). In some instances, the subject has been, or is being treated, for cancer. In some instances, the subject has not been treated with an anti-cancer therapy (or anti-cancer treatment).

[0141] In some instances, the subject (e.g., a patient) is being treated, or has been previously treated, with one or more targeted therapies. In some instances, e.g., for a patient who has been previously treated with a targeted therapy, a post-targeted therapy sample (e.g., specimen) is obtained (e.g., collected). In some instances, the post-targeted therapy sample is a sample obtained after the completion of the targeted therapy.

[0142] In some instances, the patient has not been previously treated with a targeted therapy. In some instances, e.g., for a patient who has not been previously treated with a targeted therapy, the sample comprises a resection, e.g., an original resection, or a resection following recurrence (e.g., following a disease recurrence post-therapy).Systems

[0143] Also disclosed herein are systems designed to implement any of the disclosed methods for training or deploying a generative Al model for predicting biomolecule amounts from a sample from a subject. The systems may comprise, e.g., one or more processors, and a memory unit communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: conduct an experiment on a subject from the training biological species; measure one or more amounts of a biomolecule from a sample from the experiment; create a vector based on the measured one or more amounts of the biomolecule; determine a summary value based on the created vector; and train the generative Al model to provide the generated data for a subject from the target biological species, using the determined summary value.

[0144] The systems may also comprise e.g., one or more processors, and a memory unit communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: input experiment data for a subject from a target biological species into a generative Al model trained on a training biological species, the generative Al model having been trained by a method comprising: conducting an experiment on a subject from the training biological species; measuring one or more amounts of a biomolecule from a sample from the experiment; creating a vector based on the measured one or more amounts of the biomolecule; determining a summary value based on the created vector; and training the generative Al model to provide the generated data for asubject from the target biological species, using the determined summary value; and receive the generated data. The system can comprise further instructions that, when executed by the one or more processors, cause the system to: input the generated data into an additional trained machine learning model configured to predict a clinical outcome for the subject from the target biological species; and receive from the additional trained machine learning model, the predicted clinical outcome for the subject from the target biological species.

[0145] The systems may also comprise e.g., one or more processors, and a memory unit communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: conduct an experiment on a subject; measure one or more amounts of a biomolecule from a sample from the experiment; and train the generative Al model to infer the one or more target amounts in the target anatomical region, based on the one or more measured amounts.

[0146] The systems may also comprise e.g., one or more processors, and a memory unit communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: conduct an experiment to obtain a sample from an input anatomical region of a subject; measure, from the sample, one or more amounts of the first biomolecule or a second biomolecule; input the one or more measured amounts into a trained generative Al model; and infer, from the one or more measured amounts, the one or more target amounts of the first biomolecule in the target anatomical region.

[0147] The systems may also comprise e.g., one or more processors, and a memory unit communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: conduct an experiment on a subject; measure one or more amounts of a biomolecule from a sample from the experiment; include, in a training dataset exclusionary of pre-existing structural data, a value based on the one or more measured amounts; and train the generative Al model to provide the generated data, based on the training dataset.

[0148] The systems may also comprise e.g., one or more processors, and a memory unit communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: input experiment data into agenerative Al model, the generative Al model having been trained by a method comprising: conducting an experiment on a subject; measuring one or more amounts of a biomolecule from a sample from the experiment; including, in a training dataset exclusionary of pre-existing structural data, a value based on the one or more measured amounts; and training the generative Al model to provide the generated data, based on the training dataset; and receive the generated data.

[0149] The systems may also comprise e.g., one or more processors, and a memory unit communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: conduct an experiment on a subject; measure one or more amounts of a biomolecule from a sample from the experiment; include, in a training dataset, a value based on the one or more measured amounts; exclude from the training dataset, pre-existing structural data corresponding to a biomolecule; and train the generative Al model to provide the generated data, based on the training dataset.

[0150] The systems may also comprise e.g., one or more processors, and a memory unit communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: input experiment data into a generative Al model, the generative Al model having been trained by a method comprising: conducting an experiment on a subject; measuring one or more amounts of a biomolecule from a sample from the experiment; including, in a training dataset, a value based on the one or more measured amounts; excluding from the training dataset, pre-existing structural data corresponding to a biomolecule; and training the generative Al model to provide the generated data, based on the training dataset; and receive the generated data.

[0151] The systems may also comprise e.g., one or more processors, and a memory unit communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: determine a signal strength value from the data; determine an entropy value based on the signal strength value; and output tokens based on the determined entropy value.

[0152] The systems may also comprise e.g., one or more processors, and a memory unit communicatively coupled to the one or more processors and configured to store instructions that,when executed by the one or more processors, cause the system to: exclude from the data, preexisting structural data corresponding to a biomolecule; determine a signal strength value from the data; determine an entropy value based on the signal strength value; and output tokens based on the determined entropy value.

[0153] The systems may also comprise e.g., one or more processors, and a memory unit communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: conduct an experiment on a subject from the training biological species; measure one or more amounts of a biomolecule from a sample from the experiment; create a vector based on the measured one or more amounts of the biomolecule; determine a summary value based on the created vector; and train the generative Al model to provide the generated data for a subject from the target biological species, using the determined summary value.

[0154] In some instances, the disclosed systems may be used for training or deploying a generative Al model for predicting biomolecule amounts, provided data, e.g., training data, from any of a variety of samples as described herein (e.g., a tissue sample, biopsy sample, hematological sample, or liquid biopsy sample derived from the subject).

[0155] In some instances, the training or deploying the generative Al model for predicting biomolecule amounts is further used to select, initiate, adjust, or terminate a treatment for cancer in the subject (e.g., a patient) from which the sample was derived, as described elsewhere herein.

[0156] In some instances, the disclosed systems may further comprise additional data storage modules, data communication modules (e.g., Bluetooth®, WiFi, intranet, or internet communication hardware and associated software), display modules, one or more local and / or cloud-based software packages (e.g., instrument / system control software packages, sequencing data analysis software packages), etc., or any combination thereof. In some instances, the systems may comprise, or be part of, a computer system or computer network as described elsewhere herein.Computer systems and networks

[0157] FIG. 16 illustrates an example of a computing device or system in accordance with one embodiment. Device 1600 can be a host computer connected to a network. Device 1600 can be a client computer or a server. As shown in FIG. 16, device 1600 can be any suitable type of microprocessor-based device, such as a personal computer, workstation, server or handheld computing device (portable electronic device) such as a phone or tablet. The device can include, for example, one or more processor(s) 1610, input devices 1620, output devices 1630, memory or storage devices 1640, and communication devices 1660. Software 1650 residing in memory or storage device 1640 may comprise, e.g., an operating system as well as software for executing the methods described herein. Input device 1620 and output device 1630 can generally correspond to those described herein, and can either be connectable or integrated with the computer.

[0158] Input device 1620 can be any suitable device that provides input, such as a touch screen, keyboard or keypad, mouse, or voice-recognition device. Output device 1630 can be any suitable device that provides output, such as a touch screen, haptics device, or speaker.

[0159] Storage 1640 can be any suitable device that provides storage (e.g., an electrical, magnetic or optical memory including a RAM (volatile and non-volatile), cache, hard drive, or removable storage disk). Communication device 1660 can include any suitable device capable of transmitting and receiving signals over a network, such as a network interface chip or device. The components of the computer can be connected in any suitable manner, such as via a wired media (e.g., a physical system bus 1680, Ethernet connection, or any other wire transfer technology) or wirelessly (e.g., Bluetooth®, Wi-Fi®, or any other wireless technology).

[0160] Software module 1650, which can be stored as executable instructions in storage 1640 and executed by processor(s) 1610, can include, for example, an operating system and / or the processes that embody the functionality of the methods of the present disclosure (e.g., as embodied in the devices as described herein).

[0161] Software module 1650 can also be stored and / or transported within any non-transitory computer-readable storage medium for use by or in connection with an instruction executionsystem, apparatus, or device, such as those described herein, that can fetch instructions associated with the software from the instruction execution system, apparatus, or device and execute the instructions. In the context of this disclosure, a computer-readable storage medium can be any medium, such as storage 1640, that can contain or store processes for use by or in connection with an instruction execution system, apparatus, or device. Examples of computer- readable storage media may include memory units like hard drives, flash drives and distribute modules that operate as a single functional unit. Also, various processes described herein may be embodied as modules configured to operate in accordance with the embodiments and techniques described above. Further, while processes may be shown and / or described separately, those skilled in the art will appreciate that the above processes may be routines or modules within other processes.

[0162] Software module 1650 can also be propagated within any transport medium for use by or in connection with an instruction execution system, apparatus, or device, such as those described above, that can fetch instructions associated with the software from the instruction execution system, apparatus, or device and execute the instructions. In the context of this disclosure, a transport medium can be any medium that can communicate, propagate or transport programming for use by or in connection with an instruction execution system, apparatus, or device. The transport readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic or infrared wired or wireless propagation medium.

[0163] Device 1600 may be connected to a network (e.g., network 1704, as shown in FIG. 17 and / or described below), which can be any suitable type of interconnected communication system. The network can implement any suitable communications protocol and can be secured by any suitable security protocol. The network can comprise network links of any suitable arrangement that can implement the transmission and reception of network signals, such as wireless network connections, T1 or T3 lines, cable networks, DSE, or telephone lines.

[0164] Device 1600 can be implemented using any operating system, e.g., an operating system suitable for operating on the network. Software module 1650 can be written in any suitable programming language, such as C, C++, Java or Python. In various embodiments, application software embodying the functionality of the present disclosure can be deployed in different configurations, such as in a client / server arrangement or through a Web browser as a Web-basedapplication or Web service, for example. In some embodiments, the operating system is executed by one or more processors, e.g., processor(s) 1610.

[0165] FIG. 17 illustrates an example of a computing system in accordance with one embodiment. In system 1700, device 1600 (e.g., as described above and illustrated in FIG. 16) is connected to network 1704, which is also connected to device 1706.

[0166] Devices 1600 and 1706 may communicate, e.g., using suitable communication interfaces via network 1704, such as a Local Area Network (LAN), Virtual Private Network (VPN), or the Internet. In some embodiments, network 1704 can be, for example, the Internet, an intranet, a virtual private network, a cloud network, a wired network, or a wireless network. Devices 1600 and 1706 may communicate, in part or in whole, via wireless or hardwired communications, such as Ethernet, IEEE 802.11b wireless, or the like. Additionally, devices 1600 and 1706 may communicate, e.g., using suitable communication interfaces, via a second network, such as a mobile / cellular network. Communication between devices 1600 and 1706 may further include or communicate with various servers such as a mail server, mobile server, media server, telephone server, and the like. In some embodiments, Devices 1600 and 1706 can communicate directly (instead of, or in addition to, communicating via network 1704), e.g., via wireless or hardwired communications, such as Ethernet, IEEE 802.11b wireless, or the like. In some embodiments, devices 1600 and 1706 communicate via communications 1708, which can be a direct connection or can occur via a network (e.g., network 1704).

[0167] One or all of devices 1600 and 1706 generally include logic (e.g., http web server logic) or are programmed to format data, accessed from local or remote databases or other sources of data and content, for providing and / or receiving information via network 1704 according to various examples described herein.EXAMPLES

[0168] The following examples further demonstrate to one skilled in the art how to make and use the methods and systems described herein, and are not intended to limit the scope of the claimed invention.Example 1

[0169] FIG. 11 depicts example schematics for performing and analyzing experiments relating to cross-species gene-mapping. FIG. 11A shows an example schematic depicting the experimental regimen. First, blood was drawn daily from mouse (Mus musculus) and rat (Rattus norvegicu.s) models for 3 days, and then the animals were dosed with one of three drug molecules just after the blood drawn from Day 3. Daily blood draws continued for 2 more days, before the animals were sacrificed just after the Day 5 blood draw, and then necropsy was performed on 10 common organ systems. RNA extracted from the 10 different tissue samples, as well as the 5 daily blood samples, for a total of 15 RNAseq samples per animal. For each of 3 drugs and 1 saline control treatment, 3 male rats, 3 female rats, 3 male mice, and 3 female mice were used, for a total of n=48 animals and =720 RNAseq samples. FIG. 11B shows an example schematic depicting the analyses of the RNAseq data for cross-species gene-mapping. For each gene in each species, a 30-dimensional scalar vector corresponding to the gene’s expression change following drug dosing in each organ is created (3 drugs * 10 organs, i.e., one scalar value corresponding to the mean gene expression change for each of the 3 drugs). The vector can be referred to as a gene expression reactome, i.e., a reactome. Genes across the two species can be functionally mapped to each other based on the similarity of their reactomes.Example 2

[0170] FIG. 12 provides example data indicating gene expression profiles across organs, in response to drugs.

[0171] FIG. 12A shows gene expression for the rat Bexl gene for all 10 organs. FIG. 12A shows the normalized log2 expression of the rat Bexl gene using saline control injection in the 10 different organs as a radar plot, with individual traces shown for each of the 6 animals. The reproductive organ denotes testes for male animals and ovaries for female animals, and PBMC denotes the peripheral mononuclear blood cells separated from the blood buffy coat. For rat Bexl, some organs showed highly consistent expression levels across the 6 animals (e.g. brain), and other organs showed significant individual variability (e.g. kidney and heart). The variability of gene expression in control samples can potentially arise from a combination of the individual's biological state and technical variability during the RNA extraction and next-generation sequencing (NGS) library preparation process, with the latter being a bigger contributor for low expression and the former being a bigger contributor at high expression.

[0172] FIG. 12B shows the rat Bexl gene expression in different organs after dosing with 200 mg / kg tetracycline, 500 mg / kg valproate, or 2000 mg / kg carbon tetrachloride (CC14). The rat Bexl gene was selected for illustration here because it exhibits similar tissue gene expression responses for all three drugs: up-regulation in liver and down-regulation in heart. Most genes in general do not exhibit similar expression responses to all three drugs

[0173] FIG. 12C and FIG. 12C illustrate the calculation of the rat and mouse Rassf5 gene's tetracycline reactome values. The arithmetic mean of the normalized log2 expression values of all 6 animals for the control group are subtracted from the corresponding arithmetic mean of the normalized log2 expression values for the tetracycline dosed group to produce a single scalar number for each organ. Note that for the reproductive organ, the expression levels of the 3 male testes samples and the expression levels of the 3 female ovaries samples were averaged together. These reproductive organs were treated as a single tissue group rather than two separate tissue groups because we observed that the vast majority of genes exhibit similar expression for ovaries and testes. The mouse Rassf5 gene has similar a tetracycline reactome as the rat Rassf5 gene, though the mouse shows up-regulation in PBMC that is absent in rat.

[0174] Herein, the LI norm of the reactome was used as the summary value(s) for quantifying the overall degree of expression perturbation for a gene. The LI norm is the sum of the absoluste values of all scalar value components of the reactome. For example, the LI norm of the vector (- 1, 2, 0) is 3. FIG. 12E shows the cumulative distribution function (CDF) plots of reactome LI norm values for genes in the mouse and rat genes for tetracycline, valproate, and CC14. From FIG. 12E, a wide distribution of reactome LI norm values for the mouse and rat genes were observed, and a small number of genes with large reactome LI norm values contributing to the long tail of the distribution were seen. Roughly 50% of genes exhibited reactome LI norms below about 3 for all three drugs. The LI norm was chosen as the summary value because such an analog approach avoids outsized changes in reactome values arising from slight measurement errors around cutoff thresholds that would occur when discretizing gene expression changes.

[0175] To assess the degree of overlap between genes affected by the drugs, lists of genes with gene expression changes over 4-fold (2 units of log2 expression) in any (union) of the 10 organs for each drug were constructed, and their overlap distribution is shown in FIG. 2F. For both mouse and rat, the genes impacted by the three drugs appear to be relatively independent, as all sectors of the Venn diagram exhibited a significant population of genes. This outcome suggests that more independence between genes in their reactomes to different drugs allows for more information for performing gene mapping. In contrast, if gene expression reactomes were nearly identical across the three drugs, the second and third drugs would not provide significant additional information for performing cross-species gene mapping.Example 3

[0176] A total of 18407 genes out of the 35 085 mouse genes and 29 376 rat genes share the same name, based on high protein structural homology. Given the close phylogenetic distance between mouse and rat, the same-name gene pairs across mice and rats having relatively higher reactome similarities (e.g., low reactome LI norm distances) were expected. FIG. 13 provides example data indicating that protein structure data similarities do not correlate with gene expression profiles in response to drugs. FIG. 13A depicts the structural homology for rat and mouse genes with the same gene names. Furthermore, the same-name gene pairs exhibited a distribution of protein structures reflected in both the primary structure (amino acid sequence, FIG. 13A) and the AlphaFold-folded structures. Within these same-name gene pairs, higher reactome similarity in the gene pairs with higher structural similarity were expected.

[0177] Surprisingly, the data depicted in FIG. 13B suggested no significant correlation between structural similarity and reactome distance for same-name gene pairs (FIG. 13B). In particular, none of the same-name gene pairs with over 98% a.a. identity exhibited reaction LI distance of below 2 (z.e., the lowest bin). Furthermore, the gene expression of same-name gene pairs across organs in the control animals also showed essentially no correlation with the structural similarity. These findings indicated that organ- specific gene expression, both at homeostasis and in response to drug dosing, does not translate across even closely related species. By extension, to the extent that toxicology of new drug molecules are correlated across species, such new drug molecules may manifest in very different ways to affect different sets of genes. The details of gene regulatory networks in systems biology appear to be species-specific.

[0178] Given the unexpected nature of the results, the detailed reactomes of a number of samename genes were analyzed to spot-check the conclusions. FIG. 13C shows the reactomes of the rat and mouse Dmtrc2 gene (98.4% a.a. identity). The responses were highly distinct with mouse Dmtrc2 down-regulated in response to all 3 drugs in skin and reproductive organs, and rat Dmtrc2 up-regulated in liver and pancreas. FIG. 13D shows the highly similar reactomes of the rat and mouse Cenpk gene (68.4% a.a. identity), and is an example of a same name gene pair with a low a.a. identity and a high reactome similarity.

[0179] Next, the mouse-rat gene pairs with the lowest reactome LI distances (highest reactome similarity) were analyzed, to confirm whether there was a correlation between reactome LI distance and protein a.a. identity. FIG. 13E shows these results, and again, no significant correlation was found. FIG. 13E shows an analysis of the best reactome-matched genes, for genes with no corresponding same-name gene in the other rodent species, and a lack of correlation between reactome LI distance and protein a.a. identity was observed. FIG. 13F plots the sorted distribution of reactome LI distances of the same-name mouse-rat gene pairs, and also shows for comparison, the reactome LI distance of the best reactome-matched gene. In only 400 out of the 18407 genes (2.1%) was the same-name gene also the best reactome match.Example 4

[0180] Given the data that some cross-species gene pairs exhibit significantly better reactome similarity than both same-name gene pairs, the functional mapping of genes across species based on the reactome were explored. FIG. 14 provides example data depicting cross-species functional gene mapping based on gene expression profiles. The degree of gene matching was quantified using Shannon entropy, a concept commonly used in information theory to represent the degree of uncertainty of a variable. FIG. 14A provides a schematic explaining Shannon entropy in the context of cross-species gene-mapping. As shown in FIG. 14A, when mapping the n possible rat genes to one mouse gene, no information initially exists, so every rat gene is equally like to be mapped, and the Shannon entropy on the mouse gene is computed as Iog2(n). Similarly, each rat gene starts with a Shannon entropy of log2(m), because the Shannon entropy for the two species is asymmetrical. As shown in FIG. 14A, as rat genes are excluded from matching to a particular mouse gene based on the reactome values, the Shannon entropy of the mouse gene decreases to a minimum of 0 (perfect matching). In practice, however, no gene pairs are expected to reach 0Shannon entropy because of measurement errors on expression and because of differences in biology.

[0181] The Shannon entropy of a mouse gene, Emouse(i), was determined based on the following formulas:(1) Emouse(i) = - Xj p(i,j) • log2(p(i,j))(2) p(i,j) = SA(-D(i,j)) / Z(i)(3) D(i,j) = Wc • Io Abs(Xmouse(i,O) - Xrat(j,O)) + Io Id Abs(R mouse (i,o,d) - Rrat(j,o,d))(4) R mouse (i,o,d) = X mouse (i,o,d) - X mouse (i,o)(5) Rrat (j,O,d) = Xrat(j,O,d) - Xrat(j,O)(6) z = I j SA(-D(i,j)) where: p(i,j) denotes probability of a match between mouse gene i and rat gene jS denotes the sharpness of the probability dependence on reactome distanceD(i,j) denotes the reactome distance between mouse gene i and rat gene jWcdenotes the control tissue expression weightingXmouse(i,o) denotes the expression (normalized log2 units) of mouse gene i in organ oXrat(j,o) denotes the expression of rat gene j in organ oRmouse(i,o,d) denotes the reactome (change in expression) of mouse gene i in organ o for drug dRrat(j,o,d) denotes the reactome (change in expression) of rat gene j in organ o for drug dZ denotes the partition function

[0182] The Shannon entropy calculations used depend sensitively on two hyperparameters: control expression weighting, Wc, and probability sharpness, S. Higher values of Wcprioritize the similarity of expression levels in control animals over the reactomes. Higher values of S more harshly penalize the fit probability of a gene pair based on the reactome distance, D, resulting in lower Shannon entropies but being more fragile to measurement errors. Expressedanother way, the value of sharpness S describes how strongly the method favors a marginally better match (in the form of lower reactome LI distance). In general, larger values of S lead to smaller values of Shannon Entropy E, but increases the rate of false matching between unrelated genes. In the extreme case of S = infinity, then the method would assign probability 1 to matching between a rat gene and a mouse gene with the lowest reactome LI distance, even if there’s a second mouse gene that is nearly identical in its reactome (e.g. off by 0.01). In practice, two mouse genes with nearly identical reactomes (i.e. within RNA expression measurement noise) should both be assigned probability 0.5, assuming there are no other mouse genes that have remotely similar reactomes. In contrast at the other extreme case of S = 1, then the method essentially ignores all information provides by the reactome, and assigns equal probability to all mouse genes to map to a particular rat gene. Consequently an intermediate value of S is ideal to balance the sensitivity and specificity of cross-species gene matching, with higher values of S favoring higher sensitivity but yielding lower specificity.

[0183] To quantify the degree of potential false cross-species gene mapping, we created a “shuffled” dataset in which the 35,085 mouse genes’ reactome values for each organ / drug pair is randomly permuted. For example, in the shuffled dataset, gene Ml may be randomly assigned the liver • tetracycline reactome value for M2, the liver • valproate reactome value for M2000, the lung • CC14 reactome value for Ml 2000, etc.. In this shuffled dataset, any gene mapping between the rat and mouse transcriptomes would be purely coincidental, and all gene mapping would by construction by nonspecific.

[0184] For a particular sharpness value S, we can generate a Shannon entropy E histogram distribution for the shuffled dataset. In general, this E distribution will be biased to lower values of E, compared to the real dataset. We define K as the maximum value of the scaling factor on the E histogram for the shuffled dataset that allows the scaled shuffled dataset E histogram to be circumscribed by the biological dataset E histogram. In the left panel of FIG. 4B, the brown bars shows the shuffled dataset E histogram scaled by K=0.0987 and the blue bars show the biological dataset E histogram for mapping rat genes onto mouse genes. Likewise, the right panel of FIG. 4B shows the scaled shuffled E histogram in brown and the biological E histogram in blue for mapping mouse genes onto rat genes (K = 0.0937).

[0185] One way of interpreting the value of K is the maximum false matching rate. In other words, (1-K) is the minimum specificity of the method for a given value of S. At S = infinity, the observed value of K approaches 1, but presumably a large fraction (e.g. at least 50%) of the mapped gene pairs with minimal least reactome distance are actually correct.

[0186] FIG. 14C shows the observed values of mean Shannon entropy E and maximum false matching rate K for different values of S. We find that S = 2.7 is the maximum value of S that ensures K < 10%, corresponding to at least 90% specificity, and use S=2.7 for the remaining analysis. In addition, the control tissue expression weighting is set to 1 / 3 (Wc = 1 / 3).

[0187] The final distributions of Shannon entropy for the mouse and rat genes using both control tissue expression and drug reactome data, sorted in ascending order, are shown in FIG. 14D. Roughly 60% of the genes had Shannon entropies relatively uniformly distributed across 0 and 4, which indicates varying levels of precision in gene mapping. Simultaneously, another mode group in the distribution comprised 10%-20% of genes and had high Shannon entropies above 10. Genes in this mode group are primarily genes with low expression and with minimal response to drug dosing, and separating or matching these genes based on the collected data is practically unfeasible.

[0188] One feature of the information theory -based approach to cross-species gene mapping based on reactomes is that the approach is scalable. As the amount of expression data collected in response to additional drugs and / or from additional organs increases, the dimensionality of the reactome vectors increases. FIG. 14E shows the mean Shannon entropies of mouse and rat genes, starting from just considering control tissue expression levels, and successively adding the reactomes for each drug. A steady trend of declining Shannon entropies with data from each additional drug is observed, with a marginal decrease of 0.4 to 0.5 units of Shannon entropy per additional drug. If this scaling law holds with additional drugs, the Shannon entropy for crossspecies gene mapping can likely be minimized to near 0 with 8 to 12 additional drugs (for 11 to 15 drugs in total). Critical to this assumption of continued linearity in Shannon entropy decrease is that the additional drugs must be relatively independent in their mechanisms of action. For the 3 drugs tested here, FIG. 12F showed that the drugs are relatively independent in terms of the genes that they affect.

[0189] FIG. 14F shows the scaling of the Shannon entropy based on increasing the number of analyzed organ reactomes. Like with drugs, a consistent decrease in Shannon entropy is observed, as additional tissue types are analyzed. The current already included organs cover the most common organ systems, but these organs / tissue types could be further subdivided. For example, brain tissue could be subdievided by lobe, and small intestines could include ileum and jejunum tissue, as well as duodenum tissue. Given, however, that one important application of gene mapping is to predict toxicology effects for new potential drug molecules, scaling via additional drugs would be preferable to additional tissue types.Example 5

[0190] In addition to the endpoint organ samples, RNAseq was performed also on whole blood collected on a daily basis from each animal. Whole blood is distinct from the peripheral blood mononuclear cells (PBMCs) that was analyzed for the reactomes, because roughly 98% of the whole blood RNA derives red blood cells, and PBMCs comprise only a small fraction of the remaining 2% of cells in the buffy coat. The longitudinal samples collected and analyzed are whole blood because only 20 uL of whole blood could be collected from mice on a daily basis, without severely affecting the animal’s health. Separating PBMCs from 20 uL of whole blood is not currently feasible with any commercially available instruments or solutions.

[0191] Previous US patent applications 63 / 539,082 and 63 / 602,039 and at least a portion of its systems and methods relating to longitudinal whole blood RNAseq analysis indicated that whole blood RNA expression contained significant numbers of temporally varying genes that dynamically express in response to small molecule drug dosing. FIG. 15 provides data relating to longitudinal whole blood RNAseq analysis, in relation to the methods and systems described herein. FIG. 15 provides example data depicting longitudinal whole blood gene expression profiles in response to drugs versus organ endpoint gene expression profiles in response to drugs. FIG. 15A shows that the mouse gene Pigh and the rat gene Dgat2 both show highly similar organ endpoint reactomes to CC14. FIG. 15B shows that the temporal expression patterns in whole blood of mouse Pigh and rat Dgat2 are highly distinct from each other, with rat Dgat2 showing an up-regulation response and mouse Pigh showing a down-regulation response.Furthermore, mouse Pigh exhibits a bleeding acclimation up-regulation response on Day 2 of the experiment that is absent in rat Dgat2. FIG. 15C shows that the mouse Nop53 and rat Flnb genesare observed to have no expression response to CC14 for all organ endpoints. FIG. 15D shows that in temporal whole blood samples, both genes are observed to have significant downregulation response. The mouse Nop53 further exhibits an up-regulation bleeding acclimation response. FIGS. 16A, 16B, 16C and 16D suggest that whole blood RNA expression correlates highest with PBMC and pancreas tissues. The observed RNA expression correlates for blood versus PBMC and pancreas tissues can be used to infer, e.g., predict, RNA expression amounts, e.g., dynamics, from a blood sample, provided a generative Al model such as those used in the methods and systems described herein. FIG. 15E depicts a summary of gene expression reactome responses in temporal blood samples versus gene expression reactome responses in different organs at endpoint. Whole blood reactome does not appear to correlate strongly with any of the organs we characterized. That is, a significant fraction of genes (10%-30%) with expression responses in whole blood are not reflected in any other organ. Pancreas and PBMCs showed the highest overlap in drug response genes with whole blood, but in both cases, the overlap was less than 40% of genes for all three drugs. In some sense, the data shown in FIG. 15 can be interpreted as supporting an analysis position where whole blood (and by extension, red blood cells) are practically considered a different organ or tissue type, given its own pattern of drug responses, e.g., reactomes.

[0192] It should be understood from the foregoing that, while particular implementations of the disclosed methods and systems have been illustrated and described, various modifications can be made thereto and are contemplated herein. It is also not intended that the invention be limited by the specific examples provided within the specification. While the invention has been described with reference to the aforementioned specification, the descriptions and illustrations of the preferable embodiments herein are not meant to be construed in a limiting sense. Furthermore, it shall be understood that all aspects of the invention are not limited to the specific depictions, configurations or relative proportions set forth herein which depend upon a variety of conditions and variables. Various modifications in form and detail of the embodiments of the invention will be apparent to a person skilled in the art. It is therefore contemplated that the invention shall also cover any such modifications, variations and equivalents.

Claims

CLAIMSWhat is claimed is:

1. A method for training a generative artificial intelligence (Al) model for providing generated data for a target biological species based on training data from a training biological species, comprising: conducting an experiment on a subject from the training biological species; measuring one or more amounts of a biomolecule from a sample from the experiment; creating a vector based on the measured one or more amounts of the biomolecule; determining a summary value based on the created vector; and training the generative Al model to provide the generated data for a subject from the target biological species, using the determined summary value.

2. A method for generating generated data for a subject of a target biological species, comprising: inputting experiment data for a subject from a target biological species into a generative Al model trained on a training biological species, the generative Al model having been trained by a method comprising: conducting an experiment on a subject from the training biological species; measuring one or more amounts of a biomolecule from a sample from the experiment; creating a vector based on the measured one or more amounts of the biomolecule; determining a summary value based on the created vector; and training the generative Al model to provide the generated data for a subject from the target biological species, using the determined summary value; and receiving the generated data.

3. The method of claim 2, further comprising: inputting the generated data into an additional trained machine learning model configured to predict a clinical outcome for the subject from the target biological species; and receiving from the additional trained machine learning model, the predicted clinical outcome for the subject from the target biological species.

4. The method of claim 3, wherein the predicted clinical outcome comprises determining a disease signal in the subject.

5. The method of claim 3 or 4, wherein the predicted clinical outcome comprises determining a diagnosis for the subject.

6. The method of any of claims 3-5, wherein the predicted clinical outcome comprises determining a therapy for the subject.

7. The method of claim 6, wherein the determining the therapy for the subject comprises determining a dosage of the therapy.

8. The method of claim 6, wherein the predicted clinical outcome comprises administering the therapy to the subject.

9. The method of any of claims 3-8, wherein the predicted clinical outcome comprises determining a prognosis for the subject.

10. The method of any of claims 9, wherein the prognosis comprises determining a response from the subject to a predetermined therapy.

11. The method of claim 10, wherein the response comprises an adverse response to the predetermined therapy for the subject.

12. The method of claim 10, wherein the response comprises a recovery from the predetermined therapy for the subject.

13. The method of any of claims 1-12, wherein the target biological species is a human.

14. The method of any of claims 1-13, wherein the determined summary value is based on one or more relative amounts of the biomolecule.

15. The method of claim 14, wherein the one or more relative amounts of the biomolecule comprise comparing the measured one or more amounts to measured one or more amounts from a control group.

16. The method of any of claims 1-15, wherein the summary value comprises determining a norm.

17. The method of claim 16, wherein the norm is an Lp norm.

18. The method of claim 16 or 17, wherein the norm is an LI norm, an L2 norm, or an L max norm.

19. A method for training a generative Al model for inferring one or more target amounts in a target anatomical region, comprising: conducting an experiment on a subject; measuring one or more amounts of a biomolecule from a sample from the experiment; and training the generative Al model to infer the one or more target amounts in the target anatomical region, based on the one or more measured amounts.

20. A method of inferring one or more target amounts of a first biomolecule in a target anatomical region, comprising:conducting an experiment to obtain a sample from an input anatomical region of a subject; measuring, from the sample, one or more amounts of the first biomolecule or a second biomolecule; inputting the one or more measured amounts into a trained generative Al model; and inferring, from the one or more measured amounts, the one or more target amounts of the first biomolecule in the target anatomical region.

21. The method of claim 19 or 20, wherein the obtaining the sample comprises non- invasively obtaining the sample.

22. The method of any of claims 19-21, wherein the obtaining the sample comprises drawing blood from the subject.

23. The method of claim 21 or 22, wherein the sample is obtained from a necropsy of the subject.

24. The method of any of claims 19-23, wherein the input anatomical region comprises peripheral blood mononuclear cells, red blood cells, or a combination thereof.

25. The method of any of claims 19-24, wherein the target anatomical region comprises pancreas.

26. A method for training a generative Al model for providing generated data comprising: conducting an experiment on a subject; measuring one or more amounts of a biomolecule from a sample from the experiment; including, in a training dataset exclusionary of pre-existing structural data, a value based on the one or more measured amounts; andtraining the generative Al model to provide the generated data, based on the training dataset.

27. A method for generating generated data, comprising: inputting experiment data into a generative Al model, the generative Al model having been trained by a method comprising: conducting an experiment on a subject; measuring one or more amounts of a biomolecule from a sample from the experiment; including, in a training dataset exclusionary of pre-existing structural data, a value based on the one or more measured amounts; and training the generative Al model to provide the generated data, based on the training dataset; and receiving the generated data.

28. A method for training a generative Al model for providing generated data comprising: conducting an experiment on a subject; measuring one or more amounts of a biomolecule from a sample from the experiment; including, in a training dataset, a value based on the one or more measured amounts; excluding from the training dataset, pre-existing structural data corresponding to a biomolecule; and training the generative Al model to provide the generated data, based on the training dataset.

29. A method for generating generated data, comprising: inputting experiment data into a generative Al model, the generative Al model having been trained by a method comprising: conducting an experiment on a subject;measuring one or more amounts of a biomolecule from a sample from the experiment; including, in a training dataset, a value based on the one or more measured amounts; excluding from the training dataset, pre-existing structural data corresponding to a biomolecule; and training the generative Al model to provide the generated data, based on the training dataset; and receiving the generated data.

30. The method of any of claims 1-29, wherein the experiment comprises providing a stimulus to the subject.

31. The method of claim 30, wherein the stimulus is a drug.

32. The method of claim 31, wherein the drug is isoniazid, tetracycline, carbon tetrachloride, valproate, or a combination thereof.

33. The method of any of claims 30-32, wherein the providing the stimulus comprises administering a dosage of the drug that is at or above a predetermined dosage level.

34. The method of any of claims 1-33, wherein the one or more measured amounts is based on RNA sequencing.

35. The method of any of claims 1-34, wherein the one or more measured amounts is based on mass spectrometry.

36. The method of any of claims 1-35, wherein the one or more measured amounts is based on methyl- sequencing.

37. The method of any of claims 1-36, wherein the one or more measured amounts is based on DNA sequencing.

38. The method of any of claims 1-37, wherein the sample comprises liquid biopsy samples.

39. The method of claim 38, wherein the liquid biopsy sample comprises blood, plasma, cerebrospinal fluid, sputum, stool, urine, sweat, or saliva.

40. The method of any of claims 1-39, wherein the sample comprises a tissue biopsy sample.

41. The method of any of claims 1-40, wherein the generated data comprises generated RNA sequencing data, generated mass spectrometry data, generated DNA sequencing data, generated methyl- sequencing data or a combination thereof.

42. A method of tokenizing data exclusionary of pre-existing structural data, to generate tokens for a generative Al model, comprising: determining a signal strength value from the data; determining an entropy value based on the signal strength value; and outputting tokens based on the determined entropy value.

43. A method of tokenizing data to generate tokens for a generative Al model, comprising: excluding from the data, pre-existing structural data corresponding to a biomolecule; determining a signal strength value from the data; determining an entropy value based on the signal strength value; and outputting tokens based on the determined entropy value.

44. The method of claim 43, wherein the entropy value is a Shannon entropy value.

45. The method of claim 43 or 44, wherein the entropy value is based on determining the probability of a correspondence between an amount of a first biomolecule of a firstbiological species and the amount of a second biomolecule of a second biological species.

46. The method of claim 45, wherein the determining the probability of the correspondence is based on a predetermined probability sharpness value, a predetermined control expression weighting, or a combination thereof.

47. The method of any of claims 26-46, wherein the pre-existing structural data comprises sequence data or molecular structure data.

48. The method of claim 47, wherein the sequence data comprises DNA sequences, RNA sequences, or amino acid sequences.

49. The method of claim 47 or 48, wherein the sequence data comprises sequence similarity values.

50. The method of claim 47, wherein the molecular structure data comprises protein structure data.

51. The method of any of claims 47-50, wherein the excluding the pre-existing structural data is based on a correspondence value indicating the correspondence between the preexisting structural data and the entropy value.

52. The method of any of claims 1-51, wherein the biomolecule is an RNA molecule.

53. The method of claim 52, wherein the RNA molecule is an mRNA transcript.

54. The method of any of claims 1-53, wherein the biomolecule is a protein molecule.

55. The method of any of claims 1-54, wherein the biomolecule is an epigenetic marker.

56. The method of any of claims 1-55, wherein the biomolecule is a methyl group on a nucleotide base.

57. The method of any of claims 1-56, wherein the biomolecule is a DNA molecule.

58. The method of claim 57, wherein the DNA molecule is a cell-free DNA (cfDNA).

59. The method of any of claims 1-58, wherein the biomolecule is a metabolite.

60. A system comprising: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: conduct an experiment on a subject from the training biological species; measure one or more amounts of a biomolecule from a sample from the experiment; create a vector based on the measured one or more amounts of the biomolecule; determine a summary value based on the created vector; and train the generative Al model to provide the generated data for a subject from the target biological species, using the determined summary value.

61. A system comprising: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: input experiment data for a subject from a target biological species into a generative Al model trained on a training biological species, the generative Al model having been trained by a method comprising:conducting an experiment on a subject from the training biological species; measuring one or more amounts of a biomolecule from a sample from the experiment; creating a vector based on the measured one or more amounts of the biomolecule; determining a summary value based on the created vector; and training the generative Al model to provide the generated data for a subject from the target biological species, using the determined summary value; and receive the generated data.

62. The system of claim 60, comprising further instructions that, when executed by the one or more processors, cause the system to: input the generated data into an additional trained machine learning model configured to predict a clinical outcome for the subject from the target biological species; and receive from the additional trained machine learning model, the predicted clinical outcome for the subject from the target biological species.

63. A system comprising: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: conduct an experiment on a subject; measure one or more amounts of a biomolecule from a sample from the experiment; and train the generative Al model to infer the one or more target amounts in the target anatomical region, based on the one or more measured amounts.

64. A system comprising: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: conduct an experiment to obtain a sample from an input anatomical region of a subject; measure, from the sample, one or more amounts of the first biomolecule or a second biomolecule; input the one or more measured amounts into a trained generative Al model; and infer, from the one or more measured amounts, the one or more target amounts of the first biomolecule in the target anatomical region.

65. A system comprising: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: conduct an experiment on a subject; measure one or more amounts of a biomolecule from a sample from the experiment; include, in a training dataset exclusionary of pre-existing structural data, a value based on the one or more measured amounts; and train the generative Al model to provide the generated data, based on the training dataset.

66. A system comprising: one or more processors; anda memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: input experiment data into a generative Al model, the generative Al model having been trained by a method comprising: conducting an experiment on a subject; measuring one or more amounts of a biomolecule from a sample from the experiment; including, in a training dataset exclusionary of pre-existing structural data, a value based on the one or more measured amounts; and training the generative Al model to provide the generated data, based on the training dataset; and receive the generated data.

67. A system comprising: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: conduct an experiment on a subject; measure one or more amounts of a biomolecule from a sample from the experiment; include, in a training dataset, a value based on the one or more measured amounts; exclude from the training dataset, pre-existing structural data corresponding to a biomolecule; and train the generative Al model to provide the generated data, based on the training dataset.

68. A system comprising: one or more processors; anda memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: input experiment data into a generative Al model, the generative Al model having been trained by a method comprising: conducting an experiment on a subject; measuring one or more amounts of a biomolecule from a sample from the experiment; including, in a training dataset, a value based on the one or more measured amounts; excluding from the training dataset, pre-existing structural data corresponding to a biomolecule; and training the generative Al model to provide the generated data, based on the training dataset; and receive the generated data.

69. A system comprising: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: determine a signal strength value from the data; determine an entropy value based on the signal strength value; and output tokens based on the determined entropy value.

70. A system comprising: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to:exclude from the data, pre-existing structural data corresponding to a biomolecule; determine a signal strength value from the data; determine an entropy value based on the signal strength value; and output tokens based on the determined entropy value.

71. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of a system, cause the system to: conduct an experiment on a subject from the training biological species; measure one or more amounts of a biomolecule from a sample from the experiment; create a vector based on the measured one or more amounts of the biomolecule; determine a summary value based on the created vector; and train the generative Al model to provide the generated data for a subject from the target biological species, using the determined summary value.

72. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of a system, cause the system to: input experiment data for a subject from a target biological species into a generative Al model trained on a training biological species, the generative Al model having been trained by a method comprising: conducting an experiment on a subject from the training biological species; measuring one or more amounts of a biomolecule from a sample from the experiment; creating a vector based on the measured one or more amounts of the biomolecule; determining a summary value based on the created vector; andtraining the generative Al model to provide the generated data for a subject from the target biological species, using the determined summary value; and receive the generated data.

73. The non-transitory computer-readable storage medium of claim 72, further comprising instructions that, when executed by the one or more processors, cause the system to: input the generated data into an additional trained machine learning model configured to predict a clinical outcome for the subject from the target biological species; and receive from the additional trained machine learning model, the predicted clinical outcome for the subject from the target biological species.

74. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of a system, cause the system to: conduct an experiment on a subject; measure one or more amounts of a biomolecule from a sample from the experiment; and train the generative Al model to infer the one or more target amounts in the target anatomical region, based on the one or more measured amounts.

75. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of a system, cause the system to: conduct an experiment to obtain a sample from an input anatomical region of a subject; measure, from the sample, one or more amounts of the first biomolecule or a second biomolecule; input the one or more measured amounts into a trained generative Al model; andinfer, from the one or more measured amounts, the one or more target amounts of the first biomolecule in the target anatomical region.

76. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of a system, cause the system to: conduct an experiment on a subject; measure one or more amounts of a biomolecule from a sample from the experiment; include, in a training dataset exclusionary of pre-existing structural data, a value based on the one or more measured amounts; and train the generative Al model to provide the generated data, based on the training dataset.

77. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of a system, cause the system to: input experiment data into a generative Al model, the generative Al model having been trained by a method comprising: conducting an experiment on a subject; measuring one or more amounts of a biomolecule from a sample from the experiment; including, in a training dataset exclusionary of pre-existing structural data, a value based on the one or more measured amounts; and training the generative Al model to provide the generated data, based on the training dataset; and receive the generated data.

78. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of a system, cause the system to:conduct an experiment on a subject; measure one or more amounts of a biomolecule from a sample from the experiment; include, in a training dataset, a value based on the one or more measured amounts; exclude from the training dataset, pre-existing structural data corresponding to a biomolecule; and train the generative Al model to provide the generated data, based on the training dataset.

79. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of a system, cause the system to: input experiment data into a generative Al model, the generative Al model having been trained by a method comprising: conducting an experiment on a subject; measuring one or more amounts of a biomolecule from a sample from the experiment; including, in a training dataset, a value based on the one or more measured amounts; excluding from the training dataset, pre-existing structural data corresponding to a biomolecule; and training the generative Al model to provide the generated data, based on the training dataset; and receive the generated data.

80. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of a system, cause the system to: determine a signal strength value from the data; determine an entropy value based on the signal strength value; andoutput tokens based on the determined entropy value.

81. A non- transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of a system, cause the system to: exclude from the data, pre-existing structural data corresponding to a biomolecule; determine a signal strength value from the data; determine an entropy value based on the signal strength value; and output tokens based on the determined entropy value.

Citation Information

Patent Citations

  • Methods for using artificial neural network analysis on flow cytometry data for cancer diagnosis

    US20180247195A1

  • Diagnostic Process for Disease Detection using Gene Expression based Multi Layer PCA Classifier

    US20200402660A1

  • Workflow for generating compounds with biological activity against a specific biological target

    US20210057050A1

  • Generative TNA sequence design with experiment-in-the-loop training

    US20230081439A1