Systems and methods for novel molecule generation
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- NEOCLEASE INC
- Filing Date
- 2025-07-25
- Publication Date
- 2026-08-06
AI Technical Summary
The generation of novel molecules for drug discovery and other applications is a costly and time-consuming process, often requiring numerous iterations and manual modification of pre-existing molecules, with high failure rates due to lack of in situ validation.
A computer-implemented method utilizing a language learning model (LLM) and a structure generation module (SGM) to generate novel molecules or biological systems, followed by predictive assessment tools for validation, including folding evaluation, molecular simulation, and molecular dynamics simulation, with feedback loops to refine the output and minimize off-target effects.
Facilitates the efficient and cost-effective generation of novel molecules or biological systems, reducing the need for manual intervention and improving the success rate by leveraging predictive assessment tools and feedback loops.
Smart Images

Figure US2025039305_06082026_PF_FP_ABST
Abstract
Description
SYSTEMS AND METHODS FOR NOVEL MOLECULE GENERATIONCROSS-REFERENCE
[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 676,115 filed July 26, 2024 which is incorporated by reference herein in its entirety.BACKGROUND
[0002] Generation of novel molecules that achieve a desired result, for example in drug discovery, is a notoriously expensive and lengthy process. Generating novel molecules that are effective for certain purposes, such as treatment of genetic disease, is becoming increasingly complicated and lengthy as treatment needs progress beyond simple chemical compounds to biologies with complicated structures and functions. In most cases, hundreds or thousands of iterations of molecules are developed and tested for efficacy before a molecule is discovered that adequately achieves a desired result, such as disease treatment.
[0003] Additionally, drug discovery performed in a lab is often a manual process constrained to modifying molecules readily in nature or previously synthesized. Researchers often sift through large libraries of pre-existing molecules and choose a limited number of molecular modifications from at least thousands, if not millions, of possible combinations of modifications. The validation of these modified molecules often involves expensive and lengthy rounds of in vitro testing and in vivo testing, only to reveal the ineffectiveness of the modified molecules.
[0004] Accordingly, the need exists for improved systems and methods of molecular generation, not constrained to pre-existing or natural molecules, and not subject to high rates of failure during validation testing due to lack of in situ validation and development. There is also a need for improved molecular generation in a more cost-efficient and time-efficient manner.
[0005] These and other objects, along with advantages and features of embodiments of the present invention herein disclosed, will become more apparent through reference to the following description, the figures, and the claims. Furthermore, it is to be understood that the features of the various embodiments described herein are not mutually exclusive and may exist in various combinations and permutations.SUMMARY
[0006] In some embodiments, disclosed herein is a computer-implemented method for assisting in generation of one or more novel molecules or biological systems, the methodcomprising: receiving an input comprising information related to a target function; providing the input to a language learning model (LLM) and using the LLM to output one or more sequences of the one or more novel molecules or biological systems; applying a validation model to evaluate or predict one or more predictive assessment tools of the output of the LLM; and outputting a sequence of (b) or structural representation thereof, validated for the one or more predictive assessment tools of (c), wherein the output comprises the one or more novel molecules or biological systems.
[0007] In some embodiments, the molecule is a protein. In some embodiments, the molecule is a nuclease. In some embodiments, the biological systems comprise DNA, RNA, ions, lipids, nanoparticles, small molecules, or any combination thereof. In some embodiments, the input further comprises one or more of raw nuclease sequences, structural data, simulation data, quantum calculation data, or any combination thereof. In some embodiments, the input further comprises information related to one or more phenotypic characteristics.
[0008] In some embodiments, the input is pre-processed to be configured as LLM input. In some embodiments of the method, (b) further comprises providing the input to a structure generation module (SGM), and using the SGM to output one or more structures of the one or more novel molecules or biological systems. In some embodiments, the LLM receives as feedback input the one or more structures output by the SGM to refine the LLM output of the one or more sequences of the one or more novel molecules or biological systems. In some embodiments, the SGM receives as feedback input the one or more sequences output by the LLM to refine the SGM output of the one or more structures of the one or more novel molecules or biological systems.
[0009] In some embodiments of the method, (c) further comprises applying a trained classifier module to classify into one or more groups the one or more sequences of the one or more novel molecules or biological systems output in (b). In some embodiments of the method, the one or more predictive assessment tools of (c) comprise folding evaluation, structural evaluation, molecular simulation, molecular dynamics simulation, or any combination thereof. In some embodiments of the method, the validation of (c) using the folding evaluation predictive assessment tool further comprises classifying the output of the folding evaluation tool into one or more categories based on one or more predefined parameters. In some embodiments of the method, the one or more predefined parameters comprise similarity to known molecular structures, solubility, sequence features, foldability, cleavage features, catalytic sites features, or any combination thereof. In some embodiments of the method, the validation of (c) using the molecular simulation predictive assessment tool further comprises evaluating the output of themolecular simulation predictive assessment tool using one or more parameters. In some embodiments, the one or more parameters comprise comparing the output of the molecular simulation tool to the molecular bonding of known molecules, energy for binding, catalytic sites features, active sites features, sequence features, PAM sequence features, cleavage features, charge, microenvironment features, delivery simulation without cleavage, or any combination thereof.
[0010] In some embodiments of the method, the validation of (c) using the molecular dynamics simulation tool further comprises evaluating the output of the molecular dynamics simulation tool using one or more parameters. In some embodiments, the one or more parameters comprise energy of one or more reactions, one or more energies of activation, one or more reaction pathways, one or more energies calculated using molecular mechanics, one or more energies calculated using quantum mechanics, or one or more energies calculated using both quantum mechanics and molecular mechanics, or any combination thereof. In some embodiments of the method, the output of the validation model in (c) is used as feedback input for the LLM model of (b) to refine the output of the LLM model of (b).
[0011] In some embodiments, the method further comprises a feedback loop from the output of the models of (b), (c), or both, wherein the feedback loop is used to enhance the one or more novel molecules or biological systems to minimize off-target effects. In some embodiments, the one or more novel molecules or biological systems comprise miniaturized nucleases. In some embodiments, the input further comprises genetic information associated with a specific disease.
[0012] In some embodiments, the method further comprises generating a delivery system component corresponding to the output of (d) to deliver the one or more novel molecules or biological systems to a target. In some embodiments of the method, the delivery system component is generated using the output of (b) or (c), or both. In some embodiments of the method, the output of (b), or (c), or both, is used as feedback input to the LLM of (b) to generate the delivery system component. In some embodiments, the input further comprises novel genes or mutations associated with one or more disease states.
[0013] In some embodiments of the method, the one or more predictive assessment tools of the validation model of (c) utilize quantum computing. In some embodiments, the method further comprises in vitro testing of the output of (d). In some embodiments of the method, the input comprises feedback input based on the results of in vitro testing of the output of (d).
[0014] In some embodiments, the method further comprises in vivo testing of the output of (d). In some embodiments of the method, the input comprises feedback input based on the results ofin vivo testing of the output of (d). In some embodiments, in vivo testing further comprises testing in one or more mammals. In some embodiments, the one or more mammals further comprise a human, a non-human primate, a murine mammal, or any combination thereof.
[0015] Another embodiment disclosed herein comprises a computer-implemented method for assisting in generation of one or more novel molecules or biological systems, the method comprising: receiving an input comprising information related to a target function; providing the input to a structure generation model (SGM), and using the SGM to output one or more predicted structures of the one or more novel molecules or biological systems; applying a validation model to evaluate or predict one or more predictive assessment tools of the output of the SGM; and outputting a structure of (b) validated for the one or more predictive assessment tools of (c), wherein the output comprises the one or more novel molecules or biological systems.
[0016] In some embodiments, the molecule is a protein. In some embodiments, the molecule is a nuclease. In some embodiments, the biological systems comprise DNA, RNA, ions, lipids, nanoparticles, small molecules, or any combination thereof. In some embodiments, the input further comprises one or more of raw nuclease sequences, structural data, simulation data, quantum calculation data, related input data, or any combination thereof. In some embodiments, the input further comprises information related to one or more phenotypic characteristics. In some embodiments, the input is pre-processed to be configured as SGM input.
[0017] In some embodiments of the method, (c) further comprises applying one or more structural evaluation tools to the one or more structures of the one or more novel molecules or biological systems output in (b). In some embodiments of the method, the one or more structural evaluation tools further comprise a trained classifier module to classify the output of (b) into one or more groups. In some embodiments of the method, the one or more predictive assessment tools of (c) comprise folding evaluation, structural evaluation, molecular simulation, molecular dynamics simulation, or any combination thereof. In some embodiments of the method, the validation of (c) using the folding evaluation predictive assessment tool further comprises classifying the output of the folding evaluation tool into one or more categories based on one or more predefined parameters.
[0018] In some embodiments, the one or more predefined parameters comprise similarity to known molecular structures, solubility, sequence features, foldability, cleavage features, catalytic sites features, or any combination thereof. In some embodiments for the method, the validation of (c) using the molecular simulation predictive assessment tool further comprises evaluating the output of the molecular simulation predictive assessment tool using one or moreparameters. In some embodiments, the one or more parameters comprise comparing the output of the molecular simulation tool to the molecular bonding of known molecules, energy for binding, catalytic sites features, active sites features, sequence features, PAM sequence features, cleavage features, charge, microenvironment features, delivery simulation without cleavage, or any combination thereof.
[0019] In some embodiments of the method, the validation of (c) using the molecular dynamics simulation tool further comprises evaluating the output of the molecular dynamics simulation tool using one or more parameters. In some embodiments, the one or more parameters comprise energy of one or more reactions, one or more energies of activation, one or more reaction pathways, one or more energies calculated using molecular mechanics, one or more energies calculated using quantum mechanics, or one or more energies calculated using both quantum mechanics and molecular mechanics, or any combination thereof.
[0020] In some embodiments of the method, the output of the validation model in (c) is used as feedback input for the SGM model of (b) to refine the output of the SGM model of (b).
[0021] In some embodiments, the method further comprises receiving a feedback loop from the output of the models of (b), (c), or both, wherein the feedback loop is used to enhance the one or more novel molecules or biological systems to minimize off-target effects. In some embodiments, the one or more novel molecules or biological systems comprise miniaturized nucleases. In some embodiments, the input further comprises genetic information associated with a specific disease.
[0022] In some embodiments, the method further comprises generating a delivery system component corresponding to the output of (d) to deliver the one or more novel molecules or biological systems to a target. In some embodiments of the method, the delivery system component is generated using the output of (b) or (c), or both. In some embodiments of the method, the output of (b), or (c), or both, is used as feedback input to the SGM of (b) to generate the delivery system component. In some embodiments, the input further comprises novel genes or mutations associated with one or more disease states.
[0023] In some embodiments of the method, the one or more predictive assessment tools of the validation model of (c) utilize quantum computing. In some embodiments, the method further comprises in vitro testing of the output of (d). In some embodiments of the method, the input comprises feedback input based on the results of in vitro testing of the output of (d).
[0024] In some embodiments, the method further comprises in vivo testing of the output of (d). In some embodiments of the method, the input comprises feedback input based on the results ofin vivo testing of the output of (d). In some embodiments of the method, the in vivo testing further comprises testing in one or more mammals. In some embodiments, the one or more mammals further comprise a human, a non-human primate, a murine mammal, or any combination thereof.
[0025] Yet another embodiment disclosed herein is a computer-implemented method for assisting in generation of one or more novel target-specific nucleases, the method comprising: receiving input comprising information related to one or more targets; providing the input of (a) to a language learning model (LLM), and using the LLM to generate one or more nuclease genetic sequences based at least in part on the one or more targets; using a validation model to evaluate the one or more nuclease genetic sequences of (b) based on one or more predictive assessment tools, wherein the validation model outputs a modified sequence of (b) or a structure thereof; and outputting the one or more validated novel nuclease sequences or structures of (c).
[0026] In some embodiments, the one or more predictive assessment tools comprise folding evaluation, structural evaluation, molecular simulation, or any combination thereof.
[0027] In some embodiments of the method, the molecular simulation further comprises at least one of quantum chemistry evaluation or simulated molecular dynamics evaluation.
[0028] In yet another embodiment, described herein is a computer-implemented method for assisting in generation of one or more novel molecules or biological systems, the method comprising: applying a classification model to classify one or more inputs into one or more desired molecule types, wherein the one or more inputs comprise information related to one or more target functions; applying a language learning model (LLM) to generate one or more novel sequences of the one or more novel molecules or biological systems, wherein the one or more novel molecules or biological systems are of the one or more desired molecule types classified in (a); simulating in situ one or more novel structures of each of the one or more novel sequences of the one or more novel molecules or biological systems generated in (b); and outputting the one or more novel sequences of (b) or the one or more novel structures (c), or both.
[0029] In yet another embodiment, described herein is a system for assisting in generation of one or more novel molecules or biological systems, the system comprising: an input module configured to receive instructions comprising one or more desired features, targets, or functions; a language learning model (LLM) configured to generate one or more sequences corresponding the one or more novel molecules or biological systems having the one or more desired features, targets, or functions; a validation model configured to evaluate the one or more sequences of (b),or one or more structures thereof, using predictive assessment tools; and an output module configured to output the one or more sequences or the one or more structures of the one or more novel molecules or biological systems validated in (c).
[0030] In some embodiments, the molecule is a protein. In some embodiments, the molecule is a nuclease. In some embodiments, the biological systems comprise DNA, RNA, ions, lipids, nanoparticles, small molecules, or any combination thereof. In some embodiments, the input further comprises one or more of raw nuclease sequences, structural data, simulation data, quantum calculation data, or any combination thereof. In some embodiments, the input further comprises information related to one or more phenotypic characteristics. In some embodiments, further comprising a pre-processing module configured to transform the input to LLM input.
[0031] In some embodiments, the system further comprises a structure generation module (SGM). In some embodiments of the system, the LLM of (b) is further configured to provide the input to the SGM. In some embodiments, the SGM is configured to output one or more structures of the one or more novel molecules or biological systems. In some embodiments, the LLM is further configured to receive as feedback input the one or more structures output by the SGM to refine the LLM output of the one or more sequences of the one or more novel molecules or biological systems. In some embodiments, the SGM is further configured to receive as feedback input the one or more sequences output by the LLM to refine the SGM output of the one or more structures of the one or more novel molecules or biological systems.
[0032] In some embodiments of the system, (c) further comprises a classifier module. In some embodiments of the system, the classifier module is configured to classify into one or more groups the one or more sequences of the one or more novel molecules or biological systems output in (b). In some embodiments of the system, the one or more predictive assessment tools of (c) comprise folding evaluation, structural evaluation, molecular simulation, molecular dynamics simulation, or any combination thereof. In some embodiments of the system, the validation model of (c) is further configured to classify the output of the folding evaluation tool into one or more categories based on one or more predefined parameters.
[0033] In some embodiments, the one or more predefined parameters comprise similarity to known molecular structures, solubility, sequence features, foldability, cleavage features, catalytic sites features, or any combination thereof. In some embodiments of the system, the validation model of (c) is further configured to evaluate the output of the molecular simulation predictive assessment tool using one or more parameters. In some embodiments, the one or more parameters comprise comparing the output of the molecular simulation tool to the molecular bonding of known molecules, energy for binding, catalytic sites features, active sites features,sequence features, PAM sequence features, cleavage features, charge, microenvironment features, delivery simulation without cleavage, or any combination thereof.
[0034] In some embodiments of the system, the validation model of (c) is further configured to evaluate the output of the molecular dynamics simulation tool using one or more parameters. In some embodiments, the one or more parameters comprise energy of one or more reactions, one or more energies of activation, one or more reaction pathways, one or more energies calculated using molecular mechanics, one or more energies calculated using quantum mechanics, or one or more energies calculated using both quantum mechanics and molecular mechanics, or any combination thereof.
[0035] In some embodiments of the system, the validation model of (c) is further configured to transmit output to the LLM model of (b) to refine the output of the LLM model of (b).
[0036] In some embodiments, the system further comprises a system feedback loop from the output of the models of (b), (c), or both, wherein the feedback loop is used to enhance the one or more novel molecules or biological systems to minimize off-target effects. In some embodiments, the one or more novel molecules or biological systems comprise miniaturized nucleases. In some embodiments, input further comprises genetic information associated with a specific disease. In some embodiments of the system, the output of (d) further comprises generation of one or more delivery molecules to deliver the one or more novel molecules or biological systems to a target. In some embodiments of the system, the output further comprises output derived from the output of (b) or (c), or both.
[0037] In some embodiments of the system, the LLM is further configured to generate the one or more delivery molecules using output of (b), or (c), or both. In some embodiments, the input further comprises novel genes or mutations associated with one or more disease states. In some embodiments of the system, the one or more predictive assessment tools of the validation model of (c) utilize quantum computing.
[0038] In some embodiments, the system further comprises in vitro testing of the output of (d). In some embodiments of the system, the input comprises feedback input based on the results of in vitro testing of the output of (d).
[0039] In some embodiments, the system further comprises in vivo testing of the output of (d). In some embodiments of the system, the input comprises feedback input based on the results of in vivo testing of the output of (d). In some embodiments, the in vivo testing further comprises testing in one or more mammals. In some embodiments, one or more mammals further comprise a human, a non-human primate, a murine mammal, or any combination thereof.
[0040] In yet another embodiment, disclosed herein is a system for assisting in generation of one or more novel molecules sequences, structures, dynamic simulation trajectories, or biological systems, or any combination thereof, the system comprising: an input module configured to receive instructions comprising one or more desired features, targets, or functions; a structure generation model (SGM) configured to generate one or more sequences, structures, or dynamics simulation trajectories, or any combination thereof, corresponding the one or more novel molecules or biological systems having the one or more desired features, targets, or functions; a validation model configured to evaluate the one or more sequences of (b), or one or more structures thereof, using predictive assessment tools; and an output module configured to output the one or more sequences or the one or more structures of the one or more novel molecules or biological systems validated in (c).
[0041] In some embodiments, the molecule is a protein. In some embodiments, the molecule is a nuclease. In some embodiments, the biological systems comprise DNA, RNA, ions, lipids, nanoparticles, small molecules, or any combination thereof. In some embodiments, the input further comprises one or more of raw nuclease sequences, structural data, simulation data, quantum calculation data, or any combination thereof. In some embodiments, the input further comprises information related to one or more phenotypic characteristics.
[0042] In some embodiments, the system further comprises a pre-processing module configured to format the input to an SGM input. In some embodiments of the system, the validation model of (c) further comprises one or more structural evaluation tools applied to the one or more structures of the one or more novel molecules or biological systems output by the SGM of (b). In some embodiments of the system, the one or more structural evaluation tools further comprise a trained classifier module to classify the output of (b) into one or more groups.
[0043] In some embodiments of the system, the one or more predictive assessment tools of the validation model of (c) comprise folding evaluation, structural evaluation, molecular simulation, molecular dynamics simulation, or any combination thereof. In some embodiments of the system, the validation model of (c) is further configured to classify the output of the folding evaluation tool into one or more categories based on one or more predefined parameters. In some embodiments, the one or more predefined parameters comprise similarity to known molecular structures, solubility, sequence features, foldability, cleavage features, catalytic sites features, or any combination thereof.
[0044] In some embodiments of the system, the validation model of (c) is further configured to evaluate the output of the molecular simulation predictive assessment tool using one or more parameters. In some embodiments, the one or more parameters comprise comparing the outputof the molecular simulation tool to the molecular bonding of known molecules, energy for binding, catalytic sites features, active sites features, sequence features, PAM sequence features, cleavage features, charge, microenvironment features, delivery simulation without cleavage, or any combination thereof.
[0045] In some embodiments of the system, the validation model of (c) is further configured to evaluate the output of the molecular dynamics simulation tool using one or more parameters. In some embodiments, the one or more parameters comprise energy of one or more reactions, one or more energies of activation, one or more reaction pathways, one or more energies calculated using molecular mechanics, one or more energies calculated using quantum mechanics, or one or more energies calculated using both quantum mechanics and molecular mechanics, or any combination thereof. In some embodiments of the system, the SGM model of (b) is further configured to receive the output of the validation model in (c) to refine the output of the SGM model of (b).
[0046] In some embodiments, the system further comprises a system feedback loop from the output of the models of (b), (c), or both, wherein the feedback loop is used to enhance the one or more novel molecules or biological systems to minimize off-target effects. In some embodiments, the one or more novel molecules or biological systems comprise miniaturized nucleases. In some embodiments, the input further comprises genetic information associated with a specific disease.
[0047] In some embodiments of the system, the output of (d) further comprises a delivery system component corresponding to the output of (d) to deliver the one or more novel molecules or biological systems to a target. In some embodiments of the system, the delivery system component is generated using the output of (b) or (c), or both. In some embodiments of the system, the output of (b), or (c), or both, is used as feedback input to the SGM of (b) to generate the delivery system component. In some embodiments, the input further comprises novel genes or mutations associated with one or more disease states.
[0048] In some embodiments of the system, the one or more predictive assessment tools of the validation model of (c) utilize quantum computing. In some embodiments, the system further comprises in vitro testing of the output of (d). In some embodiments of the system, the input comprises feedback input based on the results of in vitro testing of the output of (d).
[0049] In some embodiments, the system further comprises in vivo testing of the output of (d). In some embodiments of the system, the input comprises feedback input based on the results of in vivo testing of the output of (d). In some embodiments, the in vivo testing further comprisestesting in one or more mammals. In some embodiments, the one or more mammals further comprise a human, a non-human primate, a murine mammal, or any combination thereof.
[0050] In yet another embodiment, disclosed herein is a system for assisting in generation of one or more novel sequences or molecular structures, the system comprising: an input module configured to receive instructions comprising one or more desired features, targets, or functions; a language learning model (LLM) configured to generate one or more novel sequences based on the one or more desired features, targets or functions; a structure prediction model configured to generate one or more novel molecular structures from the one or more novel sequences of (b); and an output module configured to generate output comprising the one or more novel sequences of (b) or the one or more novel molecular structures of (c).INCORPORATION BY REFERENCE
[0051] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference.BRIEF DESCRIPTION OF THE DRAWINGS
[0052] The novel features of the disclosure are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present disclosure will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the disclosure are utilized, and the accompanying drawings of which:
[0053] FIG. 1 illustrates a non-limiting example of a computing device; in this case, a device with one or more processors, memory, storage, and a network interface, per one or more embodiments herein.
[0054] FIG.2 shows a non-limiting diagram overview of exemplary functional blocks for one embodiment of generation of novel molecules, novel gene editors, or other novel biological materials.
[0055] FIG.3 depicts a non-limiting diagram of exemplary functional blocks for one embodiment of generation of novel molecules, novel gene editors, or other novel biological materials and verification.
[0056] FIG. 4 illustrates a non-limiting diagram of exemplary functional blocks for one embodiment of generation novel molecules, novel gene editors, or other novel biological materials, including a custom trained language learning model (LLM) and custom trained structure generation model (SGM) including submodules of a machine learning and data training block, a generative and computational block, a computational validation block, and a lab validation block.
[0057] FIG. 5 shows a non-limiting diagram of an exemplary workflow for exemplary functional blocks for one embodiment of generation of novel molecules, novel gene editors, or other novel biological materials, utilizing a custom trained LLM and SGM not including submodules of a machine learning and data training block, a generative and computational block, and a computational validation block.
[0058] FIG. 6 depicts a non-limiting diagram of exemplary functional blocks for one embodiment of generation of novel molecules, novel gene editors, or other novel biological materials, highlighting submodules of the computational validation model.
[0059] FIG. 7 illustrates a non-limiting diagram of exemplary functional blocks for one embodiment of generation of novel molecules, novel gene editors, or other novel biological materials, using generalized structural evaluation tools in the computational validation model and highlighting submodules of the computational validation model.
[0060] FIG. 8 shows a non-limiting diagram of exemplary functional blocks for one embodiment of generation of novel molecules, novel gene editors, or other novel biological materials, including model output feedback for model training.
[0061] FIG. 9 depicts a non-limiting diagram of exemplary functional blocks for one embodiment of generation of novel molecules, novel gene editors, or other novel biological materials, including evaluative and classification input criteria for novel output for model training.
[0062] FIG. 10 illustrates a non-limiting diagram of exemplary functional blocks for one embodiment of generation of novel molecules, novel gene editors, or other novel biological materials, including evaluation and classification steps for model training.
[0063] FIG. 11 shows a non-limiting diagram of exemplary functional blocks for one embodiment of generation of novel molecules, novel gene editors, or other novel biological materials, including evaluative and classification input criteria for model training concerning target genes.
[0064] FIG. 12 depicts a non-limiting diagram of exemplary functional blocks for one embodiment of generation of novel molecules, novel gene editors, or other novel biological materials, including evaluation and classification steps for model training concerning proteins.
[0065] FIG. 13 illustrates a non-limiting diagram of exemplary functional blocks for one embodiment of generation of novel molecules, novel gene editors, or other novel biological materials, including evaluation and classification steps for model training concerning current solutions.
[0066] FIG. 14 shows a non-limiting diagram of exemplary functional blocks for one embodiment of generation of novel molecules, novel gene editors, or other novel biological materials, including a delivery learning model for model training concerning delivery solutions.
[0067] FIG. 15 depicts a non-limiting diagram of exemplary functional blocks for one embodiment of generation of novel molecules, novel gene editors, or other novel biological materials, including a structure learning model for model training concerning image structure.
[0068] FIG. 16 illustrates a non-limiting diagram of exemplary functional blocks for one embodiment of generation of novel molecules, novel gene editors, or other novel biological materials, including a predictive video simulator.
[0069] FIG. 17 shows a non-limiting diagram of exemplary functional blocks for one embodiment of generation of novel molecules, novel gene editors, or other novel biological materials, including multiple output feedback training models providing training feedback input to the LLM.
[0070] FIG. 18 depicts a non-limiting diagram of exemplary functional blocks for one embodiment of generation of novel molecules, novel gene editors, or other novel biological materials, including multiple output feedback training models providing training feedback input to the LLM and SGM.
[0071] FIG. 19 illustrates a non-limiting diagram of exemplary functional blocks for one embodiment of generation of novel molecules, novel gene editors, or other novel biological materials, including evaluation and classification steps and parameters in a computational validation model.
[0072] FIG.20 shows a non-limiting diagram of exemplary functional blocks for one embodiment of generation of novel molecules, novel gene editors, or other novel biological materials, including transformer submodules of training models.
[0073] FIG.21 illustrates a non-limiting diagram of exemplary functional blocks for one embodiment of generation of novel molecules, novel gene editors, or other novel biological materials, including a simulated prediction module and various databases containing classified data, such as classified evaluative output, classified molecular simulation evaluative output, classified quantum chemical evaluation output, and classified quantum simulation output.
[0074] FIG.22 shows a non-limiting diagram of exemplary functional blocks for one embodiment of generation of novel molecules, novel gene editors, or other novel biological materials, including various learning models for machine learning and data training.
[0075] FIG.23 shows a non-limiting diagram of a system architecture for training of a machine learning system configured to generate novel molecules, novel gene editors, or other novel biological materials.
[0076] FIG.24 depicts a non-limiting diagram of exemplary functional blocks and network connections in one embodiment between nodes of multiple models in classification training.
[0077] FIG.25 illustrates a non-limiting diagram of exemplary functional blocks for one embodiment of a digital optimization game for novel molecules, novel gene editors, or other novel biological materials.
[0078] FIG.26 illustrates exemplary data concerning the training pipeline and the generation pipeline for an exemplary molecule.
[0079] FIG.27 shows an exemplary pipeline visualization of a multi-model architecture including optimization feedback input as training data.
[0080] FIG.28 illustrates an exemplary multi-model architecture including an optimization feedback loop and data collection for feedback training using a multi -model architecture.
[0081] FIG.29 illustrates an exemplary multi-model architecture including molecular simulation trajectories and quantum computing.
[0082] FIG.30 shows an exemplary multi-model architecture for image visualization of an exemplary molecule.
[0083] FIGs.31A-31 J show exemplary data and multi-model architecture for image visualization of an exemplary molecule.
[0084] FIG.32 illustrates an exemplary multi-model architecture including an encoder and decoder for image visualization of an exemplary molecule.
[0085] FIG.33 shows an exemplary multi-model architecture including a deep learning model for unique target identification.
[0086] FIG.34A shows exemplary data relating to analysis of distance between atoms of exemplary molecules of various types.
[0087] FIG.34B shows exemplary data relating to analysis of distance between domains of exemplary protein molecules.
[0088] FIG.35 illustrates an exemplary method process for target-specific generation of an exemplary molecule.DETAILED DESCRIPTION
[0089] Described herein, in certain embodiments, are computer-implemented methods for deriving operational insights about customer infrastructure, the method comprising: (a) obtaining customer incident data reported by one or more customers, wherein the customer incident data is pre-encoded using one or more encoding techniques that are specific or unique to the one or more customers; and (b) using a decoder module to directly process the customer incident data to generate an output, wherein the output comprises a plurality of configuration item(s) (Cis) discovered, derived or extracted from the customer incident data.
[0090] Also described herein, in certain embodiments, are computer-implemented systems comprising at least one processor and instructions causing the at least one processor to perform operations for deriving operational insights about customer infrastructure, the system comprising: a decoder module configured to directly process customer incident data to generate an output, wherein the customer incident data is reported by one or more customers and is preencoded using one or more encoding techniques that are specific or unique to the one or more customers, and wherein the output comprises a plurality of configuration item(s) (Cis) discovered, derived or extracted from the customer incident data.
[0091] Also described herein, in certain embodiments, are one or more non-transitory computer-readable storage media encoded with instructions executable by one or more processors to provide an application for deriving operational insights about customer infrastructure, the application comprising: (a) obtaining customer incident data reported by one or more customers, wherein the customer incident data is pre-encoded using one or more encoding techniques that are specific or unique to the one or more customers; and (b) using a decoder module to directly process the customer incident data to generate an output, wherein the output comprises a plurality of configuration item(s) (Cis) discovered, derived or extracted from the customer incident data.Terms and Definitions
[0092] Unless otherwise defined, all technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.
[0093] As used herein, the singular forms “a,” “an,” and “the” include plural references unless the context clearly dictates otherwise. Any reference to “or” herein is intended to encompass “and / or” unless otherwise stated.
[0094] As used herein, the term “about” in some cases refers to an amount that is approximately the stated amount, in some cases near the stated amount by 10%, 5%, or 1%, including increments therein, and in some cases, in reference to a percentage, refers to an amount that is greater or less the stated percentage by 10%, 5%, or 1%, including increments therein.
[0095] As used herein, the phrases “at least one,” “one or more,” and “and / or” are open-ended expressions that are both conjunctive and disjunctive in operation. For example, each of the expressions “at least one of A, B and C,” “at least one of A, B, or C,” “one or more of A, B, and C”, “one or more of A, B, or C” and “A, B, and / or C” means A alone, B alone, C alone, A and B together, A and C together, B and C together, or A, B and C together.
[0096] Reference throughout this specification to “some embodiments,” “further embodiments,” or “a particular embodiment,” means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment. Thus, the appearances of the phrase “in some embodiments,” or “in further embodiments,” or “in a particular embodiment” in various places throughout this specification are not necessarily all referring to the same embodiment. Furthermore, the particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.Examples of Machine Learning Techniques
[0097] As disclosed throughout, in some cases, the systems, the methods, the computer-readable media, and the techniques disclosed herein may implement one or more machine learning techniques. In some cases, ML may generally involve identifying and recognizing patterns in existing data in order to facilitate making predictions for subsequent data. ML may include a ML model (which may include, for example, a ML algorithm). Machine learning, whether analytical or statistical in nature, may provide deductive or abductive inference based on real or simulated data. The ML model may be a trained model. ML techniques may comprise one or more supervised, semi-supervised, self-supervised, or unsupervised ML techniques. For example, an ML model may be a trained model that is trained through supervised learning (e.g., various parameters are determined as weights or scaling factors). ML may comprise one or more of regression analysis, regularization, classification, dimensionality reduction, ensemble learning, meta learning, association rule learning, cluster analysis, anomaly detection, deeplearning, or ultra-deep learning. ML may comprise: k-means, k-means clustering, k-nearest neighbors, learning vector quantization, linear regression, non-linear regression, least squares regression, partial least squares regression, logistic regression, stepwise regression, multivariate adaptive regression splines, ridge regression, principal component regression, least absolute shrinkage and selection operation (LASSO), least angle regression, canonical correlation analysis, factor analysis, independent component analysis, linear discriminant analysis, multidimensional scaling, non-negative matrix factorization, principal components analysis, principal coordinates analysis, projection pursuit, Sammon mapping, t-distributed stochastic neighbor embedding, AdaBoosting, boosting, gradient boosting, bootstrap aggregation, ensemble averaging, decision trees, conditional decision trees, boosted decision trees, gradient boosted decision trees, random forests, stacked generalization, Bayesian networks, Bayesian belief networks, naive Bayes, Gaussian naive Bayes, multinomial naive Bayes, hidden Markov models, hierarchical hidden Markov models, support vector machines, encoders, decoders, autoencoders, stacked auto-encoders, perceptions, multi-layer perceptions, artificial neural networks, feedforward neural networks, convolutional neural networks, recurrent neural networks, residual neural networks, physics-informed neural networks, long short-term memory, deep belief networks, deep Boltzmann machines, deep convolutional neural networks, deep recurrent neural networks, large language models, transformer models, vision transformers, or generative adversarial networks.Examples of Decision Trees and Random Forests
[0098] As described above, the machine learning model may implement a decision tree. A decision tree may be a supervised ML algorithm that can be applied to both regression and classification problems. For example, a decision tree may grow from a root (base condition), and when it meets a condition (internal node / feature), it may split into multiple branches. The end of the branch that does not split anymore may be an outcome (leaf). A decision tree can be generated using a training dataset set according to the following operations: (A) starting from a root node (the entire dataset), the algorithm may split the dataset in two branches using a decision rule or branching criterion; (B) each of these two branches may generate a new child node; (C) for each new child node, the branching process may be repeated until the dataset cannot be split any further; (D) each branching criterion may be chosen to maximize information gain (e.g., a quantification of how much a branching criterion reduces a quantification of how mixed the labels are in the children nodes). The labels may be the data or the classification that is predicted by the decision tree.
[0099] A random forest regression is an extension of the decision tree model that tends to yield more robust predictions by stretching the use of the training dataset partition. Whereas adecision tree may make a single pass through the data, a random forest regression may the training data by sampling with replacement, typically using the full dataset per tree, or at least a portion of the full dataset per tree. Rather than using all explanatory variables as candidates for splitting, a random subset of candidate variables may be used for splitting, which may enable trees that have different data and different variables (hence the term random). The predictions from the trees, which may be collectively referred to as the “forest,” may then be averaged to produce a final prediction. Many trees (e.g., ten trees, fifty trees, one hundred trees, one thousand trees, etc.) may be included in a random forest model, with a number (e.g., 3, 6, 10, etc.) of terms sampled per split, a minimum of number (e.g., 1, 2, 4, 10, etc.) of splits per tree, and a minimum split size (e.g., 16, 32, 64, 128, 256, etc.). Random forests may be trained in a similar way as decision trees. Specifically, training a random forest may include the following operations: (A) randomly select k features from the total number of features; (B) create a decision tree from these k features using the same operations as for generating a decision tree; and (C) repeat the previous two operations until a target number of trees is created.
[0100] As disclosed, a random forest classifier, which may comprise a plurality of decision trees where the output prediction may be the mode of the predicted classifications of the individual trees, can be helpful in reducing overfitting to training dataset. In some cases, an ensemble of decision trees can be constructed using a random subset of features at each split or decision node. The Gini criterion may be employed, in some cases, to choose the best partition, where decision nodes having the lowest calculated Gini impurity index are selected. The Gini impurity can be used, in some cases, as a criterion to find informative features based on which the splits in each decision tree may be constructed.
[0101] In some cases, each decision tree of a random forest may comprise one or more decision nodes, where each decision node specifies a predicate condition. For example, decision node may predicate the condition that, for a given dataset, the outcome to a question is a specific outcome. At each decision node, a decision tree can be split based on whether the predicate condition attached to the decision node holds true, leading to various prediction nodes. Each prediction node can comprise output values that represent “votes” for one or more of the classifications or conditions being evaluated by the assessment model. At prediction time, a “vote” can be taken over all of the decision trees, and the majority vote (or mode of the predicted classifications) can be output as the predicted classification.
[0102] In some cases, when the dataset being queried in the assessment model reaches a “leaf’, or a final prediction node with no further downstream splits, the output values of the leaf can be output as the votes for the particular decision tree. Since a random forest model comprises a plurality of decision trees, the final votes across all trees in the forest can be summed to yield thefinal votes and the corresponding classification of the subject. A large number of decision trees can help reduce overfitting of the assessment model to the training dataset, by reducing the variance of each individual decision tree. For example, an assessment model can comprise, for example, at least about 3 decision trees, at least about 5 decision trees, at least about 10 decision trees, at least about 20 decision trees, at least about 50 decision trees, at least about 100 decision trees, etc.Computing Systems
[0103] Referring to FIG. 1, a block diagram is shown depicting an exemplary machine that includes a computer system 100 (e.g., a processing or computing system) within which a set of instructions may execute for causing a device to perform or execute any one or more of the aspects and / or methodologies for static code scheduling of the present disclosure. The components in FIG. 1 are examples only and do not limit the scope of use or functionality of any hardware, software, embedded logic component, or a combination of two or more such components implementing particular embodiments.
[0104] Computer system 100 may include one or more processors 101, a memory 103, and a storage 108 that communicate with each other, and with other components, via a bus 140. The bus 140 may also link a display 132, one or more input devices 133 (which may, for example, include a keypad, a keyboard, a mouse, a stylus, etc.), one or more output devices 134, one or more storage devices 135, and various tangible storage media 136. All of these elements may interface directly or via one or more interfaces or adaptors to the bus 140. For instance, the various tangible storage media 136 may interface with the bus 140 via storage medium interface 126. Computer system 100 may have any suitable physical form, including but not limited to one or more integrated circuits (ICs), printed circuit boards (PCBs), mobile handheld devices (such as mobile telephones or PDAs), laptop or notebook computers, distributed computer systems, computing grids, or servers.
[0105] Computer system 100 includes one or more processor(s) 101 (e.g., central processing units (CPUs) or general-purpose graphics processing units (GPGPUs)) that carry out functions. Processor(s) 101 optionally contains a cache memory unit 102 for temporary local storage of instructions, data, or computer addresses. Processor(s) 101 are configured to assist in execution of computer readable instructions. Computer system 100 may provide functionality for the components depicted in FIG. 1 as a result of the processor(s) 101 executing non-transitory, processor-executable instructions embodied in one or more tangible computer-readable storage media, such as memory 103, storage 108, storage devices 135, and / or storage medium 136. The computer-readable media may store software that implements particular embodiments, andprocessor(s) 101 may execute the software. Memory 103 may read the software from one or more other computer-readable media (such as mass storage device(s) 135, 136) or from one or more other sources through a suitable interface, such as network interface 120. The software may cause processor(s) 101 to carry out one or more processes or one or more steps of one or more processes described or illustrated herein. Carrying out such processes or steps may include defining data structures stored in memory 103 and modifying the data structures as directed by the software.
[0106] The memory 103 may include various components (e.g., machine readable media) including, but not limited to, a random access memory component (e.g., RAM 104) (e.g., static RAM (SRAM), dynamic RAM (DRAM), ferroelectric random access memory (FRAM), phasechange random access memory (PRAM), etc.), a read-only memory component (e.g., ROM 105), and any combinations thereof. ROM 105 may act to communicate data and instructions unidirectionally to processor(s) 101, and RAM 104 may act to communicate data and instructions bidirectionally with processor(s) 101. ROM 105 and RAM 104 may include any suitable tangible computer-readable media described below. In one example, a basic input / output system 106 (BIOS), including basic routines that help to transfer information between elements within computer system 100, such as during start-up, may be stored in the memory 103.
[0107] Fixed storage 108 is connected bidirectionally to processor(s) 101, optionally through storage control unit 107. Fixed storage 108 provides additional data storage capacity and may also include any suitable tangible computer-readable media described herein. Storage 108 may be used to store operating system 109, executable(s) 110, data 111, applications 112 (application programs), and the like. Storage 108 may also include an optical disk drive, a solid-state memory device (e.g., flash-based systems), or a combination of any of the above. Information in storage 108 may, in appropriate cases, be incorporated as virtual memory in memory 103.
[0108] In one example, storage device(s) 135 may be removably interfaced with computer system 100 (e.g., via an external port connector (not shown)) via a storage device interface 125. Particularly, storage device(s) 135 and an associated machine-readable medium may provide non-volatile and / or volatile storage of machine-readable instructions, data structures, program modules, and / or other data for the computer system 100. In one example, software may reside, completely or partially, within a machine-readable medium on storage device(s) 135. In another example, software may reside, completely or partially, within processor(s) 101.
[0109] Bus 140 connects a wide variety of subsystems. Herein, reference to a bus may encompass one or more digital signal lines serving a common function, where appropriate. Bus140 may be any of several types of bus structures including, but not limited to, a memory bus, a memory controller, a peripheral bus, a local bus, and any combinations thereof, using any of a variety of bus architectures. As an example and not by way of limitation, such architectures include an Industry Standard Architecture (ISA) bus, an Enhanced ISA (EISA) bus, a Micro Channel Architecture (MCA) bus, a Video Electronics Standards Association local bus (VLB), a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, an Accelerated Graphics Port (AGP) bus, HyperTransport (HTX) bus, serial advanced technology attachment (SATA) bus, and any combinations thereof.
[0110] Computer system 100 may also include an input device 133. In one example, a user of computer system 100 may enter commands and / or other information into computer system 100 via input device(s) 133. Examples of an input device(s) 133 include, but are not limited to, an alpha-numeric input device (e.g., a keyboard), a pointing device (e.g., a mouse or touchpad), a touchpad, a touch screen, a multi-touch screen, a joystick, a stylus, a gamepad, an audio input device (e.g., a microphone, a voice response system, etc.), an optical scanner, a video or still image capture device (e.g., a camera), and any combinations thereof. In some embodiments, the input device is a Kinect, Leap Motion, or the like. Input device(s) 133 may be interfaced to bus 140 via any of a variety of input interfaces 123 (e.g., input interface 123) including, but not limited to, serial, parallel, game port, USB, FIREWIRE, THUNDERBOLT, or any combination of the above.
[0111] In particular embodiments, when computer system 100 is connected to network 130, computer system 100 may communicate with other devices, specifically mobile devices and enterprise systems, distributed computing systems, cloud storage systems, cloud computing systems, and the like, connected to network 130. Communications to and from computer system 100 may be sent through network interface 120. For example, network interface 120 may receive incoming communications (such as requests or responses from other devices) in the form of one or more packets (such as Internet Protocol (IP) packets) from network 130, and computer system 100 may store the incoming communications in memory 103 for processing. Computer system 100 may similarly store outgoing communications (such as requests or responses to other devices) in the form of one or more packets in memory 103 and communicated to network 130 from network interface 120. Processor(s) 101 may access these communication packets stored in memory 103 for processing.
[0112] Examples of the network interface 120 include, but are not limited to, a network interface card, a modem, and any combination thereof. Examples of a network 130 or network segment 130 include, but are not limited to, a distributed computing system, a cloud computingsystem, a wide area network (WAN) (e.g., the Internet, an enterprise network), a local area network (LAN) (e.g., a network associated with an office, a building, a campus or other relatively small geographic space), a telephone network, a direct connection between two computing devices, a peer-to-peer network, and any combinations thereof. A network, such as network 130, may employ a wired and / or a wireless mode of communication. In general, any network topology may be used.
[0113] Information and data may be displayed through a display 132. Examples of a display 132 include, but are not limited to, a cathode ray tube (CRT), a liquid crystal display (LCD), a thin film transistor liquid crystal display (TFT-LCD), an organic liquid crystal display (OLED) such as a passive-matrix OLED (PMOLED) or active-matrix OLED (AMOLED) display, a plasma display, and any combinations thereof. The display 132 may interface to the processor(s) 101, memory 103, and fixed storage 108, as well as other devices, such as input device(s) 133, via the bus 140. The display 132 is linked to the bus 140 via a video interface 122, and transport of data between the display 132 and the bus 140 may be controlled via the graphics control 121. In some embodiments, the display is a video projector. In some embodiments, the display is a head-mounted display (HMD) such as a VR headset. In further embodiments, suitable VR headsets include, by way of non-limiting examples, HTC Vive, Oculus Rift, Samsung Gear VR, Microsoft HoloLens, Razer OSVR, FOVE VR, Zeiss VR One, Avegant Glyph, Freefly VR headset, and the like. In still further embodiments, the display is a combination of devices such as those disclosed herein.
[0114] In addition to a display 132, computer system 100 may include one or more other peripheral output devices 134 including, but not limited to, an audio speaker, a printer, a storage device, and any combinations thereof. Such peripheral output devices may be connected to the bus 140 via an output interface 124. Examples of an output interface 124 include, but are not limited to, a serial port, a parallel connection, a USB port, a FIREWIRE port, a THUNDERBOLT port, and any combinations thereof.
[0115] In addition, or as an alternative, computer system 100 may provide functionality as a result of logic hardwired or otherwise embodied in a circuit, which may operate in place of or together with software to execute one or more processes or one or more steps of one or more processes described or illustrated herein. Reference to software in this disclosure may encompass logic, and reference to logic may encompass software. Moreover, reference to a computer-readable medium may encompass a circuit (such as an IC) storing software for execution, a circuit embodying logic for execution, or both, where appropriate. The present disclosure encompasses any suitable combination of hardware, software, or both.
[0116] Those of skill in the art will appreciate that the various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein may be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality.
[0117] The various illustrative logical blocks, modules, and circuits described in connection with the embodiments disclosed herein may be implemented or performed with a general purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general-purpose processor may be a microprocessor, but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.
[0118] The steps of a method or algorithm described in connection with the embodiments disclosed herein may be embodied directly in hardware, in a software module executed by one or more processor(s), or in a combination of the two. A software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor such the processor may read information from, and write information to, the storage medium. In the alternative, the storage medium may be integral to the processor. The processor and the storage medium may reside in an ASIC. The ASIC may reside in a user terminal. In the alternative, the processor and the storage medium may reside as discrete components in a user terminal.
[0119] In accordance with the description herein, suitable computing devices include, by way of non-limiting examples, cloud computing platforms, distributed computing platforms, server clusters, server computers, desktop computers, laptop computers, notebook computers, subnotebook computers, netbook computers, and netpad computers.
[0120] In some embodiments, the computing device includes an operating system configured to perform executable instructions. The operating system is, for example, software, including programs and data, which manages the device’s hardware and provides services for execution ofapplications. Those of skill in the art will recognize that suitable server operating systems include, by way of non-limiting examples, FreeBSD, OpenBSD, NetBSD®, Linux, Apple® Mac OS X Server®, Oracle® Solaris®, Windows Server®, and Novell® NetWare®. Those of skill in the art will recognize that suitable personal computer operating systems include, by way of nonlimiting examples, Microsoft® Windows®, Apple® Mac OS X®, UNIX®, and UNIX-like operating systems such as GNU / Linux®. In some embodiments, the operating system is provided by cloud computing. Those of skill in the art will also recognize that suitable mobile smartphone operating systems include, by way of non-limiting examples, Nokia® Symbian® OS, Apple® iOS®, Research in Motion® BlackBerry OS®, Google® Android®, Microsoft® Windows Phone® OS, Microsoft® Windows Mobile® OS, Linux®, and Palm® WebOS®.Non-transitorv Computer Readable Storage Medium
[0121] In some embodiments, the platforms, systems, media, and methods disclosed herein include one or more non-transitory computer readable storage media encoded with a program including instructions executable by the operating system of an optionally networked computing device. In further embodiments, a computer readable storage medium is a tangible component of a computing device. In still further embodiments, a computer readable storage medium is optionally removable from a computing device. In some embodiments, a computer readable storage medium includes, by way of non-limiting examples, CD-ROMs, DVDs, flash memory devices, solid state memory, magnetic disk drives, magnetic tape drives, optical disk drives, distributed computing systems including cloud computing systems and services, and the like. In some cases, the program and instructions are permanently, substantially permanently, semipermanently, or non-transitorily encoded on the media.Computer Programs
[0122] In some embodiments, the platforms, systems, media, and methods disclosed herein include at least one computer program, or use of the same. A computer program includes a sequence of instructions, executable by one or more processor(s) of the computing device’s CPU, written to perform a specified task. Computer readable instructions may be implemented as program modules, such as functions, objects, Application Programming Interfaces (APIs), computing data structures, and the like, which perform particular tasks or implement particular abstract data types. In light of the disclosure provided herein, those of skill in the art will recognize that a computer program may be written in various versions of various languages.
[0123] The functionality of the computer readable instructions may be combined or distributed as desired in various environments. In some embodiments, a computer program comprises one sequence of instructions. In some embodiments, a computer program comprises a plurality ofsequences of instructions. In some embodiments, a computer program is provided from one location. In other embodiments, a computer program is provided from a plurality of locations. In various embodiments, a computer program includes one or more software modules. In various embodiments, a computer program includes, in part or in whole, one or more web applications, one or more mobile applications, one or more standalone applications, one or more web browser plug-ins, extensions, add-ins, or add-ons, or combinations thereof.Software Modules
[0124] In some embodiments, the platforms, systems, media, and methods disclosed herein include software, server, and / or database modules, or use of the same. In view of the disclosure provided herein, software modules are created by techniques known to those of skill in the art using machines, software, and languages known to the art. The software modules disclosed herein are implemented in a multitude of ways. In various embodiments, a software module comprises a file, a section of code, a programming object, a programming structure, a distributed computing resource, a cloud computing resource, or combinations thereof. In further various embodiments, a software module comprises a plurality of files, a plurality of sections of code, a plurality of programming objects, a plurality of programming structures, a plurality of distributed computing resources, a plurality of cloud computing resources, or combinations thereof. In various embodiments, the one or more software modules comprise, by way of nonlimiting examples, a web application, a mobile application, a standalone application, and a distributed or cloud computing application. In some embodiments, software modules are in one computer program or application. In other embodiments, software modules are in more than one computer program or application. In some embodiments, software modules are hosted on one machine. In other embodiments, software modules are hosted on more than one machine. In further embodiments, software modules are hosted on a distributed computing platform such as a cloud computing platform. In some embodiments, software modules are hosted on one or more machines in one location. In other embodiments, software modules are hosted on one or more machines in more than one location.Databases
[0125] In some embodiments, the platforms, systems, media, and methods disclosed herein include one or more databases, or use of the same. In view of the disclosure provided herein, those of skill in the art will recognize that many databases are suitable for storage and retrieval of information, for example customer incident data. In various embodiments, suitable databases include, by way of non-limiting examples, relational databases, non-relational databases, object-oriented databases, object databases, entity -relationship model databases, associative databases,XML databases, document oriented databases, and graph databases. Further non-limiting examples include SQL, PostgreSQL, MySQL, Oracle, DB2, Sybase, and MongoDB. In some embodiments, a database is Internet-based. In further embodiments, a database is web-based. In still further embodiments, a database is cloud computing-based. In a particular embodiment, a database is a distributed database. In other embodiments, a database is based on one or more local computer storage devices.Novel Nuclease Generation and Validation Models
[0126] In some embodiments, the present disclosure provides a software platform, otherwise known as a software system, which may generate novel molecules, including novel nucleases. In some embodiments, the present disclosure provides a software method that may generate novel molecules, including novel nucleases.
[0127] In some embodiments, the present disclosure provides a computer-implemented method for assisting in generation of one or more novel molecules or biological systems. In some embodiments, the method may comprise receiving an input comprising information related to a target function. In some embodiments, the input may comprise data from a database. In some embodiments, the method may comprise one or more models. In some embodiments, the one or more models may be machine learning models. In some embodiments, the one or more models may be artificially intelligent models (Al models). In some embodiments, the models may comprise one or more of: diffusion models, language learning models, structural learning models, linear regression models, logistic regression models, random forest architecture models, support vector machine models, k-nearest neighbor models, naive bayes models, deep learning models, neural network models, deep learning neural network models, convolutional neural network models, recurrent neural network models, long short term memory network models, autoencoder models, generative adversarial network models (GANs), principal component analysis models, linear discriminant analysis models, gradient boosting models, extreme gradient boosting models, ridge regression models, lasso regression models, elasticnet regression models, stochastic gradient descent regression models, self-organizing map models, k-means clustering models, hierarchal clustering models, DBscan clustering models, hidden Markov models, temporal different learning models, state-action-reward-state-action learning models, factorization models, collaborative filtering models, artificial immune system models, Bayesian network models, gaussian mixture models, reinforcement learning models, multilayer perception models, radial basis function network models, Boltzmann models, fuzzy logic models, genetic models, association rule learning models, multivariate adaptive regression splines,multidimensional scaling models, learning vector quantization models, instance-based learning models, chi-squared automatic interaction detection models, gradient boosting regressor models, or another type of machine learning model, or any combination thereof.
[0128] In some embodiments, the models may comprise one or more of: taxonomic-guided diffusion models, MSA-based protein language models, equivariant denoising diffusion models, or any combination thereof. In some embodiments, the taxonomic-guided diffusion models may comprise, for example, one or more models that may incorporate taxonomic identifiers within transformer architecture to achieve, for example, controllable protein sequence generation. The MSA-based protein language models may comprise, for example, one or more multiple sequence alignment transformers that may be configured to, for example, leverage evolutionary information from protein sequences to generate new sequences. The MSA-based protein language models may be further configured to, for example, utilize iterative masking procedures for generation of optimized sequences with, for example, improved homology, coevolution, and structural properties. The equivariant denoising diffusion models may comprise, for example, generation of protein structures and sequences using a denoising process, for example encoding structural constraints and using invariant point attention transformers for production of diverse protein sequences with improved structural accuracy.
[0129] In some embodiments, the models comprising the systems and methods described herein may comprise one or more sequence-based classification models. In some embodiments, the one or more sequence-based classification models may comprise one or more of: knowledge-aided neural networks (KAN) networks, convolutional neural networks (CNNs), or recurrent neural networks (RNNs) and variants, or any combination thereof. In some embodiments, the KAN networks may comprise models configured to perform operations comprising integrating domain-specific knowledge for optimized protein sequence classification. In some embodiments, the CNNs may comprise models configured to perform operations comprising detecting local motifs and patterns in sequence data to learn hierarchical features for improved classification functions. In some embodiments, the RNNs and variants may comprise one or more models comprising, for example, long short-term memory models (LSTM models) and gated recurrent units (GRUs) configured to perform operations comprising capturing long-range dependencies in data.
[0130] In some embodiments, the systems and methods described herein may comprise one or more sampling techniques. The one or more sampling techniques may comprise one or more of: latent space annealing, Bayesian optimization-based sampling, or stochastic gradient Langevin dynamics (SGLD), or any combination thereof. In some embodiments, latent space annealingmay comprise gradually reducing temperature of the sampling process in the latent space for improved optimal sequence detection. In some embodiments, Bayesian optimization-based sampling may further comprise incorporating Bayesian optimization methods into sampling for improved sampling efficiency. In some embodiments, SGLD may comprise gradient-based optimization paired with stochastic sampling for use in generating diverse sequences with improved confidence in functional properties.
[0131] As shown in FIG. 2, in some embodiments, the input may comprise data from UniProt 210. In some embodiments, the input may comprise data from PubMed 209. In some embodiments, the input may comprise data from both UniProt 210 and PubMed 209. In some embodiments, the system or method may take place in a series of interconnected models. In some embodiments, the models may be present on the AWS cloud 211. In some embodiments, the input may be received by a raw data storage model 208. In some embodiments, the raw data storage model may be a raw data storage database 208.
[0132] In some embodiments, the raw data storage model may output data to an Al model for training 207. In some embodiments, the training model may be a third-party training model 207. In some embodiments, the training model may output data to a model storage model 206. In some embodiments, the model storage model may be a model storage database 206. In some embodiments, the model storage model or database 206 may receive additional data. In some embodiments, the additional data may comprise one or more requests to create one or more molecules. In some embodiments, the additional data may comprise one or more requests to create one or more nucleases 205. In some embodiments, the one or more requests to create one or more molecules, for example one or more nucleases, may be received by a molecule storage library, for example a nuclease library storage 204. In some embodiments, data from the molecule storage library, for example the nuclease storage library 204, may be input into a prediction model 202. In some embodiments, the prediction model may predict a structure of one or more molecules 202. In some embodiments, the structure of one or more molecules may be the structure of one or more nucleases 202.
[0133] In some embodiments, the prediction model may output data to a verification model 201. In some embodiments, the verification model may verify the structure of one or more molecules, for example one or more nucleases 201. In some embodiments, the verification model may perform verification using one or more molecular dynamics simulations 201. In some embodiments, the verification model may output data 203. In some embodiments, the output data may be output of one or more molecules, for example one or more nucleases 203. In someembodiments, the output data may comprise tier and use case for the one or more molecules, for example one or more nucleases 203.
[0134] In some embodiments, as shown in FIG. 3, for example, the present disclosure provides a computer-implemented system or method to output one or more molecules, for example nucleases. In some embodiments, the method or system may comprise receiving data from a database, for example data from UniProt 317. In some embodiments, the method or system may comprise receiving one or more requests to generate one or more molecules 311. In some embodiments, the method or system may comprise receiving one or more requests to generate one or more nucleases 311. In some embodiments, the method or system may comprise receiving input from an untrained, pretrained, or custom language learning model (LLM) 312. In some embodiments, the LLM may be a large language model. In some embodiments, the LLM may be an untrained, pretrained, or custom LLM architecture 312. In some embodiments, the LLM may be an open-source LLM 312.
[0135] In some embodiments, the input data may be separated into two or more separate databases. In some embodiments, the databases may comprise a raw sequence database, for example input raw nuclease sequences 313, or a protein sequence database 315, or both. In some embodiments, the databases may output data that undergoes data pre-processing. In some embodiments, the data in the input raw nuclease sequences 313 may undergo pre-processing 314 and may be input into an LLM model for training and tuning 302. In some embodiments, the training and tuning of the LLM model may output a trained LLM model 310. In some embodiments, data from the protein sequences database 315 may undergo data pre-processing 316 and may be input into a training model for random forest classifier training 303.
[0136] In some embodiments, the data pre-processing and the training and tuning of the LLM and random forest classifier may occur in parallel. In some embodiments, one or more requests to generate one or more molecules, for example one or more nucleases 311 may be input into the trained LLM model 310. In some embodiments, the trained LLM model may generate sequences of one or more molecules. In some embodiments, the trained LLM model may generate sequences of one or more nucleases. In some embodiments, the generated sequences output by the trained LLM model 310 may be stored in a generated molecular sequences, for example nuclease sequences, model 309. In some embodiments, the generated molecular sequences model may be a generated molecule sequences database 309, for example a generated nuclease sequences database. In some embodiments, data from the generated molecular sequences model or database may be input to a trained molecular classifier model 308. In some embodiments, data from the trained random forest classifier model 303 may be input into the trained molecularclassifier model 308. In some embodiments, the trained molecular classifier model may be a trained nuclease classifier model 308. In some embodiments, data from a random forest classifier 303 may be input into the trained molecular classifier model 308 in addition to the data from the generated molecular sequences database 309.
[0137] In some embodiments, output from the trained molecular classifier may be used to predict one or more molecular structures 307. In some embodiments, the prediction of the one or more molecular structures may be performed with a structural prediction model 307. In some embodiments, the structural prediction model 307 may be a third-party structural prediction model. In some embodiments, the structural prediction model 307 may be one or more of AlphaFold, Rosetta, or DeepMind Equivariant Neural Networks, or any combination thereof. In some embodiments, the structural prediction model 307 may be a customized, bespoke, in-house generated, or otherwise independently created structure prediction model. In some embodiments, the output from the structural prediction model 307 may be input into a verification model 306.In some embodiments, the verification model 306 may perform verification with molecular dynamics or statistical methods, or both. In some embodiments, the output from the verification model may be used for verification in vivo 305. In some embodiments, the verification in vivo 305 may be verification of one or more novel molecules. In some embodiments, the results of the verification in vivo may be used to output one or more molecules, for example nucleases 304. In some embodiments, the molecules or nucleases may be tier 1 molecules or nucleases 304
[0138] In some embodiments, the system or method for generating a novel molecule as shown in FIG. 4, for example a novel nuclease, may comprise one or more of a machine learning / data training model 408, a generative / computational model 409, a computational validation model 413, a novel nuclease library output 424, or a lab validation model 401, or any combination thereof.
[0139] In some embodiments, the machine learning / data training model 408 may comprise a protein data model 403, an input raw molecular sequences model or database 405, a cloud server 407, a data processing model 406, a train and tune model 404, an LLM / structural architecture model 402, and a training feedback model 412, or any combination thereof.
[0140] In some embodiments, the machine learning / data training model 408 may receive input comprising protein data 403. In some embodiments, the protein data 403 may comprise UniProt. In some embodiments, the protein data input may be stored in an input raw molecular sequences model 405. In some embodiments, the input raw molecular sequences model may comprise an input raw molecular database 405. In some embodiments, the raw molecular sequences may beraw nuclease sequences 405. In some embodiments, data from the input raw molecular sequences model or database 405 may be processed or uploaded to the Cloud server 407. In some embodiments, data from the input raw molecular sequences model may undergo data preprocessing 406. Data preprocessing may be performed by a data preprocessing model 406. In some embodiments, the preprocessed data may be input into the train and tune model 404 for a custom trained LLM model 410.
[0141] In some embodiments, the generative / computational model 409 may comprise a custom trained LLM model 410 or a custom trained structure model 411, or both.
[0142] In some embodiments, an LLM / structural architecture model 402 may provide input to the train and tune model 404. In some embodiments, the LLM / structural architecture model may be an open-source model 402. In some embodiments, the LLM / structural architecture model may be an untrained, pretrained, or custom model 402. In some embodiments, the train and tune model 404 may receive input from one or more training feedback machine learning models 412.In some embodiments, the train and tune model may receive feedback from a custom trained LLM model 410. In some embodiments, the train and tune model may receive input from one or more of the data processing model 406, the LLM / structural architecture untrained, pretrained, or custom model 402, the custom trained LLM model 410, or the training feedback machine learning models 412, or any combination thereof. In some embodiments, the train and tune model 404 may provide input to a custom trained structure model 411. In some embodiments, the LLM model 514 may comprise a GPT-1, GPT-2, GPT-3, GPT-4, GPT-5, GPT-6, GPT-7, GPT-8, GPT-9, GPT-10, or another GPT model. In some embodiments, the LLM model 410 may comprise a protein language model. In some embodiments, the LLM model 410 may comprise a protein language model comprising one or more of ESM-1, ESM-lv, ESM-2, ESM-2v, ESM-3, ESM-3v, ESM-4, ESM-4v, ESM-5, ESM-5v, ESMFold, ProGen, proteinBERT, another protein language model, or any combination thereof. In some embodiments, the LLM model 410 may comprise an antibody-specific protein language model. In some embodiments, the LLM model 410 may comprise an antibody-specific protein language model comprising one or more of IgLM, AbLang, AbLang2, AbLang3, AbLang4, AbLang5, AbLang6, AbLang7, AbLang8, AbLang9, AbLanglO, AntiBERTa, Bio-inspired Antibody Language Model (BALM), BALM-paired and unpaired, OAS-trained RoBERTa, IgBert, IgT5, AntiBERTy, Sapiens, FAbConSapiens, or another antibody-specific protein language model, or any combination thereof. In some embodiments, the training feedback machine learning models 412 may comprise a model output feedback for machine learning (ML) training 425, a known molecular, for example known nucleases, model for ML training 426, a gene target DNA model for MLtraining 427, another protein model for ML training 428, a current solution model for ML training 429, a delivery solution model for ML training 430, a structural model for ML training 431, or a molecular simulation model for ML training 432, or any combination thereof.
[0143] In some embodiments, a computational validation model 413 may comprise a generated molecular sequences model or database 419, a trained molecular classifier model 420, one or more folding evaluation tools, for example one or more of AlphaFold, Rosetta, or DeepMind Equivariant Neural Networks, a customized or bespoke folding evaluation tool, or any combination thereof 416, a generated molecular structure model or database 418, one or more structural evaluation tools or models 417, a full validation model 414, and a molecular simulation model 415, or any combination thereof.
[0144] In some embodiments, the custom trained LLM model 410 may provide output to a custom trained structure model 411. In some embodiments, the custom trained LLM model 410 may receive input from one or more of the train and tune model 404, the custom trained structure model 411, or a full validation model 414, or any combination thereof. In some embodiments, the custom trained LLM model 410 may provide output to one or more of a custom trained structure model 411, a model output feedback model for machine learning training 425, the train and tune model 404, a generated molecular or nuclease sequences model or database 419, the full validation model 414, or any combination thereof. In some embodiments, the custom trained structure model 411 may receive input from the custom trained LLM model 410, the train and tune model 404, the training feedback machine learning models 412, or any combination thereof. In some embodiments, the training feedback machine learning models 412 may comprise a model output feedback for machine learning (ML) training 425, a known molecular, for example known nucleases, model for ML training 426, a gene target DNA model for ML training 427, another protein model for ML training 428, a current solution model for ML training 429, a delivery solution model for ML training 430, a structural model for ML training 431, or a molecular simulation movie for ML training 432, or any combination thereof. In some embodiments, the custom trained structure model 411 may provide output to the custom trained LLM model 410 or a generated molecular structure model or database 418, or any combination thereof.
[0145] In some embodiments, the generated molecular sequences model or database 419 may be a generated nuclease sequences model or database. In some embodiments, the generated molecular or nuclease sequences model or database 419 may output data to a trained molecular classifier model, in some cases a trained nuclease classifier model 420. In some embodiments, the trained molecular classifier model 420 may output data to a molecular folding evaluationtool 416. In some embodiments, the molecular folding evaluation tool 416 may be a third-party evaluation tool. In some embodiments, the molecular folding evaluation tool 416 may be one or more of AlphaFold, Rosetta, or DeepMind Equivariant Neural Networks, or any combination thereof. In some embodiments, the molecular folding evaluation tool 416 may be a customized, bespoke, in-house generated, or otherwise independently created molecular folding evaluation tool. In some embodiments, the molecular folding evaluation tool 416 may output data to a full validation model 414. In some embodiments, the generated molecular structure model or database 418 may comprise a generated nuclease structure model or database. In some embodiments, the generated molecular structure model or database 418 may output data to one or more structural evaluation tools 417. In some embodiments, the one or more structural evaluation tools 417 may output data to a full validation model 414. In some embodiments, the full validation model 414 may output data to a train and tune model 404, a model output feedback model for machine learning and training 425, and a molecular simulation model 415, or any combination thereof.
[0146] In some embodiments, the molecular simulation model 415 may output data to a full validation model 414 or a novel molecular library output 424, or both. In some embodiments, the novel molecular library output 424 may be a novel nuclease library output. In some embodiments, the novel molecular library output 424 may be optimized for one or more gene targets. In some embodiments, the novel molecular library output 424 may be optimized for one or more variables comprising vehicle, guide RNA, micro-environment, chemical environment, charge, nearby receptors, binding energy, polarity, catalytic sites, active sites, sequences, PAM sequences, tissue type, ion concentration, oxygen concentration, efficacy, size, specificity, RNA binding, cleavage evaluation such as for cleavage efficacy or cleavage needs, delivery features, large deletions or large insertions, nearby small molecules, RNA nucleotide type, DNA nucleotide type, external electric and magnetic fields, light environment, microbial environment, ion type, pH value, single nucleotide mutations, multiple nucleotide mutations, single nucleotide deletions, multiple nucleotide deletions, single nucleotide insertions, multiple nucleotide insertions, epigenetic factors, cleavage timing, cleavage control, or any combination thereof. In some embodiments, the novel molecular library output 424 may output data to a lab validation model 401.
[0147] In some embodiments, the lab validation model 401 may comprise one or more target processes and one or more further research processes 421. The one or more target processes may comprise, for example, a verification in vitro 423 or a verification in vivo 422. In someembodiments, data from the verification in vitro 423 may be output for use in the verification in vivo 422
[0148] In some embodiments, as shown in FIG. 5, a machine learning / data training model may comprise a Uniprot database 511, an input raw molecular sequences model or database 509, a cloud server database 512, a data preprocessing model 510, an LLM architecture model 507, and a train and tune model 508, or any combination thereof. In some embodiments, a database, for example UniProt Data 511 may output data to an input raw molecular sequences model or database 509. In some embodiments, the input raw molecular sequences model or database 509 may be an input raw nuclease sequences model or database. In some embodiments, the input raw molecular sequences model or database 509 may transfer data through an AWS cloud 512 to a data preprocessing model 510. In some embodiments, the data preprocessing model 510 may preprocess the molecular data, and output the preprocessed data to a train and tune model 508.In some embodiments, the train and tune model 508 may receive input data from an LLM architecture model 507. In some embodiments, the LLM architecture model 507 may be an open source model. In some embodiments, the LLM architecture model 507 may comprise an untrained, pretrained, or custom open source model.
[0149] In some embodiments, a generative / computational model 513 may comprise an LLM model 514. In some embodiments, the LLM model 514 may comprise a GPT-1, GPT-2, GPT-3, GPT-4, GPT-5, GPT-6, GPT-7, GPT-8, GPT-9, GPT-10, or another GPT model. In some embodiments, the LLM model 514 may comprise a protein language model. In some embodiments, the LLM model 514 may comprise a protein language model comprising one or more of ESM-1, ESM-lv, ESM-2, ESM-2v, ESM-3, ESM-3v, ESM-4, ESM-4v, ESM-5, ESM-5v, ESMFold, ProGen, proteinBERT, another protein language model, or any combination thereof. In some embodiments, the LLM model 514 may comprise an antibody-specific protein language model. In some embodiments, the LLM model 514 may comprise an antibody-specific protein language model comprising one or more of IgLM, AbLang, AbLang2, AbLang3, AbLang4, AbLang5, AbLang6, AbLang7, AbLang8, AbLang9, AbLanglO, AntiBERTa, Bioinspired Antibody Language Model (BALM), BALM-paired and unpaired, OAS-trained RoBERTa, IgBert, IgT5, AntiBERTy, Sapiens, FAbConSapiens, or another antibody-specific protein language model, or any combination thereof. In some embodiments, the generative / computational model 513 may solely comprise an LLM model 514. In some embodiments, the LLM model 514 may be a custom trained LLM model. In some embodiments, the LLM model 514 may receive input data from a train and tune model 508. In someembodiments, the LLM model 514 may output data to a generated molecular sequences model or database 505.
[0150] In some embodiments, the computational validation model 501 may comprise a generated molecular sequences model or database 505, a trained molecular classifier model 504, one or more evaluation tools 503, and a molecular simulation model 502, or any combination thereof.
[0151] In some embodiments, the generated molecular sequences model or database 505 may be a generated nuclease sequences model or database 505. In some embodiments, the generated molecular sequences model or databases 505 may output data to a trained molecular classifier model 504. In some embodiments, the trained molecular classifier model 504 may comprise a trained nuclease classifier model. In some embodiments, the trained molecular classifier model 504 may output data to one or more molecular folding evaluation tools 503. In some embodiments, the molecular folding evaluation tools 503 may be third-party evaluation tools. In some embodiments, the molecular folding evaluation tools 503 may be one or more of AlphaFold, Rosetta, or DeepMind Equivariant Neural Networks, or any combination thereof. In some embodiments, the molecular folding evaluation tools 503 may be a customized, bespoke, in-house generated, or otherwise independently created molecular folding evaluation tools. In some embodiments, the one or more molecular folding evaluation tools 503 may output data to a molecular simulation model 502. In some embodiments, the molecular simulation model 502 may output data to a novel molecular library 515. In some embodiments, the novel molecular library 515 may comprise a novel nuclease library 515. In some embodiments, the novel molecular library 515 may output one or more novel molecules to a user, wherein the novel molecules may possess one or more desired characteristics.
[0152] In some embodiments, the system or method for generating one or more novel molecules, in some aspects one or more novel nucleases, is shown in, for example FIG. 6. In some embodiments a machine learning / data training model 614 may comprise a protein data input model or database 616, an input raw molecular sequences model or database 618, a cloud server database 620, a data preprocessing model 619, a structural architecture model 615, and a train and tune model 617, or any combination thereof.
[0153] In some embodiments, the method or system may receive input data from a protein model or database 616. The protein model or database may comprise, for example UniProt Data 616. In some embodiments, the protein model or database may output data to an input raw molecular sequences model or database 618. In some embodiments, the input raw molecular sequences model or database 618 may be an input raw nuclease sequences model or database. Insome embodiments, the input raw molecular sequences model or database 618 may transfer data through a cloud server 620 to a data preprocessing model 619. In some embodiments, the data preprocessing model 619 may preprocess the molecular data, and output the preprocessed data to a train and tune model 617. In some embodiments, the train and tune model 617 may receive input data from a structural architecture model 615. In some embodiments, the train and tune model 617 may receive data from one or more of an evaluative output classification model or database 609, a molecular simulation evaluative output classification model or database 608, a quantum chemistry evaluative output classification model or database 607, a quantum simulation output classification database 606, or any combination thereof. In some embodiments, the structural architecture model 615 may be an open source model. In some embodiments, the structural architecture model 615 may comprise an untrained, pretrained, or custom open source model. In some embodiments, the train and tune model 617 may output data to the structural architecture model 615.
[0154] In some embodiments, a generative / computational model 621 may comprise an LLM model 622. In some embodiments, the generative / computational model 621 may solely comprise an LLM model 622. In some embodiments, the LLM model 622 may be a custom trained LLM model. In some embodiments, the LLM model 622 may receive input data from a train and tune model 617. In some embodiments, the LLM model 622 may output data to a generated molecular sequences model or database 612. In some embodiments, the LLM model 622 may comprise a GPT-1, GPT-2, GPT-3, GPT-4, GPT-5, GPT-6, GPT-7, GPT-8, GPT-9, GPT-10, or another GPT model. In some embodiments, the LLM model 622 may comprise a protein language model. In some embodiments, the LLM model 622 may comprise a protein language model comprising one or more of ESM-1, ESM-lv, ESM-2, ESM-2v, ESM-3, ESM-3v, ESM-4, ESM-4v, ESM-5, ESM-5v, ESMFold, ProGen, proteinBERT, another protein language model, or any combination thereof. In some embodiments, the LLM model 622 may comprise an antibody-specific protein language model. In some embodiments, the LLM model 622 may comprise an antibody-specific protein language model comprising one or more of IgLM, AbLang, AbLang2, AbLang3, AbLang4, AbLang5, AbLang6, AbLang7, AbLang8, AbLang9, AbLanglO, AntiBERTa, Bio-inspired Antibody Language Model (BALM), BALM-paired and unpaired, OAS-trained RoBERTa, IgBert, IgT5, AntiBERTy, Sapiens, FAbConSapiens, or another antibody-specific protein language model, or any combination thereof.
[0155] In some embodiments, the computational validation model 601 may comprise a generated molecular sequences model or database 612, a trained molecular classifier model 613, one or more evaluation tools 611, a molecular simulation model 610, a quantum chemistryevaluation model 603, a simulated prediction model 602, an evaluative output classification model or database 609, a molecular simulation evaluative output classification model or database 608, a quantum chemistry evaluative output classification model or database 607, a quantum simulation output classification database 606, or any combination thereof.
[0156] In some embodiments, the generated molecular sequences model or database 612 may be a generated nuclease sequences model or database. In some embodiments, the generated molecular sequences model or database 612 may output data to a trained molecular classifier model 613. In some embodiments, the trained molecular classifier model 613 may comprise a trained nuclease classifier model. In some embodiments, the trained molecular classifier model 613 may output data to one or more molecular folding evaluation tools 611. In some embodiments, the molecular folding evaluation tools 611 may be third-party evaluation tools. In some embodiments, the molecular folding evaluation tools 611 may be one or more of AlphaFold, Rosetta, or DeepMind Equivariant Neural Networks, or any combination thereof. In some embodiments, the molecular folding evaluation tools 611 may be customized, bespoke, inhouse generated, or otherwise independently created molecular folding evaluation tools.
[0157] In some embodiments, the molecular folding evaluation tools may output data to a novel output evaluation model 605. In some embodiments, the novel output evaluation model 605 may apply one or more algorithms to execute one or more method steps comprising receiving novel output from the one or more molecular folding evaluation tools 611, classifying the novel output from the one or more molecular folding evaluation tools 611, compare the novel output from the one or more molecular folding evaluation tools 611 to known molecules, and evaluate the novel output from the one or more molecular folding evaluation tools 611 based on one or more parameters. The one or more parameters may comprise one or more of solubility, sequence features, foldability, cleavage evaluation, subclassification, and catalytic site, or any combination thereof. In some embodiments, the novel output evaluation model 605 may output data to a molecular simulation model 610. In some embodiments, the novel output evaluation model 605 may output data to an evaluative output classification model or database 609. In some cases, the data output to the evaluative output classification model or database 609 may comprise a folded molecular structure of one or more molecules linked to one or more numerical values, a set of one or more values representing one or more parameters of one or more molecules, or the structures thereof, or a classification scheme of one or more molecules, or the structures thereof, a matrix of values representing a classification scheme of one or more molecules, or the structures thereof, or any combination thereof. In some embodiments, the one or more numerical values may comprise one or more evaluation scores for one or more ofsolubility, domain, classification, binding, sequence similarity, or any combination thereof. In some embodiments, the evaluative output classification model or database 609 may output data to the train and tune model 617.
[0158] In some embodiments, the molecular simulation model 610 may output data to a novel output molecular simulation evaluative model 604. In some embodiments, the output data of the molecular simulation model 610 may comprise virtual modeling of the interactions comprising chemical reactions between one or more novel molecules. In some embodiments, the novel output molecular simulation evaluative model 604 may apply one or more algorithms configured to execute one or more method steps comprising receiving input data from the molecular simulation model 610, evaluating the input data from the molecular simulation model 610, comparing the input data from the molecular simulation model 610 to known molecular simulation data, simulating one or more interactions comprising one or more chemical reactions between one or more target genes and the input data from the molecular simulation model 610, and evaluating the input data based on one or more parameters. In some embodiments, the one or more parameters may comprise energy for binding, catalytic sites, active sites, sequences, PAM sequences, cleavage evaluation, charge, microenvironment, and delivery simulation without cleavage, or any combination thereof. In some embodiments, the novel output molecular simulation evaluative model 604 may output data to a molecular simulation evaluative output classification database 608. In some embodiments, the data output to the molecular simulation evaluative output classification database 608 may comprise a simulated molecular structure of one or more molecules linked to one or more numerical values, a set of one or more values representing one or more parameters of one or more molecules, or the chemical reactions thereof, or a classification scheme of one or more molecules, or the chemical reactions thereof, a matrix of values representing a classification scheme of one or more molecules, or the chemical reactions thereof, or any combination thereof. In some embodiments, the simulated interactive molecular structure may be evaluated and assigned a numerical value comprising a score based on one or more numerical values representing a simulated action or a target gene, or both. In some embodiments, the molecular simulation evaluative output classification model or database 608 may output data to the train and tune model 617.
[0159] In some embodiments, the novel output molecular simulation evaluative model 604 may output data to a quantum chemistry evaluation model 603. In some embodiments, the data output to the quantum chemistry evaluation model 603 may comprise a simulated molecular interaction comprising a chemical reaction of one or more molecules one or more other materials, linked to one or more numerical values or a set of one or more numerical valuesrepresenting one or more parameters of the one or more molecules, or the interactions thereof, or a classification scheme of one or more molecules, or the interactions thereof, a matrix of values representing a classification scheme of one or more molecules, or the interactions thereof, or any combination thereof. In some embodiments, the quantum chemistry evaluation model 603 may comprise a simulated molecular interaction comprising a chemical reaction of a biological material with a separate biological material, in some cases a gene editor, a small molecule, a macromolecule, or another biological material. In some embodiments, the quantum chemistry evaluation model 603 may comprise a simulated molecular interaction comprising a chemical reaction of two molecules, three molecules, four molecules, five molecules, six molecules, seven molecules, eight molecules, nine molecules, ten molecules, or more than ten molecules. In some embodiments, the quantum chemistry evaluation model 603 may comprise one or more mathematical values for evaluation of one or more quantum chemistry reactions or characteristics between one or more molecules output by the novel output molecular simulation evaluative model 604. In some embodiments, the mathematical models of the quantum chemistry evaluation model 603 may utilize one or more values representing one or more characteristics of the molecular simulation of the interaction comprising one or more chemical reactions of one or more novel molecules output by the novel output molecular simulation evaluative model 604. In some embodiments, the quantum chemistry evaluation model 603 may output data to a quantum chemistry evaluation output classification model or database 607. In some embodiments, the data output by the quantum chemistry evaluation model 603 to the quantum chemistry evaluation output classification model or database 607 may comprise a simulated quantum chemical reaction of one or more molecules linked to one or more numerical values, a set of one or more values representing one or more parameters of one or more quantum chemical reaction of one or more molecules, or a classification scheme of quantum chemical reaction of one or more molecules, a matrix of values representing a classification scheme of quantum chemical reaction of one or more molecules, or any combination thereof. In some embodiments, the quantum chemistry evaluative output classification model or database 607 may output data to the train and tune model 617.
[0160] In some embodiments, the characteristics of the molecular simulation of the chemical reactions of one or more novel molecules may be binding specificity, binding site, kinetics, binding energy, stoichiometry, cleavage characteristics, charge, conformational changes, or a combination thereof. In some embodiments, the characteristics of the molecular simulation of the chemical reactions of one or more novel molecules may be impacted by pH, temperature, ionic strength, solvent composition, or a combination thereof. In some embodiments, the binding specificity is measured by a binding assay, including, for example, Isothermal TitrationCalorimetry (ITC), Fluorescence Polarization (FP), and Surface Plasmon Resonance (SPR). In some embodiments, the binding specificity is measured by a binding competition assay or a computational method. In some embodiments, measurement of the kinetics is rate constant, halflife, reaction rate, turnover number, or a combination thereof. In some embodiments, the binding energy is measured by ITC, SPR, a fluorescence-based assay, NMR, molecular docking, thermodynamic analysis, or a combination thereof. In some embodiments, the stoichiometry is measured by gel electrophoresis, ITC, SPR, mass spectrometry, or a combination thereof. In some embodiments, the cleavage characteristics is the type of enzymes, cleavage specificity, or mechanism of cleavage.
[0161] In some embodiments, the quantum chemistry evaluation model 603 may output data to a simulated prediction model 602. In some embodiments, the simulated prediction model 602 may comprise one or more algorithms for conducting a future quantum computing simulation. In some embodiments, the future quantum computing simulation may be generated based on data for one or more characteristics. In some embodiments, the one or more characteristics may comprise cleavage efficacy, one or more calculations of energy speed, an exothermic action barrier, one or more quantum calculations, or any combination thereof. In some embodiments, the simulated prediction model 602 may output data to a quantum simulation output classification model or database 606. In some embodiments, the data output by the simulated prediction model 602 to the quantum simulation output classification model or database 606 may comprise a quantum computing simulation of the molecular interaction comprising one or more chemical reactions of one or more molecules linked to one or more numerical values, a set of one or more values representing one or more parameters of one or more quantum computing simulations of the molecular interactions of one or more molecules comprising one or more chemical reactions, or a classification scheme of one or more quantum computing simulations of the molecular interactions of one or more molecules comprising one or more chemical reactions, a matrix of values representing a classification scheme of one or more quantum computing simulations of the molecular interactions of one or more molecules comprising one or more chemical reactions, or any combination thereof. In some embodiments, the quantum simulation output classification model 606 may output data to the train and tune model 617.
[0162] In some embodiments, the molecular simulation model 610 may output data to a novel molecular library 623. In some embodiments, the novel molecular library 623 may comprise a novel nuclease library. In some embodiments, the novel molecular library 623 may output one or more novel molecules to a user, wherein the novel molecules may possess one or more desired characteristics.
[0163] In some embodiments, the system or method for generating one or more novel molecules, in some aspects one or more novel nucleases, is shown in, for example FIG. 7. In some embodiments a machine learning / data training model 713 may comprise a protein data input model or database 715, an input raw molecular sequences model or database 717, a cloud server database 719, a data preprocessing model 718, a structural architecture model 714, and a train and tune model 716, or any combination thereof.
[0164] In some embodiments, the method or system may receive input data from a protein model or database 715. The protein model or database may comprise, for example UniProt Data 715. In some embodiments, the protein model or database may output data to an input raw molecular sequences model or database 717. In some embodiments, the input raw molecular sequences model or database 717 may be an input raw nuclease sequences model or database. In some embodiments, the input raw molecular sequences model or database 717 may transfer data through a cloud server 719 to a data preprocessing model 718. In some embodiments, the data preprocessing model 718 may preprocess the molecular data, and output the preprocessed data to a train and tune model 716. In some embodiments, the train and tune model 716 may receive input data from a structural architecture model 714. In some embodiments, the train and tune model 716 may receive data from one or more of an evaluative output classification model or database 709, a molecular simulation evaluative output classification model or database 708, a quantum chemistry evaluative output classification model or database 707, a quantum simulation output classification database 706, or any combination thereof. In some embodiments, the structural architecture model 714 may be an open source model. In some embodiments, the structural architecture model 714 may comprise an untrained, pretrained, or custom open source model. In some embodiments, the train and tune model 716 may output data to the structural architecture model 714.
[0165] In some embodiments, a generative / computational model 720 may comprise a structure generation model (SGM) 721. In some embodiments, the generative / computational model 720 may solely comprise a SGM 721. In some embodiments, the SGM 721 may be a custom trained SGM. In some embodiments, the SGM 721 may receive input data from a train and tune model 716. In some embodiments, the SGM 721 may output data to a generated molecular structures model or database 712.
[0166] In some embodiments, the computational validation model 701 may comprise a generated molecular structures model or database 712, one or more structural evaluation tools 711, a molecular simulation model 710, a quantum chemistry evaluation model 703, a simulated prediction model 702, an evaluative output classification model or database 709, a molecularsimulation evaluative output classification model or database 708, a quantum chemistry evaluative output classification model or database 707, a quantum simulation output classification database 706, or any combination thereof.
[0167] In some embodiments, the generated molecular structures model or database 712 may be a generated nuclease structures model or database. In some embodiments, the generated molecular structures model or database 712 may output data to one or more structural evaluation tools 711.
[0168] In some embodiments, the one or more structural evaluation tools 711 may output data to a novel output evaluation model 705. In some embodiments, the novel output evaluation model 705 may apply one or more algorithms to execute one or more method steps comprising receiving novel output from the one or more structural evaluation tools 711, classifying the novel output from the one or more structural evaluation tools 711, compare the novel output from the one or more structural evaluation tools 711 to known molecules, and evaluate the novel output from the one or more structural evaluation tools 711 based on one or more parameters. The one or more parameters may comprise one or more of solubility, sequence features, foldability, cleavage evaluation, subclassification, and catalytic site, or any combination thereof. In some embodiments, the novel output evaluation model 705 may output data to a molecular simulation model 710. In some embodiments, the novel output evaluation model 705 may output data to an evaluative output classification model or database 709. In some cases, the data output to the evaluative output classification model or database 709 may comprise a molecular structure of one or more molecules linked to one or more numerical values, a set of one or more values representing one or more evaluative parameters of one or more molecules, or the structures thereof, or a classification scheme of one or more molecules, or the structures thereof, a matrix of values representing an evaluative classification scheme of one or more molecules, or the structures thereof, or any combination thereof. In some embodiments, the evaluative output classification model or database 709 may output data to the train and tune model 716.
[0169] In some embodiments, the molecular simulation model 710 may output data to a novel output molecular simulation evaluative model 704. In some embodiments, the output data of the molecular simulation model 710 may comprise virtual modeling of the chemical reactions between one or more novel molecules. In some embodiments, the novel output molecular simulation evaluative model 704 may apply one or more algorithms configured to execute one or more method steps comprising receiving input data from the molecular simulation model 710, evaluating the input data from the molecular simulation model 710, comparing the input data from the molecular simulation model 710 to known molecular simulation data, simulating one ormore chemical reactions between one or more target genes and the input data from the molecular simulation model 710, and evaluating the input data based on one or more parameters. In some embodiments, the one or more parameters may comprise energy for binding, catalytic sites, active sites, sequences, PAM sequences, cleavage evaluation, charge, microenvironment, and delivery simulation without cleavage, or any combination thereof. In some embodiments, the novel output molecular simulation evaluative model 704 may output data to a molecular simulation evaluative output classification database 708. In some embodiments, the data output to the molecular simulation evaluative output classification database 708 may comprise a simulated interactive molecular structure of one or more molecules linked to one or more numerical values, a set of one or more values representing one or more parameters of one or more molecules, or the chemical reactions thereof, or a classification scheme of one or more molecules, or the chemical reactions thereof, a matrix of values representing a classification scheme of one or more molecules, or the chemical reactions thereof, or any combination thereof. In some embodiments, the molecular simulation evaluative output classification model or database 708 may output data to the train and tune model 716.
[0170] In some embodiments, the novel output molecular simulation evaluative model 704 may output data to a quantum chemistry evaluation model 703. In some embodiments, the data output to the quantum chemistry evaluation model 703 may comprise a simulated molecular interaction of one or more molecules comprising one or more chemical reactions linked to one or more numerical values, a set of one or more values representing one or more parameters of one or more molecules, or the molecular interactions comprising one or more chemical reactions thereof, or a classification scheme of one or more molecules, or the chemical reactions thereof, a matrix of values representing a classification scheme of one or more molecules, or the chemical reactions thereof, or any combination thereof. In some embodiments, the quantum chemistry evaluation model 703 may comprise one or more mathematical values for evaluation of one or more quantum chemistry reactions or characteristics between one or more molecules output by the novel output molecular simulation evaluative model 704. In some embodiments, the mathematical models of the quantum chemistry evaluation model 703 may utilize one or more values representing one or more characteristics of the molecular simulation of the quantum chemistry reactions between novel molecules output by the novel output molecular simulation evaluative model 704. In some embodiments, the quantum chemistry evaluation model 703 may output data to a quantum chemistry evaluation output classification model or database 707. In some embodiments, the data output by the quantum chemistry evaluation model 703 to the quantum chemistry evaluation output classification model or database 707 may comprise a simulated quantum chemical reaction of one or more molecules linked to one or more numericalvalues, a set of one or more values representing one or more parameters of one or more quantum chemical reactions of one or more molecules, or a classification scheme of quantum chemistry reactions of one or more molecules, a matrix of values representing a classification scheme of quantum chemistry reactions of one or more molecules, or any combination thereof. In some embodiments, the quantum chemistry evaluative output classification model or database 707 may output data to the train and tune model 716. In some embodiments, the characteristics of the simulation of molecular interactions comprising chemical reactions between novel molecules may be binding specificity, binding site, kinetics, binding energy, stoichiometry, cleavage characteristics, charge, conformational changes, or a combination thereof. In some embodiments, the characteristics of the molecular simulation of the chemical reactions between novel molecules may be impacted by pH, temperature, ionic strength, solvent composition, or a combination thereof. In some embodiments, the binding specificity is measured by a binding assay, including, for example, Isothermal Titration Calorimetry (ITC), Fluorescence Polarization (FP), and Surface Plasmon Resonance (SPR). In some embodiments, the binding specificity is measured by a binding competition assay or a computational method. In some embodiments, measurement of the kinetics is rate constant, half-life, reaction rate, turnover number, or a combination thereof. In some embodiments, the binding energy is measured by ITC, SPR, a fluorescence-based assay, NMR, molecular docking, thermodynamic analysis, or a combination thereof. In some embodiments, the stoichiometry is measured by gel electrophoresis, ITC, SPR, mass spectrometry, or a combination thereof. In some embodiments, the cleavage characteristics is the type of enzymes, cleavage specificity, or mechanism of cleavage.
[0171] In some embodiments, the quantum chemistry evaluation model 703 may output data to a simulated prediction model 702. In some embodiments, the simulated prediction model 702 may comprise one or more algorithms for conducting a future quantum computing simulation. In some embodiments, the future quantum computing simulation may be generated based on data for one or more characteristics. In some embodiments, the one or more characteristics may comprise cleavage efficacy, one or more calculations of energy speed, an exothermic action barrier, one or more quantum calculations, or any combination thereof. In some embodiments, the simulated prediction model 702 may output data to a quantum simulation output classification model or database 706. In some embodiments, the data output by the simulated prediction model 702 to the quantum simulation output classification model or database 706 may comprise a quantum computing simulation of the chemical reactions of one or more molecules linked to one or more numerical values, a set of one or more values representing one or more parameters of one or more quantum computing simulations of the chemical reactions of one or more molecules, or a classification scheme of one or more quantum computingsimulations of the chemical reactions of one or more molecules, a matrix of values representing a classification scheme of one or more quantum computing simulations of the chemical reactions of one or more molecules, or any combination thereof. In some embodiments, the quantum simulation output classification model 706 may output data to the train and tune model 716.
[0172] In some embodiments, the molecular simulation model 710 may output data to a novel molecular library 722. In some embodiments, the novel molecular library 722 may comprise a novel nuclease library. In some embodiments, the novel molecular library 722 may output one or more novel molecules to a user, wherein the novel molecules may possess one or more desired characteristics.Novel Molecular Generation and Machine Learning Models
[0173] In one non-limiting embodiment of the invention, as illustrated in FIG. 8, the invention may comprise a system or method 800 for assisting in generation of one or more novel molecules or biological systems. In some embodiments, the invention may comprise a method for assisting in generation of one or more novel molecules or biological systems. In some embodiments, the method may utilize a machine learning / data training model 801, a generative / computational model 813, and a computational validation model 816. In some embodiments, the machine learning / data training model 801 may comprise one or more model continuous improvement and testing models 806 comprising a new LLM model evaluation model 825, or a new structural model evaluation model 826, or both, a protein model or database 808, an input raw molecular sequences model or database 810, a cloud server platform 812, a data preprocessing model 811, a train and tune model 809, an LLM architecture model 807, and a model output feedback for machine leaming / training model 802 comprising an input evaluated output model 804, an output data model learning model 803, and an output data novel sequences with classification model or database 805.
[0174] In some embodiments, the generative / computational model may comprise a trained LLM model 814 and a trained structural model 815. In some embodiments, the trained LLM model may be a custom trained LLM model 814. In some embodiments, the trained structural model may be a custom trained structural model 815.
[0175] In some embodiments, the computational validation model 816 may comprise a generated molecular sequences model or database 823, a generated molecular structure model or database 822, a trained molecular classifier model 821, one or more molecular folding evaluation tools 819, one or more structural evaluation tools 820, a full validation model 817, and a molecular simulation model 818.
[0176] In some embodiments, the method may comprise providing an input to a large language learning model (LLM), and using the LLM to output one or more sequences of the one or more novel molecules or biological systems. In some embodiments, the method may comprise applying a validation model to evaluate or predict one or more predictive assessment tools of the output of the LLM. In some embodiments, the method may comprise outputting a sequence of the one or more novel molecules or biological systems, or structural representation thereof. In some embodiments, the output sequence or structural representation may be validated for the one or more predictive assessment tools. In some embodiments, the validation output may comprise the one or more novel molecules or biological systems. In some embodiments, the molecule may be a protein. In some embodiments, the molecule may be an enzyme. In some embodiments, the molecule may be a recombinase. In some embodiments, the molecule may be a structural protein. In some embodiments, the molecule may be a hormonal protein. In some embodiments, the molecule may be a receptor protein. In some embodiments, the molecule may be a transport protein. In some embodiments, the molecule may be a hydrolase. In some embodiments, the molecule may be an oxidoreductase. In some embodiments, the molecule may be a transferase. In some embodiments, the molecule may be a lyase. In some embodiments, the molecule may be an isomerase. In some embodiments, the molecule may be a ligase. In some embodiments, the molecule may be a carbohydrate active enzyme. In some embodiments, the molecule may be a polymerase. In some embodiments, the molecule may be a topoisomerase. In some embodiments, the molecule may be an ATPase. In some embodiments, the molecule may be a phosphatase. In some embodiments, the molecule may be a ubiquitin ligase. In some embodiments, the molecule may be a nucleic acid. In some embodiments, the molecule may be a deoxyribonucleic acid (DNA). In some embodiments, the molecule may be a ribonucleic acid (RNA). In some embodiments, the molecule may be a carbohydrate. In some embodiments, the molecule may be a lipid. In some embodiments, the molecule may be an ion. In some embodiments, the molecule may be a vitamin. In some embodiments, the molecule may be a hormone. In some embodiments, the molecule may be a metabolic intermediate. In some embodiments, the molecule may be a coenzyme and / or a cofactor. In some embodiments, the molecule may be a secondary metabolite. In some embodiments, the molecule may be a nuclease. In some embodiments, the molecule may be an endonuclease. In some embodiments, the molecule may be an exonuclease. In some embodiments, the molecule may be a viral nuclease. In some embodiments, the molecule may be an RNA-guided endonuclease. In some embodiments, the molecule may be a component of a Clustered Regularly Interspaced Short Palindromic Repeats and CRISPR-associated proteins (CRISPR-Cas) system. In some embodiments, the molecule may be a meganuclease. In some embodiments, the molecule maybe a transcription activator-like effector nuclease. In some embodiments, the molecule may be a zinc finger nuclease. In some embodiments, the biological systems may comprise DNA, RNA, ions, lipids, nanoparticles, small molecules, or any combination thereof.
[0177] In some embodiments, the input may comprise one or more of raw nuclease sequences, structural data, simulation data, quantum calculation data, or any combination thereof. In some embodiments, the input may comprise protein data 808. In some embodiments, the input may comprise, for example, UniProt data 808. In some embodiments, the input may comprise information related to one or more phenotypic characteristics. In some embodiments, the input data may be received by an input raw molecular sequences model or database 810. In some embodiments, the input raw molecular sequences model or database 810 may output raw molecular sequence data through a cloud server 812 to a data preprocessing model 811. In some embodiments, the molecular sequence data may be preprocessed and output by the data preprocessing model 811 to the train and tune model 809. In some embodiments, the input may be pre-processed to be configured as LLM input. In some embodiments, the train and tune model 809 may receive input from one or more model continuous improvement and testing models 806. In some embodiments, the train and tune model 809 may receive input from a new LLM model evaluation model 825, a new structural model evaluation model 826, an LLM architecture model 807, a model output feedback for ML / Training model 802, in some cases an output data novel sequences with classification model or database 805, and an LLM model 814, a structural model 815, or both. In some embodiments, the method may further comprise providing the input to an LLM model 814. In some embodiments, the LLM model may be a trained LLM model 814. In some embodiments, the LLM model may be a custom trained LLM model 814.
[0178] In some embodiments, the method may further comprise providing the input to a structural generation simulation module (SGSM) 815. In some embodiments, the SGSM may be a trained SGSM 815. In some embodiments, the SGSM may be a custom trained SGSM 815.
[0179] In some embodiments, the method may further comprise using the LLM 814 to output one or more sequences of the one or more novel molecules or biological systems. In some embodiments, the method may further comprise using the SGSM 815 to output one or more structures of the one or more novel molecules or biological systems. In some embodiments, the LLM model 814 may receive feedback input. In some embodiments, the feedback input may comprise the one or more structures output by the SGSM 815. In some embodiments, the feedback input may modify the LLM model 814 output. In some embodiments, the modifiedLLM model 814 output may comprise the one or more sequences of the one or more novel molecules or biological systems.
[0180] In some embodiments, the SGSM 815 may receive as feedback input the one or more sequences output by the LLM model 814. In some embodiments, the SGSM 815 may generate output comprising modified output. In some embodiments, the output of the SGSM 815 is modified based on feedback input from the LLM model 814. In some embodiments, the modified SGSM 815 output may comprise the one or more structures of the one or more novel molecules or biological systems.
[0181] In some embodiments, the method may further comprise applying an output data model learning model 803 to evaluate and classify the novel output of the LLM model 814, the novel output of the SGSM 815, or both. In some embodiments, the one or more groups categorize the one or more sequences of the one or more novel molecules or biological systems output by the LLM model 814, the SGSM 815, or both. In some embodiments, the SGSM 815 may output data to an input evaluated output model or database 804. In some embodiments, the input evaluated output model or database may comprise values assigned to each SGSM 815 output based on evaluation of one or more characteristic as compared to one or more inputs. In some cases, the one or more inputs may comprise LLM model 814 data, train and tune model 809 data, sequence data, or protein data. In some embodiments, the SGSM 815 and LLM model 814 may be retrained to optimize model improvement. In some embodiments, the SGSM 815 and LLM model 814 may be merged into a single model comprising both a structural generation component and a language model component. In some embodiments, the input evaluated output model 804 may output data to the output data model learning model 803. In some embodiments, the output data model learning model 803 may comprise one or more algorithms configured to learn methods comprising evaluating and classifying novel input. In some embodiments, the output data model learning model 803 may output evaluative or classification data, or both, to an output data novel sequences with classification model or database 805. In some embodiments, the output data novel sequences with classification model or database 805 may store or process the output data of the output data model learning model 803 associated with one or more corresponding sequences. In some embodiments, the evaluative or classification data may be used for feedback training. In some embodiments, the evaluative or classification data may comprise one or more numerical values attached to one or more identified parameters. In some embodiments, the output data novel sequence with classification model or database 805 may output data of the train and tune model 809. In some embodiments, the output of the train andtune model 809 may be received as feedback input by one or more models. In some embodiments, the feedback input may be used to train or fine tune one or more models.
[0182] In some embodiments, the LLM model 814, or the SGSM 815, or both, may output data to a computational validation model 816. In some embodiments, the LLM model 814 may output data to a generated molecular sequence model or database 823. In some embodiments, the generated molecular sequence model or database 823 may output one or more molecular sequences from the LLM model 814. In some embodiments, the generated molecular sequence model or database 823 may output data to a trained molecular classifier model 821. In some embodiments, the trained molecular classifier model 821 may classify one or more sequences input from the generated molecular sequences model or database 823. In some embodiments, the classification may be performed based on one or more characteristics of each molecule. In some embodiments, the trained molecular classifier model 821 may output data to one or more predictive assessment tools. In some embodiments, the one or more predictive assessment tools may comprise folding evaluation, structural evaluation, molecular simulation, molecular dynamics simulation, or any combination thereof. In some embodiments, the trained molecular classifier model 821 may output data to one or more molecular folding evaluation tools 819. In some embodiments, the molecular folding evaluation tools 819 may be third-party evaluation tools. In some embodiments, the molecular folding evaluation tools 819 may be one or more of AlphaFold, Rosetta, or DeepMind Equivariant Neural Networks, or any combination thereof. In some embodiments, the molecular folding evaluation tools 819 may be customized, bespoke, inhouse generated, or otherwise independently created molecular folding evaluation tools. In some embodiments, the one or more molecular folding evaluation tools 819 may output data to a full validation model 817. In some embodiments, the SGSM 815 may output data to a generated molecular structure model or database 822. In some embodiments, the generated molecular structure model or database 822 may output data to one or more structural evaluation tools 820. In some embodiments, the one or more structural evaluation tools 820 may output data to a full validation model 817. In some embodiments, the full validation model 817 may output validated data to a molecular simulation model 818, or receive data from a molecular simulation model 818, or both. In some embodiments, the full validation model 817 may output validation data to the input evaluated output model 804 of the model output feedback for machine learning / training model 802. In some embodiments, the molecular simulation model 818 may output data to a novel molecular library 824. In some embodiments, the novel molecular library 824 may output one or more novel molecules or biologic materials to a user.
[0183] In another non-limiting embodiment of the invention, as illustrated in FIG. 9, the invention may comprise a system or method 900 for assisting in generation of one or more novel molecules or biological systems. In some embodiments, the invention may comprise a method for assisting in generation of one or more novel molecules or biological systems. In some embodiments, the method may utilize a machine learning / data training model 901, a generative / computational model 920, and a computational validation model 912. In some embodiments, the machine learning / data training model 901 may comprise an LLM architecture model 906, a protein model or database 907, an input raw molecular sequences model or database 909, a cloud platform 911, a data preprocessing model 910, a train and tune model 908, a model output feedback for machine leaming / training model 926, and a machine learning of known molecules model 902 comprising an input raw molecular sequences model or database 905, an evaluative and classification of novel output model 903, and an output data molecular sequences model or database 904 with criteria.
[0184] In some embodiments, the generative / computational model 920 may comprise a trained LLM model 921 and a trained structural model 922. In some embodiments, the trained LLM model may be a custom trained LLM model 921. In some embodiments, the trained structural model may be a custom trained structural model 922.
[0185] In some embodiments, the computational validation model 912 may comprise a generated molecular sequences model or database 917, a generated molecular structure model or database 918, a trained molecular classifier model 919, one or more molecular folding evaluation tools 915, one or more structural evaluation tools 916, a full validation model 913, and a molecular simulation model 914.
[0186] In some embodiments, the method may comprise providing an input to a large language learning model (LLM) 921, and using the LLM to output one or more sequences of the one or more novel molecules or biological systems. In some embodiments, the method may comprise applying a validation model 913 to evaluate or predict one or more predictive assessment tools of the output of the LLM. In some embodiments, the method may comprise outputting a sequence of the one or more novel molecules or biological systems, or structural representation thereof. In some embodiments, the output sequence or structural representation may be validated for the one or more predictive assessment tools. In some embodiments, the validation output may comprise the one or more novel molecules or biological systems. In some embodiments, the molecule may be a protein. In some embodiments, the molecule may be an enzyme. In some embodiments, the molecule may be a recombinase. In some embodiments, the molecule may be a structural protein. In some embodiments, the molecule may be a hormonal protein. In someembodiments, the molecule may be a receptor protein. In some embodiments, the molecule may be a transport protein. In some embodiments, the molecule may be a hydrolase. In some embodiments, the molecule may be an oxidoreductase. In some embodiments, the molecule may be a transferase. In some embodiments, the molecule may be a lyase. In some embodiments, the molecule may be an isomerase. In some embodiments, the molecule may be a ligase. In some embodiments, the molecule may be a carbohydrate active enzyme. In some embodiments, the molecule may be a polymerase. In some embodiments, the molecule may be a topoisomerase. In some embodiments, the molecule may be an ATPase. In some embodiments, the molecule may be a phosphatase. In some embodiments, the molecule may be a ubiquitin ligase. In some embodiments, the molecule may be a nucleic acid. In some embodiments, the molecule may be a deoxyribonucleic acid (DNA). In some embodiments, the molecule may be a ribonucleic acid (RNA). In some embodiments, the molecule may be a carbohydrate. In some embodiments, the molecule may be a lipid. In some embodiments, the molecule may be an ion. In some embodiments, the molecule may be a vitamin. In some embodiments, the molecule may be a hormone. In some embodiments, the molecule may be a metabolic intermediate. In some embodiments, the molecule may be a coenzyme and / or a cofactor. In some embodiments, the molecule may be a secondary metabolite. In some embodiments, the molecule may be a nuclease. In some embodiments, the molecule may be an endonuclease. In some embodiments, the molecule may be an exonuclease. In some embodiments, the molecule may be a viral nuclease. In some embodiments, the molecule may be an RNA-guided endonuclease. In some embodiments, the molecule may be a component of a Clustered Regularly Interspaced Short Palindromic Repeats and CRISPR-associated proteins (CRISPR-Cas) system. In some embodiments, the molecule may be a meganuclease. In some embodiments, the molecule may be a transcription activator-like effector nuclease. In some embodiments, the molecule may be a zinc finger nuclease. In some embodiments, the biological systems may comprise DNA, RNA, ions, lipids, nanoparticles, small molecules, or any combination thereof.
[0187] In some embodiments, the input may comprise one or more of raw nuclease sequences, structural data, simulation data, quantum calculation data, or any combination thereof. In some embodiments, the input may comprise protein data 907. In some embodiments, the input may comprise, for example, UniProt data 907. In some embodiments, the input may comprise information related to one or more phenotypic characteristics. In some embodiments, the input data may be received by an input raw molecular sequences model or database 909. In some embodiments, the input raw molecular sequences model or database 909 may output raw molecular sequence data through a cloud server 911 to a data preprocessing model 910. In some embodiments, the molecular sequence data may be preprocessed and output by the datapreprocessing model 910 to the train and tune model 908. In some embodiments, the input may be pre-processed to be configured as LLM input. In some embodiments, the train and tune model 908 may receive input from an LLM architecture model 906. In some embodiments, the LLM architecture model 906 may be an open source LLM architecture model. In some embodiments, the LLM architecture model 906 may be an untrained, pretrained, or custom LLM architecture model. In some embodiments, the LLM architecture model 906 may be an untrained, pretrained, or custom open source LLM architecture model. In some embodiments, the train and tune model 908 may receive input from a data preprocessing model 910, an LLM architecture model 906, a model output feedback for ML / Training model 926, an output data molecular sequences with classification model or database 904, and an LLM model 921, a structural model 922, or both. In some embodiments, the method may further comprise providing input from the train and tune model 908 to an LLM model 921. In some embodiments, the LLM model may be a trained LLM model 921. In some embodiments, the LLM model may be a custom trained LLM model 921.
[0188] In some embodiments, the method may further comprise providing input from the train and tune model 908 to a structural generation simulation module (SGSM) 922. In some embodiments, the SGSM may be a trained SGSM 922. In some embodiments, the SGSM may be a custom trained SGSM 922.
[0189] In some embodiments, the method may further comprise using the LLM 921 to output one or more sequences of the one or more novel molecules or biological systems. In some embodiments, the method may further comprise using the SGSM 922 to output one or more structures of the one or more novel molecules or biological systems. In some embodiments, the LLM model 921 may receive feedback input. In some embodiments, the feedback input may comprise the one or more structures output by the SGSM 922. In some embodiments, the feedback input may modify the LLM model 921 output. In some embodiments, the modified LLM model 921 output may comprise the one or more sequences of the one or more novel molecules or biological systems.
[0190] In some embodiments, the SGSM 922 may receive as feedback input the one or more sequences output by the LLM model 921. In some embodiments, the SGSM 922 may generate output comprising modified output. In some embodiments, the output of the SGSM 922 is modified based on feedback input from the LLM model 921. In some embodiments, the modified SGSM 922 output may comprise the one or more structures of the one or more novel molecules or biological systems.
[0191] In some embodiments, the method may further comprise applying a model output feedback for ML / training model 926 to evaluate and classify the novel output of the LLM model 921, the novel output of the SGSM 922, or both into one or more groups. In some embodiments, the one or more groups categorize the one or more sequences of the one or more novel molecules or biological systems output by the LLM model 921, the SGSM 922, or both. In some embodiments, the model output feedback for ML / training model 926 may receive input from a machine learning of known molecules model 902. In some embodiments, the machine learning of known molecules model 902 may receive input from the model output feedback for ML / training model 926. In some embodiments, the machine learning of known molecules model 903 may comprise an input raw molecular sequences model or database 905.
[0192] In some embodiments, the input raw molecular sequences model or database 905 may output data to an evaluative and classifying novel output model 903. In some embodiments, the evaluative and classifying novel output model 903 may comprise one or more algorithms configured to execute one or more methods comprising classifying and comparing one or more molecules from the input raw molecular sequences model or database 905. In some embodiments, the evaluative and classifying novel output model 903 may comprise one or more algorithms configured to execute one or more methods comprising identifying and subcategorizing molecules from the input raw molecular sequences model or database 905. In some embodiments, the evaluative and classifying novel output model 903 utilizes one or more input criteria to for machine learning. In some embodiments, the input criteria comprise one or more molecular characteristics. In some embodiments, the one or more molecular characteristics of the input criteria comprise one or more of sequence, efficacy, size, specificity, PAM sequence criteria, RNA binding, and cleavage efficacy, or any combination thereof. In some embodiments, the evaluative and classifying novel output model 903 outputs data to an output data molecular sequence model or database 904 with criteria. In some embodiments, the criteria comprises the input criteria. In some embodiments, the input criteria may comprise one or more of sequence, efficacy, size, specificity, PAM sequence criteria, RNA binding, and cleavage efficacy, or any combination thereof. In some embodiments, the input criteria may comprise training data of one or more specific features and parameters. In some embodiments, the training data may be custom training data, intentional training data, or both. In some embodiments, the one or more input criteria may be linked to the one or more molecular sequences in the output data molecular sequences model or database 904 with criteria. In some embodiments, the output data molecular sequences model or database 904 with criteria outputs data to the train and tune model 908. In some cases, the criteria may be received by one or more models as feedback input. In some cases, the feedback input may be used to train the one or more models.
[0193] In some embodiments, the LLM model 921, or the SGSM 922, or both, may output data to a computational validation model 912. In some embodiments, the LLM model 921 may output data to a generated molecular sequence model or database 917. In some embodiments, the generated molecular sequence model or database 917 may output one or more molecular sequences from the LLM model 921. In some embodiments, the generated molecular sequence model or database 917 may output data to a trained molecular classifier model 919. In some embodiments, the trained molecular classifier model 919 may classify one or more sequences input from the generated molecular sequences model or database 917. In some embodiments, the classification may be performed based on one or more characteristics of each molecule. In some embodiments, the trained molecular classifier model 919 may output data to one or more predictive assessment tools. In some embodiments, the one or more predictive assessment tools may comprise folding evaluation, structural evaluation, molecular simulation, molecular dynamics simulation, or any combination thereof. In some embodiments, the trained molecular classifier model 919 may output data to one or more molecular folding evaluation tools 915 In some embodiments, the molecular folding evaluation tools 915 may be third-party evaluation tools. In some embodiments, the molecular folding evaluation tools 915 may be one or more of AlphaFold, Rosetta, or DeepMind Equivariant Neural Networks, or any combination thereof. In some embodiments, the molecular folding evaluation tools 915 may be customized, bespoke, inhouse generated, or otherwise independently created molecular folding evaluation tools. In some embodiments, the one or more molecular folding evaluation tools 915 may output data to a full validation model 913. In some embodiments, the SGSM 922 may output data to a generated molecular structure model or database 918. In some embodiments, the generated molecular structure model or database 918 may output data to one or more structural evaluation tools 916.In some embodiments, the one or more structural evaluation tools 916 may output data to a full validation model 913. In some embodiments, the full validation model 913 may output validated data to a molecular simulation model 914, or receive data from a molecular simulation model 914, or both. In some embodiments, the full validation model 913 may output validation data to the model output feedback for ML / training model 926. In some embodiments, the molecular simulation model 914 may output data to a novel molecular library 923. In some embodiments, the novel molecular library 923 may output one or more novel molecules or biologic materials to a user.
[0194] In yet another non-limiting embodiment of the invention, as illustrated in FIG. 10, the invention may comprise a system or method 1000 for assisting in generation of one or more novel molecules or biological systems. In some embodiments, the invention may comprise a method for assisting in generation of one or more novel molecules or biological systems. Insome embodiments, the method may utilize a machine learning / data training model 1001, a generative / computational model 1014, and a computational validation model 1017. In some embodiments, the machine learning / data training model 1001 may comprise an LLM architecture model 1006, a protein model or database 1007, an input raw molecular sequences model or database 1009, a cloud platform 1011, a data preprocessing model 1010, a train and tune model 1008, a model output feedback for machine leaming / training model 1012, and a machine learning of molecular evolution model 1002 comprising an input phylogenic data mutation discoveries model or database 1005, an evaluative and classification of evolution and mutations model 1003, and an output data molecular sequences model or database 1004 with criteria.
[0195] In some embodiments, the generative / computational model 1014 may comprise a trained LLM model 1015 and a trained structural model 1016. In some embodiments, the trained LLM model may be a custom trained LLM model 1015. In some embodiments, the trained structural model may be a custom trained structural model 1016.
[0196] In some embodiments, the computational validation model 1017 may comprise a generated molecular sequences model or database 1021, a generated molecular structure model or database 1022, a trained molecular classifier model 1023, one or more molecular folding evaluation tools 1025, one or more structural evaluation tools 1024, a full validation model 1019, and a molecular simulation model 1020.
[0197] In some embodiments, the method may comprise providing an input to a large language learning model (LLM) 1015, and using the LLM to output one or more sequences of the one or more novel molecules or biological systems. In some embodiments, the method may comprise applying a validation model 1019 to evaluate or predict the effectiveness of the output of the LLM model 1015 using one or more predictive assessment tools. In some embodiments, the method may comprise outputting a sequence of the one or more novel molecules or biological systems, or structural representation thereof. In some embodiments, the output sequence or structural representation may be validated for the one or more predictive assessment tools. In some embodiments, the validation output may comprise the one or more novel molecules or biological systems. In some embodiments, the molecule may be a protein. In some embodiments, the molecule may be an enzyme. In some embodiments, the molecule may be a recombinase. In some embodiments, the molecule may be a structural protein. In some embodiments, the molecule may be a hormonal protein. In some embodiments, the molecule may be a receptor protein. In some embodiments, the molecule may be a transport protein. In some embodiments, the molecule may be a hydrolase. In some embodiments, the molecule maybe an oxidoreductase. In some embodiments, the molecule may be a transferase. In some embodiments, the molecule may be a lyase. In some embodiments, the molecule may be an isomerase. In some embodiments, the molecule may be a ligase. In some embodiments, the molecule may be a carbohydrate active enzyme. In some embodiments, the molecule may be a polymerase. In some embodiments, the molecule may be a topoisomerase. In some embodiments, the molecule may be an ATPase. In some embodiments, the molecule may be a phosphatase. In some embodiments, the molecule may be a ubiquitin ligase. In some embodiments, the molecule may be a nucleic acid. In some embodiments, the molecule may be a deoxyribonucleic acid (DNA). In some embodiments, the molecule may be a ribonucleic acid (RNA). In some embodiments, the molecule may be a carbohydrate. In some embodiments, the molecule may be a lipid. In some embodiments, the molecule may be an ion. In some embodiments, the molecule may be a vitamin. In some embodiments, the molecule may be a hormone. In some embodiments, the molecule may be a metabolic intermediate. In some embodiments, the molecule may be a coenzyme and / or a cofactor. In some embodiments, the molecule may be a secondary metabolite. In some embodiments, the molecule may be a nuclease. In some embodiments, the molecule may be an endonuclease. In some embodiments, the molecule may be an exonuclease. In some embodiments, the molecule may be a viral nuclease. In some embodiments, the molecule may be an RNA-guided endonuclease. In some embodiments, the molecule may be a component of a Clustered Regularly Interspaced Short Palindromic Repeats and CRISPR-associated proteins (CRISPR-Cas) system. In some embodiments, the molecule may be a meganuclease. In some embodiments, the molecule may be a transcription activator-like effector nuclease. In some embodiments, the molecule may be a zinc finger nuclease. In some embodiments, the biological systems may comprise DNA, RNA, ions, lipids, nanoparticles, small molecules, or any combination thereof.
[0198] In some embodiments, the input may comprise one or more of raw nuclease sequences, structural data, simulation data, quantum calculation data, or any combination thereof. In some embodiments, the input may comprise protein data 1007. In some embodiments, the input may comprise, for example, UniProt data 1007. In some embodiments, the input may comprise information related to one or more phenotypic characteristics. In some embodiments, the input data may be received by an input raw molecular sequences model or database 1009. In some embodiments, the input raw molecular sequences model or database 1009 may output raw molecular sequence data through a cloud server 1011 to a data preprocessing model 1010. In some embodiments, the molecular sequence data may be preprocessed and output by the data preprocessing model 1010 to the train and tune model 1008. In some embodiments, the input may be pre-processed to be configured as LLM input. In some embodiments, the train and tunemodel 1008 may receive input from an LLM architecture model 1006. In some embodiments, the LLM architecture model 1006 may be an open source LLM architecture model. In some embodiments, the LLM architecture model 1006 may be an untrained, pretrained, or custom LLM architecture model. In some embodiments, the LLM architecture model 1006 may be an untrained, pretrained, or custom open source LLM architecture model. In some embodiments, the train and tune model 1008 may receive input from a data preprocessing model 1010, an LLM architecture model 1006, a model output feedback for ML / Training model 1012, an output data molecular sequences with classification model or database 1004, and an LLM model 1015, a structural model 1016, or both. In some embodiments, the method may further comprise providing input from the train and tune model 1008 to an LLM model 1015. In some embodiments, the LLM model may be a trained LLM model 1015. In some embodiments, the LLM model may be a custom trained LLM model 1015. In some embodiments, the LLM model 1015 may comprise a GPT-1, GPT-2, GPT-3, GPT-4, GPT-5, GPT-6, GPT-7, GPT-8, GPT-9, GPT-10, or another GPT model. In some embodiments, the LLM model 1015 may comprise a protein language model. In some embodiments, the LLM model 1015 may comprise a protein language model comprising one or more of ESM-1, ESM-lv, ESM-2, ESM-2v, ESM-3, ESM-3v, ESM-4, ESM-4v, ESM-5, ESM-5v, ESMFold, ProGen, proteinBERT, another protein language model, or any combination thereof. In some embodiments, the LLM model 1015 may comprise an antibody-specific protein language model. In some embodiments, the LLM model 1015 may comprise an antibody-specific protein language model comprising one or more of IgLM, AbLang, AbLang2, AbLang3, AbLang4, AbLang5, AbLang6, AbLang7, AbLang8, AbLang9, AbLang 10, AntiBERTa, Bio-inspired Antibody Language Model (BALM), BALM-paired and unpaired, O AS-trained RoBERTa, IgBert, IgT5, AntiBERTy, Sapiens, FAbConSapiens, or another antibody-specific protein language model, or any combination thereof.
[0199] In some embodiments, the method may further comprise providing input from the train and tune model 1008 to a structural generation simulation module (SGSM) 1016. In some embodiments, the SGSM may be a trained SGSM 1016. In some embodiments, the SGSM may be a custom trained SGSM 1016.
[0200] In some embodiments, the method may further comprise using the LLM 1015 to output one or more sequences of the one or more novel molecules or biological systems. In some embodiments, the method may further comprise using the SGSM 1016 to output one or more structures of the one or more novel molecules or biological systems. In some embodiments, the LLM model 1015 may receive feedback input. In some embodiments, the feedback input maycomprise the one or more structures output by the SGSM 1016. In some embodiments, the feedback input may modify the LLM model 1015 output. In some embodiments, the modified LLM model 1015 output may comprise the one or more sequences of the one or more novel molecules or biological systems.
[0201] In some embodiments, the SGSM 1016 may receive as feedback input the one or more sequences output by the LLM model 1015. In some embodiments, the SGSM 1016 may generate output comprising modified output. In some embodiments, the output of the SGSM 1016 is modified based on feedback input from the LLM model 1015. In some embodiments, the modified SGSM 1016 output may comprise the one or more structures of the one or more novel molecules or biological systems.
[0202] In some embodiments, the method may further comprise applying a model output feedback for ML / training model 1012 to evaluate and classify the novel output of the LLM model 1015, the novel output of the SGSM 1016, or both into one or more groups. In some embodiments, the one or more groups categorize the one or more sequences of the one or more novel molecules or biological systems output by the LLM model 1015, the SGSM 1016, or both. In some embodiments, the model output feedback for ML / training model 1012 may receive input from a machine learning of molecular evolution model 1002. In some embodiments, the machine learning of molecular evolution model 1002 may receive input from the model output feedback for ML / training model 1012. In some embodiments, the machine learning of molecular evolution model 1002 may comprise a phylogenic data mutation discoveries model or database 1005
[0203] In some embodiments, the phylogenic data mutation discoveries model or database 1005 may output data to an evaluative and classifying evolution and mutations model 1003. In some embodiments, the evaluative and classifying evolution and mutations model 1003 may comprise one or more algorithms configured to execute one or more methods comprising classifying and comparing one or more molecules from the phylogenic data mutation discoveries model or database 1005. In some embodiments, the evaluative and classifying evolution and mutations model 1003 may comprise one or more algorithms configured to execute one or more methods comprising identifying and subcategorizing molecules from the phylogenic data mutation discoveries model or database 1005. In some embodiments, the evaluative and classifying evolution and mutations model 1003 may comprise one or more algorithms configured to execute one or more methods comprising identifying similar learning criteria for molecules from the phylogenic data mutation discoveries model or database 1005. In some embodiments, the evaluative and classifying evolution and mutations model 1003 may utilize one or more inputcriteria to for machine learning. In some embodiments, the evaluative evolution and mutations model 1003 may assess one or more features of one or more biological materials using the one or more input criteria. In some embodiments, the evaluative evolution and mutations model 1003 may assess one or more features of one or more biological materials at one or more time points. In some embodiments, the time points may be one or more past time points. In some embodiments, the time points may be one or more future time points. In some embodiments, the one or more features may be predicted features. In some embodiments, the evaluative evolution and mutations model 1003 may evaluate one or more features of one or more biological materials, in some cases one or more gene editors or one or more other molecules, over time in their evolutionary history. In some embodiments, the input criteria comprise one or more molecular characteristics. In some embodiments, the input criteria may be used to assess one or more molecular characteristics at one or more time points. In some embodiments, the one or more molecular characteristics of the input criteria comprise one or more of sequence, efficacy, size, specificity, PAM sequence criteria, RNA binding, and cleavage efficacy, or any combination thereof. In some embodiments, the evaluative and classifying evolution and mutations model 1003 outputs data to an output data molecular sequence model or database 1004 with criteria. In some embodiments, the criteria comprises the input criteria. In some embodiments, the input criteria comprises one or more of sequence, efficacy, size, specificity, PAM sequence criteria, RNA binding, and cleavage efficacy, or any combination thereof. In some embodiments, the one or more input criteria may be linked to the one or more molecular sequences in the output data molecular sequences model or database 1004 with criteria. In some embodiments, the output data molecular sequences model or database 1004 with criteria outputs data to the train and tune model 1008. In some embodiments, the output data from the evaluative and classifying evolution and mutations model 1003 stored in the sequence model or database 1004 may comprise assessments of one or more of the criteria at one or more time points. In some embodiments, the one or more time points may be past time points. In some embodiments, the one or more time points may be future time points. In some embodiments, the one or more input criteria may be used by the evaluative and classifying evolution and mutations model 1003 to predict one or more characteristics of one or more biological materials, in some cases gene editors or other molecules. In some embodiments, the one or more input criteria may be used by the evaluative and classifying evolution and mutations model 1003 to assess one or more features in the evolutionary history of one or more biological materials, in some cases gene editors or other molecules.
[0204] In some embodiments, the LLM model 1015, or the SGSM 1016, or both, may output data to a computational validation model 1017. In some embodiments, the LLM model 1015may output data to a generated molecular sequence model or database 1021. In some embodiments, the generated molecular sequence model or database 1021 may output one or more molecular sequences from the LLM model 1015. In some embodiments, the generated molecular sequence model or database 1021 may output data to a trained molecular classifier model 1023. In some embodiments, the trained molecular classifier model 1023 may classify one or more sequences input from the generated molecular sequences model or database 1021.In some embodiments, the classification may be performed based on one or more characteristics of each molecule. In some embodiments, the trained molecular classifier model 1023 may output data to one or more predictive assessment tools. In some embodiments, the one or more predictive assessment tools may comprise folding evaluation, structural evaluation, molecular simulation, molecular dynamics simulation, or any combination thereof. In some embodiments, the trained molecular classifier model 1023 may output data to one or more molecular folding evaluation tools 1025. In some embodiments, the molecular folding evaluation tools 1025 may be third-party evaluation tools. In some embodiments, the molecular folding evaluation tools 1025 may be one or more of AlphaFold, Rosetta, or DeepMind Equivariant Neural Networks, or any combination thereof. In some embodiments, the molecular folding evaluation tools 1025 may be customized, bespoke, in-house generated, or otherwise independently created molecular folding evaluation tools. In some embodiments, the one or more molecular folding evaluation tools 1025 may output data to a full validation model 1019. In some embodiments, the SGSM 1016 may output data to a generated molecular structure model or database 1022. In some embodiments, the generated molecular structure model or database 1022 may output data to one or more structural evaluation tools 1024. In some embodiments, the one or more structural evaluation tools 1024 may output data to a full validation model 1019. In some embodiments, the full validation model 1019 may output validated data to a molecular simulation model 1020, or receive data from a molecular simulation model 1020, or both. In some embodiments, the full validation model 1019 may output validation data to the model output feedback for ML / training model 1012. In some embodiments, the molecular simulation model 1020 may output data to a novel molecular library 1026. In some embodiments, the novel molecular library 1026 may output one or more novel molecules or biologic materials to a user.
[0205] In yet another non-limiting embodiment of the invention, as illustrated in FIG. 11, the invention may comprise a system or method 1100 for assisting in generation of one or more novel molecules or biological systems. In some embodiments, the invention may comprise a method for assisting in generation of one or more novel molecules or biological systems. In some embodiments, the method may utilize a machine learning / data training model 1101, a generative / computational model 1114, and a computational validation model 1117. In someembodiments, the machine leaming / data training model 1101 may comprise an LLM architecture model 1106, a protein model or database 1107, an input raw molecular sequences model or database 1109, a cloud platform 1111, a data preprocessing model 1110, a train and tune model 1108, a model output feedback for machine leaming / training model 1130, a known molecules model for ML / training model 1131, and a machine learning of target genes model 1102 comprising an input raw target DNA sequences model or database 1105 with criteria, an evaluative and classification of target genes model 1103, and an output data target DNA sequences model or database 1104 with criteria.
[0206] In some embodiments, the generative / computational model 1114 may comprise a trained LLM model 1116 and a trained structural model 1112. In some embodiments, the trained LLM model may be a custom trained LLM model 1116. In some embodiments, the trained structural model may be a custom trained structural model 1112.
[0207] In some embodiments, the computational validation model 1117 may comprise a generated molecular sequences model or database 1123, a generated molecular structure model or database 1122, a trained molecular classifier model 1121, one or more molecular folding evaluation tools 1119, one or more structural evaluation tools 1120, a full validation model 1124, and a molecular simulation model 1118.
[0208] In some embodiments, the method may comprise providing an input to a large language learning model (LLM) 1116, and using the LLM to output one or more sequences of the one or more novel molecules or biological systems. In some embodiments, the method may comprise applying a validation model 1124 to evaluate or predict the effectiveness of the output of the LLM model 1116 using one or more predictive assessment tools. In some embodiments, the method may comprise outputting a sequence of the one or more novel molecules or biological systems, or structural representation thereof. In some embodiments, the output sequence or structural representation may be validated for the one or more predictive assessment tools. In some embodiments, the validation output may comprise the one or more novel molecules or biological systems. In some embodiments, the molecule may be a protein. In some embodiments, the molecule may be an enzyme. In some embodiments, the molecule may be a recombinase. In some embodiments, the molecule may be a structural protein. In some embodiments, the molecule may be a hormonal protein. In some embodiments, the molecule may be a receptor protein. In some embodiments, the molecule may be a transport protein. In some embodiments, the molecule may be a hydrolase. In some embodiments, the molecule may be an oxidoreductase. In some embodiments, the molecule may be a transferase. In some embodiments, the molecule may be a lyase. In some embodiments, the molecule may be anisomerase. In some embodiments, the molecule may be a ligase. In some embodiments, the molecule may be a carbohydrate active enzyme. In some embodiments, the molecule may be a polymerase. In some embodiments, the molecule may be a topoisomerase. In some embodiments, the molecule may be an ATPase. In some embodiments, the molecule may be a phosphatase. In some embodiments, the molecule may be a ubiquitin ligase. In some embodiments, the molecule may be a nucleic acid. In some embodiments, the molecule may be a deoxyribonucleic acid (DNA). In some embodiments, the molecule may be a ribonucleic acid (RNA). In some embodiments, the molecule may be a carbohydrate. In some embodiments, the molecule may be a lipid. In some embodiments, the molecule may be an ion. In some embodiments, the molecule may be a vitamin. In some embodiments, the molecule may be a hormone. In some embodiments, the molecule may be a metabolic intermediate. In some embodiments, the molecule may be a coenzyme and / or a cofactor. In some embodiments, the molecule may be a secondary metabolite. In some embodiments, the molecule may be a nuclease. In some embodiments, the molecule may be an endonuclease. In some embodiments, the molecule may be an exonuclease. In some embodiments, the molecule may be a viral nuclease. In some embodiments, the molecule may be an RNA-guided endonuclease. In some embodiments, the molecule may be a component of a Clustered Regularly Interspaced Short Palindromic Repeats and CRISPR-associated proteins (CRISPR-Cas) system. In some embodiments, the molecule may be a meganuclease. In some embodiments, the molecule may be a transcription activator-like effector nuclease. In some embodiments, the molecule may be a zinc finger nuclease. In some embodiments, the biological systems may comprise DNA, RNA, ions, lipids, nanoparticles, small molecules, or any combination thereof.
[0209] In some embodiments, the input may comprise one or more of raw nuclease sequences, structural data, simulation data, quantum calculation data, or any combination thereof. In some embodiments, the input may comprise protein data 1107. In some embodiments, the input may comprise, for example, UniProt data 1107. In some embodiments, the input may comprise information related to one or more phenotypic characteristics. In some embodiments, the input data may be received by an input raw molecular sequences model or database 1109. In some embodiments, the input raw molecular sequences model or database 1109 may output raw molecular sequence data through a cloud server 1111 to a data preprocessing model 1110. In some embodiments, the molecular sequence data may be preprocessed and output by the data preprocessing model 1110 to the train and tune model 1108. In some embodiments, the input may be pre-processed to be configured as LLM input. In some embodiments, the train and tune model 1108 may receive input from an LLM architecture model 1106. In some embodiments, the LLM architecture model 1106 may be an open source LLM architecture model. In someembodiments, the LLM architecture model 1106 may be an untrained, pretrained, or custom LLM architecture model. In some embodiments, the LLM architecture model 1106 may be an untrained, pretrained, or custom open source LLM architecture model. In some embodiments, the train and tune model 1108 may receive input from a data preprocessing model 1110, an LLM architecture model 1106, a model output feedback for ML / Training model 1130, an output data target DNA sequences with classification model or database 1104, and an LLM model 1116, a structural model 1112, or both. In some embodiments, the method may further comprise providing input from the train and tune model 1108 to an LLM model 1116. In some embodiments, the LLM model may be a trained LLM model 1116. In some embodiments, the LLM model may be a custom trained LLM model 1116. In some embodiments, the LLM model 1116 may comprise a GPT-1, GPT-2, GPT-3, GPT-4, GPT-5, GPT-6, GPT-7, GPT-8, GPT-9, GPT-10, or another GPT model. In some embodiments, the LLM model 1116 may comprise a protein language model. In some embodiments, the LLM model 1116 may comprise a protein language model comprising one or more of ESM-1, ESM-lv, ESM-2, ESM-2v, ESM-3, ESM-3v, ESM-4, ESM-4v, ESM-5, ESM-5v, ESMFold, ProGen, proteinBERT, another protein language model, or any combination thereof. In some embodiments, the LLM model 1116 may comprise an antibody-specific protein language model. In some embodiments, the LLM model 1116 may comprise an antibody-specific protein language model comprising one or more of IgLM, AbLang, AbLang2, AbLang3, AbLang4, AbLang5, AbLang6, AbLang7, AbLang8, AbLang9, AbLang 10, AntiBERTa, Bio-inspired Antibody Language Model (BALM), BALM-paired and unpaired, O AS-trained RoBERTa, IgBert, IgT5, AntiBERTy, Sapiens, FAbConSapiens, or another antibody-specific protein language model, or any combination thereof.
[0210] In some embodiments, the method may further comprise providing input from the train and tune model 1108 to a structural generation simulation module (SGSM) 1112. In some embodiments, the SGSM may be a trained SGSM 1112. In some embodiments, the SGSM may be a custom trained SGSM 1112.
[0211] In some embodiments, the method may further comprise using the LLM 1116 to output one or more sequences of the one or more novel molecules or biological systems. In some embodiments, the method may further comprise using the SGSM 1112 to output one or more structures of the one or more novel molecules or biological systems. In some embodiments, the LLM model 1116 may receive feedback input. In some embodiments, the feedback input may comprise the one or more structures output by the SGSM 1112. In some embodiments, the feedback input may modify the LLM model 1116 output. In some embodiments, the modifiedLLM model 1116 output may comprise the one or more sequences of the one or more novel molecules or biological systems.
[0212] In some embodiments, the SGSM 1112 may receive as feedback input the one or more sequences output by the LLM model 1116. In some embodiments, the SGSM 1112 may generate output comprising modified output. In some embodiments, the output of the SGSM 1112 is modified based on feedback input from the LLM model 1116. In some embodiments, the modified SGSM 1112 output may comprise the one or more structures of the one or more novel molecules or biological systems.
[0213] In some embodiments, the method may further comprise applying a model output feedback for ML / training model 1130 to evaluate and classify the novel output of the LLM model 1116, the novel output of the SGSM 1112, or both into one or more groups. In some embodiments, the one or more groups categorize the one or more sequences of the one or more novel molecules or biological systems output by the LLM model 1116, the SGSM 1112, or both. In some embodiments, the model output feedback for ML / training model 1130 may receive input from a machine learning of target genes model 1102. In some embodiments, the machine learning of target genes model 1102 may receive input from the model output feedback for ML / training model 1130. In some embodiments, the model output feedback for ML / training model 1130 may receive input from a known molecules model for ML / training 1131. In some embodiments, the machine learning of target genes model 1102 may receive input from the known molecules model for ML / training model 1131. In some embodiments, the known molecules model for ML / training model 1131 may receive input from the machine learning of target genes model 1102. In some embodiments, the known molecules model for ML / training model 1131 may receive input from the model output feedback for ML / training model 1130.
[0214] In some embodiments, the machine learning of target genes model 1102 may comprise an input raw target DNA sequences model or database 1105 with criteria. In some embodiments, the input raw target DNA sequences model or database 1105 with criteria may output data to an evaluative and classifying target genes model 1103. In some embodiments, the evaluative and classifying target genes model 1103 may comprise one or more algorithms configured to execute one or more methods comprising classifying and comparing one or more molecules from the input raw target DNA sequences model or database 1105 with criteria. In some embodiments, the evaluative and classifying target genes model 1103 may comprise one or more algorithms configured to execute one or more methods comprising identifying and subcategorizing target genes data from the input raw target DNA sequences model or database 1105 with criteria. In some embodiments, the evaluative and classifying target genes model 1103 may utilize one ormore input criteria to for machine learning. In some embodiments, the input criteria comprise one or more molecular characteristics. In some embodiments, the one or more molecular characteristics of the input criteria comprise one or more of sequences, cleavage need, for example single nucleotide or double nucleotide, insertion needs, PAM sequence, binding criteria, large deletion, large insertion, tissue properties, micro-environment needs, for example pH, ion concentration, and oxygen concentration, and delivery features, sequence, efficacy, size, specificity, RNA binding, cleavage efficacy, target sequences such as target nucleotide sequences, or any combination thereof. In some embodiments, the evaluative and classifying target genes model 1103 outputs data to an output data target DNA sequences model or database 1104 with criteria. In some embodiments, the criteria comprises the input criteria. In some embodiments, the input criteria comprises sequence, cleavage need, for example single nucleotide or double nucleotide, insertion needs, PAM sequence, binding criteria, large deletion, large insertion, tissue properties, micro-environment needs, for example pH, ion concentration, and oxygen concentration, and delivery features, sequence, efficacy, size, specificity, RNA binding, cleavage efficacy, target sequences such as target nucleotide sequences, or any combination thereof. In some embodiments, the one or more input criteria may be linked to the one or more target gene sequences in the output data target DNA sequences model or database 1104 with criteria. In some embodiments, the output data target DNA sequences model or database 1104 with criteria outputs data to the train and tune model 1108.
[0215] In some embodiments, the LLM model 1116, or the SGSM 1112, or both, may output data to a computational validation model 1117. In some embodiments, the LLM model 1116 may output data to a generated molecular sequence model or database 1123. In some embodiments, the generated molecular sequence model or database 1123 may output one or more molecular sequences from the LLM model 1116. In some embodiments, the generated molecular sequence model or database 1123 may output data to a trained molecular classifier model 1121. In some embodiments, the trained molecular classifier model 1121 may classify one or more sequences input from the generated molecular sequences model or database 1123.In some embodiments, the classification may be performed based on one or more characteristics of each molecule. In some embodiments, the trained molecular classifier model 1121 may output data to one or more predictive assessment tools. In some embodiments, the one or more predictive assessment tools may comprise folding evaluation, structural evaluation, molecular simulation, molecular dynamics simulation, or any combination thereof. In some embodiments, the trained molecular classifier model 1121 may output data to one or more molecular folding evaluation tools 1119. In some embodiments, the molecular folding evaluation tools 1119 may be third-party evaluation tools. In some embodiments, the molecular folding evaluation tools1119 may be one or more of AlphaFold, Roseta, or DeepMind Equivariant Neural Networks, or any combination thereof. In some embodiments, the molecular folding evaluation tools 1119 may be customized, bespoke, in-house generated, or otherwise independently created molecular folding evaluation tools. In some embodiments, the one or more molecular folding evaluation tools 1119 may output data to a full validation model 1124. In some embodiments, the SGSM 1112 may output data to a generated molecular structure model or database 1122. In some embodiments, the generated molecular structure model or database 1122 may output data to one or more structural evaluation tools 1120. In some embodiments, the one or more structural evaluation tools 1120 may output data to a full validation model 1124. In some embodiments, the full validation model 1124 may output validated data to a molecular simulation model 1118, or receive data from a molecular simulation model 1118, or both. In some embodiments, the full validation model 1124 may output validation data to the model output feedback for ML / training model 1130. In some embodiments, the molecular simulation model 1118 may output data to a novel molecular library 1125. In some embodiments, the novel molecular library 1125 may output one or more novel molecules or biologic materials to a user.
[0216] In yet another non -limiting embodiment of the invention, as illustrated in FIG. 12, the invention may comprise a system or method 1200 for assisting in generation of one or more novel molecules or biological systems. In some embodiments, the invention may comprise a method for assisting in generation of one or more novel molecules or biological systems. In some embodiments, the method may utilize a machine learning / data training model 1201, a generative / computational model 1212, and a computational validation model 1216. In some embodiments, the machine learning / data training model 1201 may comprise an LLM architecture model 1206, a protein model or database 1207, an input raw molecular sequences model or database 1209, a cloud platform 1211, a data preprocessing model 1210, a train and tune model 1208, a model output feedback for machine leaming / training model 1224, a known molecules model for ML / training model 1225, a gene target DNA model for ML / training 1226, and a machine learning of other proteins model 1202 comprising an input raw protein sequences model or database 1205, an evaluative and classification of other proteins model 1203, and an output data protein sequences model or database 1204 with criteria.
[0217] In some embodiments, the generative / computational model 1212 may comprise atained LLM model 1214 and a trained structural model 1213. In some embodiments, the trained LLM model may be a custom trained LLM model 1214. In some embodiments, the trained structural model may be a custom trained structural model 1213. In some embodiments, the LLM model 1214 may comprise a GPT-1, GPT-2, GPT-3, GPT-4, GPT-5, GPT-6, GPT-7, GPT-8, GPT-9,GPT-10, or another GPT model. In some embodiments, the LLM model 1214 may comprise a protein language model. In some embodiments, the LLM model 1214 may comprise a protein language model comprising one or more of ESM-1, ESM-lv, ESM-2, ESM-2v, ESM-3, ESM-3v, ESM-4, ESM-4v, ESM-5, ESM-5v, ESMFold, ProGen, proteinBERT, another protein language model, or any combination thereof. In some embodiments, the LLM model 1214 may comprise an antibody-specific protein language model. In some embodiments, the LLM model 1214 may comprise an antibody-specific protein language model comprising one or more of IgLM, AbLang, AbLang2, AbLang3, AbLang4, AbLang5, AbLang6, AbLang7, AbLang8, AbLang9, AbLang 10, AntiBERTa, Bio-inspired Antibody Language Model (BALM), BALM-paired and unpaired, O AS-trained RoBERTa, IgBert, IgT5, AntiBERTy, Sapiens, FAbConSapiens, or another antibody-specific protein language model, or any combination thereof.
[0218] In some embodiments, the computational validation model 1216 may comprise a generated molecular sequences model or database 1221, a generated molecular structure model or database 1215, a trained molecular classifier model 1220, one or more molecular folding evaluation tools 1218, one or more structural evaluation tools 1219, a full validation model 1222, and a molecular simulation model 1217.
[0219] In some embodiments, the method may comprise providing an input to a large language learning model (LLM) 1214, and using the LLM to output one or more sequences of the one or more novel molecules or biological systems. In some embodiments, the method may comprise applying a validation model 1222 to evaluate or predict the effectiveness of the output of the LLM model 1214 using one or more predictive assessment tools. In some embodiments, the method may comprise outputting a sequence of the one or more novel molecules or biological systems, or structural representation thereof. In some embodiments, the output sequence or structural representation may be validated for the one or more predictive assessment tools. In some embodiments, the validation output may comprise the one or more novel molecules or biological systems. In some embodiments, the molecule may be a protein. In some embodiments, the molecule may be an enzyme. In some embodiments, the molecule may be a recombinase. In some embodiments, the molecule may be a structural protein. In some embodiments, the molecule may be a hormonal protein. In some embodiments, the molecule may be a receptor protein. In some embodiments, the molecule may be a transport protein. In some embodiments, the molecule may be a hydrolase. In some embodiments, the molecule may be an oxidoreductase. In some embodiments, the molecule may be a transferase. In some embodiments, the molecule may be a lyase. In some embodiments, the molecule may be anisomerase. In some embodiments, the molecule may be a ligase. In some embodiments, the molecule may be a carbohydrate active enzyme. In some embodiments, the molecule may be a polymerase. In some embodiments, the molecule may be a topoisomerase. In some embodiments, the molecule may be an ATPase. In some embodiments, the molecule may be a phosphatase. In some embodiments, the molecule may be a ubiquitin ligase. In some embodiments, the molecule may be a nucleic acid. In some embodiments, the molecule may be a deoxyribonucleic acid (DNA). In some embodiments, the molecule may be a ribonucleic acid (RNA). In some embodiments, the molecule may be a carbohydrate. In some embodiments, the molecule may be a lipid. In some embodiments, the molecule may be an ion. In some embodiments, the molecule may be a vitamin. In some embodiments, the molecule may be a hormone. In some embodiments, the molecule may be a metabolic intermediate. In some embodiments, the molecule may be a coenzyme and / or a cofactor. In some embodiments, the molecule may be a secondary metabolite. In some embodiments, the molecule may be a nuclease. In some embodiments, the molecule may be an endonuclease. In some embodiments, the molecule may be an exonuclease. In some embodiments, the molecule may be a viral nuclease. In some embodiments, the molecule may be an RNA-guided endonuclease. In some embodiments, the molecule may be a component of a Clustered Regularly Interspaced Short Palindromic Repeats and CRISPR-associated proteins (CRISPR-Cas) system. In some embodiments, the molecule may be a meganuclease. In some embodiments, the molecule may be a transcription activator-like effector nuclease. In some embodiments, the molecule may be a zinc finger nuclease. In some embodiments, the biological systems may comprise DNA, RNA, ions, lipids, nanoparticles, small molecules, or any combination thereof.
[0220] In some embodiments, the input may comprise one or more of raw nuclease sequences, structural data, simulation data, quantum calculation data, or any combination thereof. In some embodiments, the input may comprise protein data 1207. In some embodiments, the input may comprise, for example, UniProt data 1207. In some embodiments, the input may comprise information related to one or more phenotypic characteristics. In some embodiments, the input data may be received by an input raw molecular sequences model or database 1209. In some embodiments, the input raw molecular sequences model or database 1209 may output raw molecular sequence data through a cloud server 1211 to a data preprocessing model 1210. In some embodiments, the molecular sequence data may be preprocessed and output by the data preprocessing model 1210 to the train and tune model 1208. In some embodiments, the input may be pre-processed to be configured as LLM input. In some embodiments, the train and tune model 1208 may receive input from an LLM architecture model 1206. In some embodiments, the LLM architecture model 1206 may be an open source LLM architecture model. In someembodiments, the LLM architecture model 1206 may be an untrained, pretrained, or custom LLM architecture model. In some embodiments, the LLM architecture model 1206 may be an untrained, pretrained, or custom open source LLM architecture model. In some embodiments, the train and tune model 1208 may receive input from a data preprocessing model 1210, an LLM architecture model 1214, a model output feedback for ML / Training model 1224, an output data protein sequences model or database 1204 with criteria, and an LLM model 1214, a structural model 1213, or both. In some embodiments, the method may further comprise providing input from the train and tune model 1208 to an LLM model 1214. In some embodiments, the LLM model may be a trained LLM model 1214. In some embodiments, the LLM model may be a custom trained LLM model 1214.
[0221] In some embodiments, the method may further comprise providing input from the train and tune model 1208 to a structural generation simulation module (SGSM) 1213. In some embodiments, the SGSM may be a trained SGSM 1213. In some embodiments, the SGSM may be a custom trained SGSM 1213.
[0222] In some embodiments, the method may further comprise using the LLM 1214 to output one or more sequences of the one or more novel molecules or biological systems. In some embodiments, the method may further comprise using the SGSM 1213 to output one or more structures of the one or more novel molecules or biological systems. In some embodiments, the LLM model 1214 may receive feedback input. In some embodiments, the feedback input may comprise the one or more structures output by the SGSM 1213. In some embodiments, the feedback input may modify the LLM model 1214 output. In some embodiments, the modified LLM model 1214 output may comprise the one or more sequences of the one or more novel molecules or biological systems.
[0223] In some embodiments, the SGSM 1213 may receive as feedback input the one or more sequences output by the LLM model 1214. In some embodiments, the SGSM 1213 may generate output comprising modified output. In some embodiments, the output of the SGSM 1213 is modified based on feedback input from the LLM model 1214. In some embodiments, the modified SGSM 1213 output may comprise the one or more structures of the one or more novel molecules or biological systems.
[0224] In some embodiments, the method may further comprise applying a model output feedback for ML / training model 1224 to evaluate and classify the novel output of the LLM model 1214, the novel output of the SGSM 1213, or both into one or more groups. In some embodiments, the one or more groups categorize the one or more sequences of the one or more novel molecules or biological systems output by the LLM model 1214, the SGSM 1213, or both.In some embodiments, the model output feedback for ML / training model 1224 may receive input from a machine learning of other proteins model 1202. In some embodiments, the machine learning of other proteins model 1202 may receive input from the model output feedback for ML / training model 1224. In some embodiments, the model output feedback for ML / training model 1224 may receive input from a known molecules model for ML / training 1225. In some embodiments, the machine learning of other proteins model 1202 may receive input from the known molecules model for ML / training model 1225. In some embodiments, the known molecules model for ML / training model 1225 may receive input from the machine learning of other proteins model 1202. In some embodiments, the known molecules model for ML / training model 1225 may receive input from the model output feedback for ML / training model 1224. In some embodiments, the model output feedback for ML / training model 1224, the known molecules model for ML / training model 1225, or the machine learning of other proteins model 1202, or any combination thereof, may receive input from a gene target DNA model for ML / training model 1226. In some embodiments, a gene target DNA model for ML / training model 1226 may receive input from the model output feedback for ML / training model 1224, the known molecules model for ML / training model 1225, or the machine learning of other proteins model 1202, or any combination thereof.
[0225] In some embodiments, the machine learning of other proteins model 1202 may comprise an input raw protein sequences model or database 1205. In some embodiments, the input raw protein sequences model or database 1205 may output data to an evaluative and classifying other proteins model 1203. In some embodiments, the evaluative and classifying other proteins model 1203 may comprise one or more algorithms configured to execute one or more methods comprising classifying and comparing one or more molecules from the input raw protein sequences model or database 1205. In some embodiments, the evaluative and classifying other proteins model 1203 may comprise one or more algorithms configured to execute one or more methods comprising classifying and comparing data from the input raw protein sequences model or database 1205. In some embodiments, the evaluative and classifying other proteins model 1203 may comprise one or more algorithms configured to execute one or more methods comprising identifying and subcategorizing other proteins data from the input raw protein sequences model or database 1205. In some embodiments, the evaluative and classifying other proteins model 1203 may output data to an output data protein sequences model or database 1204 with criteria. In some embodiments, the evaluative and classifying other proteins model 1203 may receive data comprising one or more structural features of one or more gene editors, other molecules, or biological materials. In some cases, the evaluative and classifying other proteins model 1203 may receive training data comprising the one or more structural features. Insome cases, the data comprising one or more structural features may comprise data concerning one or more subparts, one or more substructures, one or more domains, one or more size characteristics, one or more binding sites, or any combination thereof. In some embodiments, the output data protein sequences model or database 1204 with criteria outputs data to the train and tune model 1208.
[0226] In some embodiments, the LLM model 1214, or the SGSM 1213, or both, may output data to a computational validation model 1216. In some embodiments, the LLM model 1214 may output data to a generated molecular sequence model or database 1221. In some embodiments, the generated molecular sequence model or database 1221 may output one or more molecular sequences from the LLM model 1214. In some embodiments, the generated molecular sequence model or database 1221 may output data to a trained molecular classifier model 1220. In some embodiments, the trained molecular classifier model 1220 may classify one or more sequences input from the generated molecular sequences model or database 1221. In some embodiments, the classification may be performed based on one or more characteristics of each molecule. In some embodiments, the trained molecular classifier model 1220 may output data to one or more predictive assessment tools. In some embodiments, the one or more predictive assessment tools may comprise folding evaluation, structural evaluation, molecular simulation, molecular dynamics simulation, or any combination thereof. In some embodiments, the trained molecular classifier model 1220 may output data to one or more molecular folding evaluation tools 1218. In some embodiments, the molecular folding evaluation tools 1218 may be third-party evaluation tools. In some embodiments, the molecular folding evaluation tools 1218 may be one or more of AlphaFold, Rosetta, or DeepMind Equivariant Neural Networks, or any combination thereof. In some embodiments, the molecular folding evaluation tools 1218 may be customized, bespoke, in-house generated, or otherwise independently created molecular folding evaluation tools. In some embodiments, the one or more molecular folding evaluation tools 1218 may output data to a full validation model 1222. In some embodiments, the SGSM 1213 may output data to a generated molecular structure model or database 1215. In some embodiments, the generated molecular structure model or database 1215 may output data to one or more structural evaluation tools 1219. In some embodiments, the one or more structural evaluation tools 1219 may output data to a full validation model 1222. In some embodiments, the full validation model 1222 may output validated data to a molecular simulation model 1217, or receive data from a molecular simulation model 1217, or both. In some embodiments, the full validation model 1222 may output validation data to the model output feedback for ML / training model 1224. In some embodiments, the molecular simulation model 1217 may output data to anovel molecular library 1223. In some embodiments, the novel molecular library 1223 may output one or more novel molecules or biologic materials to a user.
[0227] In yet another non-limiting embodiment of the invention, as illustrated in FIG. 13, the invention may comprise a system or method 1300 for assisting in generation of one or more novel molecules or biological systems. In some embodiments, the invention may comprise a method for assisting in generation of one or more novel molecules or biological systems. In some embodiments, the method may utilize a machine learning / data training model 1301, a generative / computational model 1311, and a computational validation model 1316. In some embodiments, the machine learning / data training model 1301 may comprise an LLM architecture model 1306, a protein model or database 1307, an input raw molecular sequences model or database 1309, a cloud platform 1310, a data preprocessing model 1325, a train and tune model 1308, a model output feedback for machine leaming / training model 1326, a known molecules model for ML / training model 1327, a gene target DNA model for ML / training 1328, an other protein model for ML / training model 1329, and a machine learning of current solutions model 1302 comprising an input current successes / failures model or database 1305, an evaluative and classification of current solutions model 1303, and an output data current solutions model or database 1304 with learned knowledge. In some embodiments, the evaluative and classification of current solutions model 1303 may identify one or more characteristics linked to successes of one or more biological materials, in some cases gene editors or other molecules. In some embodiments, the learned knowledge may comprise one or more features linked to successes in the success / failures model or database 1305. In some embodiments, the learned knowledge linked to successes may comprise, for example, one or more of binding efficacy, off-target editing, immunogenicity, specificity, cleavage efficacy, or any combination thereof.
[0228] In some embodiments, the generative / computational model 1311 may comprise a trained LLM model 1312 and a trained structural model 1314. In some embodiments, the trained LLM model may be a custom trained LLM model 1312. In some embodiments, the trained structural model may be a custom trained structural model 1314. In some embodiments, the LLM model 1312 may comprise a GPT-1, GPT-2, GPT-3, GPT-4, GPT-5, GPT-6, GPT-7, GPT-8, GPT-9, GPT-10, or another GPT model. In some embodiments, the LLM model 1312 may comprise a protein language model. In some embodiments, the LLM model 1312 may comprise a protein language model comprising one or more of ESM-1, ESM-lv, ESM-2, ESM-2v, ESM-3, ESM-3v, ESM-4, ESM-4v, ESM-5, ESM-5v, ESMFold, ProGen, proteinBERT, another protein language model, or any combination thereof. In some embodiments, the LLM model 1312 maycomprise an antibody-specific protein language model. In some embodiments, the LLM model 1312 may comprise an antibody-specific protein language model comprising one or more of IgLM, AbLang, AbLang2, AbLang3, AbLang4, AbLang5, AbLang6, AbLang7, AbLang8, AbLang9, AbLang 10, AntiBERTa, Bio-inspired Antibody Language Model (BALM), BALM-paired and unpaired, O AS-trained RoBERTa, IgBert, IgT5, AntiBERTy, Sapiens, FAbConSapiens, or another antibody-specific protein language model, or any combination thereof.
[0229] In some embodiments, the computational validation model 1316 may comprise a generated molecular sequences model or database 1322, a generated molecular structure model or database 1321, a trained molecular classifier model 1320, one or more molecular folding evaluation tools 1318, one or more structural evaluation tools 1319, a full validation model 1323, and a molecular simulation model 1317.
[0230] In some embodiments, the method may comprise providing an input to a large language learning model (LLM) 1312, and using the LLM to output one or more sequences of the one or more novel molecules or biological systems. In some embodiments, the method may comprise applying a validation model 1323 to evaluate or predict the effectiveness of the output of the LLM model 1312 using one or more predictive assessment tools. In some embodiments, the method may comprise outputting a sequence of the one or more novel molecules or biological systems, or structural representation thereof. In some embodiments, the output sequence or structural representation may be validated for the one or more predictive assessment tools. In some embodiments, the validation output may comprise the one or more novel molecules or biological systems. In some embodiments, the molecule may be a protein. In some embodiments, the molecule may be an enzyme. In some embodiments, the molecule may be a recombinase. In some embodiments, the molecule may be a structural protein. In some embodiments, the molecule may be a hormonal protein. In some embodiments, the molecule may be a receptor protein. In some embodiments, the molecule may be a transport protein. In some embodiments, the molecule may be a hydrolase. In some embodiments, the molecule may be an oxidoreductase. In some embodiments, the molecule may be a transferase. In some embodiments, the molecule may be a lyase. In some embodiments, the molecule may be an isomerase. In some embodiments, the molecule may be a ligase. In some embodiments, the molecule may be a carbohydrate active enzyme. In some embodiments, the molecule may be a polymerase. In some embodiments, the molecule may be a topoisomerase. In some embodiments, the molecule may be an ATPase. In some embodiments, the molecule may be a phosphatase. In some embodiments, the molecule may be a ubiquitin ligase. In someembodiments, the molecule may be a nucleic acid. In some embodiments, the molecule may be a deoxyribonucleic acid (DNA). In some embodiments, the molecule may be a ribonucleic acid (RNA). In some embodiments, the molecule may be a carbohydrate. In some embodiments, the molecule may be a lipid. In some embodiments, the molecule may be an ion. In some embodiments, the molecule may be a vitamin. In some embodiments, the molecule may be a hormone. In some embodiments, the molecule may be a metabolic intermediate. In some embodiments, the molecule may be a coenzyme and / or a cofactor. In some embodiments, the molecule may be a secondary metabolite. In some embodiments, the molecule may be a nuclease. In some embodiments, the molecule may be an endonuclease. In some embodiments, the molecule may be an exonuclease. In some embodiments, the molecule may be a viral nuclease. In some embodiments, the molecule may be an RNA-guided endonuclease. In some embodiments, the molecule may be a component of a Clustered Regularly Interspaced Short Palindromic Repeats and CRISPR-associated proteins (CRISPR-Cas) system. In some embodiments, the molecule may be a meganuclease. In some embodiments, the molecule may be a transcription activator-like effector nuclease. In some embodiments, the molecule may be a zinc finger nuclease. In some embodiments, the biological systems may comprise DNA, RNA, ions, lipids, nanoparticles, small molecules, or any combination thereof.
[0231] In some embodiments, the input may comprise one or more of raw nuclease sequences, structural data, simulation data, quantum calculation data, or any combination thereof. In some embodiments, the input may comprise protein data 1307. In some embodiments, the input may comprise, for example, UniProt data 1307. In some embodiments, the input may comprise information related to one or more phenotypic characteristics. In some embodiments, the input data may be received by an input raw molecular sequences model or database 1309. In some embodiments, the input raw molecular sequences model or database 1309 may output raw molecular sequence data through a cloud server 1310 to a data preprocessing model 1325. In some embodiments, the molecular sequence data may be preprocessed and output by the data preprocessing model 1325 to the train and tune model 1308. In some embodiments, the input may be pre-processed to be configured as LLM input. In some embodiments, the train and tune model 1308 may receive input from an LLM architecture model 1306. In some embodiments, the LLM architecture model 1306 may be an open source LLM architecture model. In some embodiments, the LLM architecture model 1306 may be an untrained, pretrained, or custom LLM architecture model. In some embodiments, the LLM architecture model 1306 may be an untrained, pretrained, or custom open source LLM architecture model. In some embodiments, the train and tune model 1308 may receive input from a data preprocessing model 1325, an LLM architecture model 1306, a model output feedback for ML / Training model 1326, an output datacurrent solutions model or database 1204 with learned knowledge, and an LLM model 1312, a structural model 1314, or both. In some embodiments, the method may further comprise providing input from the train and tune model 1308 to an LLM model 1312. In some embodiments, the LLM model may be a trained LLM model 1312. In some embodiments, the LLM model may be a custom trained LLM model 1312.
[0232] In some embodiments, the method may further comprise providing input from the train and tune model 1308 to a structural generation simulation module (SGSM) 1314. In some embodiments, the SGSM may be a trained SGSM 1314. In some embodiments, the SGSM may be a custom trained SGSM 1314.
[0233] In some embodiments, the method may further comprise using the LLM 1312 to output one or more sequences of the one or more novel molecules or biological systems. In some embodiments, the method may further comprise using the SGSM 1314 to output one or more structures of the one or more novel molecules or biological systems. In some embodiments, the LLM model 1312 may receive feedback input. In some embodiments, the feedback input may comprise the one or more structures output by the SGSM 1314. In some embodiments, the feedback input may modify the LLM model 1312 output. In some embodiments, the modified LLM model 1312 output may comprise the one or more sequences of the one or more novel molecules or biological systems.
[0234] In some embodiments, the SGSM 1314 may receive as feedback input the one or more sequences output by the LLM model 1312. In some embodiments, the SGSM 1314 may generate output comprising modified output. In some embodiments, the output of the SGSM 1314 is modified based on feedback input from the LLM model 1312. In some embodiments, the modified SGSM 1314 output may comprise the one or more structures of the one or more novel molecules or biological systems.
[0235] In some embodiments, the method may further comprise applying a model output feedback for ML / training model 1326 to evaluate and classify the novel output of the LLM model 1312, the novel output of the SGSM 1314, or both into one or more groups. In some embodiments, the one or more groups categorize the one or more sequences of the one or more novel molecules or biological systems output by the LLM model 1312, the SGSM 1314, or both. In some embodiments, the model output feedback for ML / training model 1312 may receive input from a machine learning of current solutions model 1302. In some embodiments, the machine learning of current solutions model 1302 may receive input from the model output feedback for ML / training model 1326. In some embodiments, the model output feedback for ML / training model 1326 may receive input from a known molecules model for ML / trainingmodel 1327. In some embodiments, the machine learning of current solutions model 1302 may receive input from the known molecules model for ML / training model 1327. In some embodiments, the known molecules model for ML / training model 1327 may receive input from the machine learning of current solutions model 1302. In some embodiments, the known molecules model for ML / training model 1327 may receive input from the model output feedback for ML / training model 1326. In some embodiments, the model output feedback for ML / training model 1326, the known molecules model for ML / training model 1327, or the machine learning of current solutions model 1302, or any combination thereof, may receive input from a gene target DNA model for ML / training model 1328. In some embodiments, a gene target DNA model for ML / training model 1328 may receive input from the model output feedback for ML / training model 1326, the known molecules model for ML / training model 1327, or the machine learning of current solutions model 1302, or any combination thereof. In some embodiments, an other protein model for ML / training model 1329 may receive input from one or more of the model output feedback for ML / training model 1326, the known molecules model for ML / training model 1327, the gene target DNA model for ML / training model 1328, or the machine learning of current solutions model 1302, or any combination thereof. In some embodiments, an other protein model for ML / training model 1329 may output data to one or more of the model output feedback for ML / training model 1326, the known molecules model for ML / training model 1327, the gene target DNA model for ML / training model 1328, or the machine learning of current solutions model 1302, or any combination thereof.
[0236] In some embodiments, the machine learning of current solutions model 1302 may comprise an input current successes / failures model or database 1305. In some embodiments, the input current successes / failures model or database 1305 may output data to an evaluative and classifying current solutions model 1303. In some embodiments, the evaluative and classifying current solutions model 1303 may comprise one or more algorithms configured to execute one or more methods comprising classifying and comparing one or more current solutions or external data from the input current successes / failures model or database 1305. In some cases, the one or more current solutions may comprise a successful novel molecule resulting from a previous input to the system or method. In some cases, the one or more current solutions may be processed from input data. In some embodiments, the evaluative and classifying current solutions model 1303 may comprise one or more algorithms configured to execute one or more methods comprising classifying and comparing one or more current solutions from input data originating from one or more external sources. In some cases, the external sources may comprise documents, such as one or more scientific research publications. In some embodiments, the evaluative and classifying current solutions model 1303 may comprise one or more algorithmsconfigured to execute one or more methods comprising classifying and comparing one or more current solutions from input data originating from one or more external sources comprising biological materials such as molecules, with one or more desirable cleavage, efficacy, and efficiency characteristics. In some embodiments, the input data may be used for training the classifying current solutions model 1303 or another model of the system. In some embodiments, the evaluative and classifying current solutions model 1303 may comprise one or more algorithms configured to execute one or more methods comprising classifying and comparing one or more current solutions to one or more previously generated biological materials such as gene editors or other molecules. In some cases, the previously generated biological materials such as gene editors or other molecules may be classified as successes. In some embodiments, the evaluative and classifying current solutions model 1303 may comprise one or more algorithms configured to execute one or more methods comprising identifying and subcategorizing data from the input current successes / failures model or database 1305. In some embodiments, the evaluative and classifying current solutions model 1303 may comprise one or more algorithms configured to execute one or more methods comprising quantifying successes from the input current successes / failures model or database 1305. In some embodiments, the evaluative and classifying current solutions model 1303 may comprise one or more algorithms configured to execute one or more methods comprising quantifying one or more desirable characteristics originating from one or more external sources. In some cases, the one or more desirable characteristics originating from one or more external sources may comprise data originating from one or more external sources comprising biological materials such as molecules, with one or more desirable cleavage, efficacy, and efficiency characteristics, or other characteristics. In some embodiments, the one or more desirable characteristics may comprise characteristics of one or more previously generated biological materials such as gene editors or other molecules. In some cases, the previously generated biological materials such as gene editors or other molecules may be classified as successes. In some embodiments, the evaluative and classifying current solutions model 1303 may comprise one or more algorithms configured to execute one or more methods comprising quantifying failures from the input current successes / failures model or database 1305. In some embodiments, the evaluative and classifying current solutions model 1303 may comprise one or more algorithms configured to execute one or more methods comprising learning key molecular features from the input current successes / failures model or database 1305. In some embodiments, the evaluative and classifying current solutions model 1303 may comprise one or more algorithms configured to execute one or more methods comprising comparing successful features of the input current successes / failures model or database 1305. In some embodiments, the evaluative and classifyingcurrent solutions model 1303 may comprise one or more algorithms configured to execute one or more methods comprising analyzing failures from the input current successes / failures model or database 1305. In some embodiments, the evaluative and classifying current solutions model 1303 may output data to an output data current solutions model or database 1304 with learned knowledge. In some embodiments, the learned knowledge may comprise input data form one or more external sources. . In some cases, the external sources may comprise documents, such as one or more scientific research publications. In some embodiments, the learned knowledge input data may comprise data comprising one or more values linked to one or more desirable characteristics, for example one or more desirable cleavage, efficacy, and efficiency characteristics, or other characteristics. In some embodiments, learned knowledge may comprise one or more previously generated biological materials such as gene editors or other molecules. In some cases, the previously generated biological materials such as gene editors or other molecules may be classified as successes. In some embodiments, the output data current solutions model or database 1304 with learned knowledge may output data to the train and tune model 1308.
[0237] In some embodiments, the LLM model 1312, or the SGSM 1314, or both, may output data to a computational validation model 1316. In some embodiments, the LLM model 1312 may output data to a generated molecular sequence model or database 1322. In some embodiments, the generated molecular sequence model or database 1322 may output one or more molecular sequences from the LLM model 1312. In some embodiments, the generated molecular sequence model or database 1322 may output data to a trained molecular classifier model 1320. In some embodiments, the trained molecular classifier model 1320 may classify one or more sequences input from the generated molecular sequences model or database 1322.In some embodiments, the classification may be performed based on one or more characteristics of each molecule. In some embodiments, the trained molecular classifier model 1320 may output data to one or more predictive assessment tools. In some embodiments, the one or more predictive assessment tools may comprise folding evaluation, structural evaluation, molecular simulation, molecular dynamics simulation, or any combination thereof. In some embodiments, the trained molecular classifier model 1320 may output data to one or more molecular folding evaluation tools 1318. In some embodiments, the molecular folding evaluation tools 1318 may be third-party evaluation tools. In some embodiments, the molecular folding evaluation tools 1318 may be one or more of AlphaFold, Rosetta, or DeepMind Equivariant Neural Networks, or any combination thereof. In some embodiments, the molecular folding evaluation tools 1318 may be customized, bespoke, in-house generated, or otherwise independently created molecular folding evaluation tools. In some embodiments, the one or more molecular folding evaluationtools 1318 may output data to a full validation model 1323. In some embodiments, the SGSM 1314 may output data to a generated molecular structure model or database 1321. In some embodiments, the generated molecular structure model or database 1321 may output data to one or more structural evaluation tools 1319. In some embodiments, the one or more structural evaluation tools 1319 may output data to a full validation model 1323. In some embodiments, the full validation model 1323 may output validated data to a molecular simulation model 1317, or receive data from a molecular simulation model 1317, or both. In some embodiments, the full validation model 1323 may output validation data to the model output feedback for ML / training model 1326. In some embodiments, the molecular simulation model 1317 may output data to a novel molecular library 1324. In some embodiments, the novel molecular library 1324 may output one or more novel molecules or biologic materials to a user.
[0238] In yet another non-limiting embodiment of the invention, as illustrated in FIG. 14, the invention may comprise a system or method 1400 for assisting in generation of one or more novel molecules or biological systems. In some embodiments, the invention may comprise a method for assisting in generation of one or more novel molecules or biological systems. In some embodiments, the method may utilize a machine learning / data training model 1401, a generative / computational model 1413, and a computational validation model 1416. In some embodiments, the machine learning / data training model 1401 may comprise an LLM architecture model 1406, a protein model or database 1407, an input raw molecular sequences model or database 1409, a cloud platform 1411, a data preprocessing model 1410, a train and tune model 1408, a model output feedback for machine leaming / training model 1425, a known molecules model for ML / training model 1426, a gene target DNA model for ML / training 1427, an other protein model for ML / training model 1428, a current solution model for ML / training 1429, and a machine learning of delivery solutions model 1402 comprising an input delivery knowledge model or database 1405, a delivery learning model 1403, and an output data delivery knowledge model or database 1404.
[0239] In some embodiments, the generative / computational model 1413 may comprise a trained LLM model 1412 and a trained structural model 1414. In some embodiments, the trained LLM model may be a custom trained LLM model 1412. In some embodiments, the trained structural model may be a custom trained structural model 1414. In some embodiments, the LLM model 1412 may comprise a GPT-1, GPT-2, GPT-3, GPT-4, GPT-5, GPT-6, GPT-7, GPT-8, GPT-9, GPT-10, or another GPT model. In some embodiments, the LLM model 1412 may comprise a protein language model. In some embodiments, the LLM model 1412 may comprise a protein language model comprising one or more of ESM-1, ESM-lv, ESM-2, ESM-2v, ESM-3, ESM-3v, ESM-4, ESM-4v, ESM-5, ESM-5v, ESMFold, ProGen, proteinBERT, another protein language model, or any combination thereof. In some embodiments, the LLM model 1412 may comprise an antibody-specific protein language model. In some embodiments, the LLM model 1412 may comprise an antibody-specific protein language model comprising one or more of IgLM, AbLang, AbLang2, AbLang3, AbLang4, AbLang5, AbLang6, AbLang7, AbLang8, AbLang9, AbLang 10, AntiBERTa, Bio-inspired Antibody Language Model (BALM), BALM-paired and unpaired, O AS-trained RoBERTa, IgBert, IgT5, AntiBERTy, Sapiens, FAbConSapiens, or another antibody-specific protein language model, or any combination thereof.
[0240] In some embodiments, the computational validation model 1416 may comprise a generated molecular sequences model or database 1422, a generated molecular structure model or database 1421, a trained molecular classifier model 1420, one or more molecular folding evaluation tools 1418, one or more structural evaluation tools 1419, a full validation model 1423, and a molecular simulation model 1417.
[0241] In some embodiments, the method may comprise providing an input to a large language learning model (LLM) 1412, and using the LLM to output one or more sequences of the one or more novel molecules or biological systems. In some embodiments, the method may comprise applying a validation model 1423 to evaluate or predict the effectiveness of the output of the LLM model 1412 using one or more predictive assessment tools. In some embodiments, the method may comprise outputting a sequence of the one or more novel molecules or biological systems, or structural representation thereof. In some embodiments, the output sequence or structural representation may be validated for the one or more predictive assessment tools. In some embodiments, the validation output may comprise the one or more novel molecules or biological systems. In some embodiments, the molecule may be a protein. In some embodiments, the molecule may be an enzyme. In some embodiments, the molecule may be a recombinase. In some embodiments, the molecule may be a structural protein. In some embodiments, the molecule may be a hormonal protein. In some embodiments, the molecule may be a receptor protein. In some embodiments, the molecule may be a transport protein. In some embodiments, the molecule may be a hydrolase. In some embodiments, the molecule may be an oxidoreductase. In some embodiments, the molecule may be a transferase. In some embodiments, the molecule may be a lyase. In some embodiments, the molecule may be an isomerase. In some embodiments, the molecule may be a ligase. In some embodiments, the molecule may be a carbohydrate active enzyme. In some embodiments, the molecule may be a polymerase. In some embodiments, the molecule may be a topoisomerase. In someembodiments, the molecule may be an ATPase. In some embodiments, the molecule may be a phosphatase. In some embodiments, the molecule may be a ubiquitin ligase. In some embodiments, the molecule may be a nucleic acid. In some embodiments, the molecule may be a deoxyribonucleic acid (DNA). In some embodiments, the molecule may be a ribonucleic acid (RNA). In some embodiments, the molecule may be a carbohydrate. In some embodiments, the molecule may be a lipid. In some embodiments, the molecule may be an ion. In some embodiments, the molecule may be a vitamin. In some embodiments, the molecule may be a hormone. In some embodiments, the molecule may be a metabolic intermediate. In some embodiments, the molecule may be a coenzyme and / or a cofactor. In some embodiments, the molecule may be a secondary metabolite. In some embodiments, the molecule may be a nuclease. In some embodiments, the molecule may be an endonuclease. In some embodiments, the molecule may be an exonuclease. In some embodiments, the molecule may be a viral nuclease. In some embodiments, the molecule may be an RNA-guided endonuclease. In some embodiments, the molecule may be a component of a Clustered Regularly Interspaced Short Palindromic Repeats and CRISPR-associated proteins (CRISPR-Cas) system. In some embodiments, the molecule may be a meganuclease. In some embodiments, the molecule may be a transcription activator-like effector nuclease. In some embodiments, the molecule may be a zinc finger nuclease. In some embodiments, the biological systems may comprise DNA, RNA, ions, lipids, nanoparticles, small molecules, or any combination thereof.
[0242] In some embodiments, the input may comprise one or more of raw nuclease sequences, structural data, simulation data, quantum calculation data, or any combination thereof. In some embodiments, the input may comprise protein data 1407. In some embodiments, the input may comprise, for example, UniProt data 1407. In some embodiments, the input may comprise information related to one or more phenotypic characteristics. In some embodiments, the input data may be received by an input raw molecular sequences model or database 1409. In some embodiments, the input raw molecular sequences model or database 1409 may output raw molecular sequence data through a cloud server 1411 to a data preprocessing model 1410. In some embodiments, the molecular sequence data may be preprocessed and output by the data preprocessing model 1410 to the train and tune model 1408. In some embodiments, the input may be pre-processed to be configured as LLM input. In some embodiments, the train and tune model 1408 may receive input from an LLM architecture model 1406. In some embodiments, the LLM architecture model 1406 may be an open source LLM architecture model. In some embodiments, the LLM architecture model 1406 may be an untrained, pretrained, or custom LLM architecture model. In some embodiments, the LLM architecture model 1406 may be an untrained, pretrained, or custom open source LLM architecture model. In some embodiments,the train and tune model 1408 may receive input from a data preprocessing model 1410, an LLM architecture model 1406, a model output feedback for ML / Training model 1425, an output data delivery knowledge model or database 1204, and an LLM model 1412, a structural model 1414, or both. In some embodiments, the method may further comprise providing input from the train and tune model 1408 to an LLM model 1412. In some embodiments, the LLM model may be a trained LLM model 1412. In some embodiments, the LLM model may be a custom trained LLM model 1412.
[0243] In some embodiments, the method may further comprise providing input from the train and tune model 1408 to a structural generation simulation module (SGSM) 1414. In some embodiments, the SGSM may be a trained SGSM 1414. In some embodiments, the SGSM may be a custom trained SGSM 1414.
[0244] In some embodiments, the method may further comprise using the LLM 1412 to output one or more sequences of the one or more novel molecules or biological systems. In some embodiments, the method may further comprise using the SGSM 1414 to output one or more structures of the one or more novel molecules or biological systems. In some embodiments, the LLM model 1412 may receive feedback input. In some embodiments, the feedback input may comprise the one or more structures output by the SGSM 1414. In some embodiments, the feedback input may modify the LLM model 1412 output. In some embodiments, the modified LLM model 1412 output may comprise the one or more sequences of the one or more novel molecules or biological systems.
[0245] In some embodiments, the SGSM 1414 may receive as feedback input the one or more sequences output by the LLM model 1412. In some embodiments, the SGSM 1414 may generate output comprising modified output. In some embodiments, the output of the SGSM 1414 is modified based on feedback input from the LLM model 1412. In some embodiments, the modified SGSM 1414 output may comprise the one or more structures of the one or more novel molecules or biological systems.
[0246] In some embodiments, the method may further comprise applying a model output feedback for ML / training model 1425 to evaluate and classify the novel output of the LLM model 1412, the novel output of the SGSM 1414, or both into one or more groups. In some embodiments, the one or more groups categorize the one or more sequences of the one or more novel molecules or biological systems output by the LLM model 1412, the SGSM 1414, or both. In some embodiments, the model output feedback for ML / training model 1425 may receive input from a machine learning of delivery solutions model 1402. In some embodiments, the machine learning of delivery solutions model 1402 may receive input from the model outputfeedback for ML / training model 1425. In some embodiments, the model output feedback for ML / training model 1425 may receive input from a known molecules model for ML / training model 1426. In some embodiments, the machine learning of delivery solutions model 1402 may receive input from the known molecules model for ML / training model 1426. In some embodiments, the known molecules model for ML / training model 1426 may receive input from the machine learning of delivery solutions model 1402. In some embodiments, the known molecules model for ML / training model 1426 may receive input from the model output feedback for ML / training model 1425. In some embodiments, the model output feedback for ML / training model 1425, the known molecules model for ML / training model 1426, or the machine learning of delivery solutions model 1402, or any combination thereof, may receive input from a gene target DNA model for ML / training model 1427. In some embodiments, a gene target DNA model for ML / training model 1427 may receive input from the model output feedback for ML / training model 1425, the known molecules model for ML / training model 1426, or the machine learning of delivery solutions model 1402, or any combination thereof. In some embodiments, an other protein model for ML / training model 1428 may receive input from one or more of the model output feedback for ML / training model 1425, the known molecules model for ML / training model 1426, the gene target DNA model for ML / training model 1427, or the machine learning of delivery solutions model 1402, or any combination thereof. In some embodiments, an other protein model for ML / training model 1428 may output data to one or more of the model output feedback for ML / training model 1425, the known molecules model for ML / training model 1426, the gene target DNA model for ML / training model 1427, or the machine learning of delivery solutions model 1402, or any combination thereof. In some embodiments, a current solution model for ML / training model 1429 may receive input from one or more of the model output feedback for ML / training model 1425, the known molecules model for ML / training model 1426, the gene target DNA model for ML / training model 1427, the other protein model for ML / training model 1428, or the machine learning of delivery solutions model 1402, or any combination thereof. In some embodiments, a current solution model for ML / training model 1429 may output data to one or more of the model output feedback for ML / training model 1425, the known molecules model for ML / training model 1426, the gene target DNA model for ML / training model 1427, the other protein model for ML / training model 1428, or the machine learning of delivery solutions model 1402, or any combination thereof.
[0247] In some embodiments, the machine learning of delivery solutions model 1402 may comprise an input delivery knowledge model or database 1405. In some embodiments, the input delivery knowledge model or database 1405 may output data to a delivery learning model 1403.In some embodiments, the delivery learning model 1403 may comprise one or more algorithmsconfigured to execute one or more methods comprising learning to classify and compare one or more delivery methods from the input delivery knowledge model or database 1405. In some embodiments, the delivery learning model 1403 may comprise one or more algorithms configured to execute one or more methods comprising identifying and quantifying successful delivery methods and materials from the input delivery knowledge model or database 1405. In some embodiments, the delivery methods and materials may comprise delivery vehicles, for example lipids, nanoparticles, lipid nanoparticles, viral vectors, protein envelopes, micelles, magnetic nanoparticles, nanoemulsions, lipoplexes, or other delivery vehicles. In some embodiments, the delivery methods and materials may comprise data input from one or more external sources. In some cases, the one or more external sources may comprise one or more documents such as scientific publications. In some embodiments, the delivery learning model 1403 may comprise one or more algorithms configured to execute one or more methods comprising identifying and quantifying failures from the input delivery knowledge model or database 1405. In some embodiments, the failures may comprise one or more of immunogenic reactions, off target effects, lack of efficacy, lack of binding affinity, other undesirable characteristics, or any combination thereof. In some embodiments, the delivery learning model 1403 may output data to an output data delivery knowledge model or database 1404. In some embodiments, the output data delivery knowledge model or database 1404 may output data to the train and tune model 1408. In some embodiments, the delivery method uses a viral vector. In some embodiments, the viral vector is an adeno-associated virus. In some embodiments, the viral vector is a lentivirus. In some embodiments, the delivery method uses a non-viral delivery system. In some embodiments, non-viral delivery system is a nanoparticle, electroporation, a cell-penetrating peptide, or a combination thereof. In some embodiments, the delivery method uses cellular transplantation. In some embodiments, the delivery method uses a CRISPR ribonucleoprotein complex. In some embodiments, the delivery method is injection. In some embodiments, the delivery method uses a lipid. In some embodiments, the delivery method uses a nanoemulsion or microemulsion. In some embodiments, the delivery method uses one or more vesicles. In some embodiments, the delivery method uses one or more lipid vesicles. In some embodiments, the delivery method uses one or more nanoparticles. In some embodiments, the delivery method uses a capsid. In some embodiments, the delivery method uses an AAV capsid. In some embodiments, the delivery method uses a bacteriophage. In some embodiments, the delivery method uses one or more CPMV capsids. In some embodiments, the delivery methods uses one or more CCMV capsids.
[0248] In some embodiments, the LLM model 1412, or the SGSM 1414, or both, may output data to a computational validation model 1416. In some embodiments, the LLM model 1412may output data to a generated molecular sequence model or database 1422. In some embodiments, the generated molecular sequence model or database 1422 may output one or more molecular sequences from the LLM model 1412. In some embodiments, the generated molecular sequence model or database 1422 may output data to a trained molecular classifier model 1420. In some embodiments, the trained molecular classifier model 1420 may classify one or more sequences input from the generated molecular sequences model or database 1422.In some embodiments, the classification may be performed based on one or more characteristics of each molecule. In some embodiments, the trained molecular classifier model 1420 may output data to one or more predictive assessment tools. In some embodiments, the one or more predictive assessment tools may comprise folding evaluation, structural evaluation, molecular simulation, molecular dynamics simulation, or any combination thereof. In some embodiments, the trained molecular classifier model 1420 may output data to one or more molecular folding evaluation tools 1418. In some embodiments, the molecular folding evaluation tools 1418 may be third-party evaluation tools. In some embodiments, the molecular folding evaluation tools 1418 may be one or more of AlphaFold, Rosetta, or DeepMind Equivariant Neural Networks, or any combination thereof. In some embodiments, the molecular folding evaluation tools 1418 may be customized, bespoke, in-house generated, or otherwise independently created molecular folding evaluation tools. In some embodiments, the one or more molecular folding evaluation tools 1318 may output data to a full validation model 1423. In some embodiments, the SGSM 1414 may output data to a generated molecular structure model or database 1421. In some embodiments, the generated molecular structure model or database 1421 may output data to one or more structural evaluation tools 1419. In some embodiments, the one or more structural evaluation tools 1419 may output data to a full validation model 1423. In some embodiments, the full validation model 1423 may output validated data to a molecular simulation model 1417, or receive data from a molecular simulation model 1417, or both. In some embodiments, the full validation model 1423 may output validation data to the model output feedback for ML / training model 1425. In some embodiments, the molecular simulation model 1417 may output data to a novel molecular library 1424. In some embodiments, the novel molecular library 1424 may output one or more novel molecules or biologic materials to a user.
[0249] In yet another non-limiting embodiment of the invention, as illustrated in FIG. 15, the invention may comprise a system or method 1500 for assisting in generation of one or more novel molecules or biological systems. In some embodiments, the invention may comprise a method for assisting in generation of one or more novel molecules or biological systems. In some embodiments, the method may utilize a machine learning / data training model 1501, a generative / computational model 1513, and a computational validation model 1516. In someembodiments, the machine leaming / data training model 1501 may comprise an LLM architecture model 1506, a protein model or database 1507, an input raw molecular sequences model or database 1509, a cloud platform 1511, a data preprocessing model 1510, a train and tune model 1508, a model output feedback for machine leaming / training model 1525, a known molecules model for ML / training model 1526, a gene target DNA model for ML / training 1527, an other protein model for ML / training model 1528, a current solution model for ML / training 1529, a delivery solution model for ML / training model 1530 and a machine learning of image structure model 1502 comprising an input raw structure model or database 1505, a structure learning model 1503, and an output data structure knowledge model or database 1504.
[0250] In some embodiments, the generative / computational model 1513 may comprise a trained LLM model 1512 and a trained structural model 1514. In some embodiments, the trained LLM model may be a custom trained LLM model 1512. In some embodiments, the trained structural model may be a custom trained structural model 1514. In some embodiments, the LLM model 1512 may comprise a GPT-1, GPT-2, GPT-3, GPT-4, GPT-5, GPT-6, GPT-7, GPT-8, GPT-9, GPT-10, or another GPT model. In some embodiments, the LLM model 1512 may comprise a protein language model. In some embodiments, the LLM model 1512 may comprise a protein language model comprising one or more of ESM-1, ESM-lv, ESM-2, ESM-2v, ESM-3, ESM-3v, ESM-4, ESM-4v, ESM-5, ESM-5v, ESMFold, ProGen, proteinBERT, another protein language model, or any combination thereof. In some embodiments, the LLM model 1512 may comprise an antibody-specific protein language model. In some embodiments, the LLM model 1512 may comprise an antibody-specific protein language model comprising one or more of IgLM, AbLang, AbLang2, AbLang3, AbLang4, AbLang5, AbLang6, AbLang7, AbLang8, AbLang9, AbLang 10, AntiBERTa, Bio-inspired Antibody Language Model (BALM), BALM-paired and unpaired, O AS-trained RoBERTa, IgBert, IgT5, AntiBERTy, Sapiens, FAbConSapiens, or another antibody-specific protein language model, or any combination thereof.
[0251] In some embodiments, the computational validation model 1516 may comprise a generated molecular sequences model or database 1522, a generated molecular structure model or database 1521, a trained molecular classifier model 1520, one or more molecular folding evaluation tools 1518, one or more structural evaluation tools 1519, a full validation model 1523, and a molecular simulation model 1517.
[0252] In some embodiments, the method may comprise providing an input to a large language learning model (LLM) 1512, and using the LLM to output one or more sequences of the one or more novel molecules or biological systems. In some embodiments, the method may compriseapplying a validation model 1523 to evaluate or predict the effectiveness of the output of the LLM model 1512 using one or more predictive assessment tools. In some embodiments, the method may comprise outputting a sequence of the one or more novel molecules or biological systems, or structural representation thereof. In some embodiments, the output sequence or structural representation may be validated for the one or more predictive assessment tools. In some embodiments, the validation output may comprise the one or more novel molecules or biological systems. In some embodiments, the molecule may be a protein. In some embodiments, the molecule may be an enzyme. In some embodiments, the molecule may be a recombinase. In some embodiments, the molecule may be a structural protein. In some embodiments, the molecule may be a hormonal protein. In some embodiments, the molecule may be a receptor protein. In some embodiments, the molecule may be a transport protein. In some embodiments, the molecule may be a hydrolase. In some embodiments, the molecule may be an oxidoreductase. In some embodiments, the molecule may be a transferase. In some embodiments, the molecule may be a lyase. In some embodiments, the molecule may be an isomerase. In some embodiments, the molecule may be a ligase. In some embodiments, the molecule may be a carbohydrate active enzyme. In some embodiments, the molecule may be a polymerase. In some embodiments, the molecule may be a topoisomerase. In some embodiments, the molecule may be an ATPase. In some embodiments, the molecule may be a phosphatase. In some embodiments, the molecule may be a ubiquitin ligase. In some embodiments, the molecule may be a nucleic acid. In some embodiments, the molecule may be a deoxyribonucleic acid (DNA). In some embodiments, the molecule may be a ribonucleic acid (RNA). In some embodiments, the molecule may be a carbohydrate. In some embodiments, the molecule may be a lipid. In some embodiments, the molecule may be an ion. In some embodiments, the molecule may be a vitamin. In some embodiments, the molecule may be a hormone. In some embodiments, the molecule may be a metabolic intermediate. In some embodiments, the molecule may be a coenzyme and / or a cofactor. In some embodiments, the molecule may be a secondary metabolite. In some embodiments, the molecule may be a nuclease. In some embodiments, the molecule may be an endonuclease. In some embodiments, the molecule may be an exonuclease. In some embodiments, the molecule may be a viral nuclease. In some embodiments, the molecule may be an RNA-guided endonuclease. In some embodiments, the molecule may be a component of a Clustered Regularly Interspaced Short Palindromic Repeats and CRISPR-associated proteins (CRISPR-Cas) system. In some embodiments, the molecule may be a meganuclease. In some embodiments, the molecule may be a transcription activator-like effector nuclease. In some embodiments, the molecule may be azinc finger nuclease. In some embodiments, the biological systems may comprise DNA, RNA, ions, lipids, nanoparticles, small molecules, or any combination thereof.
[0253] In some embodiments, the input may comprise one or more of raw nuclease sequences, structural data, simulation data, quantum calculation data, or any combination thereof. In some embodiments, the input may comprise protein data 1507. In some embodiments, the input may comprise, for example, UniProt data 1507. In some embodiments, the input may comprise information related to one or more phenotypic characteristics. In some embodiments, the input data may be received by an input raw molecular sequences model or database 1509. In some embodiments, the input raw molecular sequences model or database 1509 may output raw molecular sequence data through a cloud server 1511 to a data preprocessing model 1510. In some embodiments, the molecular sequence data may be preprocessed and output by the data preprocessing model 1510 to the train and tune model 1508. In some embodiments, the input may be pre-processed to be configured as LLM input. In some embodiments, the train and tune model 1508 may receive input from an LLM architecture model 1506. In some embodiments, the LLM architecture model 1506 may be an open source LLM architecture model. In some embodiments, the LLM architecture model 1506 may be an untrained, pretrained, or custom LLM architecture model. In some embodiments, the LLM architecture model 1506 may be an untrained, pretrained, or custom open source LLM architecture model. In some embodiments, the train and tune model 1508 may receive input from a data preprocessing model 1510, an LLM architecture model 1506, a model output feedback for ML / Training model 1525, an output data structure knowledge model or database 1504, and an LLM model 1512, a structural model 1514, or both. In some embodiments, the method may further comprise providing input from the train and tune model 1508 to an LLM model 1512. In some embodiments, the LLM model may be a trained LLM model 1512. In some embodiments, the LLM model may be a custom trained LLM model 1512.
[0254] In some embodiments, the method may further comprise providing input from the train and tune model 1508 to a structural generation simulation module (SGSM) 1514. In some embodiments, the SGSM may be a trained SGSM 1514. In some embodiments, the SGSM may be a custom trained SGSM 1514.
[0255] In some embodiments, the method may further comprise using the LLM 1512 to output one or more sequences of the one or more novel molecules or biological systems. In some embodiments, the method may further comprise using the SGSM 1514 to output one or more structures of the one or more novel molecules or biological systems. In some embodiments, the LLM model 1512 may receive feedback input. In some embodiments, the feedback input maycomprise the one or more structures output by the SGSM 1514. In some embodiments, the feedback input may modify the LLM model 1512 output. In some embodiments, the modified LLM model 1512 output may comprise the one or more sequences of the one or more novel molecules or biological systems.
[0256] In some embodiments, the SGSM 1514 may receive as feedback input the one or more sequences output by the LLM model 1512. In some embodiments, the SGSM 1514 may generate output comprising modified output. In some embodiments, the output of the SGSM 1514 is modified based on feedback input from the LLM model 1512. In some embodiments, the modified SGSM 1514 output may comprise the one or more structures of the one or more novel molecules or biological systems.
[0257] In some embodiments, the method may further comprise applying a model output feedback for ML / training model 1525 to evaluate and classify the novel output of the LLM model 1512, the novel output of the SGSM 1514, or both into one or more groups. In some embodiments, the one or more groups categorize the one or more sequences of the one or more novel molecules or biological systems output by the LLM model 1512, the SGSM 1514, or both. In some embodiments, the model output feedback for ML / training model 1525 may receive input from a machine learning of image structure model 1502. In some embodiments, the machine learning of image structure model 1502 may receive input from the model output feedback for ML / training model 1525. In some embodiments, the model output feedback for ML / training model 1525 may receive input from a known molecules model for ML / training model 1526. In some embodiments, the machine learning of image structure model 1502 may receive input from the known molecules model for ML / training model 1526. In some embodiments, the known molecules model for ML / training model 1526 may receive input from the machine learning of image structure model 1502. In some embodiments, the known molecules model for ML / training model 1526 may receive input from the model output feedback for ML / training model 1525. In some embodiments, the model output feedback for ML / training model 1525, the known molecules model for ML / training model 1526, or the machine learning of image structure model 1502, or any combination thereof, may receive input from a gene target DNA model for ML / training model 1527. In some embodiments, a gene target DNA model for ML / training model 1527 may receive input from the model output feedback for ML / training model 1525, the known molecules model for ML / training model 1526, or the machine learning of image structure model 1502, or any combination thereof. In some embodiments, an other protein model for ML / training model 1528 may receive input from one or more of the model output feedback for ML / training model 1525, the known molecules modelfor ML / training model 1526, the gene target DNA model for ML / training model 1527, or the machine learning of image structure model 1502, or any combination thereof. In some embodiments, an other protein model for ML / training model 1528 may output data to one or more of the model output feedback for ML / training model 1525, the known molecules model for ML / training model 1526, the gene target DNA model for ML / training model 1527, or the machine learning of image structure model 1502, or any combination thereof. In some embodiments, a current solution model for ML / training model 1529 may receive input from one or more of the model output feedback for ML / training model 1525, the known molecules model for ML / training model 1526, the gene target DNA model for ML / training model 1527, the other protein model for ML / training model 1528, or the machine learning of image structure model 1502, or any combination thereof. In some embodiments, a current solution model for ML / training model 1529 may output data to one or more of the model output feedback for ML / training model 1525, the known molecules model for ML / training model 1526, the gene target DNA model for ML / training model 1527, the other protein model for ML / training model 1528, or the machine learning of image structure model 1502, or any combination thereof. In some embodiments, a delivery solution model for ML / training model 1530 may receive input from one or more of the model output feedback for ML / training model 1525, the known molecules model for ML / training model 1526, the gene target DNA model for ML / training model 1527, the other protein model for ML / training model 1528, the delivery solution model for ML / training model 1529, or the machine learning of image structure model 1502, or any combination thereof. In some embodiments, a delivery solution model for ML / training model 1530 may output data to one or more of the model output feedback for ML / training model 1525, the known molecules model for ML / training model 1526, the gene target DNA model for ML / training model 1527, the other protein model for ML / training model 1528, the current solution model for ML / training model 1529, or the machine learning of image structure model 1502, or any combination thereof.
[0258] In some embodiments, the machine learning of image structure model 1502 may comprise an input raw structure model or database 1505. In some embodiments, the input raw structure model or database 1505 may output data to a structure learning model 1503. In some embodiments, the structure learning model 1503 may comprise one or more algorithms configured to execute one or more methods comprising learning to classify and compare structures and imaging for one or more molecules from the input raw structure model or database 1505. In some embodiments, the structure learning model 1503 may comprise one or more algorithms configured to execute one or more methods comprising learning structure from sequence of one or more molecules from the input raw structure model or database 1505. Insome embodiments, the input received by the structure learning model 1503 from the input raw structure model or database 1505 may comprise one or more three-dimensional vectors or structures, or both. In some embodiments, the three-dimensional vectors or structures, or both may comprise identifiable shapes. In some embodiments, the structure learning model 1503 may identify the shapes from the input data. In some embodiments, the input data may comprise one or more characteristics of one or more structures. In some cases, the one or more characteristics of one or more input structures may comprise, one or more areas of the one or more structures being tightly bound, compact, expansive, having one or more angles of various degrees, a distance from another point of the structures, other characteristics, or any combination thereof. In some embodiments, the structure learning model 1503 may comprise one or more algorithms configured to execute one or more methods comprising learning sequence from structure of one or more molecules from the input raw structure model or database 1505. In some embodiments, the structure learning model 1503 may output data to an output data structure knowledge model or database 1504. In some embodiments, the output may comprise one or more shape identifications or one or more structural simulations. In some embodiments, the one or more structural simulations may comprise overlaying structures having differing gene targets. In some embodiments, the structural overlay may be used by the structure learning model 1503 to detect an appropriate match between the structural areas for binding or cleaving activities. In some embodiments, the output may comprise images including one or more of protein binding DNA, guide RNA, transcription factors, genetic enhancers, genetic silencers, genetic repressors, genetic activators, microRNAs, promoters, receptors, or any combination thereof. In some embodiments, the output data structure knowledge model or database 1504 may output data to the train and tune model 1508. In some embodiments, the structures and imaging for one or more molecules may be characterized using mass spectrometry imaging, electron microscopy, atomic force microscopy, X-ray crystallography, nuclear magnetic resonance spectroscopy, infrared spectroscopy, X-ray computed tomography imaging, optical imaging, radionucleotide imaging, ultrasound imaging, or any combination thereof.
[0259] In some embodiments, the LLM model 1512, or the SGSM 1514, or both, may output data to a computational validation model 1516. In some embodiments, the LLM model 1512 may output data to a generated molecular sequence model or database 1522. In some embodiments, the generated molecular sequence model or database 1522 may output one or more molecular sequences from the LLM model 1512. In some embodiments, the generated molecular sequence model or database 1522 may output data to a trained molecular classifier model 1520. In some embodiments, the trained molecular classifier model 1520 may classify one or more sequences input from the generated molecular sequences model or database 1522.In some embodiments, the classification may be performed based on one or more characteristics of each molecule. In some embodiments, the trained molecular classifier model 1520 may output data to one or more predictive assessment tools. In some embodiments, the one or more predictive assessment tools may comprise folding evaluation, structural evaluation, molecular simulation, molecular dynamics simulation, or any combination thereof. In some embodiments, the trained molecular classifier model 1520 may output data to one or more molecular folding evaluation tools 1518. In some embodiments, the molecular folding evaluation tools 1518 may be third-party evaluation tools. In some embodiments, the molecular folding evaluation tools 1518 may be one or more of AlphaFold, Rosetta, or DeepMind Equivariant Neural Networks, or any combination thereof. In some embodiments, the molecular folding evaluation tools 1518 may be customized, bespoke, in-house generated, or otherwise independently created molecular folding evaluation tools. In some embodiments, the one or more molecular folding evaluation tools 1518 may output data to a full validation model 1523. In some embodiments, the SGSM 1514 may output data to a generated molecular structure model or database 1521. In some embodiments, the generated molecular structure model or database 1521 may output data to one or more structural evaluation tools 1519. In some embodiments, the one or more structural evaluation tools 1519 may output data to a full validation model 1523. In some embodiments, the full validation model 1523 may output validated data to a molecular simulation model 1517, or receive data from a molecular simulation model 1517, or both. In some embodiments, the full validation model 1523 may output validation data to the model output feedback for ML / training model 1525. In some embodiments, the molecular simulation model 1517 may output data to a novel molecular library 1524. In some embodiments, the novel molecular library 1524 may output one or more novel molecules or biologic materials to a user.
[0260] In yet another non-limiting embodiment of the invention, as illustrated in FIG. 16, the invention may comprise a system or method 1600 for assisting in generation of one or more novel molecules or biological systems. In some embodiments, the invention may comprise a method for assisting in generation of one or more novel molecules or biological systems. In some embodiments, the method may utilize a machine learning / data training model 1601, a generative / computational model 1613, and a computational validation model 1616. In some embodiments, the machine learning / data training model 1601 may comprise an LLM architecture model 1606, a protein model or database 1607, an input raw molecular sequences model or database 1609, a cloud platform 1611, a data preprocessing model 1610, a train and tune model 1608, a model output feedback for machine leaming / training model 1625, a known molecules model for ML / training model 1626, a gene target DNA model for ML / training 1627, an other protein model for ML / training model 1628, a current solution model for ML / training1629, a delivery solution model for ML / training model 1630, a structural model for ML / training model 1631, and a predictive molecular simulation movie generator model 1602 comprising a predictive visual molecular simulations model 1603, and a generated molecular simulation movie 1604
[0261] In some embodiments, the generative / computational model 1613 may comprise a trained LLM model 1612 and a trained structural model 1614. In some embodiments, the trained LLM model may be a custom trained LLM model 1612. In some embodiments, the trained structural model may be a custom trained structural model 1614. In some embodiments, the LLM model 1612 may comprise a GPT-1, GPT-2, GPT-3, GPT-4, GPT-5, GPT-6, GPT-7, GPT-8, GPT-9, GPT-10, or another GPT model. In some embodiments, the LLM model 1612 may comprise a protein language model. In some embodiments, the LLM model 1612 may comprise a protein language model comprising one or more of ESM-1, ESM-lv, ESM-2, ESM-2v, ESM-3, ESM-3v, ESM-4, ESM-4v, ESM-5, ESM-5v, ESMFold, ProGen, proteinBERT, another protein language model, or any combination thereof. In some embodiments, the LLM model 1612 may comprise an antibody-specific protein language model. In some embodiments, the LLM model 1612 may comprise an antibody-specific protein language model comprising one or more of IgLM, AbLang, AbLang2, AbLang3, AbLang4, AbLang5, AbLang6, AbLang7, AbLang8, AbLang9, AbLang 10, AntiBERTa, Bio-inspired Antibody Language Model (BALM), BALM-paired and unpaired, O AS-trained RoBERTa, IgBert, IgT5, AntiBERTy, Sapiens, FAbConSapiens, or another antibody-specific protein language model, or any combination thereof.
[0262] In some embodiments, the computational validation model 1616 may comprise a generated molecular sequences model or database 1622, a generated molecular structure model or database 1621, a trained molecular classifier model 1620, one or more molecular folding evaluation tools 1618, one or more structural evaluation tools 1619, a full validation model 1623, and a molecular simulation model 1617.
[0263] In some embodiments, the method may comprise providing an input to a large language learning model (LLM) 1612, and using the LLM to output one or more sequences of the one or more novel molecules or biological systems. In some embodiments, the method may comprise applying a validation model 1623 to evaluate or predict the effectiveness of the output of the LLM model 1612 using one or more predictive assessment tools. In some embodiments, the method may comprise outputting a sequence of the one or more novel molecules or biological systems, or structural representation thereof. In some embodiments, the output sequence or structural representation may be validated for the one or more predictive assessment tools. Insome embodiments, the validation output may comprise the one or more novel molecules or biological systems. In some embodiments, the molecule may be a protein. In some embodiments, the molecule may be an enzyme. In some embodiments, the molecule may be a recombinase. In some embodiments, the molecule may be a structural protein. In some embodiments, the molecule may be a hormonal protein. In some embodiments, the molecule may be a receptor protein. In some embodiments, the molecule may be a transport protein. In some embodiments, the molecule may be a hydrolase. In some embodiments, the molecule may be an oxidoreductase. In some embodiments, the molecule may be a transferase. In some embodiments, the molecule may be a lyase. In some embodiments, the molecule may be an isomerase. In some embodiments, the molecule may be a ligase. In some embodiments, the molecule may be a carbohydrate active enzyme. In some embodiments, the molecule may be a polymerase. In some embodiments, the molecule may be a topoisomerase. In some embodiments, the molecule may be an ATPase. In some embodiments, the molecule may be a phosphatase. In some embodiments, the molecule may be a ubiquitin ligase. In some embodiments, the molecule may be a nucleic acid. In some embodiments, the molecule may be a deoxyribonucleic acid (DNA). In some embodiments, the molecule may be a ribonucleic acid (RNA). In some embodiments, the molecule may be a carbohydrate. In some embodiments, the molecule may be a lipid. In some embodiments, the molecule may be an ion. In some embodiments, the molecule may be a vitamin. In some embodiments, the molecule may be a hormone. In some embodiments, the molecule may be a metabolic intermediate. In some embodiments, the molecule may be a coenzyme and / or a cofactor. In some embodiments, the molecule may be a secondary metabolite. In some embodiments, the molecule may be a nuclease. In some embodiments, the molecule may be an endonuclease. In some embodiments, the molecule may be an exonuclease. In some embodiments, the molecule may be a viral nuclease. In some embodiments, the molecule may be an RNA-guided endonuclease. In some embodiments, the molecule may be a component of a Clustered Regularly Interspaced Short Palindromic Repeats and CRISPR-associated proteins (CRISPR-Cas) system. In some embodiments, the molecule may be a meganuclease. In some embodiments, the molecule may be a transcription activator-like effector nuclease. In some embodiments, the molecule may be a zinc finger nuclease. In some embodiments, the biological systems may comprise DNA, RNA, ions, lipids, nanoparticles, small molecules, or any combination thereof.
[0264] In some embodiments, the input may comprise one or more of raw nuclease sequences, structural data, simulation data, quantum calculation data, or any combination thereof. In some embodiments, the input may comprise protein data 1607. In some embodiments, the input may comprise, for example, UniProt data 1607. In some embodiments, the input may compriseinformation related to one or more phenotypic characteristics. In some embodiments, the input data may be received by an input raw molecular sequences model or database 1609. In some embodiments, the input raw molecular sequences model or database 1609 may output raw molecular sequence data through a cloud server 1611 to a data preprocessing model 1610. In some embodiments, the molecular sequence data may be preprocessed and output by the data preprocessing model 1610 to the train and tune model 1608. In some embodiments, the input may be pre-processed to be configured as LLM input. In some embodiments, the train and tune model 1608 may receive input from an LLM architecture model 1606. In some embodiments, the LLM architecture model 1606 may be an open source LLM architecture model. In some embodiments, the LLM architecture model 1606 may be an untrained, pretrained, or custom LLM architecture model. In some embodiments, the LLM architecture model 1606 may be an untrained, pretrained, or custom open source LLM architecture model. In some embodiments, the train and tune model 1608 may receive input from a data preprocessing model 1610, an LLM architecture model 1606, a model output feedback for ML / Training model 1625, a predictive molecular simulation movie generator model 1604, and an LLM model 1612, a structural model 1614, or both. In some embodiments, the method may further comprise providing input from the train and tune model 1608 to an LLM model 1612. In some embodiments, the LLM model may be a trained LLM model 1612. In some embodiments, the LLM model may be a custom trained LLM model 1612.
[0265] In some embodiments, the method may further comprise providing input from the train and tune model 1608 to a structural generation simulation module (SGSM) 1614. In some embodiments, the SGSM may be a trained SGSM 1614. In some embodiments, the SGSM may be a custom trained SGSM 1614.
[0266] In some embodiments, the method may further comprise using the LLM 1612 to output one or more sequences of the one or more novel molecules or biological systems. In some embodiments, the method may further comprise using the SGSM 1614 to output one or more structures of the one or more novel molecules or biological systems. In some embodiments, the LLM model 1612 may receive feedback input. In some embodiments, the feedback input may comprise the one or more structures output by the SGSM 1614. In some embodiments, the feedback input may modify the LLM model 1612 output. In some embodiments, the modified LLM model 1612 output may comprise the one or more sequences of the one or more novel molecules or biological systems.
[0267] In some embodiments, the SGSM 1614 may receive as feedback input the one or more sequences output by the LLM model 1612. In some embodiments, the SGSM 1614 maygenerate output comprising modified output. In some embodiments, the output of the SGSM 1614 is modified based on feedback input from the LLM model 1612. In some embodiments, the modified SGSM 1614 output may comprise the one or more structures of the one or more novel molecules or biological systems. In some embodiments, the SGSM 1614 may output data to a novel molecular library model or database 1624. In some embodiments, the SGSM 1614 may output data to a novel molecular library model or database 1624 without processing by either or both of the machine leaming / data training model 1601 or the computational validation model 1616
[0268] In some embodiments, the method may further comprise applying a model output feedback for ML / training model 1625 to evaluate and classify the novel output of the LLM model 1612, the novel output of the SGSM 1614, or both into one or more groups. In some embodiments, the one or more groups categorize the one or more sequences of the one or more novel molecules or biological systems output by the LLM model 1612, the SGSM 1614, or both. In some embodiments, the model output feedback for ML / training model 1625 may receive input from a predictive molecular simulation movie generator model 1602. In some embodiments, the predictive molecular simulation movie generator model 1602 may receive input from the model output feedback for ML / training model 1625. In some embodiments, the model output feedback for ML / training model 1625 may receive input from a known molecules model for ML / training model 1626. In some embodiments, the predictive molecular simulation movie generator model 1602 may receive input from the known molecules model for ML / training model 1626. In some embodiments, the known molecules model for ML / training model 1626 may receive input from the predictive molecular simulation movie generator model 1602. In some embodiments, the known molecules model for ML / training model 1626 may receive input from the model output feedback for ML / training model 1625. In some embodiments, the model output feedback for ML / training model 1625, the known molecules model for ML / training model 1626, or the predictive molecular simulation movie generator model 1602, or any combination thereof, may receive input from a gene target DNA model for ML / training model 1627. In some embodiments, a gene target DNA model for ML / training model 1627 may receive input from the model output feedback for ML / training model 1625, the known molecules model for ML / training model 1626, or the predictive molecular simulation movie generator model 1602, or any combination thereof. In some embodiments, an other protein model for ML / training model 1628 may receive input from one or more of the model output feedback for ML / training model 1625, the known molecules model for ML / training model 1626, the gene target DNA model for ML / training model 1627, or the predictive molecular simulation movie generator model 1602, or any combination thereof. In someembodiments, an other protein model for ML / training model 1628 may output data to one or more of the model output feedback for ML / training model 1625, the known molecules model for ML / training model 1626, the gene target DNA model for ML / training model 1527, or the predictive molecular simulation movie generator model 1602, or any combination thereof. In some embodiments, a current solution model for ML / training model 1629 may receive input from one or more of the model output feedback for ML / training model 1625, the known molecules model for ML / training model 1626, the gene target DNA model for ML / training model 1627, the other protein model for ML / training model 1628, or the predictive molecular simulation movie generator model 1602, or any combination thereof. In some embodiments, a current solution model for ML / training model 1629 may output data to one or more of the model output feedback for ML / training model 1625, the known molecules model for ML / training model 1626, the gene target DNA model for ML / training model 1627, the other protein model for ML / training model 1628, or the predictive molecular simulation movie generator model 1602, or any combination thereof. In some embodiments, a delivery solution model for ML / training model 1630 may receive input from one or more of the model output feedback for ML / training model 1625, the known molecules model for ML / training model 1626, the gene target DNA model for ML / training model 1627, the other protein model for ML / training model 1628, the delivery solution model for ML / training model 1629, or the predictive molecular simulation movie generator model 1602, or any combination thereof. In some embodiments, a delivery solution model for ML / training model 1630 may output data to one or more of the model output feedback for ML / training model 1625, the known molecules model for ML / training model 1626, the gene target DNA model for ML / training model 1627, the other protein model for ML / training model 1628, the current solution model for ML / training model 1629, or the predictive molecular simulation movie generator model 1602, or any combination thereof. In some embodiments, a structural model for ML / training model 1631 may receive input from one or more of the model output feedback for ML / training model 1625, the known molecules model for ML / training model 1626, the gene target DNA model for ML / training model 1627, the other protein model for ML / training model 1628, the delivery solution model for ML / training model 1629, the delivery solution model for ML / training model 1630, or the predictive molecular simulation movie generator model 1602, or any combination thereof. In some embodiments, a structural model for ML / training model 1631 may output data to one or more of the model output feedback for ML / training model 1625, the known molecules model for ML / training model 1626, the gene target DNA model for ML / training model 1627, the other protein model for ML / training model 1628, the current solution model for ML / training model1629, a delivery solution model for ML / training model 1630, or the predictive molecular simulation movie generator model 1602, or any combination thereof.
[0269] In some embodiments, the predictive molecular simulation movie generator model 1602 may comprise a predictive visual molecular simulations model 1603. In some embodiments, the predictive visual molecular simulations model 1603 may predict a future chemical reactions of one or more molecules, for example the molecules output by the LLM model 1612, or the SGSM 1614, or both. In some embodiments, the predictive visual molecular simulations model 1603 may use one or more stable diffusion models to generate one or more predictive future chemical reactions of the one or more molecules. In some embodiments, the predictive visual molecular simulations model 1603 may receive input comprising one or more simulated movies of one or more biological materials, for example gene editors or other molecules. In some embodiments, the predictive visual molecular simulations model 1603 may generate one or more video simulations of the one or more predictive future chemical reactions of the one or more molecules. In some embodiments, the predictive visual molecular simulations model 1603 may comprise a generated molecular simulation movie 1604. In some embodiments, the molecular simulation movie may comprise one or more of a sequence of 3D images, sequence data, one or more additional characteristics, or any combination thereof. In some embodiments, the 3D images of the molecular simulation movie 1604 may comprise 3D images of one or more biological materials, for example gene editors or other molecules. In some embodiments, the one or more additional characteristics may comprise features of the one or more biological materials, for example gene editors or other molecules. In some embodiments, the generated molecular simulation movie 1604 may simulate one or more chemical interactions of the one or more biological materials, for example gene editors or other molecules. In some embodiments, the predictive molecular simulation movie generator model 1602 may output data to the SGSM 1614. In some cases, the data output by the predictive molecular simulation movie generator model 1602 to the SGSM 1614 may comprise the generated molecular simulation movie 1604.
[0270] In some embodiments, the LLM model 1612, or the SGSM 1614, or both, may output data to a computational validation model 1616. In some embodiments, the LLM model 1612 may output data to a generated molecular sequence model or database 1622. In some embodiments, the generated molecular sequence model or database 1622 may output one or more molecular sequences from the LLM model 1612. In some embodiments, the generated molecular sequence model or database 1622 may output data to a trained molecular classifier model 1620. In some embodiments, the trained molecular classifier model 1620 may classify one or more sequences input from the generated molecular sequences model or database 1622.In some embodiments, the classification may be performed based on one or more characteristics of each molecule. In some embodiments, the trained molecular classifier model 1620 may output data to one or more predictive assessment tools. In some embodiments, the one or more predictive assessment tools may comprise folding evaluation, structural evaluation, molecular simulation, molecular dynamics simulation, or any combination thereof. In some embodiments, the trained molecular classifier model 1620 may output data to one or more molecular folding evaluation tools 1618. In some embodiments, the molecular folding evaluation tools 1518 may be third-party evaluation tools. In some embodiments, the molecular folding evaluation tools 1518 may be one or more of AlphaFold, Rosetta, or DeepMind Equivariant Neural Networks, or any combination thereof. In some embodiments, the molecular folding evaluation tools 1518 may be customized, bespoke, in-house generated, or otherwise independently created molecular folding evaluation tools. In some embodiments, the one or more molecular folding evaluation tools 1618 may output data to a full validation model 1623. In some embodiments, the SGSM 1614 may output data to a generated molecular structure model or database 1621. In some embodiments, the generated molecular structure model or database 1621 may output data to one or more structural evaluation tools 1619. In some embodiments, the one or more structural evaluation tools 1519 may output data to a full validation model 1623. In some embodiments, the full validation model 1623 may output validated data to a molecular simulation model 1617, or receive data from a molecular simulation model 1617, or both. In some embodiments, the full validation model 1623 may output validation data to the model output feedback for ML / training model 1625. In some embodiments, the molecular simulation model 1617 may output data to a novel molecular library 1624. In some embodiments, the novel molecular library 1624 may receive input data from the SGSM 1614. In some embodiments, the novel molecular library 1624 may output one or more novel molecules or biologic materials to a user.
[0271] In yet another non-limiting embodiment of the invention, as illustrated in FIG. 17, the invention may comprise a system or method 1700 for assisting in generation of one or more novel molecules or biological systems. In some embodiments, the invention may comprise a method for assisting in generation of one or more novel molecules or biological systems. In some embodiments, the method may utilize a machine learning / data training model 1701, a generative / computational model 1713, and a computational validation model 1716. In some embodiments, the machine learning / data training model 1701 may comprise an LLM architecture model 1706, a protein model or database 1707, an input raw molecular sequences model or database 1709, a cloud platform 1711, a data preprocessing model 1710, a train and tune model 1708, a model output feedback for machine leaming / training model 1725, a known molecules model for ML / training model 1726, a gene target DNA model for ML / training 1727,an other protein model for ML / training model 1728, a current solution model for ML / training 1729, a delivery solution model for ML / training model 1730, a structural model for ML / training model 1731, and a predictive molecular simulation movie generator ML / training model 1732.
[0272] In some embodiments, the generative / computational model 1713 may comprise a trained LLM model 1712 and a trained structural model 1714. In some embodiments, the trained LLM model may be a custom trained LLM model 1712. In some embodiments, the trained structural model may be a custom trained structural model 1714. In some embodiments, the LLM model 1712 may comprise a GPT-1, GPT-2, GPT-3, GPT-4, GPT-5, GPT-6, GPT-7, GPT-8, GPT-9, GPT-10, or another GPT model. In some embodiments, the LLM model 1712 may comprise a protein language model. In some embodiments, the LLM model 1712 may comprise a protein language model comprising one or more of ESM-1, ESM-lv, ESM-2, ESM-2v, ESM-3, ESM-3v, ESM-4, ESM-4v, ESM-5, ESM-5v, ESMFold, ProGen, proteinBERT, another protein language model, or any combination thereof. In some embodiments, the LLM model 1712 may comprise an antibody-specific protein language model. In some embodiments, the LLM model 1712 may comprise an antibody-specific protein language model comprising one or more of IgLM, AbLang, AbLang2, AbLang3, AbLang4, AbLang5, AbLang6, AbLang7, AbLang8, AbLang9, AbLang 10, AntiBERTa, Bio-inspired Antibody Language Model (BALM), BALM-paired and unpaired, O AS-trained RoBERTa, IgBert, IgT5, AntiBERTy, Sapiens, FAbConSapiens, or another antibody-specific protein language model, or any combination thereof.
[0273] In some embodiments, the computational validation model 1716 may comprise a generated molecular sequences model or database 1722, a generated molecular structure model or database 1721, a trained molecular classifier model 1720, one or more molecular folding evaluation tools 1718, one or more structural evaluation tools 1719, a full validation model 1723, and a molecular simulation model 1717.
[0274] In some embodiments, the method may comprise providing an input to a large language learning model (LLM) 1712, and using the LLM to output one or more sequences of the one or more novel molecules or biological systems. In some embodiments, the method may comprise applying a validation model 1723 to evaluate or predict the effectiveness of the output of the LLM model 1712 using one or more predictive assessment tools. In some embodiments, the method may comprise outputting a sequence of the one or more novel molecules or biological systems, or structural representation thereof. In some embodiments, the output sequence or structural representation may be validated for the one or more predictive assessment tools. In some embodiments, the validation output may comprise the one or more novel molecules orbiological systems. In some embodiments, the molecule may be a protein. In some embodiments, the molecule may be an enzyme. In some embodiments, the molecule may be a recombinase. In some embodiments, the molecule may be a structural protein. In some embodiments, the molecule may be a hormonal protein. In some embodiments, the molecule may be a receptor protein. In some embodiments, the molecule may be a transport protein. In some embodiments, the molecule may be a hydrolase. In some embodiments, the molecule may be an oxidoreductase. In some embodiments, the molecule may be a transferase. In some embodiments, the molecule may be a lyase. In some embodiments, the molecule may be an isomerase. In some embodiments, the molecule may be a ligase. In some embodiments, the molecule may be a carbohydrate active enzyme. In some embodiments, the molecule may be a polymerase. In some embodiments, the molecule may be a topoisomerase. In some embodiments, the molecule may be an ATPase. In some embodiments, the molecule may be a phosphatase. In some embodiments, the molecule may be a ubiquitin ligase. In some embodiments, the molecule may be a nucleic acid. In some embodiments, the molecule may be a deoxyribonucleic acid (DNA). In some embodiments, the molecule may be a ribonucleic acid (RNA). In some embodiments, the molecule may be a carbohydrate. In some embodiments, the molecule may be a lipid. In some embodiments, the molecule may be an ion. In some embodiments, the molecule may be a vitamin. In some embodiments, the molecule may be a hormone. In some embodiments, the molecule may be a metabolic intermediate. In some embodiments, the molecule may be a coenzyme and / or a cofactor. In some embodiments, the molecule may be a secondary metabolite. In some embodiments, the molecule may be a nuclease. In some embodiments, the molecule may be an endonuclease. In some embodiments, the molecule may be an exonuclease. In some embodiments, the molecule may be a viral nuclease. In some embodiments, the molecule may be an RNA-guided endonuclease. In some embodiments, the molecule may be a component of a Clustered Regularly Interspaced Short Palindromic Repeats and CRISPR-associated proteins (CRISPR-Cas) system. In some embodiments, the molecule may be a meganuclease. In some embodiments, the molecule may be a transcription activator-like effector nuclease. In some embodiments, the molecule may be a zinc finger nuclease. In some embodiments, the biological systems may comprise DNA, RNA, ions, lipids, nanoparticles, small molecules, or any combination thereof.
[0275] In some embodiments, the input may comprise one or more of raw nuclease sequences, structural data, simulation data, quantum calculation data, or any combination thereof. In some embodiments, the input may comprise protein data 1707. In some embodiments, the input may comprise, for example, UniProt data 1707. In some embodiments, the input may comprise information related to one or more phenotypic characteristics. In some embodiments, the inputdata may be received by an input raw molecular sequences model or database 1709. In some embodiments, the input raw molecular sequences model or database 1709 may output raw molecular sequence data through a cloud server 1711 to a data preprocessing model 1710. In some embodiments, the molecular sequence data may be preprocessed and output by the data preprocessing model 1710 to the train and tune model 1708. In some embodiments, the input may be pre-processed to be configured as LLM input. In some embodiments, the train and tune model 1708 may receive input from an LLM architecture model 1706. In some embodiments, the LLM architecture model 1706 may be an open source LLM architecture model. In some embodiments, the LLM architecture model 1706 may be an untrained, pretrained, or custom LLM architecture model. In some embodiments, the LLM architecture model 1706 may be an untrained, pretrained, or custom open source LLM architecture model. In some embodiments, the train and tune model 1708 may receive input from a data preprocessing model 1710, an LLM architecture model 1706, a model output feedback for ML / Training model 1725, an LLM model 1712, a structural model 1714, or both. In some embodiments, the method may further comprise providing input from the train and tune model 1708 to an LLM model 1712. In some embodiments, the LLM model may be a trained LLM model 1712. In some embodiments, the LLM model may be a custom trained LLM model 1712.
[0276] In some embodiments, the method may further comprise providing input from the train and tune model 1708 to a structural generation simulation module (SGSM) 1714. In some embodiments, the SGSM may be a trained SGSM 1714. In some embodiments, the SGSM may be a custom trained SGSM 1714.
[0277] In some embodiments, the method may further comprise using the LLM 1712 to output one or more sequences of the one or more novel molecules or biological systems. In some embodiments, the method may further comprise using the SGSM 1714 to output one or more structures of the one or more novel molecules or biological systems. In some embodiments, the LLM model 1712 may receive feedback input. In some embodiments, the feedback input may comprise the one or more structures output by the SGSM 1714. In some embodiments, the feedback input may modify the LLM model 1712 output. In some embodiments, the modified LLM model 1712 output may comprise the one or more sequences of the one or more novel molecules or biological systems.
[0278] In some embodiments, the SGSM 1714 may receive as feedback input the one or more sequences output by the LLM model 1712. In some embodiments, the SGSM 1714 may generate output comprising modified output. In some embodiments, the output of the SGSM 1714 is modified based on feedback input from the LLM model 1712. In some embodiments, themodified SGSM 1714 output may comprise the one or more structures of the one or more novel molecules or biological systems.
[0279] In some embodiments, the method may further comprise applying a model output feedback for ML / training model 1725 to evaluate and classify the novel output of the LLM model 1712, the novel output of the SGSM 1714, or both into one or more groups. In some embodiments, the one or more groups categorize the one or more sequences of the one or more novel molecules or biological systems output by the LLM model 1712, the SGSM 1714, or both. In some embodiments, the model output feedback for ML / training model 1725 may receive input from a known molecules model for ML / training model 1726. In some embodiments, the known molecules model for ML / training model 1726 may receive input from the model output feedback for ML / training model 1725. In some embodiments, the model output feedback for ML / training model 1725, the known molecules model for ML / training model 1726, or both, may receive input from a gene target DNA model for ML / training model 1727. In some embodiments, a gene target DNA model for ML / training model 1727 may receive input from the model output feedback for ML / training model 1725, the known molecules model for ML / training model 1726, or both. In some embodiments, ...
Claims
CLAIMSWhat is claimed is:
1. A computer-implemented method for assisting in generation of one or more novel molecules or biological systems, the method comprising:(a) receiving an input comprising information related to a target function; (b) providing the input to a language learning model (LLM) and using the LLM to output one or more sequences of the one or more novel molecules or biological systems;(c) applying a validation model to evaluate or predict one or more predictive assessment tools of the output of the LLM; and(d) outputting a sequence of (b) or structural representation thereof, validated for the one or more predictive assessment tools of (c), wherein the output comprises the one or more novel molecules or biological systems.
2. The method of claim 1, wherein the molecule is a protein.
3. The method of claim 1, wherein the molecule is a nuclease.
4. The method of claim 1, wherein the biological systems comprise DNA, RNA, ions, lipids, nanoparticles, small molecules, or any combination thereof.
5. The method of claim 1, wherein the input further comprises one or more of raw nuclease sequences, structural data, simulation data, quantum calculation data, or any combination thereof.
6. The method of claim 1, wherein the input further comprises information related to one or more phenotypic characteristics.
7. The method of claim 1, wherein the input is pre-processed to be configured as LLM input.
8. The method of claim 1, wherein (b) further comprises providing the input to a structure generation module (SGM), and using the SGM to output one or more structures of the one or more novel molecules or biological systems.
9. The method of claim 8, wherein the LLM receives as feedback input the one or more structures output by the SGM to refine the LLM output of the one or more sequences of the one or more novel molecules or biological systems.
10. The method of claim 8, wherein the SGM receives as feedback input the one or more sequences output by the LLM to refine the SGM output of the one or more structures of the one or more novel molecules or biological systems.
11. The method of claim 1, wherein (c) further comprises applying a trained classifier module to classify into one or more groups the one or more sequences of the one or more novel molecules or biological systems output in (b).
12. The method of claim 1, wherein the one or more predictive assessment tools of (c) comprise folding evaluation, structural evaluation, molecular simulation, molecular dynamics simulation, or any combination thereof.
13. The method of claim 12, wherein the validation of (c) using the folding evaluation predictive assessment tool further comprises classifying output of the folding evaluation tool into one or more categories based on one or more predefined parameters.
14. The method of claim 13, wherein the one or more predefined parameters comprise similarity to known molecular structures, solubility, sequence features, foldability, cleavage features, catalytic sites features, or any combination thereof.
15. The method of claim 12, wherein the validation of (c) using the molecular simulation predictive assessment tool further comprises evaluating the output of the molecular simulation predictive assessment tool using one or more parameters.
16. The method of claim 15, wherein the one or more parameters comprise comparing the output of the molecular simulation tool to molecular bonding of known molecules, energy for binding, catalytic sites features, active sites features, sequence features, PAM sequence features, cleavage features, charge, microenvironment features, delivery simulation without cleavage, or any combination thereof.
17. The method of claim 12, wherein the validation of (c) using the molecular dynamics simulation tool further comprises evaluating output of the molecular dynamics simulation tool using one or more parameters.
18. The method of claim 17, wherein the one or more parameters comprise energy of one or more reactions, one or more energies of activation, one or more reaction pathways, one or more energies calculated using molecular mechanics, one or more energies calculated using quantum mechanics, one or more energies calculated using both quantum mechanics and molecular mechanics, or any combination thereof.
19. The method of claim 1, wherein the output of the validation model in (c) is used as feedback input for the LLM model of (b) to refine the output of the LLM model of (b).
20. The method of claim 1, further comprising a feedback loop from the output of the models of (b), (c), or both, wherein the feedback loop is used to enhance the one or more novel molecules or biological systems to minimize off-target effects.
21. The method of claim 1, wherein the one or more novel molecules or biological systems comprise miniaturized nucleases.
22. The method of claim 1, wherein the input further comprises genetic information associated with a specific disease.
23. The method of claim 1, further comprising generating a delivery system component corresponding to the output of (d) to deliver the one or more novel molecules or biological systems to a target.
24. The method of claim 23, wherein the delivery system component is generated using the output of (b) or (c), or both.
25. The method of claim 24, wherein the output of (b), or (c), or both, is used as feedback input to the LLM of (b) to generate the delivery system component.
26. The method of claim 1, wherein the input further comprises novel genes or mutations associated with one or more disease states.
27. The method of claim 1, wherein the one or more predictive assessment tools of the validation model of (c) utilize quantum computing.
28. The method of claim 1, further comprising in vitro testing of the output of (d).
29. The method of claim 1, wherein the input comprises feedback input based on the results of in vitro testing of the output of (d).
30. The method of claim 1, further comprising in vivo testing of the output of (d).
31. The method of claim 1, wherein the input comprises feedback input based on the results of in vivo testing of the output of (d).
32. The method of claim 31, wherein the in vivo testing further comprises testing in one or more mammals.
33. The method of claim 32, wherein the one or more mammals further comprise a human, a non-human primate, a murine mammal, or any combination thereof.
34. A computer-implemented method for assisting in generation of one or more novel molecules or biological systems, the method comprising:(a) receiving an input comprising information related to a target function; (b) providing the input to a structure generation model (SGM), and using the SGM to output one or more predicted structures of the one or more novel molecules or biological systems;(c) applying a validation model to evaluate or predict one or more predictive assessment tools of the output of the SGM; and(d) outputting a structure of (b) validated for the one or more predictive assessment tools of (c), wherein the output comprises the one or more novel molecules or biological systems.
35. The method of claim 34, wherein the molecule is a protein.
36. The method of claim 34, wherein the molecule is a nuclease.
37. The method of claim 34, wherein the biological systems comprise DNA, RNA, ions, lipids, nanoparticles, small molecules, or any combination thereof.
38. The method of claim 34, wherein the input further comprises one or more of raw nuclease sequences, structural data, simulation data, quantum calculation data, related input data, or any combination thereof.
39. The method of claim 34, wherein the input further comprises information related to one or more phenotypic characteristics.
40. The method of claim 34, wherein the input is pre-processed to be configured as SGM input.
41. The method of claim 34, wherein (c) further comprises applying one or more structural evaluation tools to the one or more structures of the one or more novel molecules or biological systems output in (b).
42. The method of claim 41, wherein the one or more structural evaluation tools further comprise a trained classifier module to classify the output of (b) into one or more groups.
43. The method of claim 34, wherein the one or more predictive assessment tools of (c) comprise folding evaluation, structural evaluation, molecular simulation, molecular dynamics simulation, or any combination thereof.
44. The method of claim 43, wherein the validation of (c) using the folding evaluation predictive assessment tool further comprises classifying output of the folding evaluation tool into one or more categories based on one or more predefined parameters.
45. The method of claim 44, wherein the one or more predefined parameters comprise similarity to known molecular structures, solubility, sequence features, foldability, cleavage features, catalytic sites features, or any combination thereof.
46. The method of claim 43, wherein the validation of (c) using the molecular simulation predictive assessment tool further comprises evaluating output of the molecular simulation predictive assessment tool using one or more parameters.
47. The method of claim 46, wherein the one or more parameters comprise comparing the output of the molecular simulation tool to molecular bonding of known molecules, energy for binding, catalytic sites features, active sites features, sequence features, PAM sequence features, cleavage features, charge, microenvironment features, delivery simulation without cleavage, or any combination thereof.
48. The method of claim 43, wherein the validation of (c) using the molecular dynamics simulation tool further comprises evaluating output of the molecular dynamics simulation tool using one or more parameters.
49. The method of claim 48, wherein the one or more parameters comprise energy of one or more reactions, one or more energies of activation, one or more reaction pathways, one or more energies calculated using molecular mechanics, one or more energies calculated using quantum mechanics, one or more energies calculated using both quantum mechanics and molecular mechanics, or any combination thereof.
50. The method of claim 34, wherein the output of the validation model in (c) is used as feedback input for the SGM model of (b) to refine the output of the SGM model of (b).
51. The method of claim 34, further comprising a feedback loop from the output of the models of (b), (c), or both, wherein the feedback loop is used to enhance the one or more novel molecules or biological systems to minimize off-target effects.
52. The method of claim 34, wherein the one or more novel molecules or biological systems comprise miniaturized nucleases.
53. The method of claim 34, wherein the input further comprises genetic information associated with a specific disease.
54. The method of claim 34, further comprising generating a delivery system component corresponding to the output of (d) to deliver the one or more novel molecules or biological systems to a target.
55. The method of claim 54, wherein the delivery system component is generated using the output of (b) or (c), or both.
56. The method of claim 55, wherein the output of (b), or (c), or both, is used as feedback input to the SGM of (b) to generate the delivery system component.
57. The method of claim 34, wherein the input further comprises novel genes or mutations associated with one or more disease states.
58. The method of claim 34, wherein the one or more predictive assessment tools of the validation model of (c) utilize quantum computing.
59. The method of claim 34, further comprising in vitro testing of the output of (d).
60. The method of claim 34, wherein the input comprises feedback input based on the results of in vitro testing of the output of (d).
61. The method of claim 34, further comprising in vivo testing of the output of (d).
62. The method of claim 34, wherein the input comprises feedback input based on the results of in vivo testing of the output of (d).
63. The method of claim 62, wherein the in vivo testing further comprises testing in one or more mammals.
64. The method of claim 63, wherein the one or more mammals further comprise a human, a non-human primate, a murine mammal, or any combination thereof.
65. A computer-implemented method for assisting in generation of one or more novel target-specific nucleases, the method comprising:(a) receiving input comprising information related to one or more targets; (b) providing the input of (a) to a language learning model (LLM), and using the LLM to generate one or more nuclease genetic sequences based at least in part on the one or more targets;(c) using a validation model to evaluate the one or more nuclease genetic sequences of (b) based on one or more predictive assessment tools, wherein the validation model outputs a modified sequence of (b) or a structure thereof; and(d) outputting the one or more validated novel nuclease sequences or structures of (c).
66. The method of claim 65, wherein the one or more predictive assessment tools comprise folding evaluation, structural evaluation, molecular simulation, or any combination thereof.
67. The method of claim 66, wherein molecular simulation further comprises at least one of quantum chemistry evaluation or simulated molecular dynamics evaluation.
68. A computer-implemented method for assisting in generation of one or more novel molecules or biological systems, the method comprising:(a) applying a classification model to classify one or more inputs into one or more desired molecule types, wherein the one or more inputs comprise information related to one or more target functions;(b) applying a language learning model (LLM) to generate one or more novel sequences of the one or more novel molecules or biological systems, wherein the one or more novel molecules or biological systems are of the one or more desired molecule types classified in (a);(c) simulating in situ one or more novel structures of each of the one or more novel sequences of the one or more novel molecules or biological systems generated in (b); and (d) outputting the one or more novel sequences of (b) or the one or more novel structures (c), or both.
69. A system for assisting in generation of one or more novel molecules or biological systems, the system comprising:(a) an input module configured to receive instructions comprising one or more desired features, targets, or functions;(b) a language learning model (LLM) configured to generate one or more sequences corresponding the one or more novel molecules or biological systems having the one or more desired features, targets, or functions;(c) a validation model configured to evaluate the one or more sequences of (b), or one or more structures thereof, using predictive assessment tools; and(d) an output module configured to output the one or more sequences or the one or more structures of the one or more novel molecules or biological systems validated in (c).
70. The system of claim 69, wherein the molecule is a protein.
71. The system of claim 69, wherein the molecule is a nuclease.
72. The system of claim 69, wherein the biological systems comprise DNA, RNA, ions, lipids, nanoparticles, small molecules, or any combination thereof.
73. The system of claim 69, wherein the input further comprises one or more of raw nuclease sequences, structural data, simulation data, quantum calculation data, or any combination thereof.
74. The system of claim 69, wherein the input further comprises information related to one or more phenotypic characteristics.
75. The system of claim 69, further comprising a pre-processing module configured to transform the input to LLM input.
76. The system of claim 69, wherein (b) further comprises a structure generation module (SGM).
77. The system of claim 76, wherein the LLM of (b) is further configured to provide the input to the SGM.
78. The system of claim 77, wherein the SGM is configured to output one or more structures of the one or more novel molecules or biological systems.
79. The system of claim 78, wherein the LLM is further configured to receive as feedback input the one or more structures output by the SGM to refine the LLM output of the one or more sequences of the one or more novel molecules or biological systems.
80. The system of claim 78, wherein the SGM is further configured to receive as feedback input the one or more sequences output by the LLM to refine the SGM output of the one or more structures of the one or more novel molecules or biological systems.
81. The system of claim 69, wherein (c) further comprises a classifier module.
82. The system of claim 81, wherein the classifier module is configured to classify into one or more groups the one or more sequences of the one or more novel molecules or biological systems output in (b).
83. The system of claim 69, wherein the one or more predictive assessment tools of (c) comprise folding evaluation, structural evaluation, molecular simulation, molecular dynamics simulation, or any combination thereof.
84. The system of claim 83, wherein the validation model of (c) is further configured to classify output of the folding evaluation tool into one or more categories based on one or more predefined parameters.
85. The system of claim 84, wherein the one or more predefined parameters comprise similarity to known molecular structures, solubility, sequence features, foldability, cleavage features, catalytic sites features, or any combination thereof.
86. The system of claim 83, wherein the validation model of (c) is further configured to evaluate the output of the molecular simulation predictive assessment tool using one or more parameters.
87. The system of claim 86, wherein the one or more parameters comprise comparing the output of the molecular simulation tool to molecular bonding of known molecules, energy for binding, catalytic sites features, active sites features, sequence features, PAM sequence features, cleavage features, charge, microenvironment features, delivery simulation without cleavage, or any combination thereof.
88. The system of claim 83, wherein the validation model of (c) is further configured to evaluate output of the molecular dynamics simulation tool using one or more parameters.
89. The system of claim 88, wherein the one or more parameters comprise energy of one or more reactions, one or more energies of activation, one or more reaction pathways, one or more energies calculated using molecular mechanics, one or more energies calculated using quantum mechanics, one or more energies calculated using both quantum mechanics and molecular mechanics, or any combination thereof.
90. The system of claim 69, wherein the validation model of (c) is further configured to transmit output to the LLM model of (b) to refine the output of the LLM model of (b).
91. The system of claim 69, further comprising a system feedback loop from the output of the models of (b), (c), or both, wherein the feedback loop is used to enhance the one or more novel molecules or biological systems to minimize off-target effects.
92. The system of claim 69, wherein the one or more novel molecules or biological systems comprise miniaturized nucleases.
93. The system of claim 69, wherein the input further comprises genetic information associated with a specific disease.
94. The system of claim 69, wherein the output of (d) further comprises generation of one or more delivery molecules to deliver the one or more novel molecules or biological systems to a target.
95. The system of claim 94, wherein the output further comprises output derived from the output of (b) or (c), or both.
96. The system of claim 94, wherein the LLM is further configured to generate the one or more delivery molecules using output of (b), or (c), or both.
97. The system of claim 69, wherein the input further comprises novel genes or mutations associated with one or more disease states.
98. The system of claim 69, wherein the one or more predictive assessment tools of the validation model of (c) utilize quantum computing.
99. The system of claim 69, further comprising in vitro testing of the output of (d).
100. The system of claim 69, wherein the input comprises feedback input based on the results of in vitro testing of the output of (d).
101. The system of claim 69, further comprising in vivo testing of the output of (d).
102. The system of claim 69, wherein the input comprises feedback input based on the results of in vivo testing of the output of (d).
103. The system of claim 102, wherein the in vivo testing further comprises testing in one or more mammals.
104. The system of claim 103, wherein the one or more mammals further comprise a human, a non-human primate, a murine mammal, or any combination thereof.
105. A system for assisting in generation of one or more novel sequences, structures, dynamic simulation trajectories, or biological systems, or any combination thereof, the system comprising:(a) an input module configured to receive instructions comprising one or more desired features, targets, or functions;(b) a structure generation model (SGM) configured to generate one or more sequences, structures, or dynamics simulation trajectories, or any combination thereof, corresponding the one or more novel molecules or biological systems having the one or more desired features, targets, or functions;(c) a validation model configured to evaluate the one or more sequences of (b), or one or more structures thereof, using predictive assessment tools; and(d) an output module configured to output the one or more sequences or the one or more structures of the one or more novel molecules or biological systems validated in (c).
106. The system of claim 105, wherein the molecule is a protein.
107. The system of claim 105, wherein the molecule is a nuclease.
108. The system of claim 105, wherein the biological systems comprise DNA, RNA, ions, lipids, nanoparticles, small molecules, or any combination thereof.
109. The system of claim 105, wherein the input further comprises one or more of raw nuclease sequences, structural data, simulation data, quantum calculation data, or any combination thereof.
110. The system of claim 105, wherein the input further comprises information related to one or more phenotypic characteristics.
111. The system of claim 105, further comprising a pre-processing module configured to format the input to an SGM input.
112. The system of claim 105, wherein the validation model of (c) further comprises one or more structural evaluation tools applied to the one or more structures of the one or more novel molecules or biological systems output by the SGM of (b).
113. The system of claim 112, wherein the one or more structural evaluation tools further comprise a trained classifier module to classify the output of (b) into one or more groups.
114. The system of claim 105, wherein the one or more predictive assessment tools of the validation model of (c) comprise folding evaluation, structural evaluation, molecular simulation, molecular dynamics simulation, or any combination thereof.
115. The system of claim 114, wherein the validation model of (c) is further configured to classify output of the folding evaluation tool into one or more categories based on one or more predefined parameters.
116. The system of claim 115, wherein the one or more predefined parameters comprise similarity to known molecular structures, solubility, sequence features, foldability, cleavage features, catalytic sites features, or any combination thereof.
117. The system of claim 114, wherein the validation model of (c) is further configured to evaluate the output of the molecular simulation predictive assessment tool using one or more parameters.
118. The system of claim 117, wherein the one or more parameters comprise comparing the output of the molecular simulation tool to molecular bonding of known molecules, energy for binding, catalytic sites features, active sites features, sequence features, PAM sequence features, cleavage features, charge, microenvironment features, delivery simulation without cleavage, or any combination thereof.
119. The system of claim 114, wherein the validation model of (c) is further configured to evaluate output of the molecular dynamics simulation tool using one or more parameters.
120. The system of claim 119, wherein the one or more parameters comprise energy of one or more reactions, one or more energies of activation, one or more reaction pathways, one or more energies calculated using molecular mechanics, one or more energies calculated using quantum mechanics, one or more energies calculated using both quantum mechanics and molecular mechanics, or any combination thereof.
121. The system of claim 105, wherein the SGM model of (b) is further configured to receive the output of the validation model in (c) to refine the output of the SGM model of (b).
122. The system of claim 105, further comprising a system feedback loop from the output of the models of (b), (c), or both, wherein the feedback loop is used to enhance the one or more novel molecules or biological systems to minimize off-target effects.
123. The system of claim 105, wherein the one or more novel molecules or biological systems comprise miniaturized nucleases.
124. The system of claim 105, wherein the input further comprises genetic information associated with a specific disease.
125. The system of claim 105, wherein the output of (d) further comprises a delivery system component corresponding to the output of (d) to deliver the one or more novel molecules or biological systems to a target.
126. The system of claim 125, wherein the delivery system component is generated using the output of (b) or (c), or both.
127. The system of claim 126, wherein the output of (b), or (c), or both, is used as feedback input to the SGM of (b) to generate the delivery system component.
128. The system of claim 105, wherein the input further comprises novel genes or mutations associated with one or more disease states.
129. The system of claim 105, wherein the one or more predictive assessment tools of the validation model of (c) utilize quantum computing.
130. The system of claim 105, further comprising in vitro testing of the output of (d).
131. The system of claim 105, wherein the input comprises feedback input based on the results of in vitro testing of the output of (d).
132. The system of claim 105, further comprising in vivo testing of the output of (d).
133. The system of claim 132, wherein the input comprises feedback input based on the results of the in vivo testing of the output of (d).
134. The system of claim 132, wherein the in vivo testing further comprises testing in one or more mammals.
135. The system of claim 134, wherein the one or more mammals further comprise a human, a non-human primate, a murine mammal, or any combination thereof.
136. A system for assisting in generation of one or more novel sequences or molecular structures, the system comprising:(a) an input module configured to receive instructions comprising one or more desired features, targets, or functions;(b) a language learning model (LLM) configured to generate one or more novel sequences based on the one or more desired features, targets or functions;(c) a structure prediction model configured to generate one or more novel molecular structures from the one or more novel sequences of (b); and(d) an output module configured to generate output comprising the one or more novel sequences of (b) or the one or more novel molecular structures of (c).