Techniques for artificial intelligence (AI) based protein engineering using natural language prompting

The system addresses the limitations of conventional protein engineering by using LLMs and conditional generative models to generate accurate protein sequences from human-readable prompts, enhancing flexibility and scalability in protein design.

WO2026015314A1PCT designated stage Publication Date: 2026-01-15UNIVERSITY OF CHICAGO
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/035833
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-11
Filing Date
2025-06-30
Publication Date
2026-01-15

AI Technical Summary

Technical Problem

Conventional protein engineering techniques lack flexibility and scalability, failing to generate viable protein sequences with desired characteristics and often require additional information beyond human-readable prompts.

Method used

A system utilizing large language models (LLMs) and conditional generative models to align text embeddings with protein embeddings, enabling the generation of protein sequences through natural language prompts, incorporating multilayer perceptron models and evolutionary scaling models for accurate protein sequence design.

Benefits of technology

Enables flexible and scalable protein engineering, allowing non-experts to design proteins with desired characteristics, improving accuracy and efficiency by integrating textual descriptions into the design process and enhancing model adaptability through multi-modal contrastive learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025035833_15012026_PF_FP_ABST
    Figure US2025035833_15012026_PF_FP_ABST
Patent Text Reader

Abstract

Techniques for artificial intelligence (Al) based protein engineering are disclosed herein. An example system includes one or more processors and one or more memories storing computer-executable instructions thereon. When executed, the instructions cause the processors to receive an input prompt including characteristics associated with a protein and input the input prompt into a large language model (LLM) that is configured to output a text embedding that is aligned with one or more protein embeddings output by a protein language model (pLM). The instructions further cause the processors to determine an adjusted text embedding based on the text embedding and the one or more protein embeddings, generate, using a conditional generative model, a protein sequence based on the adjusted text embedding, and cause at least one of: (i) the protein sequence or (ii) a graphic visualization of the protein sequence to be displayed for viewing by a user.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNIQUES FOR ARTIFICIAL INTELLIGENCE (Al) BASED PROTEIN ENGINEERING USING NATURAL LANGUAGE PROMPTINGRELATED APPLICATIONS

[0001] This application claims the benefit of U.S. Provisional Application No63 / 669,836 (filed on July 11, 2024), the entirety of which is incorporated by reference herein.STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT

[0002] This invention was made with government support under GM141697 awarded by the National Institutes of Health. The government has certain rights in the invention.TECHNICAL FIELD

[0003] The present disclosure generally relates to protein engineering techniques, and more particularly, to utilizing (1) large language models (LLMs) to interpret human-readable prompts and (2) a conditional generative (e.g., transformer) model to generate a protein sequence based, in part, on the LLM outputs.BACKGROUND

[0004] Engineered functional proteins have important applications in diverse areas, including clean energy, human health, and biotechnology. Conventional techniques for engineering such proteins suffer from several drawbacks. For example, conventional Potts models, autoregressive (AR) models, and Variational Autoencoders (VAEs) typically lack the flexibility and / or scalability to generate viable / optimal proteins. Namely, such conventional techniques arc frequently incapable of integrating sufficiently large, functional datasets to handle protein sequence generation tasks requiring high protein sequence characteristic specificity and / or are otherwise unable to generate viable / optimal protein sequences possessing characteristics that deviate from known proteins.

[0005] Therefore, in general, flexible and scalable protein engineering is an area of great interest, and conventional techniques can be insufficient for providing such protein sequences flexibly, at-scale, and possessing desirable functional properties. Accordingly, a need exists for techniques that enable users to flexibly engineer proteins at-scale that possess desirablefunctional properties, and thereby mitigate the negative effects stemming from inaccurate, inefficient, and inflexible conventional techniques.SUMMARY

[0006] In some aspects, the techniques described herein relate to a system for Al-bascd protein engineering, the system including: one or more processors; and one or more memories communicatively coupled with the one or more processors and storing computer-executable instructions thereon, that when executed, cause the one or more processors to: receiving an input prompt including characteristics associated with a protein, inputting the input prompt into a large language model (LLM) that is configured to output a text embedding that is aligned with one or more protein embeddings output by a protein language model (pLM), determining an adjusted text embedding based on the text embedding and the one or more protein embeddings, generating, using a conditional generative model, a protein sequence based on the adjusted text embedding, and causing at least one of: (i) the protein sequence or (ii) a graphic visualization of the protein sequence to be displayed for viewing by a user.

[0007] In some aspects, the techniques described herein relate to a system, wherein the computer-executable instructions, when executed, further cause the one or more processors to generate the protein sequence by: analyzing, by a multilayer perceptron model, the adjusted text embedding to determine a modified text representation; and inputting the adjusted text embedding into the conditional generative model to generate the protein sequence.

[0008] In some aspects, the techniques described herein relate to a system, wherein the computer-executable instructions, when executed, further cause the one or more processors to generate the protein sequence by: analyzing, with a multilayer perceptron model, a time encoding to determine a denoising time value; and inputting the denoising time value into the conditional generative model to generate the protein sequence.

[0009] In some aspects, the techniques described herein relate to a system, wherein the computer-executable instructions, when executed, further cause the one or more processors to generate the protein sequence by: converting, by a tokenizer module, a corrupted protein sequence and one or more starting tokens into continuous vector tokens; and inputting the continuous vector tokens and a positional encoding into the conditional generative model to generate the protein sequence.

[0010] In some aspects, the techniques described herein relate to a system, wherein the computer-executable instructions, when executed, further cause the one or more processors to: determine a similarity value between the protein sequence and at least one homolog protein sequence; and update the conditional generative model based on the similarity value.

[0011] In some aspects, the techniques described herein relate to a system, wherein the LLM is a Bidirectional Encoder Representations from Transformers (BERT)-based model, and wherein the computer-executable instructions, when executed, further cause the one or more processors to: train the LLM to output text embeddings based on a training loss associated with reconstructing masked tokens.

[0012] In some aspects, the techniques described herein relate to a system, wherein the pLM is an evolutionary scaling model (ESM2) configured to reconstruct masked amino acid tokens.

[0013] In some aspects, the techniques described herein relate to a system, wherein the computer-executable instructions, when executed, further cause the one or more processors to: train the LLM and the pLM in tandem for the LLM to output text embeddings aligned with protein embeddings output by the pLM based on (i) a global contrastive loss, (ii) a protein family contrastive loss, (iii) one or more masked language losses.

[0014] In some aspects, the techniques described herein relate to a system, wherein the computer-executable instractions, when executed, further cause the one or more processors to: train the conditional generative model to receive adjusted text embeddings as input and output a protein sequence by denoising a corrupted protein sequence.

[0015] In some aspects, the techniques described herein relate to a system, wherein the computer-executable instructions, when executed, further cause the one or more processors to determine the adjusted text embedding using a multilayer perceptron model that is configured to more closely align the text embedding with the one or more protein embeddings.

[0016] In some aspects, the techniques described herein relate to a system, wherein the conditional generative model is an order-agnostic autoregressive diffusion model.

[0017] In some aspects, the techniques described herein relate to a system, wherein the computer-executable instructions, when executed, further cause the one or more processors to:transmit a control instruction to a protein manufacturing system to manufacture a protein represented by the protein sequence.

[0018] In some aspects, the techniques described herein relate to a computer-implemented method for Al-based protein engineering, the computer-implemented method including: receiving, at one or more processors, an input prompt including characteristics associated with a protein; inputting, by one or more processors, the input prompt into a large language model (LLM) that is configured to output a text embedding that is aligned with one or more protein embeddings output by a protein language model (pLM); determining, by one or more processors, an adjusted text embedding based on the text embedding and the one or more protein embeddings; generating, by the one or more processors using a conditional generative model, a protein sequence based on the adjusted text embedding; and causing, by one or more processors, at least one of: (i) the protein sequence or (ii) a graphic visualization of the protein sequence to be displayed for viewing by a user.

[0019] In some aspects, the techniques described herein relate to a computer-implemented method, wherein generating the protein sequence further includes: analyzing, by a multilayer perceptron model, the adjusted text embedding to determine a modified text representation; and inputting, by the one or more processors, the adjusted text embedding into the conditional generative model to generate the protein sequence.

[0020] In some aspects, the techniques described herein relate to a computer-implemented method, wherein generating the protein sequence further includes: analyzing, with a multilayer perceptron model, a time encoding to determine a denoising time value; and inputting, by the one or more processors, the denoising time value into the conditional generative model to generate the protein sequence.

[0021] In some aspects, the techniques described herein relate to a computer-implemented method, wherein generating the protein sequence further includes: converting, by a tokenizer module, a corrupted protein sequence and one or more starting tokens into continuous vector tokens; and inputting, by the one or more processors, the continuous vector tokens and a positional encoding into the conditional generative model to generate the protein sequence.

[0022] In some aspects, the techniques described herein relate to a computer-implemented method, further including: determining, by the one or more processors, a similarity valuebetween the protein sequence and at least one homolog protein sequence; and updating, by the one or more processors, the conditional generative model based on the similarity value.

[0023] In some aspects, the techniques described herein relate to a computer-implemented method, wherein the LLM is a Bidirectional Encoder Representations from Transformers (BERT)-based model, and wherein the computer-implemented method further includes: training, by the one or more processors, the LLM to output text embeddings based on a training loss associated with reconstructing masked tokens.

[0024] In some aspects, the techniques described herein relate to a computer-implemented method, further including: transmitting, by the one or more processors, a control instruction to a protein manufacturing system to manufacture a protein represented by the protein sequence.

[0025] In some aspects, the techniques described herein relate to a tangible, non-transitory computer-readable medium storing instructions for Al-based protein engineering, that when executed by one or more processors cause the one or more processors to: receive an input prompt including characteristics associated with a protein; input the input prompt into a large language model (LLM) that is configured to output a text embedding that is aligned with one or more protein embeddings output by a protein language model (pLM); determine an adjusted text embedding based on the text embedding and the one or more protein embeddings; generate, using a conditional generative model, a protein sequence based on the adjusted text embedding; and cause at least one of: (i) the protein sequence or (ii) a graphic visualization of the protein sequence to be displayed for viewing by a user.BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The Figures described below depict preferred embodiments for purposes of illustration only. One skilled in the art will readily recognize from the following discussion that alternative embodiments of the systems and methods illustrated herein may be employed without departing from the principles of the disclosure described herein.

[0027] Figure 1 depicts an example computing system in which various embodiments of the present disclosure may be implemented.

[0028] Figure 2 depicts an example Al-based protein engineering workflow overview, in accordance with various embodiments described herein.

[0029] Figure 3A depicts an example multimodal pre-training sequence for a first stage of an Al-based protein engineering system, in accordance with various embodiments described herein.

[0030] Figure 3B depicts an example evolution-based contrastive learning workflow for a first stage of an Al-based protein engineering system, in accordance with various embodiments described herein.

[0031] Figure 3C depicts an example forward pass through training process for a first stage of an Al-based protein engineering system, in accordance with various embodiments described herein.

[0032] Figure 4 depicts example training effects for a second stage of an Al-based protein engineering system, in accordance with various embodiments described herein.

[0033] Figure 5A depicts an example text-guided diffusion process using a conditional generative model for a third stage of an Al-based protein engineering system, in accordance with various embodiments described herein.

[0034] Figure 5B depicts an example feedback loop that includes utilizing a predicted artificial protein sequence generated by an Al-based protein engineering system to re-train a model of the system, in accordance with various embodiments described herein.

[0035] Figure 6A depicts an example protein manufacturing sequence using generated protein sequences output by an Al-based protein engineering system, in accordance with various embodiments described herein.

[0036] Figure 6B depicts an example protein generation sequence to generate a protein sequence for manufacture, in accordance with various embodiments described herein.

[0037] Figure 7 depicts a flow diagram representing an example computer- implemented method, in accordance with various embodiments described herein.DETAILED DESCRIPTION

[0038] Broadly speaking, the Al-based protein engineering techniques of the present disclosure establish a new approach to protein engineering via a novel three-stage neural network architecture that enables the guided design of functional proteins via natural language prompts. These techniques fundamentally redefine the approach to protein design by adeptly mergingnatural language processing (NLP) with advanced generative Al models. Consequently, the present techniques enable the design of novel synthetic proteins with desired characteristics by non-expert users providing a text prompt (e. ., “Design a high-activity and thermostable S1A serine protease”). In doing so, this model democratizes protein design to make it accessible to a broad user base in a manner that was simply unachievable using conventional techniques.

[0039] More specifically, the present techniques provide flexible and scalable design and development of proteins across a diverse array of families, enzymes, and signaling proteins, all directed through user-input natural language prompts. These prompts include characteristics associated with a protein, and the present techniques interpret the prompts to guide Al models in creating specific protein sequences with desired characteristics. Namely, an input prompt is input into a large language model (LLM) that is configured to output a text embedding that is aligned with one or more protein embeddings output by a protein language model (pLM). The techniques of the present disclosure then adjust this text embedding and generate, using a conditional generative model (also referenced herein as a “transformer model” and / or utilizing a “transformer architecture”), a protein sequence based on the adjusted text embedding. At least these aspects of the present disclosure overcome the flexibility and scalability issues experienced by conventional techniques, demonstrate experimentally validated protein designs, and / or otherwise improve over such conventional techniques.

[0040] In particular, at least the generative Al component(s) (z.e., generating a protein sequence using a conditional generative model) and the multiple language models directly improve the flexibility and scalability of the protein generation / manufacturing processes, as compared to conventional techniques. For example, the multiple language models (LLMs) improve over conventional techniques by enabling significantly greater flexibility in the protein generation process than was previously possible using conventional techniques. As mentioned, the LLMs of the present disclosure analyze input (human-readable) prompts to generate text embeddings that are aligned with protein embeddings output by a pLM. The protein embeddings are sufficiently accurate and granular that they can be used for competitive structure predictions as part of the embedding alignment process without requiring multiple sequence alignment (MSA), unlike conventional techniques. Thus, unlike conventional techniques, the present techniques are readily scalable to Metagenome scales because the multiple language models (andcorresponding embeddings) enable significantly faster inferences than were possible using conventional techniques.

[0041] Further, these aligned text embeddings are analyzed / adjusted to generate (and subsequently manufacture) accurate protein sequences, such that only a human-readable input prompt is required to achieve a highly accurate protein sequence. Conventional techniques either struggle to analyze such prompts accurately / completely or cannot interpret / analyze such prompts entirely. Thus, the LLMs of the present disclosure improve over conventional techniques at least in that they dramatically expand the universe of usable inputs to generate highly accurate protein sequences (z.e., increase the flexibility / tolerance of the system to varied inputs) than is available for conventional techniques.

[0042] Additionally, the conditional generative model receives these outputs from the LLMs and generates protein sequences in a manner that improves over conventional techniques.Broadly speaking, the conditional generative model is crucial in enabling the present techniques to generate / construct detailed and accurate protein sequences. The conditional generative model achieves this through a process known as “inpainting”, where the model inserts specific motifs into a protein sequence based on the existing structural framework and guided by textual descriptions of the desired protein functions (e.g., an adjusted text embedding). This text-based conditioning enables the conditional generative model to accurately inpaint specific regions in multimeric proteins or antibodies, such as complementarity-determining regions (CDRs) or particular scaffolding, as guided by textual descriptions of the desired functions. As a result, the conditional generative model is particularly adept at targeting and refining complex regions within multimeric proteins and antibodies. Conventional techniques generally lack the capability to integrate text-based conditioning into inpainting processes, and as such suffer from an inability to accurately inpaint specific regions of proteins entirely or without additional information, which increases the demand on processing resources (processing cycles, bandwidth, time) to correct or refine such inaccuracies.

[0043] Moreover, in certain embodiments, the present techniques are integrated as pail of a protein manufacturing process (e.g., with protein manufacturing systems, high throughput selection assays) that further improves over conventional techniques. Such a manufacturing integration amplifies the utility of the present techniques by enabling the inclusion ofexperimental readouts directly into the conditional generative model and / or the language models. By integrating textual prompts that detail assay outcomes and functional performance into the design cycle (e.g., conditional generative model, LLMs), the present techniques iteratively refine both the model(s) and the resulting protein sequences / manufactured proteins. This synthesis of linguistic input and empirical data (i.e., multi-modal contrastive learning) empowers our model(s) with unprecedented levels of adaptability and user-directed customization without requiring alterations to the neural architecture of the underlying model(s), which was not possible in many conventional systems. Such a protein manufacturing integration therefore iteratively improves the accuracy and functionality of the designed proteins and model(s) to a degree that is not possible using conventional systems that lack such an integration and / or an iterative / re-training architecture.

[0044] Still further, in certain embodiments, the present techniques utilize one or more training processes that leverage several novel loss terms / functions, leading to more accurate model(s) than was possible using conventional training techniques. For example, the LLMs are trained using two additional loss terms (e.g., a biomedical masked language loss and protein masked language loss) that significantly enhance the performance of both the LLM (e.g., a biomedical language model (bLM)) and the pLM by enabling the expansion of the available dataset from 1 to 21 keyword headers. Such a substantial increase in keyword headers enables the model(s) to intake significantly more data than was previously possible using conventional techniques, leading to substantially richer, well-informed embeddings output by the model(s). Both models are also trained using a novel molecular evolution-based loss (e.g., a protein family contrastive loss) and a global contrastive loss, which in combination with the masked language losses, (1) accurately structures the joint embedding space (e.g., between the bLM and pLM) and (2) expands the pretraining dataset for the model(s) to significantly more data (e.g., ~ 45 million samples or text-protein pairs) than was previously available to train conventional models (e.g., ~ 560,000 samples).

[0045] The techniques of the present disclosure thus also improve the functionality of a computing device (e.g., a hosting server / device such as a central server) at least by analyzing data in a particular way to enhance the accuracy and efficiency of the computing device. The large language models, facilitator model, and conditional generative model, executing on the computing device, generate protein sequences with flexibility, scalability, and accuracy notachieved using conventional techniques. That is, the present disclosure describes improvements in the functioning of the computer itself because the computing device has increased flexibility to accurately analyze more data than was possible using conventional techniques (e.g., increased keyword headers) as a direct result of the large language models, facilitator model, and conditional generative model. This improves over the prior art at least because existing systems cannot accurately analyze human-readable input prompts, are incapable of accurately generating complex protein sequences at-scale, and / or are otherwise unable to analyze data with the flexibility, scalability, and accuracy resulting from the disclosed large language models, facilitator model, and conditional generative model.

[0046] Moreover, the present disclosure includes effecting a transformation or reduction of a particular article to a different state or thing, e.g., transforming or reducing the analytical flexibility of a computing system (and associated subsystems / components / devices) from a non- optimal or error state (e.g., highly inflexible) to an optimal (or closer to optimal) state by utilizing LLMs to interpret human-readable input prompts and other Al-based models to generate protein sequences based on such input prompts, and thereby substantially reducing the inflexibility of conventional protein engineering techniques.

[0047] Still further, the present disclosure includes specific features other than what is well- understood, routine, conventional activity in the field, or adding unconventional steps that demonstrate, in various embodiments, particular useful applications, e.g., inputting an input prompt into a large language model (LLM) that is configured to output a text embedding that is aligned with one or more protein embeddings output by a protein language model (pLM), determining an adjusted text embedding based on the text embedding and the one or more protein embeddings, and / or generating, using a conditional generative model, a protein sequence based on the adjusted text embedding, among others.

[0048] Of course, it should be appreciated that the advantages and technical improvements described above and elsewhere herein are not the only advantages and / or technical improvements that may be realized as a result of the techniques described herein. Other advantages and / or technical improvements to the functioning of a computer itself or other technologies or technical fields may be apparent to one of ordinary skill in the art.

[0049] To provide a better understanding of the techniques described herein, Figure 1 depicts an example computing environment overview, in which techniques of the present disclosure may be implemented. Figures 2, 3A-3C, 4, and 5A and 5B depict training and / or functionality associated with a first stage, a second stage, and / or a third stage of an Al-based protein engineering system, respectively. Figures 6A and 6B depict example protein manufacturing processes / sequences using the protein sequence generation processes / systems described herein. Figure 7 illustrates an example computer-implemented method for flexible and scalable Al-based protein engineering.EXAMPLE COMPUTING SYSTEMS

[0050] Figure 1 depicts an example computing system 100 in which various embodiments of the present disclosure may be implemented. Depending on the embodiment, the example computing system 100 may determine / generate text embeddings, adjusted text embeddings, protein embeddings, protein sequences, graphic visualizations of protein sequences, and / or any related values or combinations thereof. Of course, it should be appreciated that, while the various components of the example computing system 100 (e.g., central server 102, computing device 104, protein manufacturing system 106, external server 108, etc.) are illustrated in Figure 1 as single components, the example computing system 100 may include multiple (e.g., dozens, hundreds, thousands) of computing devices 104 and external servers 108 that are simultaneously connected to the network 110 at any given time.

[0051] Generally, the example computing system 100 includes a central server 102, a computing device 104, a protein manufacturing system 106, and an external server 108. Each of the central server 102, the computing device 104, the protein manufacturing system 106, and the external server 108 may communicate with the other devices e.g., transmit data, instructions, etc.) across the network 110. As an example, the computing device 104 may belong to an external user and the central server 102, protein manufacturing system 106, and / or the external server 108 may belong to a protein synthesis laboratory configured to receive protein sequence generation requests from the device 104 to generate / manufacture a protein sequence. In this example, the user using the computing device 104 may transmit data e.g., an input prompt) to the central server 102, and the server 102 may execute a protein engineering application 102b 1 to generate one or more protein sequences based on the input prompt. The central server 102 mayalso make the protein sequence accessible to the user via the computing device 104, so the user may review the protein sequence to analyze correlations between the sequence and the characteristics indicated in the input prompt, update the input prompt to include additional characteristics / parameters, and / or any other suitable actions or combinations thereof. Further, in certain embodiments, the user may transmit an input prompt indicating a request to manufacture a generated protein sequence, which the central server 102 interprets and transmits to the protein manufacturing system 106 for subsequent manufacturing of the protein indicated by the generated protein sequence.

[0052] More specifically, the central server 102 includes one or more processors 102a, the memory 102b, and a networking interface 102c. The memory 102b stores computer-executable instructions that are configured to, when executed by the one or more processors 102a, cause the one or more processors 102a to analyze data (e.g., input prompts, protein database 108a) received at the central server 102 and output various values (e.g., protein sequences, text / protein embeddings). The protein engineering application 102bl, the large language models 102b2, the facilitator model 102b3, the conditional generative model 102b4, and the application data 102b5 may all include such computer-executable instructions, as well as other data. The memory 102b may also store additional data and / or databases. It should be appreciated that the central server 102 can include one or multiple computing devices that are co-located or distributed.Additionally, in certain embodiments, the protein engineering application 102b 1 includes one or more of the large language models 102b2, the facilitator model 102b3, the conditional generative model 102b4, and / or the application data 102b5.

[0053] The central server 102 receives input prompts from the computing device 104 connected to the server 102 through a network 110 and processes the input prompts in accordance with one or more sets of instructions stored in a memory 102b to output any of the values described herein. The central server 102 executes the protein engineering application 102b 1, which in turn, accesses and applies the large language models 102b2, the facilitator model 102b3, the conditional generative model 102b4, and / or the application data 102b5 to the input prompts. The input prompts generally include or indicate characteristics associated with a protein. For example, the input prompt may include a text prompt stating that the user desires a “key transmembrane SH3 domain protein in osmotic sensing” or a “class-A beta-lactamase family enzyme,” in which the user further indicates they are “interested in TEM-1 beta-lactamase’s role in breaking down penicillin-related antibiotics.” Thus, the characteristics associated with the protein included and / or indicated in the input prompts may be directly stated and extracted and / or inferred based on the context provided by the user. Some / all of this information may eventually be stored as part of the application data 102b5 and / or stored in an external storage location {e.g., external server 108).

[0054] Generally, the data included in the input prompt include a text string of alphanumerical characters. In certain embodiments, the data included as part of the input prompt is or includes an audio stream, a video stream, a file, a document, and / or any other suitable data / datatype(s) or combinations thereof. Accordingly, in these embodiments, the input prompt is or includes a set of such text strings, audio streams, video streams, files, documents, and / or any other suitable dataZdatatype(s) or combinations thereof.

[0055] The protein engineering application 102bl receives the input prompt and generates protein sequences based on the characteristics included and / or indicated by the input prompt by accessing / applying the large language models 102b2, the facilitator model 102b3, and the conditional generative model 102b4 to the input prompt. The protein sequences generally indicate / represent a real-world protein that likely possesses the characteristics indicated / included in the input prompt, such that proteins manufactured in accordance with the protein sequences should also possess such characteristics. The large language models 102b2 generally include two large language models (e.g., a bLM and a pLM) configured to output embeddings of respective inputs (e.g., text and protein sequences, respectively). The facilitator model 102b3 analyzes a text embedding output by the large language models 102b2 {e.g., the bLM) to align the text embedding with one or more protein embeddings output by the large language models 102b2 {e.g., pLM). The conditional generative model 102b4 then analyzes this aligned text embedding, along with other inputs, to generate a protein sequence that includes the protein characteristics represented by the aligned text embedding. In certain embodiments, the conditional generative model 102b4 is or includes a transformer model and / or otherwise utilizes a transformer architecture.

[0056] The application 102b 1 then causes a computing device (e.g., computing device 104) to display the protein sequence and / or a graphic visualization of the protein sequence for viewing by the user that provided the input prompt and / or any other suitable user(s). In certainembodiments, the application 102b 1 causes the computing device to only display the protein sequence. In other embodiments, the application 102b 1 causes the computing device to only display a graphical representation of the protein sequence.

[0057] In certain embodiments, one or more of the large language models 102b2, the facilitator model 102b3, and / or the conditional generative model 102b4 is stored in a remote location from the central server 102 (e.g., a cloud-based server). In these embodiments, the protein engineering application 102bl accesses the trained large language models 102b2, the trained facilitator model 102b3, and / or the trained conditional generative model 102b4 by transmitting inputs e.g., input prompts) to the cloud-based server. The trained large language models 102b2, the trained facilitator model 102b3, and / or the trained conditional generative model 102b4 then analyzes the input prompts, generates outputs e.g., text embeddings, aligned text embeddings, protein sequences), and the cloud-based server returns these outputs to the protein engineering application 102bl.

[0058] As mentioned, components of the example computing system 100 (e.g., large language models 102b2, facilitator model 102b3, conditional generative model 102b4) may utilize machine learning as part of their operation. Generally speaking, machine learning may be implemented through machine learning methods and algorithms. In certain embodiments, the machine learning model(s) utilized as part of the example computing system 100 is or includes multiple trained large language models, order-agnostic autoregressive diffusion models, and / or multilayer perceptron models configured / trained to determine text / protein embeddings, adjusted text embeddings, and / or protein sequences based on the input prompts.

[0059] In certain embodiments, the machine learning models described herein (e.g., models 102b2, 102b3, 102b4) employ supervised learning, which involves identifying patterns in existing data to make predictions about subsequently received data. Specifically, the machine learning models may be “trained” using training data, which includes example inputs and associated example outputs. Based upon the training data, the machine learning models generate a predictive function which maps outputs to inputs and utilize the predictive function to generate machine learning outputs based upon data inputs. The example inputs and example outputs of the training data may include any of the data inputs or machine learning outputs described above. In the exemplary embodiment, a processing element may be trained by providing it with a largesample of data with known characteristics or features. In various embodiments, the implemented machine learning methods and algorithms arc directed toward at least one of a plurality of categorizations of machine learning, such as supervised learning.

[0060] For example, the large language models 102b2 described herein utilize or include natural language processing (NLP) functionality. The input prompts may be or include text prompts provided by a user at computing device 104, and the large language models 102b2 may implement NLP algorithms / models to interpret the text included therein when determining the text embeddings. Additionally, one or more of the large language models 102b2 (e.g., the pLM) may be trained to implement NLP algorithms / models on text characters that represent protein sequences, instead of human-interpretable words or phrases, when determining the protein embeddings.

[0061] In some embodiments, the ML models described herein may employ unsupervised learning, which involves finding meaningful relationships in unorganized data. Unlike supervised learning, unsupervised learning does not involve user-initiated training based upon example inputs with associated outputs / labels. Rather, in unsupervised learning, the machine learning model organizes unlabeled data according to a relationship determined by at least one machine learning method / algorithm employed by the machine learning model. Unorganized data may include any combination of data inputs and / or machine learning outputs, as described above.

[0062] It is to be understood that supervised machine learning and / or unsupervised machine learning may also comprise retraining, relearning, or otherwise updating models with new, or different, information, which may include information received, ingested, generated, or otherwise used over time. Further, it should be appreciated that, as previously mentioned, the machine learning model described herein may be used to output text / protein embeddings, adjusted embeddings, protein sequences, and / or any other values, responses, or combinations thereof using artificial intelligence (e.g., a machine learning model of the large language models 102b2, facilitator model 102b3, conditional generative model 102b5) or, in alternative aspects, without using artificial intelligence.

[0063] To train the machine learning models described herein, the central server 102 and / or the protein engineering application 102bl may utilize data stored in the protein database 108a ofthe external server 108. The protein database 108a may generally include data pairs of protein sequences and corresponding textual descriptions (e.g., characteristics, families, functions, etc.) of the protein represented by the protein sequences. In certain embodiments, the protein database 108a includes multiple databases, such as the UniProtKB / Swiss-Prot database, the Pfam database, and / or any other suitable databases or combinations thereof. These databases are data- rich repositories of such text-protein sequence pairs and thereby provide substantial training data corpuses the server 102 and / or application 102bl may leverage to train the machine learning models described herein. For example, utilizing both the UniProtKB / Swiss-Prot database and the Pfam database of the protein database 108a to train the machine learning models (e.g., large language models 102b2) described herein yields a training dataset of approximately 45 million such text-protein sequence pairs, resulting in a robust training process that was previously unaccomplished by conventional techniques. In certain embodiments, the central server 102 and / or other computing resources may re-train any of the machine learning models described herein (e.g., LLMs 102b2, facilitator model 102b3, conditional generative model 102b4) using any of the outputs and / or other data used by or associated with execution of any of the actions described herein (e.g., as part of the execution of the protein engineering application 102b 1). It should be appreciated that the external server 108 can include one or multiple computing devices that are co-located or distributed, and / or that the protein database 108a may be one or more databases that are co-located or distributed on one or multiple computing devices.

[0064] More generally, the computing device 104 is or includes any device that is associated with (e.g., owned and / or operated by) a particular entity that may provide data (e.g., input prompts) that is transmitted to and / or is otherwise accessible by the central server 102, the protein manufacturing system 106, and / or the external server 108 through the network 110. In certain embodiments, the input prompt(s) transmitted to and / or otherwise accessible by the central server 102, the protein manufacturing system 106, and / or the external server 108 is a set of human-interpretable text including characteristics associated with a protein to be evaluated by the central server 102, the protein manufacturing system 106, and / or the external server 108. In some embodiments, the computing device 104 is a server or collection of servers. However, in certain embodiments, the computing device 104 is a personal computing device of an entity / user, such as a smartphone, a tablet, smart glasses, or any other suitable device or combination of devices (e.g., a smart watch plus a smartphone) with wireless communication capability. In theembodiment of Figure 1 , the computing device 104 includes a processor 104a, a memory 104b, a networking interface 104c, and a display 104d.

[0065] The computing device 104 is communicatively coupled to the central server 102, the protein manufacturing system 106, and / or the external server 108. For example, the computing device 104, the central server 102, the protein manufacturing system 106, and / or the external server 108 may communicate via USB, Bluetooth, Wi-Fi Direct, Near Field Communication (NFC), etc. For example, the central server 102 may transmit a protein sequence or graphical indication of the protein sequence indicating extracted protein characteristics, embeddings, and / or any other values or combinations thereof to the computing device 104 via the networking interface 102c, which the computing device 104 may receive via the networking interface 104c. In another example, the central server 102 may also generate and transmit communications to the protein manufacturing system 106 to manufacture one or more proteins based on the protein sequences output by the machine learning models described herein. These communications may include the protein sequences as well as control instructions configured to control the manufacturing devices 106d included as part of the protein manufacturing system 106.

[0066] The protein manufacturing system 106 generally is a system configured to receive protein sequences from the central server 102 and manufacture proteins based on the protein sequences. To do this, the protein manufacturing system 106 receives communications from the central server 102, the computing device 104, and / or the external server 108 via the networking interface 106c and analyzes such communications using the processors 106a. The protein manufacturing system 106 may subsequently interpret the received protein sequence and / or any control instructions included in the communication from the server 102 to instruct the manufacturing devices 106d to manufacture one or more proteins based on the received protein sequence(s). In certain embodiments, the protein manufacturing system 106 may generate control instructions based on the protein sequence received from the server 102 to control the manufacturing devices 106d to manufacture one or more proteins.

[0067] Each of the processors 102a, 104a, 106a may include any suitable number of processors and / or processor types. For example, the processors 102a, 104a, 106a may each include one or more CPUs and one or more graphics processing units (GPUs). Generally, each of the processors 102a, 104a, 106a may be configured to execute software instructions stored ineach of the corresponding memories 102b, 104b, 106b. The memories 102b, 104b, 106b may each include one or more persistent memories (e.g., a hard drive and / or solid state memory) and may store one or more applications, modules, and / or models, such as the protein engineering application 102bl, the large language models 102b2, the facilitator model 102b3, and / or the conditional generative model 102b4. It should be appreciated that the external server 108 may also include one or more processors and one or more memories (both not shown for simplicity).

[0068] The networking interface 102c may enable the central server 102 to communicate with the computing device 104, the protein manufacturing system 106, the external server 108, and / or any other suitable devices or combinations thereof. More specifically, the networking interface 102c enables the central server 102 to communicate with each component of the example computing system 100 across the network 110 through their respective networking interfaces 104c, 106c, 108c. The networking interfaces 102c, 104c, 106c, 108c may support wired or wireless communications, such as USB, Bluetooth, Wi-Fi Direct, Near Field Communication (NFC), etc. The networking interface 102c may enable the central server 102 to communicate with the various components of the example computing system 100 via a wireless communication network such as a fifth-, fourth-, or third-generation cellular network (5G, 4G, or 3G, respectively), a Wi-Fi network (802.11 standards), a WiMAX network, or any other suitable wide area network (WAN), local area network (LAN), or personal area network (PAN), etc.

[0069] Moreover, the network 1 10 may be a single communication network, or may include multiple communication networks of one or more types (e.g., one or more wired and / or PANs or LANs, and / or one or more WANs such as the Internet). In some embodiments, the network 110 includes multiple, entirely distinct networks (e.g., one or more networks for communications between central server 102 and computing device 104, and a separate, Bluetooth or wireless LAN (WLAN) network for communications between central server 102 and protein manufacturing system 106, and so on).

[0070] It will be understood that the above disclosure is one example and does not necessarily describe every possible embodiment. As such, it will be further understood that alternate embodiments may include fewer, alternate, and / or additional steps or elements.EXAMPLE TRAINING AND EXECUTION OF A MULTI-STAGE, AI-BASEDPROTEIN ENGINEERING SYSTEM

[0071] Figure 2 depicts an example Al-based protein engineering workflow 200 overview, in accordance with various embodiments described herein. Generally speaking, the example AI- based protein engineering workflow 200 includes three stages, including a first stage 202, a second stage 204, and a third stage 206 that each perform specific functions to ultimately generate a protein sequence 206c. The first stage 202 generally includes training and executing two large language models to ensure that a text-based language model (e.g., bLM 202a) outputs aligned text embeddings that are aligned to one or more protein embeddings output by a proteinbased language model (e.g., pLM 202b). The second stage 204 generally includes training and executing a facilitator model 204a to further align text embeddings (e.g., text embedding 204b) output by the bLM 202a with protein embeddings (e.g., protein embedding 204c) output by the pLM 202b. The third stage 206 generally includes training and executing a conditional generative model 206a to generate protein sequences (e.g., protein sequence 206c) based on aligned text embeddings (e.g., aligned text embedding 206b) output by the facilitator model 204a.

[0072] As illustrated in Figure 2, the example Al-based protein engineering workflow 200 begins at the first stage 202 with the bLM 202a receiving an input text prompt 202a 1 and the pLM 202b receiving an input protein sequence prompt 202b 1. The models 202a, 202b analyze these input prompts 202al, 202b 1 and generate embeddings 202a2, 202b2, which are vector representations of the original input prompts 202al, 202b 1. As part of the training process for the bLM 202a and / or the pLM 202b, the applications (e.g., protein engineering application 102bl) and / or other processing components described herein (e.g., as part of central server 102) proceed to align these embeddings 202a2, 202b2 in a joint embedding space 202c to produce an aligned text embedding 202a3 and an aligned protein embedding 202b3. This alignment generally refers to adjusting parameters / weights of the bLM 202a and / or the pLM 202b and / or embeddings in a manner that minimizes several loss terms reflecting the overall differences / dis similarities that can exist between text-protein embedding pairs, groups, etc. when the data is embedded. Thus, to ensure that text embeddings output by the bLM 202a align accurately with corresponding protein embeddings output by the pLM 202b, the processors / applications analyze these loss terms in the context of the joint embedding space 202c, where larger scale trends or patterns of differences / dis similarities between many embedding pairs can be identified and adjusted, in addition to individual embedding pairdifferences / dissimilarities. These first stage 202 functions are further discussed herein in reference to Figures 3A-3C.

[0073] The example Al-based protein engineering workflow 200 continues with the second stage 204 and the third stage 206. The facilitator model 204a receives a text embedding 204b and a protein embedding 204c as outputs from the first stage 202 and further aligns the text embedding 204b with the protein embedding 204c to generate an aligned text embedding 206b. In certain embodiments, the protein embedding 204c is multiple protein embeddings that the facilitator model 204a uses to further align the text embedding 204b. Regardless, in the third stage 206, the conditional generative model 206a receives the aligned text embedding 206b as input to generate the protein sequence 206c. These second stage 204 and third stage 206 functions are further discussed herein in reference to Figures 4, 5A, and 5B.

[0074] Figure 3A depicts an example multimodal pre-training sequence 300 for a first stage of an Al-based protein engineering system, in accordance with various embodiments described herein. The example multimodal pre-training sequence 300 broadly illustrates a sequence of actions, which may be performed by central server 102 (e.g., processor 102a and / or other components of central server 102) of Figure 1, for example, to train large language models (e.g., bLM 202a and pLM 202b). The example multimodal pre-training sequence 300 illustrated in Figure 3A is for the purposes of discussion only, and additional / alternative training sequences may also, or instead, be utilized.

[0075] At a high level, and as previously mentioned, the first stage training process begins with a dual-language model framework, incorporating both a bLM 302 and a pLM 304. These models 302, 304 synergistically infer latent representations from natural language descriptions (e.g., text input prompt 302a) paired with protein sequences (e.g., protein input prompt 304a), aiming to create a joint embedding space 306. This space 306 aligns the bLM 302 output embeddings 302b with the pLM 304 output embeddings 304b, resulting in more aligned text embeddings 304c and protein embeddings 302c, and is also responsible for translating user-input natural language descriptions (i.e., input prompts) into actionable protein sequence design parameters. As such, within the joint embedding space 306, embeddings sharing similar features and / or that otherwise reference similar proteins “attract” while those that reference dissimilar proteins “repel”.

[0076] As illustrated in Figure 3A, the bLM 302 and the pLM 304 are trained using their own masked language losses (LbML, LpML, respectively), along with a novel Global Contrastive loss (LGC) and a protein family contrastive loss (LPFC). The global contrastive loss, LGC, is generally the main loss to align protein language and biomedical language loss, and the protein family contrastive loss LPFC generally clusters protein homology and thereby induces a “molecular evolution” bias during training. These losses enhance model training and scalability in several ways. For example, minimizing the LPFC efficiently aligns homologs within the same protein family. Further, both the bLM 302 and the pLM 304 receive specialized training using the masked language losses LbML, LpML to refine their understanding of their respective languages. The cumulative loss for the first stage training, Lstagei, is thus an amalgamation of these four loss components, for example, as a weighted sum, a simple sum (e.g., Lstagei = LGC + LPFC + LbML + LPML), and / or any other suitable composition or combinations thereof.:.

[0077] More specifically, the first stage is independently trained relative to the second and third stage pre-trainings. In certain embodiments, both the bLM and pLM are transformer-based models, and the losses, LbML, LpML, are known as “BERT” losses (i.e., masked language losses). In some embodiments, the bLM 302 corresponds to a PubMed-BERT model, which is a one- hundred million parameter model with a BERT-based architecture that is pre-trained on PubMed articles. The training loss for the bLM 302, LbML, is based on reconstructing masked tokens, and an advantage of this architecture is that the tokenization (i.e., vocabulary) is more particularly suited for words / terms featured in PubMed articles, resulting in the bLM 302 establishing a fluency with relevant science topics. Further, the bLM 302 is fine-tuned as part of the training represented by Figure 3A to account for the multimodality task(s) described herein.

[0078] In some embodiments, the pLM 304 is an ESM2 (evolutionary scaling model), which is pre-trained using meta-research on protein sequences. The pLM 304 generally learns data / characteristics of amino acids in terms of biophysical properties, and the pLM’s 304 objective is to mask amino acid tokens and reconstruct those tokens. As the pLM 304 is trained on numerous protein sequences, the pLM 304 has knowledge of protein features which can be used to conduct effective protein structure prediction. Similar to the bLM 302, the pLM 304 is fine-tuned as part of the training represented by Figure 3A to account for the multimodality task(s) described herein. In certain embodiments, training of the pLM 304 is performed in parallel (i.e., simultaneously) with the bLM 302, as described in reference to Figure 3C.

[0079] Figure 3B depicts an example evolution-based contrastive learning workflow 330 for a first stage of an Al-based protein engineering system, in accordance with various embodiments described herein. The example evolution-based contrastive learning workflow 330 broadly illustrates a sequence of actions, which may be performed by central server 102 (e.g., processor 102a and / or other components of central server 102) of Figure 1, for example, to train one or more language models e.g., bLM 302 and pLM 304). The example evolution-based contrastive learning workflow 330 illustrated in Figure 3B is for the purposes of discussion only, and additional / altemative training sequences may also, or instead, be utilized.

[0080] Generally, the example evolution-based contrastive learning workflow 330 includes leveraging data from one or more databases 332, 334 to train the pLM (e.g., pLM 304) and / or the bLM (e.g., bLM 302). The data taken from these databases 332, 334 marries textual descriptions with protein sequences (i.e., text-protein sequence pairs), and covers key protein properties such as function, catalytic activity, and subcellular location. These databases 332, 334 may generally be any suitable database(s) that include or reference text-protein sequence pairs. In some embodiments, the first database 332 is the UniProtKB / Swiss-Prot database and the second database 334 is the Pfam database. In these embodiments, the processing / training components (e.g., components of central server 102) incorporate additional descriptors and keywords / keyword headers from the databases 332, 334 into the training process for at least the bLM, such as "Family Description" and "Gene Ontology." Each keyword header may correspond to keys that the training components (e.g., protein engineering application) retrieve and mine within the databases 332, 334.

[0081] The four losses considered as part of the first stage training process are also shown in the example evolution-based contrastive learning workflow 330. The protein family contrastive loss LPFC and the global contrastive loss LGC both represent a joint embedding loss objective (e.g., within the joint embedding space 306) with respect to training the language models (i.e., pLM and bLM), and the masked language losses LbLM, LPLM represent a masked language loss objective with respect to training the language models. More specifically, the LPFC impacts the protein sequences from the databases 332, 334, as represented by the protein prompts 332a, 334a, and the LGC impacts the protein and text prompts, as represented by the protein prompts 332a, 334a and text prompts 332b, 334b. The LPML impacts the masked tokens from masked protein prompts 336a and reconstructs them, and the Lnvi. impacts the masked tokens frommasked text prompts 336b and reconstructs them. Each of these losses can be described mathematically as follows:where T is a temperature parameter that scales the distribution of the dot products, I is an indicator function, N is the batch size, and zpis the protein representation.

[0082] With these losses, the processing components of the Al-based systems described herein generally perform a “forward pass” training sequence that incorporates each loss value. Such a forward pass technique generally involves predicting embeddings / values at a given training step during the training process, and Figure 3C depicts an example forward pass through training process 360 for a first stage of an Al-based protein engineering system, in accordance with various embodiments described herein. The example forward pass-through training process 360 broadly illustrates a sequence of actions, which may be performed by central server 102 (e.g. , processor 102a and / or other components of central server 102) of Figure 1, for example, to train the language models (e.g. , bLM 202a, pLM 202b) described herein. The example forward pass- through training process 360 illustrated in Figure 3C is for the purposes of discussion only, and additional / altemative training sequences may also, or instead, be utilized.

[0083] Generally, the forward pass-through training process 360 illustrates how the bLM and pLM arc trained and process data from conncctcd / training databases (e.g., databases 332, 334). As mentioned, the language models are generally trained in parallel and their outputs aligned by increasing their correlation based on tensor products shown by the matrices 366, 368. Namely, the models receive input prompts 362a, 364a including data from the connected / training databases and output embeddings zpand zt, as illustrated in the bLM training input / output sequence 362 and the pLM training input / output sequence 364.

[0084] These training input / output sequences 362, 364 may occur any suitable number of iterations to roughly train and / or fine-tune the respective models. Ultimately, the training components (e.g., components of the central server 102) evaluate the losses previously described to determine when the language models have reached a suitable level of accuracy. For example, the global contrastive loss LGC allows the processing components to optimize correlation along the diagonal between zpand zt, while the protein family contrastive loss LPFC enables the processing components to optimize the correlation between homologs along the off diagonal within the known protein sequences. Further, it should be appreciated that after the pre-training discussed herein at least in reference to Figures 3A-3C is completed, the input prompt to the bLM can be written and submitted by a user with no necessary / required structure (e.g., capitalized keyword headers are not required).

[0085] Figure 4 depicts example training effects 400 for a second stage of an Al-based protein engineering system, in accordance with various embodiments described herein. The example training effects 400 broadly represent the impacts / effects of a sequence of actions (e.g., training / applying the facilitator model 402), which may be performed by central server 102 (e.g., processor 102a and / or other components of central server 102) of Figure 1. The example training effects 400 illustrated in Figure 4 is for the purposes of discussion only, and additional / altemative training sequences may also, or instead, be utilized.

[0086] Generally speaking, the second stage (e.g., second stage 204) of the protein engineering workflow described herein leverages a multi-layer perception model (i.e., facilitator model 402) with an autoencoder architecture and utilizes a max -mean discrepancy loss for training. As previously mentioned, this second stage focuses on aligning the bLM text embedding 404 (e.g., zt) with the pLM protein embedding 406 (e.g., zp). As part of thisalignment process, the processing components described herein (e.g., as part of central server 102) produce latent variables that align the textual descriptions with the protein sequences, for example, based on the dataset leveraged at the first stage.

[0087] Overall, the facilitator model 402 is configured / trained to improve the bLM text embedding 404 by aligning it further with pLM protein embedding 406 to output the aligned text embedding 408. The facilitator model 402 may be or utilize a multilayer perceptron model with a novel pretraining loss (z.e., max-mean discrepancy), which may be formulaically represented as follows:This second stage training loss associated with the facilitator model 402 represented by equation (6) generally includes calculating the maximum mean discrepancy (MMD) loss between the protein sequence representation zpand the produced protein sequence representation zc. In particular, the protein sequence representation can be represented as zc= ffacUitator(zt - The _||Xi-x.|i2kernel function / <(■,■) is a gaussian kernel, such that k( Xi, Xj) = exp ( — — ), and <7 is a hyperparameter. Additionally, or alternatively, the processing components described herein may calculate the mean-squared error (MSE) loss between zpand zc.

[0088] Based on this loss value, the facilitator model 402 accurately aligns the text embedding 404 to the protein embedding 406 to output the aligned text embedding 408. The aligned text embedding 408 is generally representative of the facilitator model’s 402 inference of the text embedding 404, represented mathematically as ffacintator(.zt)- This inference causes the text embedding 404 to have a matching magnitude with the protein embedding 406. More specifically, the facilitator model 402 matches the vector norms of the embeddings 404, 406 based on the pretraining described herein using the MMD loss represented by equation (6) and / or the MSE loss. This vector norm matching is illustrated in the MMD loss graph 410 and the MSE loss graph 412 of Figure 4, wherein the text embedding norms 410a, 412a, the protein embedding norms 410b, 412b, and the aligned text embedding norms 410c, 412c are shown.

[0089] Thus, after pre-training of the first and second stages of the protein engineering workflow described herein, the trained bLM may infer a text embedding 404 based on an input text prompt and that is aligned with one or more protein prompt embeddings. Then thefacilitator model 402 further infers an adjusted text embedding 408 which is then inserted to the third stage model (e.g., conditional generative model 102b4) to generate protein sequences that are compatible with the input text prompt.

[0090] Figure 5A depicts an example text-guided diffusion process 500 using a conditional generative model 502 for a third stage of an Al-based protein engineering system, in accordance with various embodiments described herein. The example text-guided diffusion process 500 broadly illustrates a sequence of actions, which may be performed by central server 102 (e.g., processor 102a and / or other components of central server 102) of Figure 1, for example, to generate protein sequences, and / or to train / re-train the conditional generative model 502. The example text-guided diffusion process 500 illustrated in Figure 5A is for the purposes of discussion only, and additional / altemative text-guided diffusion sequences may also, or instead, be utilized.

[0091] In the example text-guided diffusion process 500 representing the third stage of the protein engineering system described herein, the conditional generative model 502 iteratively reconstructs corrupted protein sequences over time to output a protein sequence. At each timestep iteration of the protein sequence reconstruction, the conditional generative model 502 reconstructs an amino acid position, as guided by previously refined amino acids and the adjusted textual embedding 514. The conditional generative model’s 502 transformer-based backbone architecture (e.g., order-agnostic autoregressive diffusion model) ensures scalability of the protein engineering processes described herein by handling extensive datasets and integrating with pre-trained models of up to 1.3 billion parameters.

[0092] Generally speaking, the conditional generative model 502 is a “conditional” orderagnostic autoregressive diffusion model, which differs from conventional diffusion models, and enables control over the generated outputs. The corresponding “conditionals” are determined and inferred based on the natural language prompts / embeddings following the first and second stages, as described herein.

[0093] Regardless, the conditional generative model 502 is a diffusion model, such that the training process broadly has two steps: a noising / corruption process and a denoising process, collectively represented by the sequence corruption / denoising 506. The denoising process isgenerally the process that the conditional generative model 502 is trained to perform and the corresponding loss can be represented as follows: logp(xk|xCT(<t),zc) (7),where zcis the adjusted text embedding, prepresents the conditional generative model 502, and the goal indicated by the loss represented by equation (7) is to optimize the model’s 502 ability to reconstruct missing amino acids xkwithin a corrupted protein sequence (e.g., corrupted sequence 508).

[0094] The noising / corruption process generally is or includes an absorbing state corruption process, meaning that at each timestep of the training process, the processing components mask and absorb an amino acid until the entire sequence is absorbed (e.g., hyphens within the sequence corruption / denoising 506). The denoising process generally converts a sequence of absorbing states (e.g., hyphens) into a novel generated protein sequence that is compatible with one or more natural language prompts based on the adjusted text embedding 514. Such a novel generated protein sequence is illustrated in Figure 5A as partially completed protein sequence 504 that includes one more amino acid than the corrupted sequence 508.

[0095] As illustrated in Figure 5A, the example text-guided diffusion process 500 further includes an input embedding module 510, a positional encoding 512, a first MLP module 516, a time encoding 518, and a second MLP module 520. The input embedding module 510 is generally a tokenizer module that converts amino acid and special labels (e.g., starting tokens) into continuous vector tokens. The positional encoding 512 and the time encoding 518 generally is or include one or more features used to augment the conditional generative model 502 analysis based on a known position and / or time, respectively, within the completely / partially corrupted sequence 508 during the denoising process. The MLP modules 516, 520 are multilayer perceptron layer models, the first MLP module 516 processes the adjusted text embedding 514 to provide a meaningful representation for conditional generative model 502 to process, and the second MLP module 520 processes the time encoding 518 to inform the conditional generative model 502 of timing parameters along the denoising time trajectory.

[0096] Figure 5B depicts an example feedback loop 530 that includes utilizing a predicted artificial protein sequence 534a generated by an Al-based protein engineering system to rc-train one or more of the models (e. g., conditional generative model 532) described herein, in accordance with various embodiments described herein. The example feedback loop 530 broadly illustrates a sequence of actions, which may be performed by central server 102 (e.g., processor 102a and / or other components of central server 102) of Figure 1, for example, to retrain a conditional generative model 532. The example feedback loop 530 illustrated in Figure 5B is for the purposes of discussion only, and additional / alternative feedback loops may also, or instead, be utilized.

[0097] More generally, integrating the Al-based protein engineering systems described herein with high-throughput selection assays, the systems and techniques described herein facilitate large-scale protein function readouts. These readouts, assay description(s), and / or any other suitable experimental / manufacturing results can then be reincorporated into the models (e.g., conditional generative model 532), thereby enriching the learning cycle with natural language inputs from the assays. For example, the processing components of the Al-based protein engineering systems described herein may evaluate similarity values and / or other values that represent the similarities / dis similarities between the predicted artificial protein sequence 534a and the natural homolog 534b, and these similarity values may be used to adjust parameters / weights of the conditional generative model 532 and / or any other models described herein.

[0098] In certain embodiments, the natural homolog 534b may be an artificial protein sequence, to which, the predicted artificial protein sequence 534a is compared. For example, the natural homolog 534b may be a known artificial protein sequence or may be another predicted artificial protein sequence that has been experimentally validated as possessing similar characteristics to the desired characteristics of the predicted artificial protein sequence 534a.

[0099] Additionally, or alternatively, the processing components of the Al-based protein engineering systems described herein may re-train the models described herein (e.g., conditional generative model 532) based on how well a manufactured protein based on the predicted artificial protein sequence 534a includes / provides the desired characteristics indicated in the input prompt(s). For example, the processing components described herein may determine oneor more similarity values representing the similarities between the functional aspects of a manufactured protein resulting from the predicted artificial protein sequence 534a and the desired functional characteristics indicated in an input prompt to the various models described herein. This similarity value may serve as feedback the processing components described herein may use to re-train and / or otherwise update the models (e.g., conditional generative model 532).EXAMPLE PROTEIN ENGINEERING INTEGRATIONS WITH PROTEIN MANUFACTURING SYSTEMS

[0100] Figure 6A depicts an example end-to-end protein manufacturing sequence 600 using generated protein sequences output by an Al-based protein engineering system, in accordance with various embodiments described herein. The example protein manufacturing sequence 600 generally represents an end-to-end protein manufacturing sequence including receiving input prompts 602, generating optimized protein sequences 604 utilizing the techniques described in the present disclosure, preparing a DNA library 606 based on the optimized protein sequences 604, performing a specialized selection assay 608, performing an advanced deep sequencing protocol 610, and measuring (e.g., quantitative evaluation 612) the relative abundance of the resulting proteins. Generally speaking, the specialized selection assay 608 may be or include a selection-type assay, a high-throughput fluorescence-based assay (e.g., microfluidic arrays or droplet microfluidics), high-throughput cell-based assays, and / or any readout that measures the biological activity of the designed molecules or combinations thereof.

[0101] More specifically, the example end-to-end protein manufacturing sequence 600 includes receiving input prompts 602 and advancing through the Al-driven processes for generating optimized protein sequences 604, as described in the present disclosure. With these optimized protein sequences 604 the sequence 600 further includes meticulously preparing the DNA library 606, which encapsulates the generated optimized protein sequences 604 and ensures a high-fidelity template for subsequent synthetic biological applications. The example end-to-end protein manufacturing sequence 600 further includes executing the specialized selection assay 608 that is tailored to identify highly functional protein variants, followed by the advanced deep sequencing protocol 610 designed to ensure exhaustive sequence coverage. The sequence 600 further includes a quantitative evaluation 612 of the relative abundance of theresultant proteins by employing cutting-edge analytical techniques to ascertain their expression levels and functional efficacy.

[0102] As an example, the Al-based protein engineering systems described herein may be tasked with generating synthetic sequences that complement function in vivo similar to three natural protein families: Shol-Sh3 domain, Chorismate mutase (CM) protein, and TEM-1 lactamase protein. The text-based language model (e.g., bLM) is then prompted with a textual description of each protein family to generate text embeddings. For example, an input prompt corresponding to the Shol-SH3 protein family may include: “Key transmembrane SH3 domain protein in osmotic sensing.. After the protein sequences are generated, the systems described herein (e.g., central server 102 and protein manufacturing system 106) then assemble the genes (e.g., 606), test the protein sequences 604 using high-throughput assays (e.g., 608), and perform deep sequencing 610.

[0103] Using the results of such manufacturing, the processing components described herein may re-train and / or update the various models described herein. Namely, the example end-to- end protein manufacturing sequence 600 results in a manufactured protein, from which the processing components described herein may acquire and utilize experimental data to update / re- train the models (e.g., conditional generative model 532) described herein.

[0104] Figure 6B depicts an example protein generation sequence 630 to generate one or more protein sequences for manufacture, in accordance with various embodiments described herein. The example protein generation sequence 630 generally represents the specific analyses / outputs provided by the three-stage neural network architecture described herein to achieve a generated / predicted protein sequence (e.g., predicted artificial protein sequence 534a).

[0105] More specifically, the example protein generation sequence 630 includes a set of input prompts 632, that are used as input to the bLM at the first stage 634. The bLM then infers a text embedding based on the input prompt(s) that is sent to the facilitator model, which further infers an adjusted text embedding based on the text embedding at the second stage 636. This adjusted text embedding is then used as an input to the conditional generative model at the third stage 638, and the conditional generative model generates protein sequences that are compatible to the initial text prompt. In other words, the protein sequences generated by the conditionalgenerative model at the third stage 638 generally include sequences that possess characteristics that arc similar / idcntical to those indicated by the input text prompt(s).

[0106] In any event, after the third stage 638, the protein sequences are developed into one or more protein designs 640, which may be manufactured by a protein manufacturing system (e.g., protein manufacturing system 106). Consequently, the proteins manufactured as a result of the protein sequences generated by the present techniques may find practical applications across a wide range of diverse industries, such as pharmaceuticals, sustainability efforts, and / or any other molecular-based fields.EXAMPLE COMPUTER-IMPLEMENTED METHODS

[0107] Figure 7 depicts a flow diagram representing an example computer-implemented method 700, in accordance with various embodiments described herein. The method 700 may be implemented by one or more processors of the example computing system 100, such as the processor 102a of central server 102 (e.g., by protein engineering application 102bl), for example.

[0108] The method 700 includes receiving an input prompt including characteristics associated with a protein (block 702). The method 700 further includes inputting the input prompt into an LLM that is configured to output a text embedding that is aligned with one or more protein embeddings output by a pLM (block 704). The method 700 further includes determining an adjusted text embedding based on the text embedding and the one or more protein embeddings (block 706). The method 700 further includes generating, using a conditional generative model, a protein sequence based on the adjusted text embedding (block 708). The method 700 further includes causing at least one of: (i) the protein sequence and / or (ii) a graphic visualization of the protein sequence to be displayed for viewing by a user (block 710).

[0109] In certain embodiments, method 700 further includes generating the protein sequence by: analyzing, by a multilayer perceptron model, the adjusted text embedding to determine a modified text representation; and inputting the adjusted text embedding into the conditional generative model to generate the protein sequence.

[0110] In some embodiments, the method 700 further includes generating the protein sequence by: analyzing, with a multilayer perceptron model, a time encoding to determine a denoising time value; and inputting the denoising time value into the conditional generative model to generate the protein sequence.

[0111] In certain embodiments, the method 700 further includes generating the protein sequence by: converting, by a tokenizer module, a corrupted protein sequence and one or more starting tokens into continuous vector tokens; and inputting the continuous vector tokens and a positional encoding into the conditional generative model to generate the protein sequence.

[0112] In some embodiments, the method 700 further includes: determining a similarity value between the protein sequence and at least one homolog protein sequence; and updating the conditional generative model based on the similarity value. In certain embodiments, the similarity value additionally or alternatively indicates a degree of similarity between the functional aspects of a manufactured protein resulting from a predicted artificial protein sequence (e.g., predicted artificial protein sequence 534a) and the desired functional characteristics indicated in an input prompt to the various models described herein. Accordingly, in these embodiments, updating / re-training the conditional generative model and / or other models described herein as part of a semi-supervised ML process may generally include similarities and / or other comparisons between the output / generated protein sequences and (1) natural homologs and / or (2) experimental / manufacturing data resulting from the manufacture of a protein based on the output / generated protein sequences.

[0113] In certain embodiments, the LLM is a Bidirectional Encoder Representations from Transformers (BERT)-based model, and the method 700 further includes: training the LLM to output text embeddings based on a training loss associated with reconstructing masked tokens.

[0114] In some embodiments, the pLM is an evolutionary scaling model (ESM2) configured to reconstruct masked amino acid tokens.

[0115] In certain embodiments, the method 700 further includes: training the LLM and the pLM in tandem for the LLM to output text embeddings aligned with protein embeddings output by the pLM based on (i) a global contrastive loss, (ii) a protein family contrastive loss, (iii) one or more masked language losses.

[0116] In some embodiments, the method 700 further includes: training the conditional generative model to receive adjusted text embeddings as input and output a protein sequence by denoising a corrupted protein sequence.

[0117] In certain embodiments, the method 700 further includes determining the adjusted text embedding using a multilayer perceptron model that is configured to more closely align the text embedding with the one or more protein embeddings.

[0118] In some embodiments, the conditional generative model is an order- agnostic autoregressive diffusion model.

[0119] In certain embodiments, the method 700 further includes transmitting a control instruction to a protein manufacturing system to manufacture a protein represented by the protein sequence.

[0120] Of course, it is to be appreciated that the actions of the method 700 may be performed any suitable number of times, and that the actions described in reference to the method 700 may be performed in any suitable order.ADDITIONAL CONSIDERATIONS

[0121] Throughout this specification, plural instances may implement components, operations, or structures described as a single instance. Although individual operations of one or more methods arc illustrated and described as separate operations, one or more of the individual operations may be performed concurrently, and nothing requires that the operations be performed in the order illustrated. Structures and functionality presented as separate components in example configurations may be implemented as a combined structure or component.Similarly, structures and functionality presented as a single component may be implemented as separate components. These and other variations, modifications, additions, and improvements fall within the scope of the subject matter herein.

[0122] The systems and methods described herein are directed to an improvement to computer functionality, and improve the functioning of conventional computers. Additionally, certain embodiments are described herein as including logic or a number of routines, subroutines, applications, or instructions. These may constitute either software (e.g., code embodied on a non-transitory, computer-readable medium) or hardware. In hardware, theroutines, etc., are tangible units capable of performing certain operations and may be configured or arranged in a certain manner. In example embodiments, one or more computer systems (e.g., a standalone, client or server computer system) or one or more hardware modules of a computer system e.g., a processor or a group of processors) may be configured by software (e.g., an application or application portion) as a hardware module that operates to perform certain operations as described herein.

[0123] In various embodiments, a hardware module may be implemented mechanically or electronically. For example, a hardware module may comprise dedicated circuitry or logic that is permanently configured (e.g., as a special-purpose processor, such as a field programmable gate array (FPGA) or an application-specific integrated circuit (ASIC)) to perform certain operations. A hardware module may also comprise programmable logic or circuitry (e.g., as encompassed within a general-purpose processor or other programmable processor) that is temporarily configured by software to perform certain operations. It will be appreciated that the decision to implement a hardware module mechanically, in dedicated and permanently configured circuitry, or in temporarily configured circuitry (e.g., configured by software) may be driven by cost and time considerations.

[0124] Accordingly, the term “hardware module” should be understood to encompass a tangible entity, be that an entity that is physically constructed, permanently configured (e.g., hardwired), or temporarily configured (e.g., programmed) to operate in a certain manner or to perform certain operations described herein. Considering embodiments in which hardware modules are temporarily configured (e.g., programmed), each of the hardware modules need not be configured or instantiated at any one instance in time. For example, where the hardware modules include a general-purpose processor configured using software, the general-purpose processor may be configured as respective different hardware modules at different times. Software may accordingly configure a processor, for example, to constitute a particular hardware module at one instance of time and to constitute a different hardware module at a different instance of time.

[0125] Hardware modules can provide information to, and receive information from, other hardware modules. Accordingly, the described hardware modules may be regarded as being communicatively coupled. Where multiple of such hardware modules exist contemporaneously,communications may be achieved through signal transmission (e.g., over appropriate circuits and buses) that connect the hardware modules. In embodiments in which multiple hardware modules are configured or instantiated at different times, communications between such hardware modules may be achieved, for example, through the storage and retrieval of information in memory structures to which the multiple hardware modules have access. For example, one hardware module may perform an operation and store the output of that operation in a memory device to which it is communicatively coupled. A further hardware module may then, at a later time, access the memory device to retrieve and process the stored output. Hardware modules may also initiate communications with input or output devices, and can operate on a resource (e.g., a collection of information).

[0126] The various operations of example methods described herein may be performed, at least partially, by one or more processors that are temporarily configured (e.g., by software) or permanently configured to perform the relevant operations. Whether temporarily or permanently configured, such processors may constitute processor-implemented modules that operate to perform one or more operations or functions. The modules referred to herein may, in some example embodiments, comprise processor- implemented modules.

[0127] Similarly, the methods or routines described herein may be at least partially processor- implemented. For example, at least some of the operations of a method may be performed by one or more processors or processor-implemented hardware modules. The performance of certain of the operations may be distributed among the one or more processors, not only residing within a single machine, but deployed across a number of machines. In some example embodiments, the processor or processors may be located in a single location (e.g., within a home environment, an office environment or as a server farm), while in other embodiments the processors may be distributed across a number of locations.

[0128] The performance of certain of the operations may be distributed among the one or more processors, not only residing within a single machine, but deployed across a number of machines. In some example embodiments, the one or more processors or processor-implemented modules may be located in a single geographic location (e.g., within a home environment, an office environment, or a server farm). In other example embodiments, the one or moreprocessors or processor-implemented modules may be distributed across a number of geographic locations.

[0129] It should also be understood that, unless a term is expressly defined in this patent using the sentence “As used herein, the term ‘ ’ is hereby defined to mean...” or a similar sentence, there is no intent to limit the meaning of that term, either expressly or by implication, beyond its plain or ordinary meaning, and such term should not be interpreted to be limited in scope based upon any statement made in any section of this patent (other than the language of the claims). To the extent that any term recited in the claims at the end of this disclosure is referred to in this disclosure in a manner consistent with a single meaning, that is done for sake of clarity only so as to not confuse the reader, and it is not intended that such claim term be limited, by implication or otherwise, to that single meaning.

[0130] Unless specifically stated otherwise, discussions herein using words such as “processing,” “computing,” “calculating,” “determining,” “presenting,” “displaying,” or the like may refer to actions or processes of a machine (e.g., a computer) that manipulates or transforms data represented as physical (e.g., electronic, magnetic, or optical) quantities within one or more memories (e.g., volatile memory, non-volatile memory, or a combination thereof), registers, or other machine components that receive, store, transmit, or display information.

[0131] As used herein any reference to “one embodiment” or “an embodiment” means that a particular element, feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment. The appearances of the phrase “in one embodiment” in various places in the specification are not necessarily all referring to the same embodiment.

[0132] As used herein, the terms “comprises,” “comprising,” “includes,” “including,” “has,” “having” or any other variation thereof, are intended to cover a non-exclusive inclusion. For example, a process, method, article, or apparatus that comprises a list of elements is not necessarily limited to only those elements but may include other elements not expressly listed or inherent to such process, method, article, or apparatus. Further, unless expressly stated to the contrary, “or” refers to an inclusive or and not to an exclusive or. For example, a condition A or B is satisfied by any one of the following: A is true (or present) and B is false (or not present), A is false (or not present) and B is true (or present), and both A and B are true (or present).

[0133] In addition, use of the “a” or “an” are employed to describe elements and components of the embodiments herein. This is done merely for convenience and to give a general sense of the description. This description, and the claims that follow, should be read to include one or at least one and the singular also may include the plural unless it is obvious that it is meant otherwise.

[0134] Upon reading this disclosure, those of skill in the art will appreciate still additional alternative structural and functional designs through the principles disclosed herein. Therefore, while particular embodiments and applications have been illustrated and described, it is to be understood that the disclosed embodiments are not limited to the precise construction and components disclosed herein. Various modifications, changes and variations, which will be apparent to those skilled in the art, may be made in the arrangement, operation and details of the method and apparatus disclosed herein without departing from the spirit and scope defined in the appended claims.

[0135] The patent claims at the end of this patent application are not intended to be construed under 35 U.S.C. § 112(f) unless traditional means-plus-function language is expressly recited, such as “means for” or “step for” language being explicitly recited in the claim(s).

Claims

CLAIMSWhat is claimed is:

1. A system for Al-based protein engineering, the system comprising: one or more processors; and one or more memories communicatively coupled with the one or more processors and storing computer-executable instructions thereon, that when executed, cause the one or more processors to: receiving an input prompt including characteristics associated with a protein, inputting the input prompt into a large language model (LLM) that is configured to output a text embedding that is aligned with one or more protein embeddings output by a protein language model (pLM), determining an adjusted text embedding based on the text embedding and the one or more protein embeddings, generating, using a conditional generative model, a protein sequence based on the adjusted text embedding, and causing at least one of: (i) the protein sequence or (ii) a graphic visualization of the protein sequence to be displayed for viewing by a user.

2. The system of claim 1, wherein the computer-executable instructions, when executed, further cause the one or more processors to generate the protein sequence by: analyzing, by a multilayer perceptron model, the adjusted text embedding to determine a modified text representation; and inputting the adjusted text embedding into the conditional generative model to generate the protein sequence.

3. The system of claim 1, wherein the computer-executable instructions, when executed, further cause the one or more processors to generate the protein sequence by: analyzing, with a multilayer perceptron model, a time encoding to determine a denoising time value; andinputting the denoising time value into the conditional generative model to generate the protein sequence.

4. The system of claim 1, wherein the computer-executable instructions, when executed, further cause the one or more processors to generate the protein sequence by: converting, by a tokenizer module, a corrupted protein sequence and one or more starting tokens into continuous vector tokens; and inputting the continuous vector tokens and a positional encoding into the conditional generative model to generate the protein sequence.

5. The system of claim 1, wherein the computer-executable instructions, when executed, further cause the one or more processors to: determine a similarity value between the protein sequence and at least one homolog protein sequence; and update the conditional generative model based on the similarity value.

6. The system of claim 1, wherein the LLM is a Bidirectional Encoder Representations from Transformers (BERT)-based model, and wherein the computer-executable instructions, when executed, further cause the one or more processors to: train the LLM to output text embeddings based on a training loss associated with reconstructing masked tokens.

7. The system of claim 1, wherein the pLM is an evolutionary scaling model (ESM2) configured to reconstruct masked amino acid tokens.

8. The system of claim 1, wherein the computer-executable instructions, when executed, further cause the one or more processors to: train the LLM and the pLM in tandem for the LLM to output text embeddings aligned with protein embeddings output by the pLM based on (i) a global contrastive loss, (ii) a protein family contrastive loss, (iii) one or more masked language losses.

9. The system of claim 1 , wherein the computer-executable instructions, when executed, further cause the one or more processors to: train the conditional generative model to receive adjusted text embeddings as input and output a protein sequence by denoising a corrupted protein sequence.

10. The system of claim 1, wherein the computer-executable instructions, when executed, further cause the one or more processors to determine the adjusted text embedding using a multilayer perceptron model that is configured to more closely align the text embedding with the one or more protein embeddings.

11. The system of claim 1, wherein the conditional generative model is an orderagnostic autoregressive diffusion model.

12. The system of claim 1, wherein the computer-executable instructions, when executed, further cause the one or more processors to: transmit a control instruction to a protein manufacturing system to manufacture a protein represented by the protein sequence.

13. A computer-implemented method for Al-based protein engineering, the computer- implemented method comprising: receiving, at one or more processors, an input prompt including characteristics associated with a protein; inputting, by one or more processors, the input prompt into a large language model (LLM) that is configured to output a text embedding that is aligned with one or more protein embeddings output by a protein language model (pLM); determining, by one or more processors, an adjusted text embedding based on the text embedding and the one or more protein embeddings; generating, by the one or more processors using a conditional generative model, a protein sequence based on the adjusted text embedding; and causing, by one or more processors, at least one of: (i) the protein sequence or (ii) a graphic visualization of the protein sequence to be displayed for viewing by a user.

14. The computer-implemented method of claim 13, wherein generating the protein sequence further comprises: analyzing, by a multilayer perceptron model, the adjusted text embedding to determine a modified text representation; and inputting, by the one or more processors, the adjusted text embedding into the conditional generative model to generate the protein sequence.

15. The computer-implemented method of claim 13, wherein generating the protein sequence further comprises: analyzing, with a multilayer perceptron model, a time encoding to determine a denoising time value; and inputting, by the one or more processors, the denoising time value into the conditional generative model to generate the protein sequence.

16. The computer-implemented method of claim 13, wherein generating the protein sequence further comprises: converting, by a tokenizer module, a corrupted protein sequence and one or more starting tokens into continuous vector tokens; and inputting, by the one or more processors, the continuous vector tokens and a positional encoding into the conditional generative model to generate the protein sequence.

17. The computer-implemented method of claim 13, further comprising: determining, by the one or more processors, a similarity value between the protein sequence and at least one homolog protein sequence; and updating, by the one or more processors, the conditional generative model based on the similarity value.

18. The computer- implemented method of claim 13, wherein the LLM is a Bidirectional Encoder Representations from Transformers (BERT)-based model, and wherein the computer-implemented method further comprises:training, by the one or more processors, the LLM to output text embeddings based on a training loss associated with reconstructing masked tokens.

19. The computer-implemented method of claim 13, further comprising: transmitting, by the one or more processors, a control instruction to a protein manufacturing system to manufacture a protein represented by the protein sequence.

20. A tangible, non-transitory computer-readable medium storing instructions for AI- based protein engineering, that when executed by one or more processors cause the one or more processors to: receive an input prompt including characteristics associated with a protein; input the input prompt into a large language model (LLM) that is configured to output a text embedding that is aligned with one or more protein embeddings output by a protein language model (pLM); determine an adjusted text embedding based on the text embedding and the one or more protein embeddings; generate, using a conditional generative model, a protein sequence based on the adjusted text embedding; and cause at least one of: (i) the protein sequence or (ii) a graphic visualization of the protein sequence to be displayed for viewing by a user.

Citation Information

Patent Citations

  • Embedding-based generative model for protein design

    US20220375538A1

  • Diffusion model for generative protein design

    US20240161864A1

  • Systems and methods for language modeling of protein engineering

    US20240203532A1