Systems and methods for predicting protein expression titers

A machine learning-based system predicts protein expression titer, addressing inefficiencies in candidate protein selection by providing real-time, cost-effective, and accurate analysis for therapeutic protein development.

WO2025160393A1PCT designated stage Publication Date: 2025-07-31REGENERON PHARMACEUTICALS INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/012949
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-26
Filing Date
2025-01-24
Publication Date
2025-07-31

AI Technical Summary

Technical Problem

The current process for selecting candidate proteins for therapeutic applications is time-consuming, labor-intensive, and costly, with limited data availability and qualitative/quantitative analysis, leading to inefficiencies in the manufacturability assessment.

Method used

A system utilizing supervised machine learning models predicts protein expression titer based on primary structure data, isotype, and manufacturability parameters, integrating binary space representations and historical data to classify and re-train models for accurate protein selection.

Benefits of technology

Enables real-time estimation of protein expression titer, improves qualitative and quantitative analysis, reduces processing time and costs, and facilitates effective candidate protein selection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025012949_31072025_PF_FP_ABST
    Figure US2025012949_31072025_PF_FP_ABST
Patent Text Reader

Abstract

Systems and methods for predicting expression titer for a protein are disclosed. A system in accordance with the present disclosure comprises a memory and a processor configured to receive structure data for the protein, the structure data comprising a primary structure of the protein. The processor is configured to generate a binary space representation of the structure data, the binary space representation based on a classification of each element of the primary structure, and apply the binary space representation to a supervised machine learning model. The processor is further configured receive one or more outputs from the model, at least one of the one or more outputs including an expression titer classification for the protein. The processor is further configured to update the historical data to include the structure data and the one or more outputs and re-train the model using the updated historical data.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Systems And Methods For Predicting Protein Expression Titers

[0002] Field

[0003] The present disclosure generally relates to predictive modeling and, more particularly, to systems and methods for predicting protein expression titers.

[0004] Background

[0005] Therapeutic protein drugs are an important class of medicine which are used to treat a wide variety of clinical indications, including cancers, autoimmunity / inflammation, exposure to infectious agents, and genetic disorders. More particularly, therapeutic protein drugs replace a protein that is deficient or abnormal, augment an existing pathway, provide a novel function or activity, interfere with a molecule or organism, or deliver a payload such as a radionuclide, cytotoxic drug, or protein effector.

[0006] In general, potential candidate proteins are identified and studied, and the best candidate proteins progress though preclinical and clinical research phases. To proactively minimize potential clinical failure, candidate proteins must undergo a manufacturability assessment. A manufacturability assessment offers the opportunity to evaluate the viability of a candidate protein at the industrial scale. More particularly, a manufacturability assessment may evaluate a candidate protein's physical and chemical parameters, including, but not limited to, viscosity, solubility, critical residuals, aggregate, immunogenicity, and degradation.

[0007] Therefore, the identification and selection of candidate proteins is an important aspect of the development of effective treatments for a wide variety of diseases. Currently, the candidate protein selection process is informed predominately by experimental development activities, require extensive resources, and is very costly. As such, more efficient, accurate, and cost-effective systems and methods for candidate protein selection are desirable.

[0008] Summary

[0009] The present disclosure provides systems and methods for predicting expression titer for selection of candidate proteins. In some embodiments, a system for predicting expression titer for a protein is provided. The system may include at least one memory storing computer-executable instructions and at least one processor in communication with the at least one memory. The at least one processor is configured to execute the computerexecutable instructions to: (1) receive structure data for the protein, the structure data comprising a primary structure of the protein; (2) generate a binary space representation of the structure data, the binary space representation based on a classification of each element of the primary structure; (3) apply one or more inputs to a supervised machine learning model, the one or more inputs comprising the binary space representation, the model being previously trained using historical data, the historical data comprising structure data for a plurality of proteins and their corresponding expression titers; (4) receive one or more outputs from the model, at least one of the one or more outputs including an expression titer classification for the protein; (5) update the historical data to include the structure data of the protein and the corresponding one or more outputs; and (6) re-train the model using the updated historical data. If the expression titer classification for the protein is above a predetermined threshold, the protein may be selected for further development, production, optimization, purification, and / or other type of processing. If the expression titer classification does not meet the predetermined threshold, it may not be pursued further. The system may include additional, less, or alternate functionality, including that discussed elsewhere herein.

[0010] The present disclosure also provides a computer-implemented method for predicting expression titer for a protein. The method may be implemented using a system including a computing device including a processor communicatively coupled to a memory device. Additionally, or alternatively, the computer-implemented method may be implemented via one or more local or remote processors, servers, transceivers, memory units, mobile devices, wearables, smart watches, smart contact lenses, smart glasses, augmented reality glasses, virtual reality headsets, mixed or extended reality glasses or headsets, voice or chat bots, Artificial intelligence (Al) bots (including generative Al bots), and / or other electronic or electrical components, which may be in wired or wireless communication with one another. The method comprises (1) receiving structure data for the protein, the structure data comprising a primary structure of the protein; (2) generating a binary space representation of the structure data, the binary space representation based on a classification of each element of the primary structure; (3) applying one or more inputs to a supervised machine learning model, the one or more inputs comprising the binary space representation, the model being previously trained using historical data, the historical data comprising structure data for a plurality of proteins and their corresponding expression titers; (4) receiving one or more outputs from the model, at least one of the one or more outputs including an expression titer classification for the protein; (5) updating the historical data to include the structure data of the protein and the corresponding one or more outputs; and (6) re-training the model using the updated historical data. The method may include additional, less, or alternate functionality, including those discussed elsewhere herein.

[0011] The present disclosure provides at least one non-transitory computer-readable storage media having computer-executable instructions embodied thereon. The computerexecutable instructions, when executed by at least one processor, cause the at least one processor to: (1) receive structure data for the protein, the structure data comprising a primary structure of the protein; (2) generate a binary space representation of the structure data, the binary space representation based on a classification of each element of the primary structure; (3) apply one or more inputs to a supervised machine learning model, the one or more inputs comprising the binary space representation, the model being previously trained using historical data, the historical data comprising structure data for a plurality of proteins and their corresponding expression titers; (4) receive one or more outputs from the model, at least one of the one or more outputs including an expression titer classification for the protein; (5) update the historical data to include the structure data of the protein and the corresponding one or more outputs; and (6) re-train the model using the updated historical data. The storage medium may include additional, less, or alternate actions, including those discussed elsewhere herein. If the expression titer classification for the protein is above a predetermined threshold, the protein may be selected for further development, production, optimization, purification, and / or other type of processing. If the expression titer classification does not meet the predetermined threshold, it may not be pursued further and / or may be discarded. Advantages will become more apparent to those skilled in the art from the following description of the preferred embodiments which have been shown and described by way of illustration. As will be realized, the present embodiments may be capable of other and different embodiments, and their details are capable of modification in various respects. Accordingly, the drawings and description are to be regarded as illustrative in nature and not as restrictive.

[0012] Brief Description Of The Drawings

[0013] The Figures described below depict various aspects of the systems and methods disclosed therein. It should be understood that each Figure depicts an embodiment of a particular aspect of the disclosed systems and methods, and that each of the Figures is intended to accord with a possible embodiment thereof. Further, wherever possible, the following description refers to the reference numerals included in the following Figures, in which features depicted in multiple Figures are designated with consistent reference numerals.

[0014] There are shown in the drawings arrangements which are presently discussed, it being understood, however, that the present embodiments are not limited to the precise arrangements and are instrumentalities shown, wherein:

[0015] Figure 1 illustrates a block diagram of an example computer system in accordance with an embodiment of the present disclosure.

[0016] Figure 2 illustrates a block diagram of an example titer prediction computing device that may be used with the example computer system illustrated in Figure 1.

[0017] Figure 3 illustrates a block diagram of an artificial intelligence (Al) / deep learning (DL) module that may be used with the expression titer prediction computing device illustrated in Figure 2.

[0018] Figure 4 illustrates an example process for generating a binary space representation of a primary amino acid sequence in accordance with an embodiment of the present disclosure.

[0019] Figure 5 illustrates a flow chart of an example computer-implemented process that may be carried out by the example computer system shown in Figure 1. Figure 6 illustrates an example server computing device that may be used with the computer system shown in Figure 1.

[0020] Figure 7 illustrates an example user computing device that may be used with the computer system shown in Figure 1.

[0021] The Figures depict preferred embodiments for purposes of illustration only. One skilled in the art will readily recognize from the following discussion that alternative embodiments of the systems and methods illustrated herein may be employed without departing from the principles of the aspects described herein.

[0022] Description Of Embodiments

[0023] The present embodiments may relate to, inter alia, systems and methods for predicting expression titer for a protein. For example, in some embodiments, the system for predicting expression titer for a protein comprises at least one memory storing computerexecutable instructions and at least one processor in communication with the at least one memory. The at least one processor is configured to: (1) receive structure data for the protein, the structure data comprising a primary structure ( / .e., a primary amino acid sequence) of the protein; (2) generate a binary space representation of the structure data, the binary space representation based on a classification of each element of the primary structure; (3) apply one or more inputs to a supervised machine learning model, the one or more inputs comprising the binary space representation, the model being previously trained using historical data, the historical data comprising structure data for a plurality of proteins and their corresponding expression titers; (4) receive one or more outputs from the model, at least one of the one or more outputs including an expression titer classification for the protein; (5) update the historical data to include the structure data of the protein and the corresponding one or more outputs; and (6) re-train the model using the updated historical data. If the expression titer classification for the protein is above a predetermined threshold, the protein may be selected for further development, production, optimization, purification, and / or other type of processing. If the expression titer classification does not meet the predetermined threshold, it may not be pursued further and / or may be discarded. The systems and methods described herein may include additional, less, or alternate functionality, including that discussed elsewhere herein. Although the present disclosure discusses predicting expression titer for a protein, the systems and methods disclosed herein may be used to predict expression titer in other therapeutic contexts. For example, the systems and methods disclosed herein may be used for nucleic acid-based therapeutics, such as mRNA therapeutics or adeno-associated viruses (AAVs). In such embodiments, the structure data may comprise nucleic acid data.

[0024] "Expression titer," as used herein, may refer generally to a measurement of the quantity, yield, amount, or concentration of a substance, such as a protein (e.g., a recombinant therapeutic protein), virus, or nucleic acid, in a sample. In some cases, the substance may be a product of an engineered cell line.

[0025] "Protein expression titer", as used herein, may refer to the expression titer of a specific protein in a sample. In various aspects, the protein may be any protein that is identified as a therapeutic candidate. In some aspects, the protein may be an antibody, such as a monoclonal antibody, and can encompass a human monoclonal antibody.

[0026] "Monoclonal antibody," as used herein, is not limited to antibodies produced through hybridoma technology. A monoclonal antibody can be derived from a single clone, including any eukaryotic, prokaryotic, or phage clone, by any means available or known in the art. Monoclonal antibodies useful with the present disclosure can be prepared using a wide variety of techniques known in the art including the use of hybridoma, recombinant, and phage display technologies, or a combination thereof.

[0027] "Amino acid sequence," as used herein, may refer generally to the order of amino acids as they occur in a polypeptide chain. A "primary structure" or a "primary amino acid sequence" may refer to the linear sequence of amino acids beginning at the amino-terminal (N) end and ending at the carboxyl-terminal (C). The sequence of a protein is usually notated as a string of letters, according to the order of the amino acids from the aminoterminal to the carboxyl-terminal of the protein. Either a single or three-letter code may be used to represent each amino acid in the sequence. There are 20 amino acids that occur naturally in nature, which can be represented by a three or single letter code as follows: Alanine (Ala, A); Arginine (Arg, R); Asparagine (Asn, N); Aspartic acid (Asp, D); Cysteine (Cys, C); Glutamic acid (Glu, E); Glutamine (Gin, Q); Glycine (Gly, G); Histidine (His, H); Isoleucine (lie, I); Leucine (Leu, L); Lysine (Lys, K); Methionine (Met, M); Phenylalanine (Phe, F); Proline (Pro, P); Serine (Ser, S); Threonine (Thr, T); Tryptophan (Trp, W); Tyrosine (Tyr, Y); Valine (Vai, V).

[0028] "Isotype," as used herein, may refer generally to a class of antibody that is determined by its respective heavy-chain constant region. Isotypes are defined by slight variations in structure, most predominantly number and location of disulfide bonds. Isotypes may additionally have distinct "signatures" of amino acid sequences, typically in the Fc region. In general, there are five antibody isotypes that each have a unique heavychain constant region: IgM, IgD, IgG, IgE, and IgA. An antibody's isotype may be important in determining which immune cells and molecules are recruited by the antibody to help destroy and remove a pathogen. In general, IgG is the most prevalent isotype in blood, but is also found in tissues. IgG is the predominant isotype during a secondary immune response, which is when the immune system encounters an antigen for a second or subsequent time. It is expressed as a monomer with a valency of two. There are four IgG subclasses, numbered based on their abundance in blood: IgGl, lgG2, lgG3, and lgG4.

[0029] "Cell line," as used herein, may refer generally to a defined population of cells that may be maintained in culture for an extended period of time, retaining stability of certain phenotypes and functions. In general, cell lines are clonal, meaning that the entire population originated from a single common ancestor cell. "Host cell line," as used herein, may refer to a cell line produced or encoded by host organisms used to produce recombinant therapeutic proteins. A specific host cell line may be designated by an identifier, referred to as a "host cell line identifier" herein.

[0030] "App," as used herein, may refer generally to a software application installed and downloaded on a user computing device and executed to provide an interactive graphical user interface at the user computing device. An app associated with the computer system, as described herein, may be understood to be maintained by the computer system and / or one or more components thereof. Accordingly, a "maintaining party" of the app may be understood to be responsible for any functionality of the app and may be considered to instruct other parties / components to perform such functions via the app. In addition to the structure data, the computer system may receive and analyze additional data, such as isotype data, host cell line identifier data, and / or manufacturability and / or developability data. In the example embodiment, the computer system may capture and synthesize this data, leveraging machine learning and / or artificial intelligence, to predict expression titers. The computer system may include any suitable data storage capabilities, such as cloud storage, to access and / or store any of the above data. In that way, the computer system may access and analyze historical and / or current (e.g., real-time or near real-time) data. In some embodiments, the computer system may include or be implemented via one or more local or remote processors, servers, transceivers, memory units, mobile devices, wearables, smart watches, smart contact lenses, smart glasses, augmented reality glasses, virtual reality headsets, mixed or extended reality glasses or headsets, voice or chat bots, Al bots (including generative Al bots), and / or other electronic or electrical components, which may be in wired or wireless communication with one another.

[0031] In the example embodiment, the computer system includes at least one computing device. The computing device is configured to perform the functions that may be more generally described herein as being performed by and / or attributed to the overall computer system.

[0032] In particular, in the example embodiment, the computing device may be in communication with one or more computing devices associated with a user. These computing devices may include a personal mobile computing device, such as a smart phone, tablet, and the like.

[0033] The computing device may receive data from the computing device(s), including structure data, isotype data, host cell line identifier data, manufacturability and / or developability data, and / or any other type of data associated with a protein. The structure data may comprise a primary structure (i.e., a primary amino acid sequence) of the protein. In some embodiments, the isotype data may comprise the isotype of the protein (e.g., IgGl, lgG4, etc.), which may be used as a categorical feature (e.g., 1 or 4) to inform on the titer value or class prediction. Titer may vary broadly based on isotype. Additionally, or alternatively, the entire amino acid sequence of the protein is received and analyzed to determine the protein's sequence signatures. A specific host cell line may be engineered to produce a relatively higher level of titer with each iteration, and therefore, the host cell line identifier data may be used to inform the titer value or class prediction. For example, one host cell line may produce around half the titer, on average, than another engineered host cell line. The manufacturability and / or developability data may include one or more of the following parameters: viscosity, melting transition temperatures, isoelectric point, hydrophobic character, net charge, dipole moment, molecular purity by size analysis, molecular purity by charge analysis, particle size, and / or aggregation propensity. However, this list is meant to be illustrative and not exhaustive.

[0034] The computing device may receive portions of structure data, isotype data, host cell line identifier data, manufacturability and / or developability data, and / or any other type of data associated with a protein from alternative computing devices, such as sequencing devices. Additionally, or alternatively, the computing device may access portions of such data from one or more databases or other memory devices. The computing device may be configured to aggregate, combine, synthesize, parse, compare, and / or otherwise process this data, as described in more detail herein, in order to predict an expression titer and / or expression titer classification for a protein.

[0035] The computing device may store any received, retrieved, and / or accessed data in one or more databases, and may store the expression titer and / or expression titer classification and / or other generated data in the one or more databases. A database may be any suitable storage location, and may in some embodiments include a cloud storage device such that the database may be accessed by a plurality of computing devices (e.g., a plurality of user analytics computing devices, third-party computing devices, etc.). The database may be integral to the computing device or may be remotely located with respect thereto.

[0036] The resulting expression titer and / or expression titer classification for a protein may be used in the process of selecting a candidate protein, such as a candidate antibody. More particularly, the expression titer data may provide important insights into the developability and manufacturability of a specific protein, and / or specific antibody. As an example, the model may predict that one isotype (e.g., IgGl), in a particular host cell line, is associated with a lower titer as compared to other isotypes (e.g., lgG4). In one embodiment, if the resulting expression titer is above a predetermined threshold and / or if the resulting expression titer classification is a certain classification (e.g., "high titer class" and / or "medium titer class"), the protein may undergo further development, production, optimization, purification, and / or other type of processing.

[0037] The technical effect of the systems and processes described herein may be achieved by performing at least one of the following steps: (i) receive structure data for the protein, the structure data comprising a primary structure of the protein; (ii) generate a binary space representation of the structure data, the binary space representation based on a classification of each element of the primary structure; (iii) apply one or more inputs to a supervised machine learning model, the one or more inputs comprising the binary space representation, the model being previously trained using historical data, the historical data comprising structure data for plurality of proteins and their corresponding expression titers; (iv) receive one or more outputs from the model, at least one of the one or more outputs including an expression titer classification for the protein; (v) update the historical data to include the structure data of the protein and the corresponding one or more outputs; and (vi) re-train the model using the updated historical data.

[0038] At least one of the technical problems addressed by the systems and methods disclosed herein may include: (i) time-consuming, labor-intensive, and costly process for selecting a candidate protein; (ii) limited available data in the candidate protein selection process; and (iii) limited qualitative and quantitative analysis of proteins in the candidate protein selection process.

[0039] The resulting technical effects may include, for example: (i) ability to estimate expression titer of a protein in real-time; (ii) selection of candidate proteins with characteristics likely to result in an effective treatment; (iii) improved qualitative and quantitative analysis of candidate proteins; (iv) ability to receive data from a plurality of different data sources and translating the data into a format for further analysis; and (v) reduced processing time and costs associated with the candidate protein selection process.

[0040] Example computer systems for predicting expression titer for a protein are disclosed herein. For example, Figure 1 depicts a schematic diagram of an example computer system 100. Computer system 100 is configured to predict expression titer for a protein. In one embodiment, computer system 100 may include and / or facilitate communication between a titer prediction computing device 110 and one or more user computing devices 130 (which may also be referred to as "mobile devices") and / or between titer prediction computing device 110 and one or more of third-party devices 140 and / or titer prediction servers 150.

[0041] Titer prediction computing device 110 may be implemented as a server computing device with artificial intelligence and deep learning functionality. Alternatively, titer prediction computing device 110 (and / or user computing devices 130) may be implemented as any device capable of interconnecting to the Internet, including mobile computing device or "mobile device," such as a smartphone, a "phablet," or other web-connectable equipment or mobile devices (such as one or more local or remote processors, servers, transceivers, memory units, mobile devices, wearables, smart watches, smart contact lenses, smart glasses, augmented reality glasses, virtual reality headsets, mixed or extended reality glasses or headsets, voice or chat bots, Al bots (including generative Al bots), and / or other electronic or electrical components, which may be in wired or wireless communication with one another).

[0042] Titer prediction computing device 110 may be in communication with one or more user computing devices 130, third party devices 140, and / or titer prediction server 150, such as via wireless communication or data transmission over one or more radio frequency links or wireless communication channels. In the example embodiment, components of computer system 100 may be communicatively coupled to the Internet through many interfaces including, but not limited to, at least one of a network, such as the Internet, a local area network (LAN), a wide area network (WAN), or an integrated services digital network (ISDN), a dial-up-connection, a digital subscriber line (DSL), a cellular telecommunications connection (e.g., a 3G, 4G, 5G, etc., connection), a cable modem, and a BLUETOOTH connection.

[0043] Computer system 100 also includes one or more database(s) 120 containing information on a variety of matters. For example, database 120 may include such information as protein sequence data, isotype data, host cell line identifier data, manufacturability and / or developability data, expression titer data and / or any other information used, received, and / or generated by computer system 100 and / or any component thereof, including such information as described herein. In one embodiment, database 120 may include a cloud storage device, such that information stored thereon may be securely stored but still accessed by one or more components of computer system 100, such as, for example, titer prediction computing device 110, user computing devices 130, and / or titer prediction servers 150. In some embodiments, database 120 may be stored on titer prediction computing device 110. In an alternative embodiment, database 120 may be stored remotely from titer prediction computing device 110 and may be non-centralized.

[0044] In some embodiments, user computing devices 130 may be computers that include a web browser or a software application to enable user computing devices 130 to access the functionality of titer prediction computing device 110 using the Internet or a direct connection, such as a cellular network connection. User computing devices 130 may be any device capable of accessing the Internet including, but not limited to, a desktop computer, a mobile device (e.g., a laptop computer, a personal digital assistant (PDA), a cellular phone, a smartphone, a tablet, a phablet, netbook, notebook, smart watches or bracelets, smart glasses, wearable electronics, pagers, virtual reality headsets, augmented reality glasses, voice or chat bots, wearables, etc.), or other web-based connectable equipment.

[0045] User computing devices 130 may be used to access a data management app 112 maintained by titer prediction computing device 110, for example, via a user interface 132 when data management app 112 is executed on user computing device 130. A user may use data management app 112 to provide inputs to titer prediction computing device 110, view predictions generated by titer prediction computing device 110, and perform other actions, including those described elsewhere herein.

[0046] Third party devices 140 may be computing devices associated with external sources of data. Titer prediction computing device 110 may request, receive, and / or otherwise access data from third party devices 140. Third party devices 140 may be any devices capable of interconnecting to the Internet, including a server computing device, a mobile computing device or "mobile device," such as a smartphone, or other web-connectable equipment or mobile devices. Example user analytics computing devices are disclosed herein. For example, Figure 2 depicts titer prediction computing device 110 (as shown in Figure 1), according to an embodiment. In some embodiments, titer prediction computing device 110 may include a processor 202, a memory 204 (which may be similar to database 120, also shown in Figure 1), a communication interface 206, and a storage interface 208. Processor 202 is configured to execute instructions, which may be stored in memory 204. Processor 202 includes one or more processing units (e.g., in a multi-core configuration) and may be configured to execute a plurality of modules.

[0047] In some embodiments, processor 202 is operable to execute an artificial intelligence / deep learning (AI / DL) module 210, an expression titer module 212, and a module 214 that maintains functionality for data management app 112 (shown in Figure 1). Modules 210, 212, and 214 may include specialized instruction sets, and / or coprocessors. Database 120 and / or memory 204 may store any data and / or instructions necessary for modules 210, 212, and 214 to function as described herein. In the example embodiment, database 120 may store protein sequence data 220, isotype data 222, host cell line identifier data 224, manufacturability and / or developability data 226, expression titer data 228 and / or any other information used, received, and / or generated by titer prediction computing device 110.

[0048] AI / DL module 210 may execute artificial intelligence and / or deep learning functionality on behalf of expression titer module 212. Specifically, AI / DL module 210 may include any rules, algorithms, training data sets / programs, and / or any other suitable data and / or executable instructions that enable titer prediction computing device 110 employ artificial intelligence and / or deep learning to predict expression titer for proteins.

[0049] Figure 3 depicts an AI / DL module 210 (as shown in Figure 2), according to an embodiment. In some embodiments, AI / DL module 210 includes a training set builder module 302 programmed to submit one or more queries to database 120 (shown in Figures 1 and 2) to retrieve data and / or subsets of data, and to use those subsets of data to build training data sets for generating predictive model 308.

[0050] In example embodiments, training set builder module 302 is programmed to retrieve training data sets from the retrieved subsets of data. Each training data set corresponds to historical data, which may include one or more of structure data (e.g., a primary amino acid sequence of the protein), isotype data, host cell line identifier data, and / or manufacturability and / or developability data (e.g., sequence data 220, isotype data 222, host cell line identifier data 224, and manufacturability and / or developability data 226, respectively) for a protein, and the corresponding expression titer (e.g., expression titer data 228) which was previously determined, as opposed to completed in real-time with respect to the time of retrieval by training set builder module 302. Each training data set can include model input data along with result data representing an expression titer and / or an expression titer classification. The model input data can represent factors that may be expected to, or unexpectedly be found during model training to, have some correlation with the expression titer. In some embodiments, the model input data may comprise one or more of structure data (e.g., a primary amino acid sequence of the protein), isotype data, host cell line identifier data, and / or manufacturability and / or developability data (e.g., sequence data 220, isotype data 222, host cell line identifier data 224, and manufacturability and / or developability data 226) for a protein. For example, a specific host cell line may be engineered to produce a relatively higher level of titer with each iteration, and therefore, the host cell line identifier data may be used to inform the titer value or class prediction.

[0051] After training set builder module 302 generates training data sets, it passes the training data sets to model trainer module 304, which is programmed to apply the model input data fields of each training data set as inputs to one or more machine learning models. Each of the one or more machine learning models is programmed to produce, for each training data set, at least one output intended to correspond to, or "predict," a value of the at least one result data field of the training data set. Machine learning techniques may be used to train the model to identify and recognize patterns in existing data in order to facilitate making predictions for subsequent new input data. For example, support vector machines (SVM), neural networks, and multilayer perceptron (MLP) classifiers may be used to train and optimize the model.

[0052] Model trainer module 304 is programmed to compare, for each training data set, the at least one output of the model to the at least one result data field of the training data set, and apply a machine learning algorithm to adjust parameters of the model in order to reduce the difference or "error" between the at least one output and the corresponding at least one result data field. In this way, model trainer module 304 trains the machine learning model to accurately predict expression titer for inputs. In other words, model trainer module 304 cycles the one or more machine learning models through the training data sets, causing adjustments in the model parameters, until the error between the at least one output and the expression titer falls below a suitable threshold, and then uploads at least one trained machine learning model to predictive model module 308 for application to new structure data (e.g., a new primary amino acid sequence). For example, model trainer module 304 may be programmed to compare, for each training data set, an expression titer for a protein as determined by the model to an expression titer for the protein as previously determined using traditional testing methods. Model trainer module 304 may then adjust one or more weight values using a machine learning algorithm, such as a backpropagation algorithm, in order to reduce the difference between the expression titer as determined by the model and the expression titer as determined using traditional testing methods.

[0053] In some embodiments, the one or more machine learning models may include one or more neural networks, such as a convolutional neural network, a deep learning neural network, or the like. The neural network may have one or more layers of nodes, and the model parameters adjusted during training may be respective weight values applied to one or more inputs to each node to produce a node output. In other words, the nodes in each layer may receive one or more inputs and apply a weight to each input to generate a node output. The node inputs to the first layer may correspond to the model input data fields and the node outputs of the final layer may correspond to the at least one output of the model, intended to predict the at least one result data field. For example, the node inputs to the first layer may correspond to structure data, isotype data, host cell line identifier data, and / or manufacturability and / or developability data for at least one protein and the node outputs of the final layer may correspond to at least one expression titer for the at least one protein. One or more intermediate layers of nodes may be connected between the nodes of the first layer and the nodes of the final layer. As model trainer module 304 cycles through the training data sets, model trainer module 304 applies a suitable backpropagation algorithm to adjust the weights in each node layer to minimize the error between the at least one output (e.g., an expression titer for a protein as determined by the model) and the corresponding result data field (e.g., the previously determined expression titer for the protein). In this fashion, the machine learning model is trained to produce one or more outputs which reliably predicts protein expression titer. Alternatively, the machine learning model has any suitable structure. In some embodiments, model trainer module 304 provides an advantage by automatically discovering and properly weighting complex, second- or third-order, and / or otherwise nonlinear interconnections between the model input data fields and the at least one output. Absent the machine learning model, such connections are unexpected and / or undiscoverable by human analysts.

[0054] Additionally, or alternatively, the one or more machine learning models may include one or more multilayer perceptron (MLP) classifiers. A MLP classifier may comprise input and output layers, and one or more hidden layers with many neurons stacked together. For example, an MLP classifier in accordance with the present disclosure may comprise an input layer comprising structure data, isotype data, host cell line identifier data, and / or manufacturability and / or developability data and an output layer comprising an expression titer prediction.

[0055] Additionally, or alternatively, the one or more machine learning models may include one or more support vector machines (SVMs). SVMs are supervised learning models with associated learning algorithms that analyze data for classification and regression analysis. More particularly, a SVM constructs a hyperplane or set of hyperplanes in a high or infinite-dimensional space, which can be used for classification, regression, or other tasks like outlier detection. For example, in some embodiments, training data sets may each be marked as belonging to one of two categories based on the previously determined expression titer (e.g., expression titers less than or equal to 200 pg / mL may be classified as "low titer class" and expression titers greater than 200 pg / mL be classified as "high titer class"). The SVM maps training data sets to points in space so as to maximize the width of the gap between the two categories. New data sets are then mapped into that same space and predicted to belong to a category based on which side of the gap they fall. For example, new structure data is then mapped into the space and predicted to belong to a "high titer class" or a "low titer class."

[0056] In some embodiments, one output may be formatted as a predicted expression titer class (e.g., low, medium, high, etc.). For example, in some embodiments, expression titers less than or equal to 200 pg / mL may be classified as "low titer class" and expression titers greater than 200 pg / mL be classified as "high titer class" and at least one output of the trained model is "low titer class" or "high titer class." In one embodiment, if the resulting expression titer is above a predetermined threshold and / or if the resulting expression titer classification is a certain classification (e.g., "high titer class" and / or "medium titer class"), the protein may undergo further development, production, optimization, purification, and / or other type of processing.

[0057] In some embodiments, predictive model module 308 compares the known expression titer for the two proteins with the output from the trained model, and routes the comparison result to a model updater module 306 of AI / DL module 210. Model updater module 210 is programmed to derive a correction signal from the comparison results, and to provide correction signal to model trainer module 304 to enable updating or "re-training" of the at least one machine learning model to improve performance. For example, one or more new weight values may be derived from the comparison results, and the correction signal may adjust weight values applied to one or more inputs. The re-trained machine learning model may be periodically re-uploaded to predictive model module 308.

[0058] In some embodiments, model trainer module 304 may update the training dataset by creating one or more new historical records which includes new data and re-training the operator model using the updated training dataset, further improving the accuracy of the operator model.

[0059] Expression titer module 212 may employ AI / DL module 210 to use the trained model to predict an expression titer for a protein. More particularly, expression titer module 212 may use the output from the trained model to predict an expression titer for a protein. The predicted expression titer and other data may be viewable via data management app 112. App module 214 is configured to facilitate maintaining data management app 112 and providing the functionality thereof to users. App module 214 may store instructions that enable the download and / or execution of data management app 112 at user computing devices 130. App module 214 may store instructions regarding user interfaces, controls, commands, settings, and the like, and may format data into a format suitable for transmitting to user computing devices 130 for display thereof.

[0060] In some embodiments, processor 202 is operatively coupled to communication interface 206 such that titer prediction computing device 110 is capable of communicating with remote device(s) such as user computing devices 130, third party devices 140, and / or titer prediction servers 150 (all shown in Figure 1) over a wired or wireless connection. For example, communication interface 206 may receive sequence data, isotype data, host cell line identifier data, and / or manufacturability and / or developability data, expression titer data, and the like, from user computing devices 130, and / or third-party devices 140. Communication interface 206 may include, for example, a wired or wireless network adapter and / or a wireless data transceiver for use with a mobile telecommunications network.

[0061] Processor 202 may also be operatively coupled to database 120 (and / or any other storage device) via storage interface 208. Database 120 may be any computer-operated hardware suitable for storing and / or retrieving data. In some embodiments, database 120 may be integrated in titer prediction computing device 110. For example, titer prediction computing device 110 may include one or more hard disk drives as database 120. In other embodiments, database 120 is external to titer prediction computing device 110 and is accessed by a plurality of computer devices. For example, database 120 may include a storage area network (SAN), a network attached storage (NAS) system, multiple storage units such as hard disks and / or solid-state disks in a redundant array of inexpensive disks (RAID) configuration, cloud storage devices, and / or any other suitable storage device.

[0062] Storage interface 208 may be any component capable of providing processor 202 with access to database 120. Storage interface 208 may include, for example, an Advanced Technology Attachment (ATA) adapter, a Serial ATA (SATA) adapter, a Small Computer System Interface (SCSI) adapter, a RAID controller, a SAN adapter, a network adapter, and / or any component providing processor 202 with access to database 120.

[0063] Processor 202 may execute computer-executable instructions for implementing aspects of the disclosure. In some embodiments, processor 202 may be transformed into a special purpose microprocessor by executing computer-executable instructions or by otherwise being programmed. For example, processor 202 may be programmed with the instructions such as those illustrated in Figure 5.

[0064] Memory 204 may include, but is not limited to, random access memory (RAM) such as dynamic RAM (DRAM) or static RAM (SRAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), and non-volatile RAM (NVRAM). The above memory types are example only, and are thus not limiting as to the types of memory usable for storage of a computer program.

[0065] Example binary space representations of primary amino acid sequences are disclosed herein. In some embodiments, processor 202 (shown in Figure 2) is operable to receive structure data for a protein (e.g., a primary amino acid sequence of a protein) and generate a binary space representation of the structure data, as described in "Predicting protein-protein interactions based only on sequences information" (Shen et al., The Proceedings of the National Academy of Sciences (PNAS), March 13, 2007), which is incorporated by reference in its entirety. In some embodiments, the binary space representation may be based on a classification of each element (e.g., amino acid) of the structure data (e.g., primary amino acid sequence) according to a) electrostatic properties and b) hydrophobic properties of the respective element, as described in more detail below.

[0066] In some embodiments, the amino acids may be classified according to their respective dipole scale (e.g., electrostatic interactions) and / or their respective volume scale (e.g., hydrophobic interactions). In the embodiment illustrated in Figure 4, each amino acid is classified according to seven (7) classes of amino acids. For example, a first class 431 comprises amino acids Alanine (A), Glycine (G), Valine (V); second class 432 comprises amino acids Isoleucine (I), Leucine (L), Phenylalanine (F), Proline (P); a third class 433 comprises amino acids Tyrosine (Y), Methionine (M), Threonine (T), Serine (S); a fourth class 434 comprises amino acids Histidine (H), Asparagine (N), Glutamine (Q), Tryptophan (W); a fifth class 435 comprises amino acids Arginine (R), Lysine (K); a sixth class comprises amino acids Aspartic acid (D), Glutamic acid (E); and a seventh class comprises amino acid Cysteine (C), as described in Table 1, below. The classes shown below are by way of example only, and the amino acids may be classified according to various other classifications. For example, in some embodiments, each amino acid is classified according to three (3) classes of amino acids. The number of classes may be defined by a user.

[0067] Table 1. Amino Acid Classifications

[0068] In Table 1, the dipole scale is as follows: refers to a dipole of less than 1.0 Debye (D); "+" refers to a dipole of less than 2.0 D; "++" refers to a dipole of less than 3.0 D; "+++" refers to a dipole of greater than 3.0 D; refers to a dipole of greater than 3.0 with opposite orientation. In Table 1, the volume scale is as follows: refers to a volume of less than 50 A3; and "+" refers to a volume of greater than 50 A3. In Table 1, Cysteine is not part of Class 3 because of its ability to form disulfide bonds.

[0069] Figure 4 depicts a process of generating a binary space representation of a primary amino acid sequence. In the embodiment illustrated in Figure 4, any three continuous amino acids is classified as a unit ( / .e., a triad). In some embodiments, some triads may overlap. Stated another way, an amino acid may be a part of more than one triad. For example, in the embodiment illustrated in Figure 4, triad 410 comprises amino acids 412, 414, and 416 and triad 420 comprises amino acids 414, 416, and 418. Therefore, amino acids 414 and 416 are part of both triad 410 and triad 420. A triad type may be determined for each triad. For example, a first triad type 441 is comprised of three amino acids belonging to first class 431, a second type 442 is comprised of a first amino acid belonging to second class 432 and second and third amino acids belonging to first class 431, and so forth. Triads composed of three amino acids belonging to the same class (e.g., triads having the same triad type) may be treated the same, or substantially the same, as they may be considered to play similar roles for predicting expression titer). This information may be projected into a homogeneous vector space by counting the frequency of each triad type. For example, in the embodiment illustrated in Figure 4, first triad type 441 has a frequency 461 of five (5) in primary amino acid sequence 400.

[0070] The process of generating descriptor vectors first comprises using a binary space (V, F) to represent a protein sequence, where V 450 is the vector space of the sequence features, and each feature (Vi) represents a triad type, F 470 is the frequency vector corresponding to V 450, and the value of the / th dimension of F (f) is the frequency of type V; appearing in the protein sequence. For example, in the embodiment illustrated in Figure 4, where amino acids have been catalogued into seven classes, the size of V 450 should be 343 (7 x 7 x 7).

[0071] Example computer-implemented methods for predicting expression titer for a protein is disclosed herein. For example, Figure 5 depicts a process 500 for predicting expression titer in real-time. Process 500 may be executed by one or more processors, such as processor 202 (shown in Figure 2) of titer prediction computing device 110 (shown in Figures 1 and 2). At 502, structure data for the protein is received, the structure data comprising a primary structure of the protein. At 504, a binary space representation of the primary amino acid sequence is generated. In some embodiments, the binary space representation may be based on a classification of each element (e.g., amino acid) of the primary structure according to electrostatic properties and / or hydrophobic properties of the respective amino acid, as discussed with regards to Figure 4. Next, at 506, the binary space representation is applied to a supervised machine learning model, the model being previously trained using historical data, the historical data comprising structure data for plurality of proteins and their corresponding expression titers. In some embodiments, the isotype and / or host cell line identifier of the protein may also be applied to the supervised machine learning model, in addition to the binary space representation. At 508, one or more outputs may be received from the model. At least one of the one or more outputs includes an expression titer classification for the protein. At 510, the historical data is updated to include the structure data of the protein and the corresponding one or more outputs. At 512, the model is re-trained using the updated historical data. A candidate protein may then be selected from a set of proteins for development, the set of proteins including the candidate protein. The selection is based at least in part on expression titer of the candidate protein. In one embodiment, if the resulting expression titer is above a predetermined threshold and / or if the resulting expression titer classification is a certain classification (e.g., "high titer class" and / or "medium titer class"), the protein may undergo further development, production, optimization, purification, and / or other type of processing.

[0072] Example server computing devices are disclosed herein. For example, Figure 6 is a schematic diagram of an example configuration of a server computing device 600, in accordance with some embodiments of the present disclosure. Server computing devices having an architecture similar to server computing device 600 may be used to implement one or more of the computing systems shown in Figure 1, such as titer prediction computing device 110. In the example embodiment, server computing device 600 includes processor 605 for executing instructions (not shown) stored in a memory 610. In an embodiment, processor 605 may include one or more processing units (e.g., in a multi-core configuration). The instructions may be executed within various different operating systems, such as UNIX®, LINUX® (LINUX is a registered trademark of Linus Torvalds), Microsoft Windows®, etc. It should also be appreciated that upon initiation of a computer-based method, various instructions may be executed during initialization. Some operations may be required in order to perform one or more processes described herein, while other operations may be more general or specific to a particular programming language (e.g., C, C#, C++, Java, or other suitable programming languages, etc.).

[0073] In the example embodiment, processor 605 is operatively coupled to a communication interface 615 such that server computing device 600 is capable of communicating with a remote device, such as a user or system administrator computing system (not shown) or another server computing device 600.

[0074] In the example embodiment, processor 605 is also operatively coupled to a storage device 630, which may be, for example, a computer-operated hardware unit suitable for storing or retrieving data. In some embodiments, storage device 630 is integrated into server computing device 600. For example, device 600 may include one or more hard disk drives as storage device 630. In other embodiments, storage device 630 is external to device 600 and may be accessed by a plurality of server computing devices 600. For example, storage device 630 may include multiple storage units such as hard disks or solid-state disks in a redundant array of inexpensive disks (RAID) configuration. Storage device 630 may include a storage area network (SAN) or a network attached storage (NAS) system. Storage device 630 may be used as a repository for one or more databases or other data structures for storing various data elements received, processed, and / or generated by titer prediction computing device 110 (shown in Figures 1 and 2). For example, storage device 630 may be used to implement database 120 (shown in Figures 1 and 2).

[0075] In some embodiments, processor 605 is operatively coupled to storage device 630 via an optional storage interface 620. Storage interface 620 may include, for example, a component capable of providing processor 605 with access to storage device 630. In one embodiment, storage interface 620 further includes one or more of an Advanced Technology Attachment (ATA) adapter, a Serial ATA (SATA) adapter, a Small Computer System Interface (SCSI) adapter, a RAID controller, a SAN adapter, a network adapter, or a similarly capable component providing processor 605 with access to storage device 630.

[0076] Memory area 610 may include, but is not limited to, random-access memory (RAM) such as dynamic RAM (DRAM) or static RAM (SRAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), non-volatile RAM (NVRAM), and magneto-resistive random-access memory (MRAM). The above memory types are for example only, and are thus not limiting as to the types of memory usable for storage of a computer program.

[0077] Example user computing devices are disclosed herein. For example, Figure 7 illustrates an example configuration of a user computing device 702. User computing device 702 includes a processor 704 for executing instructions. In some embodiments, executable instructions are stored in a memory area 706. Processor 704 may include one or more processing units (e.g., in a multi-core configuration). Memory area 706 is any device allowing information such as executable instructions and / or other data to be stored and retrieved. Memory area 706 may include one or more computer-readable media.

[0078] User computing device 702 also includes at least one media output component 708 for presenting information to a user. Media output component 708 is any component capable of conveying information to user. In some embodiments, media output component 708 includes an output adapter such as a video adapter and / or an audio adapter. An output adapter is operatively coupled to processor 704 and operatively couplable to an output device such as a display device (e.g., a liquid crystal display (LCD), organic light emitting diode (OLED) display, cathode ray tube (CRT), or "electronic ink" display) or an audio output device (e.g., a speaker or headphones). For example, a user may view expression titer information via media output component 708.

[0079] In some embodiments, user computing device 702 includes an input device 710 for receiving input from user. Input device 710 may include, for example, a keyboard, a pointing device, a mouse, a stylus, a touch sensitive panel (e.g., a touch pad or a touch screen), a camera, a gyroscope, an accelerometer, a position detector, and / or an audio input device. A single component such as a touch screen may function as both an output device of media output component 708 and input device 710.

[0080] User computing device 702 may also include a communication interface 712, which is communicatively couplable to a remote device such as a server system or a web server operated by a bioregistry. Communication interface 712 may include, for example, a wired or wireless network adapter or a wireless data transceiver for use with a mobile phone network (e.g., Global System for Mobile communications (GSM), 3G, 4G or Bluetooth) or other mobile data network (e.g., Worldwide Interoperability for Microwave Access (WIMAX)).

[0081] Machine learning methods are disclosed herein. The computer-implemented methods discussed herein may include additional, less, or alternate actions, including those discussed elsewhere herein. The methods may be implemented via one or more local or remote processors, transceivers, servers and / or via computer-executable instructions stored on non-transitory computer-readable media or medium.

[0082] Additionally, the computer systems discussed herein may include additional, less, or alternate functionality, including that discussed elsewhere herein. The computer systems discussed herein may include or be implemented via computer-executable instructions stored on non-transitory computer-readable media or medium.

[0083] A processor or a processing element may be trained using supervised or unsupervised machine learning, and the machine learning program may employ a neural network, which may be a convolutional neural network, a deep learning neural network, a reinforced or reinforcement learning module or program, or a combined learning module or program that learns in two or more fields or areas of interest. Machine learning may involve identifying and recognizing patterns in existing data in order to facilitate making predictions for subsequent data. Models may be created based upon example inputs in order to make valid and reliable predictions for novel inputs.

[0084] Additionally, or alternatively, the machine learning programs may be trained by inputting sample data sets or certain data into the programs, such as structural data for the protein (e.g., primary amino acid sequence of the protein) and / or a binary space representation of the structural data, isotype data, host cell line identifier data, and / or manufacturability and / or developability data, image data, and / or other data. The machine learning programs may utilize deep learning algorithms that may be primarily focused on pattern recognition, and may be trained after processing multiple examples. The machine learning programs may include Bayesian Program Learning (BPL), voice recognition and synthesis, image or object recognition, optical character recognition, and / or natural language processing - either individually or in combination. The machine learning programs may also include natural language processing, semantic analysis, automatic reasoning, and / or machine learning.

[0085] Supervised and unsupervised machine learning techniques may be used. In supervised machine learning, a processing element may be provided with example inputs and their associated outputs, and may seek to discover a general rule that maps inputs to outputs, so that when subsequent novel inputs are provided the processing element may, based upon the discovered rule, accurately predict the correct output. In unsupervised machine learning, the processing element may be required to find its own structure in unlabeled example inputs. In some embodiments, machine learning techniques may be used to extract data about a particular protein from binary space representation of the structural data, isotype data, host cell line identifier data, and / or manufacturability and / or developability data, image data, and / or other data.

[0086] In some embodiments, the voice bots or chatbots discussed herein may be configured to utilize ML and / or Al techniques. For instance, the voice bot or chatbot may be an Al chatbot. The voice bot or chatbot may employ supervised or unsupervised machine learning techniques, which may be followed by, and / or used in conjunction with, reinforced or reinforcement learning techniques.

[0087] As will be appreciated based upon the foregoing specification, the above-described embodiments of the disclosure may be implemented using computer programming or engineering techniques including computer software, firmware, hardware or any combination or subset thereof. Any such resulting program, having computer-readable code means, may be embodied or provided within one or more computer-readable media, thereby making a computer program product, i.e., an article of manufacture, according to the discussed embodiments of the disclosure. The computer-readable media may be, for example, but is not limited to, a fixed (hard) drive, diskette, optical disk, magnetic tape, semiconductor memory such as read-only memory (ROM), SD card, memory device and / or any transmitting / receiving medium, such as the Internet or other communication network or link. The article of manufacture containing the computer code may be made and / or used by executing the code directly from one medium, by copying the code from one medium to another medium, or by transmitting the code over a network.

[0088] These computer programs (also known as programs, software, software applications, "apps", or code) include machine instructions for a programmable processor, and can be implemented in a high-level procedural and / or object-oriented programming language, and / or in assembly / machine language. As used herein, the terms "machine- readable medium" and "computer-readable medium" refer to any computer program product, apparatus and / or device (e.g., magnetic discs, optical disks, memory, Programmable Logic Devices (PLDs)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The "machine-readable medium" and "computer- readable medium," however, do not include transitory signals. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0089] As used herein, a processor may include any programmable system including systems using micro-controllers, reduced instruction set circuits (RISC), application specific integrated circuits (ASICs), logic circuits, and any other circuit or processor capable of executing the functions described herein. The above examples are example only, and are thus not intended to limit in any way the definition and / or meaning of the term "processor."

[0090] As used herein, the terms "software" and "firmware" are interchangeable, and include any computer program stored in memory for execution by a processor, including RAM memory, ROM memory, EPROM memory, EEPROM memory, and non-volatile RAM (NVRAM) memory. The above memory types are example only, and are thus not limiting as to the types of memory usable for storage of a computer program.

[0091] In some embodiments, a computer program is provided, and the program is embodied on a computer readable medium. In an example embodiment, the system is executed on a single computer system, without requiring a connection to a server computer. In a further embodiment, the system is being run in a Windows® environment (Windows is a registered trademark of Microsoft Corporation, Redmond, Washington). In yet another embodiment, the system is run on a mainframe environment and a UNIX® server environment (UNIX is a registered trademark of X / Open Company Limited located in Reading, Berkshire, United Kingdom). The application is flexible and designed to run in various different environments without compromising any major functionality.

[0092] In some embodiments, the system includes multiple components distributed among a plurality of computing devices. One or more components may be in the form of computer-executable instructions embodied in a computer-readable medium. The systems and processes are not limited to the specific embodiments described herein. In addition, components of each system and each process can be practiced independent and separate from other components and processes described herein. Each component and process can also be used in combination with other assembly packages and processes. The present embodiments may enhance the functionality and functioning of computers and / or computer systems.

[0093] As used herein, an element or step recited in the singular and preceded by the word "a" or "an" should be understood as not excluding plural elements or steps, unless such exclusion is explicitly recited. Furthermore, references to "example embodiment" or "one embodiment" of the present disclosure are not intended to be interpreted as excluding the existence of additional embodiments that also incorporate the recited features.

[0094] The patent claims at the end of this document are not intended to be construed under 35 U.S.C. § 112(f) unless traditional means-plus-function language is expressly recited, such as "means for" or "step for" language being expressly recited in the claim(s).

[0095] This written description uses examples to disclose the disclosure, including the best mode, and also to enable any person skilled in the art to practice the disclosure, including making and using any devices or systems and performing any incorporated methods. The patentable scope of the disclosure is defined by the claims, and may include other examples that occur to those skilled in the art. Such other examples are intended to be within the scope of the claims if they have structural elements that do not differ from the literal language of the claims, or if they include equivalent structural elements with insubstantial differences from the literal language of the claims.

Claims

What Is Claimed Is:

1. A system for predicting expression titer for a protein, the system comprising: at least one memory storing computer-executable instructions; and at least one processor in communication with the at least one memory, wherein the at least one processor is configured to execute the computer-executable instructions to: receive structure data for the protein, the structure data comprising a primary structure of the protein; generate a binary space representation of the structure data, the binary space representation based on a classification of each element of the primary structure; apply one or more inputs to a supervised machine learning model, the one or more inputs comprising the binary space representation, the model being previously trained using historical data, the historical data comprising structure data for plurality of proteins and their corresponding expression titers; receive one or more outputs from the model, at least one of the one or more outputs including an expression titer classification for the protein; update the historical data to include the structure data of the protein and the corresponding one or more outputs; and re-train the model using the updated historical data.

2. The system of Claim 1, wherein the at least one processor is further configured to execute the computer-executable instructions to: create a plurality of groups of elements of the structure data, wherein each of the plurality of groups comprises three contiguous elements of the primary structure; determine a triad class for each of the plurality of groups, the triad class based on the classification of each element; create a vector space representing the triad class of each of the plurality of groups; determine a frequency of each triad class within the primary structure; and create a frequency vector representing the frequency of each triad class,wherein the binary space representation comprises the vector space and the frequency vector.

3. The system of Claim 2, wherein each element is represented by a single or three- letter code.

4. The system of Claim 3, wherein each element is classified based on at least a dipole classification and a volume classification of the respective element.

5. The system of Claims 1 or 2, wherein the one or more inputs further include an isotype of the primary structure, the isotype comprising immunoglobulin G (IgG) antibodies.

6. The system of Claim 5, wherein the isotype comprises at least one of IgGl or lgG4.

7. The system of Claims 1 or 2, wherein the one or more inputs further include a host cell line identifier.

8. The system of Claim 1, wherein the model comprises at least one of a support vector machine or a neural network.

9. The system of Claim 8, wherein the model comprises a multilayer perceptron (MLP) classifier comprising at least one hidden layer.

10. The system of Claim 1, wherein the model is configured for a binary classification, the binary classification comprising a high value classification and a low value classification.

11. A computer-based method for predicting expression titer for a protein, the method implemented using a system including a computing device including a processor communicatively coupled to memory device, the method comprising:receiving structure data for the protein, the structure data comprising a primary structure of the protein; generating a binary space representation of the structure data, the binary space representation based on a classification of each element of the primary structure; applying one or more inputs to a supervised machine learning model, the one or more inputs comprising the binary space representation, the model being previously trained using historical data, the historical data comprising structure data for plurality of proteins and their corresponding expression titers; receiving one or more outputs from the model, at least one of the one or more outputs including an expression titer classification for the protein; updating the historical data to include the structure data of the protein and the corresponding one or more outputs; and re-training the model using the updated historical data.

12. The computer-based method of Claim 11, further comprising: creating a plurality of groups of elements of the structure data, wherein each of the plurality of groups comprises three contiguous elements of the primary structure; determining a triad class for each of the plurality of groups, the triad class based on the classification of each element; creating a vector space representing the triad class of each of the plurality of groups; determining a frequency of each triad class within the primary structure; and creating a frequency vector representing the frequency of each triad class, wherein the binary space representation comprises the vector space and the frequency vector.

13. The computer-based method of Claim 12, wherein each element is represented by a single or three-letter code.

14. The computer-based method of Claim 13, wherein each element is classified based on at least a dipole classification and a volume classification of the respective element.

15. The computer-based method of Claim 11, wherein the model comprises at least one of a support vector machine or a neural network.

16. The computer-based method of Claim 15, wherein the model comprises a multilayer perceptron (MLP) classifier comprising at least one hidden layer.

17. At least one non-transitory computer-readable storage media having computerexecutable instructions embodied thereon, wherein when executed by at least one processor, the computer-executable instructions cause the at least one processor to: receive structure data for a protein, the structure data comprising a primary structure of the protein; generate a binary space representation of the structure data, the binary space representation based on a classification of each element of the primary structure; apply one or more inputs to a supervised machine learning model, the one or more inputs comprising the binary space representation, the model being previously trained using historical data, the historical data comprising structure data for plurality of proteins and their corresponding expression titers; receive one or more outputs from the model, at least one of the one or more outputs including an expression titer classification for the protein; update the historical data to include the structure data of the protein and the corresponding one or more outputs; and re-train the model using the updated historical data.

18. The least one non-transitory computer-readable storage media of Claim 17, wherein the computer-executable instructions further cause the at least one processor to: create a plurality of groups of elements of the structure data, wherein each of the plurality of groups comprises three contiguous elements of the primary structure;determine a triad class for each of the plurality of groups, the triad class based on the classification of each element; create a vector space representing the triad class of each of the plurality of groups; determine a frequency of each triad class within the primary structure; and create a frequency vector representing the frequency of each triad class, wherein the binary space representation comprises the vector space and the frequency vector.

19. The least one non-transitory computer-readable storage media of Claim 18, wherein each element is represented by a single or three-letter code.

20. The least one non-transitory computer-readable storage media of Claim 19, wherein each element is classified based on at least a dipole classification and a volume classification of the respective element.

Citation Information

Patent Citations

  • Machine learning and / or image processing for spectral object classification

    US20210182635A1

  • Methods and systems for biotherapeutic development

    US20220139558A1

  • Method and process for predicting and analyzing patient cohort response, progression, and survival

    US20220270763A1

  • Systems and methods for creating biomolecule embeddings

    US20230253113A1

  • HLA-ii immunopeptidome methods and systems for antigen discovery

    WO2024015892A1