Systems and methods for intelligent construction of antibody libraries - Patents.com
Patent Information
- Application Number
- JP2024525610
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-11-01
- Filing Date
- 2022-10-26
- Publication Date
- 2025-10-30
AI Technical Summary
Identifying useful antibodies for therapeutic applications is difficult due to their immunogenicity and the need for significant redesign, and existing libraries lack diversity while maintaining human sequence representation.
Utilizing machine learning models to predict biophysical and biochemical properties of antibody sequences, enabling the construction of libraries with controlled diversity and reduced immunogenicity, allowing for the synthesis of antibodies that can recognize a wide range of antigens.
The method generates antibody libraries with improved human sequence representation and desirable properties, facilitating the development of therapeutic agents with enhanced specificity and reduced immunogenicity.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to and the benefit of U.S. Provisional Patent Application No. 63 / 274,394, filed November 1, 2021, which is incorporated by reference in its entirety. [Background technology]
[0002] Antibodies are extremely important as research tools and in diagnostic and therapeutic applications, however, identifying useful antibodies can be difficult and, once identified, antibodies often require significant redesign before they are suitable for human therapeutic use.
[0003] Thus, there is a need for libraries of smaller antibodies (i.e., synthetically and physically feasible antibodies) with directed diversity that are non-immunogenic (e.g., more human) and systematically represent candidate antibodies with desirable properties (e.g., the ability to recognize a broad range of antigens). Obtaining such libraries requires balancing the competing objectives of limiting the diversity of sequences represented in the library (e.g., limiting the introduction of non-human sequences while allowing for possible oversampling to allow for synthetic and physical feasibility) while maintaining a level of diversity sufficient to recognize a broad range of different antigens.
[0004] Thus, there is a need for methods to construct antibody libraries populated with antibodies that: (a) can be readily synthesized; (b) are physically feasible and, in some cases, can be oversampled; (c) are diverse enough to recognize all of the antigens recognized by the pre-immune (i.e., pre-negative selection) human repertoire; (d) are non-immunogenic in humans (i.e., contain sequences of human origin); and / or (e) have CDR length and sequence diversity, as well as framework diversity, representative of naturally occurring human antibodies. Summary of the Invention [Means for solving the problem]
[0005] Described herein are systems and methods for constructing antibody libraries using machine learning to inform sequence selection for inclusion in the library. The techniques include (i) training and using machine learning and statistical models to predict biophysical and biochemical properties from sequences, and (ii) training and using machine learning models to predict developability and generate novel sequences from sequences. In certain embodiments, the systems and methods generate libraries of antibodies (and / or antibody-encoding polynucleotides) by individually designing libraries with specified sequence and / or length diversity. The resulting libraries are useful, for example, in the development of therapeutics.
[0006] In one aspect, the present invention relates to a system for constructing (e.g., designing) an antibody library. The system comprises a processor of a computing device and a memory having instructions stored thereon, which, when executed by the processor, cause the processor to perform one or more of the following: (i), (ii), (iii), (iv), (v), (vi), and (vii): (i) develop (e.g., train) a first machine learning model using input sequence and characterization data (e.g., (a) train a logistic regression model that derives amino acid coefficients to predict polyspecificity and hydrophobicity of individual complementarity determining regions (CDRs) and / or framework regions (FRs), and / or (b) train a tree model (e.g., Random Forest or XGBoost) to predict one or more biophysical properties and / or one or more chemical stability properties from sequences, and / or (c) train one or more bioinformatics models (e.g., genomic DNA models) to predict one or more biophysical properties and / or one or more chemical stability properties from sequences, and / or (d) train a logistic regression model that derives amino acid coefficients to predict one or more biophysical properties and / or one or more chemical stability properties from sequences, and / or (e) train a logistic regression model that derives amino acid coefficients to predict one or more biophysical properties and / or one or more chemical stability properties from sequences, and / or (f) train a logistic regression model that derives amino acid coefficients to predict one or more biophysical properties and / or one or more chemical stability properties from sequences, and / or (g) train a logistic regression model that derives amino acid coefficients to predict one or more biophysical properties and / or one or more chemical stability properties from sequences, and / or (h) train a logistic regression model that derives amino acid coefficients to predict one or more biophysical properties and / or one or more chemical stability properties from sequences, and / or (i) train a logistic regression model that derives amino acid coefficients to predict one or more biophysical properties and / or one or more chemical stability properties from sequences (d) train a deep learning model including a neural network to predict biophysical properties and / or one or more chemical stability properties from the sequence (e.g., the model includes an input layer, multiple intermediate feature extraction layers, and a final output layer); and / or (d) create a statistical model to assess bias to select sequences with low bias; and / or (e) develop hierarchical statistics to predict risk of chemical modification as a function of sequence motifs at specific positions and regions (e.g., H1, H2, H3, L1, L2, L3, HFR, LFR). (ii) use the first machine learning model in (i) to predict desirable segments (e.g., segments of favorable predicted expression enrichment), allowing selection of segments from a pool of novel and / or pre-generated segments.(iii) processing the set of input sequences prior to selection and / or use in training the first machine learning model in (i), where processing the set of input sequences includes one or more of: (a) modifying the sequences to remove chemical liability sites; (b) for CDR H3, splitting the sequence into segments to mimic VDJ recombination; (c) for CDR L3, splitting the sequence into segments to mimic VJ recombination; and (d) annotating the V-regions and CDRs (H1, H2, L3) with the number of mutations from germline. (iv) training a machine learning model for prediction of biophysical and / or biochemical properties, e.g., using data on the set of input sequences sorted for favorable biophysical properties (e.g., low polyspecificity, low hydrophobicity, and / or high expression). (v) using the machine learning model for prediction of biophysical and / or biochemical properties in (iv) to predict one or more biophysical and / or biochemical properties (e.g., polyspecificity, hydrophobicity, melting temperature, SEC monomer percentage, retention time, chemical stability data, and / or measures of sequence enrichment or depletion) from the sequences; (vi) developing (e.g., training) an autoregressive deep learning neural network model to learn joint sequence probability distributions across sequences of interest for a particular germline for different species; (vii) using the neural network model in (vi) to capture sequence composition and / or correlations from the input set of sequences to generate novel sequences or segments for consideration in the synthetic library.
[0007] In another aspect, the invention relates to a system for constructing an antibody library, comprising a processor of a computing device and a memory having instructions stored thereon that, when executed by the processor, cause the processor to process a set of input sequences through one or more machine learning models to generate a collection of final antibody library sequences.
[0008] In certain embodiments, the instructions cause the processor to (i) process each input sequence from a set of input sequences, and (ii) for each of the input sequences, a per-residue prediction of one or more structurally important properties of the sequence as predicted by a first model (e.g., a graph convolutional network (GCN)), and the instructions cause the processor to process (i) and (ii) as inputs in a second model to predict, as an output of the second model, (iii) one or more biophysical properties (e.g., hydrophobic interaction chromatography retention time (HIC RT) and / or multispecific reagent (PSR) score and / or PSR binding category) and / or (iv) one or more chemical stability properties (e.g., Asn deamidation, Asp isomerization, and / or Met oxidation) for each of the input sequences, wherein the inclusion or exclusion of each sequence from the final antibody library is based at least in part on the output of the second model.
[0009] In certain embodiments, the per-residue predictions predicted by the first model include one or more selected from the group consisting of: (i) a solvent accessibility (SASA) measure, (ii) a charge patch measure, (iii) a hydrophobic patch measure, and (iv) a Cα / Cβ coordinate prediction.
[0010] In certain embodiments, the second model includes deep convolutional and / or recurrent networks (eg, for prediction of biophysical properties).
[0011] In certain embodiments, the second model comprises a tree-based classification model (eg, for predicting chemical stability).
[0012] In one aspect, the present invention relates to a method for constructing (e.g., designing) an antibody library, the method comprising using a processor of a computing device to perform one or more of the following (i), (ii), (iii), (iv), (v), (vi), and (vii): (i) developing (e.g., training) a first machine learning model using input sequences and characterization data (e.g., (a) training a logistic regression model to derive amino acid coefficients to predict polyspecificity and hydrophobicity of individual complementarity determining regions (CDRs) and / or framework regions (FRs), and / or (b) training a tree model (e.g., Random Forest or XGBoost) to predict one or more biophysical properties and / or one or more chemical stability properties from sequences, and / or (c) training one or more bioinspired models (e.g., chromosome 1, chromosome 2, chromosome 3, chromosome 4, chromosome 5, chromosome 6, chromosome 7, chromosome 8, chromosome 9, chromosome 10, chromosome 11, chromosome 12, chromosome 13, chromosome 14, chromosome 15, chromosome 16, chromosome 17, chromosome 18, chromosome 19, chromosome 20, chromosome 21, chromosome 22, chromosome 23, chromosome 24, chromosome 25, chromosome 26, chromosome 27, chromosome 28, chromosome 29, chromosome 30, chromosome 31, chromosome 32, chromosome 33, chromosome 34, chromosome 35, chromosome 36, chromosome 37, chromosome 38, chromosome 39, chromosome 40, chromosome 41, chromosome 42, chromosome 43, chromosome 44, (d) training a deep learning model including a neural network to predict biophysical properties and / or one or more chemical stability properties from the sequence (e.g., the model includes an input layer, multiple intermediate feature extraction layers, and a final output layer); and / or (d) creating a statistical model to assess bias to select sequences with low bias; and / or (e) developing hierarchical statistics to predict risk of chemical modification as a function of sequence motifs at specific positions and regions (e.g., H1, H2, H3, L1, L2, L3, HFR, LFR). (ii) using the first machine learning model in (i) to predict desirable segments (e.g., segments of preferred predicted expression enrichment), allowing selection of segments from a pool of novel and / or pre-generated segments. (iii) processing the set of input sequences prior to selection and / or use in training the first machine learning model in (i), where processing the set of input sequences includes one or more of: (a) modifying the sequences to remove chemical liability sites; (b) for CDR H3, splitting the sequence into segments to mimic VDJ recombination; (c) for CDR L3, splitting the sequence into segments to mimic VJ recombination; and (d) assigning a number of mutations from germline to the V-regions and CDRs (H1, H2, L3).(iv) training a machine learning model for prediction of biophysical and / or biochemical properties (e.g., using data on a set of input sequences sorted for favorable biophysical properties (e.g., low polyspecificity, low hydrophobicity, and / or high expression)); (v) using the machine learning model for prediction of biophysical and / or biochemical properties in (iv) to predict one or more biophysical and / or biochemical properties from a sequence (e.g., polyspecificity, hydrophobicity, melting temperature, SEC monomer percentage, retention time, chemical stability data, and / or measures of sequence enrichment or depletion); (vi) developing (e.g., training) an autoregressive deep learning neural network model to learn joint sequence probability distributions across sequences of interest for a particular germline for different species; (vii) using the neural network model in (vi) to capture sequence composition and / or correlations from the input set of sequences to generate novel sequences or segments for consideration in synthetic libraries.
[0013] In another aspect, the present invention relates to a method for constructing (e.g., designing) an antibody library, the method comprising processing a set of input sequences by a processor of a computing device with one or more machine learning models to generate a collection of final antibody library sequences.
[0014] In certain embodiments, the method includes (i) processing each input sequence from the set of input sequences as an input in a second model, and further includes (ii) processing, for each of the input sequences, a per-residue prediction of one or more structurally important properties of the sequence as predicted by the first model (e.g., a graph convolutional network (GCN)), and predicting, as an output of the second model, (iii) one or more biophysical properties (e.g., hydrophobic interaction chromatography retention time (HIC RT) and / or multispecific reagent (PSR) score and / or PSR binding category) and / or (iv) one or more chemical stability properties (e.g., Asn deamidation, Asp isomerization, and / or Met oxidation) for each of the input sequences, wherein the inclusion or exclusion of each sequence in the final antibody library is based at least in part on the output of the second model.
[0015] In certain embodiments, the per-residue predictions predicted by the first model include one or more selected from the group consisting of: (i) a solvent accessibility (SASA) measure, (ii) a charge patch measure, (iii) a hydrophobic patch measure, and (iv) a Cα / Cβ coordinate prediction.
[0016] In certain embodiments, the second model includes deep convolutional and / or recurrent networks (eg, for prediction of biophysical properties).
[0017] In certain embodiments, the second model comprises a tree-based classification model (eg, for predicting chemical stability).
[0018] The above and other objects, aspects, features, and advantages of the present disclosure will become more apparent and better understood by referring to the following description taken in conjunction with the accompanying drawings. [Brief description of the drawings]
[0019] [Figure 1]FIG. 1 is a block flow diagram of an exemplary method for information-based construction of an antibody sequence library, according to an exemplary embodiment.
[0020] [Diagram 2] FIG. 1 is a schematic diagram of a deep learning module for predicting developability from sequences, according to an exemplary embodiment.
[0021] [Diagram 3] 1 is a chart showing examples for matching specific CDR H3 sequences in a VHH H3 library design, according to an exemplary embodiment.
[0022] [Figure 4A] FIG. 1 shows an exemplary sequence generation procedure used in CDR H1 and H2 library design, according to an exemplary embodiment.
[0023] [Figure 4B] FIG. 1 shows an exemplary sequence generation procedure used in Vλ L3 library design, according to an exemplary embodiment.
[0024] [Diagram 5] According to an exemplary embodiment,
number
number
[0025] [Figure 6] FIG. 1 is a schematic diagram of a network environment for use in providing the systems, methods, and architectures described herein.
[0026] [Figure 7]1A-1C are schematic diagrams illustrating computing and mobile computing devices that can be used to implement the techniques described herein.
[0027] [Figure 8A] FIG. 1 is a schematic showing steps in an exemplary machine learning method for predicting structural properties from sequence data. [Figure 8B] FIG. 1 is a schematic showing steps in an exemplary machine learning method for predicting structural properties from sequence data.
[0028] [Figure 9] FIG. 1 is a block diagram of a method for predicting important developable properties for therapeutics using prediction of structurally significant metrics of biophysical properties in a model, according to an exemplary embodiment.
[0029] [Figure 10A] FIG. 1 is a schematic diagram illustrating the use of a graph convolution model for residue-level prediction of structural descriptors, according to an exemplary embodiment. [Figure 10B] FIG. 1 is a schematic diagram illustrating the use of a graph convolution model for residue-level prediction of structural descriptors, according to an exemplary embodiment.
[0030] [Figure 11] 1 is a graph showing the overall SAP score calculated by summing individual residue predictions according to an exemplary embodiment (where predictions are comparable to those obtained from the AlphaFold2 model for the same input sequence).
[0031] [Figure 12] 1 is a graph showing the overall SCM score calculated by summing the individual residue predictions according to an exemplary embodiment.
[0032] [Figure 13] FIG. 13 is a schematic diagram of a Mollweide projection of a predicted net charge patch, in accordance with an exemplary embodiment;
[0033] [Figure 14] FIG. 1 is a schematic diagram illustrating the use of convolutional and recursive models for the prediction of hydrophobicity and polyspecificity, according to an exemplary embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0034] Features and advantages of the present disclosure will become more apparent from the following detailed description taken in conjunction with the drawings in which like reference characters indicate corresponding elements throughout. In the drawings, like reference numbers typically indicate identical, functionally similar, and / or structurally similar elements.
[0035] The systems, architectures, apparatus, methods, and processes of the claimed invention are believed to encompass variations and modifications developed using information from the embodiments described herein, and modifications and variations of the systems, architectures, apparatus, methods, and processes described herein are contemplated herein.
[0036] Throughout this specification, where articles, devices, systems, and architectures are described as having, including, or comprising particular elements, or processes and methods are described as having, including, or comprising particular steps, it is intended that there are further articles, devices, systems, and architectures of the invention that consist essentially of or consist of the recited elements, and that there are further processes and methods of the invention that consist essentially of or consist of the recited processing steps.
[0037] It should be understood that the order of steps or order for performing certain actions is not essential so long as the invention remains operable. Moreover, two or more steps or actions may be conducted simultaneously.
[0038] The citation of any publication in this specification (e.g., in the "Background" section) is not an admission that such publication is prior art to any claim presented herein. The "Background" section is presented for purposes of clarity and is not intended as a description of prior art with respect to any claim.
[0039] Documents referenced herein are incorporated herein by reference. In the event of a conflict in the meaning of certain terms, the meaning set forth in the Detailed Description of the Invention shall prevail.
[0040] Headings are provided for the convenience of the reader, and the presence and / or placement of headings is not intended to limit the scope of the subject matter described herein.
[0041] Section I. A set of training sequences for learning sequence patterns, correlations, and frequency of use in natural or synthetic repertoires The sequences used in the systems and methods presented herein can come from internal discovery (naive, LCBS (light chain batch shuffle), pre-made AFFMAT (affinity maturation), oligo-based read-specific AFFMAT), patent and clinical sequences, or literature NGS (next generation sequencing) data sets. Once a starting set of sequences is obtained, they can be post-processed depending on the CDRs (complementarity determining regions) and the nature of the library. A list of exemplary post-processing steps includes the following: 1) Eliminate chemical liability sites by modifying the sequence in the following ways: a.Replace exposed Met with Leu. bN(G,S,T) is replaced by Q(G,S,T), thus eliminating a potential Asn deamidation motif. cD(G,S,T) is replaced by E(G,S,T) (this will eliminate the Asp isomerization motif). Asn in the dN-gly sites was replaced with Asp, thus eliminating the N-linked glycosylation motif (these are considered potential negative factors due to host cell dependency among other factors). e. replacing the fragmentation motif DP with EP; and f. Substitute N, D, or M amino acids that are predicted (by the trained machine learning model) to be at high risk of modification, or mutate their surrounding sequence context to reduce the risk of modification (as also predicted by the machine learning model). 2) For CDR H3, the sequence is divided into segments to mimic VDJ recombination. a. Segments can be based on matching to pre-generated libraries of segments from known V, D, and J genes. b. Inferred from analysis and analysis of the output of programs such as IgBlast, Immuncantation (Vander Heiden JA, Yaari G, Bioinformatics, 30, 1930, 2014 PMID: 24618469, Gupta NT, Vander Heiden JA, Bioinformatics, 31, 3356, 2015, PMID: 26069265). And, c. De novo inferred from in-house software. 3) For CDR L3, the sequence is divided into segments to mimic VJ recombination. a. Segments are based on matching with known V- and J-genes from the literature or the IMGT database. 4) Additionally, the V-regions and CDRs (H1, H2, L3) are annotated as follows: a. the number of mutations from the germline; and b. The number of mutations from germline for preferential antigen contacting or exposed residues based on analysis of published crystal structures.
[0042] Examples of sequences used to train models for different library designs include: 1) human germline variable region sequence data, obtained from: a) OAS (Observed Antibody Space) sequence database (Kovaltsuk A, The Journal of Immunology, 201, 2502, 2018, PMID: 30217829), b) internal data derived from human germline sequences from primary discovery and affinity maturation of paired heavy and light chain sequences; c) Clinical antibody data from literature, patents, etc. for paired heavy and light chain sequences; 2) human V, D, and J-gene information from IMGT; 3) variable region sequence data for camelids from NGS; a) Llama sequence from McCoy LE, PLoS Pathogens, 10, e1004552, 2014, PMID: 25522326; b) Bactrian camel sequence from Li X, PLoS ONE, 11, e0161801, 2016, PMID: 27588755, and c) Clinical antibody data from literature, patents, etc. 4) The V, D and J gene information of camelids is as follows: a) V genes from IMGT for alpaca, llama, and Bactrian camel; b) J-genes from IMGT for alpaca, llama, and J-genes for Bactrian camel from Liang Z, Frontiers of Agricultural Science and Engineering, 2, 249, 2015; and c) D-gene from IMGT for alpaca, and D-genes for llama and Bactrian camel inferred internally from NGS data using IgScout (Safonova Y, Frontiers in Immunology, 10, 1, 2019, PMID: 31134072), and D-gene for Bactrian camel from Liang Z, Frontiers of Agricultural Science and Engineering, 2, 249, 2015.
[0043] Section II Machine learning for predicting structural properties directly from sequences Because obtaining 3D structures from experimental techniques such as X-ray crystallography, cryo-EM, or from homology modeling software such as AlphaFold, IgFold, or Schrodinger Discovery Studio can be time-consuming, we present herein machine learning models to predict structural features important for downstream developability prediction from sequence input. 1) 3D structural data for developing the mechanical model(s) can be obtained from: a) publicly deposited or internally obtained Protein Data Bank (PDB) structures, and / or b) Homology models from publicly available sources or generated via internal software pipelines and algorithms. 2) Using internally developed proprietary algorithms or internal implementations of published methods (e.g., SAP Chennamsetty et al., J. Phys. Chem., 2014; SCM Agrawal et al., mAbs, 2016) with the above 3D structure data, descriptors are generated for each residue in the input structure. Furthermore, these descriptor values can be aggregated based on residue type, antibody region, or a combination thereof to generate higher level descriptors.
[0044] The sequence of the 3D structures in item 1) serves as input data to a machine learning model that aims to predict a set of descriptors in item 2) above.
[0045] 8A and 8B show steps in an exemplary machine learning method for predicting structural features from sequence data.
[0046] Since protein structures can be represented as graphs, the graph convolutional network (GCN) architecture can be used to predict structural properties from sequences. GCN involves the following steps: 1) network weights W are learned separately for each amino acid type, position, or combination thereof (so-called node weights); and 2) Further weights are learned to represent the influence of neighboring residues on the central residue Cij (so-called edge weights). 3) The node and edge weights can be combined by a mathematical operation, denoted f (including a bias term b), to generate a descriptor or additional feature for each residue in the sequence (denoted by x). 4) To improve the network's learning ability and allow it to learn properties across multiple length scales, an independent set of the above parameters can be learned at each step. 5) Deep learning models can be built by stacking multiple such layers to train the network to learn complex relationships, with each layer shown as an "attention block" in the schematic diagram in Figure 8B. 6) A densely connected layer with nonlinear activation is implemented, where position-specific weights are learned to finally predict structural descriptors for each residue.
[0047] Examples of structural descriptors that can be learned by the model and then predicted using only sequence as input are: 1) the solvent accessibility for each residue, 2) the degree of hydrophobicity around each residue over multiple length scales, calculated using a set of published or determined hydrophobicity / hydrophilicity trends; 3) the degree of positive, negative, and total charge around each residue, calculated using charges assigned from different force fields such as CHARMM and AMBER, calculated at different pHs, and across multiple length scales; and 4) Calculation of structural coordinates for the main chain and side chains obtained after aligning the input training data into a common reference frame.
[0048] Predictions of these descriptors from sequence can then serve as input to downstream tasks and other machine learning models to predict experimentally observed developability properties for antibodies.
[0049] Sequences for the input structures were aligned using a consistent numbering scheme and converted to numerical values using a one-hot encoding scheme, addition of biophysical and biochemical features using an amino acid property scale, a position-specific scoring matrix, and pre-trained sequence embedding.
[0050] Models were trained using 25-fold Monte Carlo cross-validation with training and validation splits of 80% and 20%, respectively. Models were trained for up to 200 epochs with early stopping if no improvement was seen on the test set over 10 epochs.
[0051] Since the predicted output descriptors have different scaled values, a preprocessing step can be performed so that the distribution for each descriptor is centered by subtracting the mean per residue. Additionally, various strategies for scaling the magnitude were used, such as dividing by the variance or interquartile range of the original distribution.
[0052] Example pseudocode for a deep learning model for predicting structural descriptors from sequences is as follows:
number
number
[0053] Section III Development Capabilities Input Training Data for Machine Learning Input sequences for training machine learning models for predicting biophysical and biochemical properties can be derived from the following illustrative examples: 1) Data relating to an individual sequence, the sequence being: a) a discovery effort using an internal library; b) Data on sequences generated from the literature, such as clinical antibodies, sequences from patents, and data on sequences derived from; 2) Data regarding pools or collections of sequences, e.g. a) NGS sequencing data for libraries sorted for favorable biophysical properties (e.g., low polyspecificity, low hydrophobicity, high expression, etc.); and b) Data regarding polyclonal assessment of libraries with known input sequence or compositional differences.
[0054] The biophysical and biochemical data for the sequence may include, for example: 1) Multispecific measurements using PSR (multispecific reagents) and AC-SINS (affinity capture self-interacting nanoparticle spectroscopy); 2) hydrophobicity as measured using HIC (hydrophobic interaction chromatography) retention time; 3) Melting temperature, 4) SEC (size exclusion chromatography) monomer percentage and retention time; 5) chemical stability data under different stress conditions to identify deamidation, isomerization, oxidation, and fragmentation using tryptic peptide mapping; and 6) Sequence enrichment or depletion in positively or negatively selected populations compared to each other or to the input frequency below.
number
[0055] Section IV Machine Learning and Statistical Models for Sequence Developability In certain embodiments, the following machine learning models are developed based on the input sequence and characterization data as described above: a. A logistic regression model to derive amino acid coefficients to predict polyspecificity and hydrophobicity for individual positions, CDRs and FRs; b. Tree models such as Random Forest and XGBoost for predicting biophysical properties from sequences; c. A deep learning model that uses neural networks to predict biophysical properties from sequences; d. A statistical model to assess bias to select sequences with low bias; and e. Hierarchical statistics to predict risk of chemical modification as a function of sequence motifs at specific positions and regions (CDRH1, CDRH2, CDRH3, CDRL1, CDRL2, CDRL3, HFR, LFR) based on prior experimentally observed rate(s) of modification of that position, region, or motif anywhere in the antibody sequence (wherein a "motif" is defined as an amino acid that may be modified and the N+1 amino acid immediately following it). The statistics are hierarchical because they use the most specific statistic to predict for subjects for which prior observation is sufficient.
[0056] Each of these machine learning models is described in more detail below.
[0057] Logistic regression Logistic regression can be performed using methods such as those described in Jain T, Bioinformatics, 33, 3758, 2017, PMID: 28961999. The results from these models are region-specific amino acid coefficients for predicting low developability traits such as delayed retention time in HIC, high polyspecificity, and expression. The following formula is used:
number
number
number
number
[0058] coefficient
number
[0059] b. Tree-based regression and classification models Given properties or metrics of interest, regression and classification methods (e.g., tree-based methods such as Random Forest or XGBoost) are used as inputs to train neural networks or other machine learning models to predict such properties for novel sequences, segments of sequences, and / or individual amino acids within sequences. In this example,
number
[0060] An exemplary calculation of the hydrophobicity score and the polyspecificity score is performed as follows. 1) Calculate the solvent accessibility for each residue in the sequence from the graph folding model described above, or pre-calculated values from a database generated computationally for a set of known structures.
number
number
number
number
number
[0061] c. Deep learning model for predicting developability from sequences Deep learning neural network methods are trained to take into account properties or metrics of interest and predict such properties for novel sequences or segments. These models can include an input layer, multiple intermediate feature extraction layers, and a final output layer, as shown in the schematic diagram of FIG.
[0062] Input sequences of different lengths are processed to the same length for input to the neural network. This can be done by aligning the sequences using a consistent numbering scheme or by light padding the sequences with an appropriate number of insertions. The sequences are then converted to numerical values using a one-hot encoding scheme and the addition of biophysical and biochemical features using amino acid property scales, position-specific scoring matrices, and pre-trained sequence embedding. Descriptors calculated as output from upstream machine learning models / modules (such as, for example, the graph convolution model described herein) can also be calculated from the sequences and added to the model input.
[0063] The model input can be adapted to different modalities by adding or subtracting chain information in the input layer. The feature extraction layer can include one or more of convolutional, regression (using long-short-term memory (LSTM) units, gated recurrent units (GRUs)), self-attention, and / or densely connected layers.
[0064] In one exemplary embodiment, the model was trained using 10-fold cross-validation, where the model was trained for up to 300 epochs with early stopping if no improvement was observed on the test set over 10 epochs.
[0065] Example pseudocode for a deep learning model to predict exploitability is as follows:
number
[0066] d. Statistical models for identifying sequences with low bias Statistical and data mining approaches are presented here to identify sequences or sequence motifs that pair equally well across multiple sources of diversity in a library. Given an ideal or targeted distribution of different diversity, proposed motifs are evaluated for their ability to match that distribution with low bias, for example, using the Kullback-Leibler divergence metric. A motif can be a single amino acid at a position, a combination of amino acids at different positions, or an entire sequence. The Kullback-Leibler divergence metric for a given motif can be calculated as follows:
number
[0067] Section V Machine learning models of sequence patterns and composition An autoregressive deep learning neural network model can be implemented to learn sequence patterns in a set of curated input training sequences. The set can be composed of sequences classified according to desired criteria or characteristics, such as species and germline, favorable developability profile, etc. The goal is to learn a joint sequence probability distribution over the sequences of interest as follows:
number
[0068] In one exemplary embodiment, the input sequence data was split into a training:validation set in a ratio of 3:1. In this embodiment, the model was trained for up to 300 epochs with early stopping if no improvement was observed on the test set over 10 epochs.
[0069] The trained model can then be run in a generation mode by starting with an input seed sequence, e.g., an alphanumeric character (other than an amino acid symbol), a hyphen, or other symbol of length 1. For example, in the exemplary schematic of FIG. 4, this is a hyphen "-", which serves as an artificial construct to teach the model when H1 begins and ends. This seed sequence is updated by appending sampled amino acids from the probabilities predicted by the model. The generation process ends when an insertion is predicted, which marks the end of the sampled sequence. In addition, the probabilities of the generated sequences can also be stored for use in prioritizing sequences for a synthetic library. This results in a set of sequences specific to the germline of interest.
[0070] Section VI Obtaining segments and estimating their frequency of use in the repertoire Using the collection or subset of V-, D-, and J-genes detailed above, candidate segments can be generated using methods such as nucleotide deletion, nucleotide addition, nibbling, etc. To infer segments de novo from the data, wildcard sequences can be added as placeholders that match any sequence of length 0 to L.
[0071] A tree-based pruning algorithm (e.g., a match to design algorithm) can be used to match the pool of segments with natural repertoire sequences. Examples are provided in International Patent Application Publications WO2009 / 036379 and WO2012 / 009568, the contents of which are incorporated herein by reference in their entirety. The frequency of usage of each segment is updated based on its use in matching the target pool of sequences.
[0072] For combinations of multiple segments that maximally match a subject sequence in a repertoire, the frequency of usage of each segment can be increased, for example, by the inverse of the number of matching combinations.
[0073] If some segment types are wildcards, the portions of the subject sequence that match the wildcards may be extracted as new segments and their frequency of use may be updated as described above.
[0074] Section VII: Using both (i) the identified sequence and (ii) one or more structurally significant features of the sequence predicted by the first model as inputs to a second model (e.g., to predict chemical stability, polyspecificity, and hydrophobicity of the composition in lieu of structural information (e.g., without software-generated structures)). What is found here is that it is possible to use (i) an identified sequence and (ii) one or more structurally significant properties of the sequence predicted by a first model to predict developability properties (such as chemical stability, polyspecificity, and hydrophobicity of a sequence composition) and thereby generate novel sequences or segments for consideration for inclusion in a synthetic library. In certain embodiments, the predicted structurally significant properties can be used in place of the structure itself (e.g., in place of a software-determined structure such as predicted by AlphaFold or similar software). Below is an illustrative example showing how this "model-to-model" concept can be used to improve the ability to predict chemical stability, polyspecificity, and hydrophobicity of a sequence composition.
[0075] 9 is a block diagram of a method for predicting important developable properties for therapeutics using prediction of structurally significant metrics of biophysical properties in a model according to an exemplary embodiment. On the left is a deep graph convolutional network that provides structural descriptors from sequences, such as per-residue predictions of SASA, charge patch, hydrophobic patch, and Cα / Cβ coordinates. On the top right is a deep convolutional and recurrent network for predicting biophysical properties, such as prediction of hydrophobic interaction chromatography retention time (HIC RT) and multispecific reagent (PSR) binding category (e.g., high vs. low) from Fv sequences. Structural descriptors from sequences are shown as inputs in the deep convolutional and recurrent network for predicting biophysical properties. On the bottom right is a tree-based classification model for predicting chemical stability properties, such as Asn deamidation, Asp isomerization, and Met oxidation. Again, structural descriptors from the sequences are presented as input in a tree-based classification model for prediction of chemical stability properties.
[0076] 10A and 10B are schematic diagrams illustrating the use of graph convolutional networks (GCNs) for residue-level prediction of structural descriptors, according to an exemplary embodiment. Sequences are used as input for prediction of structural descriptors (e.g., per-residue prediction of SASA, charge patch, hydrophobic patch, and Cα / Cβ coordinates).
[0077] Figure 10A represents a molecule as a graph with residues as nodes and edges between spatial neighbors. A graph convolutional network (GCN) can be trained to learn features and predict residue-level structural / biophysical properties as a combination of self-residue features and neighboring residue features. The learned feature for the central residue is concatenated with the weighted learned neighboring features, and then a downsampling convolution is performed nonlinearly to generate the feature output for the next layer.
[0078] Figure 10B is an overview of an exemplary graph convolution architecture and training data. In this example, there are four attention weight matrices shared between blocks. The final layer learns a different set of weights on the learned features to predict the output (e.g., Solvent Accessibility Assessment (SASA), which is an important feature for determining protein folding and stability).
[0079] FIG. 11 is a graph showing the overall spatial aggregation propensity (SAP) score, calculated by summing the individual residue predictions according to the method described above, which is comparable to that obtained from the AlphaFold2 model for the same input sequence.
[0080] FIG. 12 is a graph showing the overall scoring card method (SCM) score calculated by summing the individual residue predictions according to the method described above.
[0081] Figure 13 is a schematic of the Mollweide projection of the net charge patches predicted using the method described above. The GCN model generates both properties and Cα / Cβ coordinate predictions in a single model. The example shown in Figure 13 shows that the presence of large negative patches correlates with poor solubility.
[0082] FIG. 14 is a schematic diagram illustrating the use of convolutional and recurrent models for hydrophobicity and multispecificity prediction, according to an exemplary embodiment. Patterns in N-mer peptide sequences represent local information or "features". Interactions between N-mer peptides capture information over longer length scales, and also capture information between peptides separated along the sequence. In the schematic diagram of FIG. 14, the input layer I uses "input Fv" (antibody fragment sequence), which performs one-hot encoding, e.g., existing amino acid property scales, and prediction of residue-level structural / biophysical properties from GCNs as described above. The next step is feature extraction, where the convolutional layer of the network architecture learns separate features for, e.g., penta-peptides, and the recurrent layer learns features for the entire linear sequence. The next step is to combine the extracted features across HC and LC. The densely connected layer learns to globally combine patterns from previous layers. The output layer generates a prediction, e.g., HIC RT or PSR score / category, as described in more detail herein. EXAMPLES
[0083] Section VIII a.VHH H3 library design 1. Development feasibility model for segment selection The Fc-linker-VH library, synthesized such that H3 diversity reflects the human pre-immune repertoire, was sorted for expression and multispecificity using FACS (fluorescence activated cell sorting). The input library, high and low expressers, and high and low multispecificity groups were sequenced using NGS. The frequency of segment observations in the NGS sequence was used to generate an enrichment score for the segments, as outlined in Section II above.
number
number
[0084] Pre-generated segments were obtained based on V, D, and J-gene data as described in Section I. For de novo segment inference, the collection of sequences as detailed in Section I was used in conjunction with the matching algorithm of Section VI above.
[0085] The procedure for matching camelid H3 sequences is generally as follows. 1) A collection of D- and J-genes was used to generate candidate D- and J-segments. 2) The wildcard N1 segment matches any sequence of length 1 through 9. 3) The wildcard N2 segment matches any sequence of length 0 to 7.
[0086] CDR H3 sequences from McCoy LE, PLoS Pathogens, 10, e1004552, 2014, PMID: 25522326 and Li X, PLoS ONE, 11, e0161801, 2016, PMID: 27588755 were matched to the above pool of segments using the methods described in Section VI above, maximizing the following metrics for matching to D- and J-segments:
number
[0087] The following criteria were used to eliminate matches resulting from the above process. 1) the number of mismatches for the D- and J-segments is greater than 25% of the length of the D- or J-segment; 2) the total number of mismatches is greater than 5; and 3) Maximum matches resulting in Asp or Asn in the last position or Asn in the penultimate position are eliminated until a suitable match is found subject to the other constraints listed above.
[0088] In the case of multiple viable D- and J-segments that maximize S, each segment identified in a match was inversely weighted relative to the number N of matches.
[0089] The result from this procedure is a list of usage weights P for candidate D- and J-segments generated from the collection of D- and J-genes. In addition, this procedure generates a list of new candidate N1 and N2 segments along with their usage weights P.
number
[0090] An example for matching the CDR H3 sequence AAEPSGGSWPRYEYNF is shown in Figure 3, which uses a value of x=2 for the score S. 3. Segment Selection for the Final Library
[0091] The following steps were taken to complete segment selection for the final library. 1) Input the candidate segments from the previous step into a machine learning model for segment development potential and obtain their predicted enrichment scores.
number
number
number
number
number
number
number
[0092] After segments were selected in this manner, a representative combinatorial library was sampled in silico. This library was evaluated for predicted biophysical characteristics (e.g., polyspecificity and hydrophobicity) using the model described above. In addition, sequences from the natural repertoire were also evaluated for these properties and compared with de novo synthetic designs.
[0093] It should be noted that the principles and examples disclosed herein for VHH antibody (or nanobody) library design can be applied to library design including other antibody portions (e.g., light chain framework regions (LC FRs), light chain complementarity determining regions (LC CDRs), heavy chain framework regions (HC FRs), and others). As used herein, the term VHH refers to the antigen-binding fragment (i.e., variable domain) of a heavy chain-only antibody, e.g., a camelid heavy chain-only antibody.
[0094] b. Example of VλL3 library design As detailed in Section I, human Vλ germline sequences were collected from internal databases and external sources (eg, the OAS sequence database, literature, and patent applications).
[0095] The observed sequences were split into left and right fragments to mimic VJ recombination for CDR L3. Models were constructed using the methods outlined above for the collection of CDR L3, individual left sequences, and right sequences.
[0096] After generating a model to capture the composition and correlation of sequences in the input sequence set, the generated model was run in generation mode to generate new sequences or segments for consideration in the synthetic library. An exemplary sequence generation is shown in Figure 4B. The value of T in the sampling process of Figure 4B can be increased (decreased) to generate sequences that are closer (more distant) from the sequence set used to train the model.
[0097] The final selection of germline-specific sequences was performed in the following manner. 1) Obtain the probability of the generated sequence from the generative model. 2) As detailed in Section IV above, polyspecificity and hydrophobicity scores are evaluated based on the CDR-specific amino acid coefficients obtained using a logistic regression model or a neural network model. The polyspecificity and hydrophobicity scores were converted to percentile ranks in 5% increments (lower numbers indicate more favorable properties). 3) Evaluate the probability of chemical modifications in the sequence from a neural network or tree-based regression or classification model as described in Section IV. 4) For the generated sequences, calculate the number of mutations from the germline across the entire sequence and across the preferred antigen contact residues. 5) converting sequence probabilities, mutation information, and predicted polyspecificity, hydrophobicity, and chemical stability scores (e.g., developability rankings) from the generative model into sequence priority scores; and 6) Select the top sequences based on their priority scores or draw a random sample.
[0098] Based on factors such as germline diversity, length distribution, etc., required in the final library, prioritized, differing proportions of sequences from the germline-specific libraries can be pooled together for the final synthetic library.
[0099] c. Example of CDR H1 H2 library design Sequences from the human IGHV3 germline were collected from external sources (e.g., the OAS sequence database), internal databases, literature, and patent applications, as detailed above in Section I. Sequences from llama and camelid NGS datasets were processed from literature studies.
[0100] These sequences were renumbered and the CDR H1 and H2 sequences were extracted. The process of VλL3 library design is followed by a subsequent process to train a pattern learning model and run the model in production mode.
[0101] For example, germline-specific models were constructed for a collection of CDRs H1 and H2 using the methods outlined above.
[0102] After generating a model to capture the composition and correlation of sequences in the input sequence set, the generated model was run in generation mode to generate new sequences for consideration in the synthetic library. An exemplary sequence generation is shown in Figure 4A. The value of T in the sampling process of Figure 4A can be increased (decreased) to generate sequences that are closer (farther) from the sequence set used to train the model.
[0103] The final selection of germline-specific sequences was performed in the following manner. 7) Obtain the probability of the generated sequence from the generative model. 8) As detailed in Section III above, polyspecificity and hydrophobicity scores are evaluated based on the CDR-specific amino acid coefficients obtained using a logistic regression model or a neural network model. The polyspecificity and hydrophobicity scores were converted to percentile ranks in 5% increments (lower numbers indicate more favorable properties). 9) For the generated sequences, calculate the number of mutations from the germline across the entire sequence and across the preferred antigen contact residues. 10) Convert the sequence probabilities, mutation information, and exploitability rankings from the generative model into sequence priority scores; and 11) Select the top sequences based on their priority scores or draw a random sample.
[0104] Based on factors such as germline diversity, length distribution, etc., required in the final library, prioritized, different percentages of sequences from the germline-specific libraries can be pooled together for the final synthetic library.
[0105] The subsequent process of training the model to learn the patterns and running the model in production mode follows the process of VλL3 library design described above.
[0106] d. Examples of VκL3 sequence design Data on antibodies with suitable developability properties and known heavy and light chain sequences were collected, aligned, renumbered, and annotated with germline information. CDR L3 was then extracted and the amino acids at positions L89-L97 were tabulated.
[0107] Referring to the notation in Section IVd above, each diversity set i is a set of human heavy chain germline
number
number
number
[0108]
number
number
[0109] A KL calculation for a single amino acid choice at every position results in a two-dimensional table, where the rows indicate the positions, the columns indicate the amino acids, and the KL metric is a numerical value. An additional two-dimensional table was also constructed, with the same rows and columns, but including the number of amino acids found at each position.
[0110] These tables were used to select sequences from a larger set of CDR L3 sequences by the following procedure. 1. From the tabulated counts and calculated KL scores, filter sequences where the selected individual amino acids are in low occurrence or high KL score positions, e.g., filter rare or highly biased selections; and 2. The remaining sequences are scored by summing the KL score and the number of tabulated position-specific amino acids in the sequence. Sequences are prioritized by two calculated metrics: a. The top order by descending count; and b. Top ordering of total KL scores in ascending order.
[0111] Different proportions of sequences resulting from criteria 2a and 2b can be used to select the desired number of sequences in the library.
[0112] Software, computer systems, and network environments Certain embodiments described herein utilize computer algorithms in the form of software instructions executed by a computer processor. In certain embodiments, the software instructions include a machine learning module, also referred to herein as artificial intelligence software. As used herein, a machine learning module refers to a computer-implemented process (e.g., software function) that executes one or more specific machine learning algorithms (e.g., artificial neural networks (ANN), convolutional neural networks (CNN), random forests, decision trees, support vector machines, etc.) to determine one or more output values for a given input. In certain embodiments, the input includes alphanumeric data, which may include, for example, numbers, words, phrases, or long strings of characters. In certain embodiments, the one or more output values include values that represent numbers, words, phrases, or other alphanumeric strings. In certain embodiments, the one or more output values include those that identify one or more response strings (e.g., selected from a database).
[0113] For example, a machine learning module can receive as input a text string (e.g., entered by a human user) and generate various outputs. For example, a machine learning module can automatically analyze input alphanumeric string(s) and determine an output value that classifies the content (e.g., intent) of the text, e.g., as in natural language understanding (NLU). In certain embodiments, the text string is analyzed to generate and / or derive an output alphanumeric string. For example, the machine learning module can be (or include) natural language processing (NLP) software.
[0114] In certain embodiments, a machine learning module that performs machine learning methods is trained, e.g., using a dataset that includes the categories of data described herein. Such training can be used to determine various parameters (e.g., weights associated with layers in a neural network, etc.) of the machine learning algorithm executed by the machine learning module. In certain embodiments, once the machine learning module is trained to accomplish a particular task, e.g., identifying a particular response string, the determined parameter values are fixed, and the (e.g., immutable, static) machine learning module is used to process new data (e.g., different from the training data) to accomplish the trained task without further updates to its parameters (e.g., the machine learning module does not receive feedback and / or updates). In certain embodiments, the machine learning module may receive feedback, e.g., based on a user review of accuracy, and such feedback may be used as additional training data to dynamically update the machine learning module. In certain embodiments, two or more machine learning modules may be combined and executed as a single module and / or a single software application. In certain embodiments, two or more machine learning modules may also be executed separately, e.g., as separate software applications. The machine learning modules may be software and / or hardware. For example, the machine learning module may be implemented entirely as software, or certain functions of the ANN module (e.g., a CNN) may be performed via dedicated hardware (e.g., via an application specific integrated circuit (ASIC)).
[0115] FIG. 6 illustrates and describes the implementation of a network environment 600 for providing the systems, methods, and architectures as described herein. Referring now to FIG. 6, a block diagram of an exemplary cloud computing environment 600 is illustrated and generally described. The cloud computing environment 600 may include one or more resource providers 602a, 602b, 602c (collectively 602). Each resource provider 602 may include computing resources. In some embodiments, the computing resources may include any hardware and / or software used to process data. For example, the computing resources may include hardware and / or software capable of executing algorithms, computer programs, and / or computer applications. In some embodiments, exemplary computing resources may include application servers and / or databases with storage and retrieval capabilities. Each resource provider 602 may be connected to any other resource provider 602 in the cloud computing environment 600. In some embodiments, the resource providers 602 may be connected through a computer network 608. Each resource provider 602 may be connected to one or more computing devices 604 a , 604 b , 604 c (collectively 604 ) through a computer network 608 .
[0116] The cloud computing environment 600 may include a resource manager 606. The resource manager 606 may be connected to the resource providers 602 and the computing devices 604 through a computer network 608. In one embodiment, the resource manager 606 may facilitate the provision of computing resources by one or more resource providers 602 to one or more computing devices 604. The resource manager 606 may receive a request for a computing resource from a particular computing device 604. The resource manager 606 may identify one or more resource providers 602 that have the capability to provide the computing resource requested by the computing device 604. The resource manager 606 may select a resource provider 602 that provides the computing resource. The resource manager 606 may facilitate a connection between the resource provider 602 and a particular computing device 604. In one embodiment, the resource manager 606 may establish a connection between a particular resource provider 602 and a particular computing device 604. In one embodiment, the resource manager 606 may redirect a particular computing device 604 to a particular resource provider 602 that has the requested computing resource.
[0117] 7 illustrates examples of a computing device 700 and a mobile computing device 750 that can be used to perform the techniques described in this disclosure. The computing device 700 is intended to represent various forms of digital computers (e.g., laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers). The mobile computing device 750 is intended to represent various forms of mobile devices (e.g., personal digital assistants, mobile phones, smartphones, and other similar computing devices). The components illustrated herein, their connections and relationships, and their functions are meant to be exemplary only and not limiting.
[0118] The computing device 700 comprises a processor 702, a memory 704, a storage device 706, a high-speed interface 708 connecting to the memory 704 and multiple high-speed expansion ports 710, and a low-speed interface 712 connecting to a low-speed expansion port 714 and the storage device 706. Each of the processor 702, the memory 704, the storage device 706, the high-speed interface 708, the high-speed expansion port 710, and the low-speed interface 712 are interconnected by various buses, which may be mounted on a common motherboard or in other formats as desired. The processor 702 is capable of processing instructions for execution within the computing device 700. The instructions include instructions stored in the memory 704 or the storage device 706, which display graphical information for a GUI on an external input / output device (such as, for example, a display 716 connected to the high-speed interface 708). In other implementations, multiple processors and / or multiple buses may be used, as desired, along with multiple memories and multiple types of memories. Also, multiple computing devices may be connected, with each device providing a portion of the required operations (e.g., as a bank of servers, a group of blade servers, or a multi-processor system). Thus, as the terms are used herein, when functions are described as being performed by a "processor," this encompasses embodiments in which the functions are performed by any number of processor(s) in any number of computing device(s). Furthermore, when a function is described as being performed by a "processor," this encompasses embodiments in which the function is performed by any number of processor(s) in any number of computing device(s) (e.g., in a distributed computing system).
[0119] The memory 704 stores information within the computing device 700. In some implementations, the memory 704 is a volatile memory unit or units. In some implementations, the memory 704 is a non-volatile memory unit or units. The memory 704 may also be another form of computer-readable medium, such as a magnetic disk or an optical disk.
[0120] The storage device 706 has mass storage capability for the computing device 700. In an embodiment, the storage device 706 may be or may include a computer-readable medium. The medium may be, for example, a floppy disk device, a hard disk device, an optical disk device, or a tape device, an array of devices including flash memory or other similar solid-state memory devices, devices in a storage area network or other configuration, and the like. The instructions may be stored on an information carrier. The instructions, when executed by one or more processing devices (e.g., the processor 702), perform one or more methods, such as those described above. The instructions may also be stored in one or more storage devices (e.g., the memory 704, the storage device 706, or a memory on the processor 702), such as a computer-readable medium or a machine-readable medium.
[0121] The high-speed interface 708 manages bandwidth-intensive operations for the computing device 700, while the low-speed interface 712 manages less bandwidth-intensive operations. Such an allocation of functions is merely an example. In one embodiment, the high-speed interface 708 is connected to the memory 704, the display 716 (e.g., through a graphics processor or accelerator), and a high-speed expansion port 710 that may accept various expansion cards (not shown). In this embodiment, the low-speed interface 712 is connected to the storage device 706 and the low-speed expansion port 714. The low-speed expansion port 714, which may include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), may be connected, for example, through a network adapter to one or more input / output devices such as a keyboard, pointing device, scanner, or a networking device such as a switch or router.
[0122] Computing device 700 may be implemented in a number of different forms, as shown in the figure. For example, it may be implemented as a standard server 720, or multiple times in a group of such servers. In addition, it may be implemented in a personal computer, such as a laptop computer 722. It may also be implemented as part of a rack server system 724. Alternatively, components from computing device 700 may be combined with other components in a mobile device (not shown), such as mobile computing device 750. Each such device may include one or more of computing device 700 and mobile computing device 750, and the entire system may be composed of multiple computing devices in communication with each other.
[0123] The mobile computing device 750 includes, among other components, a processor 752, a memory 764, an input / output device such as a display 754, a communication interface 766, and a transceiver 768. The mobile computing device 750 may also be provided with a storage device, such as a microdrive or other device, to provide additional storage. Each of the processor 752, memory 764, display 754, communication interface 766, and transceiver 768 are interconnected by various buses, and some of the components may be mounted on a common motherboard or in other forms, as appropriate.
[0124] The processor 752 can execute instructions within the mobile computing device 750, including instructions stored in the memory 764. The processor 752 can be implemented as a chipset of chips including multiple separate analog and digital processors. The processor 752 can be responsible for coordinating other components of the mobile computing device 750, such as controlling a user interface, applications run by the mobile computing device 750, and wireless communication by the mobile computing device 750.
[0125] The processor 752 may communicate with a user through a control interface 758 and a display interface 756 connected to a display 754. The display 754 may be, for example, a TFT display (thin film-transistor liquid crystal display) or an OLED (organic light-emitting diode) display, or other suitable display technology. The display interface 756 may include appropriate circuitry for driving the display 754 to provide graphical and other information to the user. The control interface 758 may receive commands from the user and convert them for sending to the processor 752. Additionally, an external interface 762 may enable communication with the processor 752 to enable short-range communication of the mobile computing device 750 with other devices. The external interface 762 may provide, for example, wired communication in some implementations and wireless communication in other implementations, and multiple interfaces may also be used.
[0126] The memory 764 stores information within the mobile computing device 750. The memory 764 may be implemented as one or more of a computer-readable medium or media, a volatile memory unit or units, or a non-volatile memory unit or units. An expansion memory 774 may also be provided, which may be connected to the mobile computing device 750 through an expansion interface 772, which may include, for example, a SIM (single in-line memory module) card interface. The expansion memory 774 may provide additional storage space for the mobile computing device 750 or may also store applications or other information for the mobile computing device 750. In particular, the expansion memory 774 may include instructions that perform or complement the processes described above, and may also include secure information. Thus, for example, the expansion memory 774 may be provided as a security module for the mobile computing device 750 and may be programmed with instructions that enable secure use of the mobile computing device 750. Additionally, secure applications may be provided via the SIM card along with additional information, such as placing identifying information on the SIM card in a manner that cannot be hacked.
[0127] The memory may include, for example, flash memory and / or NVRAM memory (non-volatile random access memory), as described below. In an embodiment, the instructions are stored on an information carrier. The instructions, when executed by one or more processing devices (e.g., processor 752), perform one or more methods, such as those described above. The instructions may also be stored in one or more storage devices, such as one or more computer-readable or machine-readable media (e.g., memory 764, expansion memory 774, or memory on processor 752). In an embodiment, the instructions may be received by a propagated signal through transceiver 768 or external interface 762.
[0128] The mobile computing device 750 may communicate wirelessly through a communications interface 766, which may include digital signal processing circuitry, if necessary. The communications interface 766 may support communications under various modes or protocols, such as GSM voice calls (Global System for Mobile Communications), SMS (Short Message Service), EMS (Enhanced Messaging Service), or MMS messaging (Multimedia Messaging Service), CDMA (Code Division Multiple Access), TDMA (Time Division Multiple Access), PDC (Personal Digital Cellular), WCDMA (Wideband Code Division Multiple Access), CDMA2000, or GPRS (General Packet Radio Service), among others. Such communications may be performed, for example, by radio frequencies through a transceiver 768. Additionally, short-range communications may be performed, such as by Bluetooth, Wi-Fi, or other such transceivers (not shown). In addition, a GPS (Global Positioning System) receiver module 770 can transmit additional navigational and location-related radio data to the mobile computing device 750, such data can be used by applications running on the mobile computing device 750 as desired.
[0129] The mobile computing device 750 may also communicate audibly using an audio codec 760, which may receive spoken information from a user and convert it into usable digital information. The audio codec 760 may also generate audible sounds for the user, such as through a speaker in a handset of the mobile computing device 750. Such sounds may include sounds from voice telephone calls, may include recorded sounds (e.g., voice messages, music files, etc.), and may also include sounds generated by applications running on the mobile computing device 750.
[0130] The mobile computing device 750 may be implemented in a number of different forms, as shown in the figure. For example, it may be implemented as a mobile phone 760. It may also be implemented as part of a smartphone 782, a personal digital assistant, or other similar mobile device.
[0131] Various implementations of the systems and techniques described herein can be realized in digital electronic circuitry, integrated circuits, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementation in one or more computer programs that can be executed and / or interpreted on a programmable system, such a system comprising at least one programmable processor, which may be special purpose or general purpose, and which may be coupled to receive data and instructions from and send data and instructions to a storage system, at least one input device, and at least one output device.
[0132] These computer programs (also known as programs, software, software applications, or code) include machine instructions for a programmable processor and may be implemented in high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. As used herein, the terms machine-readable medium and computer-readable medium refer to any computer program product, apparatus, and / or device (e.g., magnetic disks, optical disks, memory, programmable logic circuits (PLDs)) used to provide machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as a machine-readable signal. The term machine-readable signal refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0133] To interact with a user, the systems and techniques described herein can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user, as well as a keyboard and a pointing device (e.g., a mouse or trackball) by which the user can provide input to the computer. Similarly, other types of devices can be used to interact with the user. For example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback), and input from the user can be received in any form, including acoustic input, speech input, or tactile input.
[0134] The systems and techniques described herein can be implemented in a computing system. Such a computing system may include a back-end component (e.g., as a data server), a middleware component (e.g., an application server), a front-end component (e.g., a client computer having a graphical user interface or a web browser through which a user can interact with an embodiment of the systems and techniques described herein), or any combination of back-end, middleware, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communications network). Examples of communications networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0135] A computing system may include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
[0136] In some implementations, certain modules described herein may be separated, combined, or incorporated into a single or combined module. Any modules shown in the figures are not intended to limit the systems described herein to the illustrated software architecture.
[0137] Elements of different embodiments described herein may be combined to create other embodiments not specifically described above. Elements may be removed from the processes, computer programs, databases, etc. described herein without adversely affecting their operation. Furthermore, the illustrated logic flows do not require the particular order or sequential order shown to achieve desirable results. Various individual elements may be combined into one or more individual elements to perform the functions described herein.
[0138] While the present invention has been particularly shown and described with reference to certain preferred embodiments, it will be apparent to those skilled in the art that various changes in form and details may be made therein without departing from the spirit and scope of the invention as defined in the appended claims.
Claims
1. A system for constructing an antibody library, comprising: a processor of a computing device; a memory having instructions stored therein; The instructions, when executed by the processor, cause the processor to: (i) developing a first machine learning model using the input sequence and characterization data; (ii) using the first machine learning model in (i) to predict desirable segments and enable selection of segments from a pool of novel and / or pre-generated segments; (iii) processing a set of input sequences prior to selection and / or use in training the first machine learning model in (i), the processing comprising one or more of: (a) modifying the sequences to remove chemical liability sites; (b) for CDR H3, dividing the sequences into segments to mimic VDJ recombination; (c) for CDR L3, dividing the sequences into segments to mimic VJ recombination; and (d) annotating V-regions and CDRs (H1, H2, L3) with the number of mutations from germline; (iv) training a machine learning model for the prediction of biophysical and / or biochemical properties; (v) predicting one or more biophysical and / or biochemical properties from the sequence using the machine learning model for predicting biophysical and / or biochemical properties in (iv); (vi) developing an autoregressive deep learning neural network model to learn joint sequence probability distributions across sequences of interest for specific germline sequences for different species; and (vii) using the neural network model in (vi) to capture sequence composition and / or correlation from an input set of sequences to generate novel sequences or segments for consideration in a synthetic library.
2. A system for constructing an antibody library, comprising: a processor of a computing device; a memory having instructions stored therein; The instructions, when executed by the processor, cause the processor to process a set of input sequences through one or more machine learning models to generate a collection of final antibody library sequences.
3. 3. The system of claim 2, wherein the instructions cause the processor to (i) process each input sequence from the set of input sequences and (ii) for each of the input sequences, a per-residue prediction of one or more structurally important properties of the sequence as predicted by a first model, and the instructions cause the processor to process (i) and (ii) as inputs in a second model and predict as outputs of the second model (iii) one or more biophysical properties and / or (iv) one or more chemical stability properties for each of the input sequences, wherein the inclusion or exclusion of each sequence in the final antibody library is based at least in part on the output of the second model.
4. 4. The system of claim 3, wherein the per-residue predictions predicted by the first model include one or more selected from the group consisting of: (i) a measure of solvent accessibility (SASA), (ii) a measure of charge patching, (iii) a measure of hydrophobic patching, and (iv) a Cα / Cβ coordinate prediction.
5. The system of claim 3 , wherein the second model comprises a deep convolutional and / or recurrent network (e.g., for predicting biophysical properties).
6. The system of claim 3 , wherein the second model comprises a tree-based classification model.
7. 1. A method for constructing an antibody library, comprising: using a processor of a computing device to perform the following (i), (ii), (iii), (iv), (v), (vi), and (vii): (i) developing a first machine learning model using the input sequence and characterization data; (ii) using the first machine learning model in (i) to predict desirable segments and enable selection of segments from a pool of novel and / or pre-generated segments; (iii) processing a set of input sequences prior to selection and / or use in training the first machine learning model in (i), by (a) modifying the sequences; (b) for CDR H3, dividing the sequence into segments to mimic VDJ recombination; (c) for CDR L3, dividing the sequence into segments to mimic VJ recombination; and (d) annotating the V-regions and CDRs (H1, H2, L3) with the number of mutations from germline. (iv) training a machine learning model for the prediction of biophysical and / or biochemical properties; (v) predicting one or more biophysical and / or biochemical properties from the sequence using the machine learning model for predicting biophysical and / or biochemical properties in (iv); (vi) developing an autoregressive deep learning neural network model to learn joint sequence probability distributions across sequences of interest for specific germline sequences for different species; and (vii) using the neural network model in (vi) to capture sequence composition and / or correlation from an input set of sequences to generate novel sequences or segments for consideration in a synthetic library.
8. 1. A method for constructing an antibody library, comprising: processing the set of input sequences with a processor of a computing device using one or more machine learning models to generate a collection of final antibody library sequences.
9. 9. The method of claim 8, comprising: (i) processing each input sequence from the set of input sequences as an input in a second model; and (ii) for each of the input sequences, processing a per-residue prediction of one or more structurally important properties of the sequence as predicted by the first model; and predicting, as an output of the second model, (iii) one or more biophysical properties and / or (iv) one or more chemical stability properties for each of the input sequences, wherein the inclusion or exclusion of each sequence in the final antibody library is based at least in part on the output of the second model.
10. 10. The method of claim 9, wherein the per-residue predictions predicted by the first model include one or more selected from the group consisting of: (i) a measure of solvent accessibility (SASA), (ii) a measure of charge patching, (iii) a measure of hydrophobic patching, and (iv) a Cα / Cβ coordinate prediction.
11. The method of claim 9 , wherein the second model comprises a deep convolutional and / or recurrent network.
12. The method of claim 9 , wherein the second model comprises a tree-based classification model.
13. The system of claim 1, wherein the instructions, when executed by the processor, cause the processor to develop the first machine learning model using input sequences and characterization data. (a) training a logistic regression model to derive amino acid coefficients to predict the polyspecificity and hydrophobicity of individual complementarity determining regions (CDRs) and / or framework regions (FRs); (b) training a tree model to predict one or more biophysical properties and / or one or more chemical stability properties from the sequence; (c) training a deep learning model, including a neural network, to predict one or more biophysical properties and / or one or more chemical stability properties from the sequence; (d) creating a statistical model to assess bias to select sequences with low bias; and (e) Developing hierarchical statistics to predict the risk of chemical modification as a function of sequence motifs at specific positions and regions. comprising one or more selected from the group consisting of: The system.
14. The method of claim 7, comprising using the processor to develop the first machine learning model using input sequences and characterization data, wherein developing the first machine learning model comprises: (a) training a logistic regression model to derive amino acid coefficients to predict the polyspecificity and hydrophobicity of individual complementarity determining regions (CDRs) and / or framework regions (FRs); (b) training a tree model (e.g., Random Forest or XGBoost) to predict one or more biophysical properties and / or one or more chemical stability properties from the sequence; (c) training a deep learning model, including a neural network, to predict one or more biophysical properties and / or one or more chemical stability properties from the sequence; (d) creating a statistical model to assess bias to select sequences with low bias; and (e) Developing hierarchical statistics to predict the risk of chemical modification as a function of sequence motifs at specific positions and regions. comprising one or more selected from the group consisting of: The method.