Artificial intelligence (AI) validation dataset generation systems and methods for generating ai validation datasets for validating ai-based protein sequence models
The described method generates AI validation datasets by filtering and clustering protein sequences to create a validation dataset independent of the training dataset, addressing the unreliability of conventional methods and enhancing model accuracy and deployment efficiency.
Patent Information
- Application Number
- PCT/US2025/035786
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-09-19
- Filing Date
- 2025-06-27
- Publication Date
- 2026-01-02
AI Technical Summary
Conventional methods for creating validation sets for training large language models on protein sequences result in unreliable validation metrics due to data manipulation and clustering from the same sequences, leading to inaccurate model performance assessment.
A method for generating AI validation datasets that includes filtering a reference proteome dataset to create a high completeness proteome data subset, clustering protein sequences, and randomly allocating a validation cluster subset to form a protein sequence validation dataset, independent of the training dataset, allowing for accurate model validation and assessment.
This approach enables more accurate model assessment with lower perplexity, loss, and higher accuracy, facilitating the improvement of AI-based protein sequence models by providing reliable validation metrics and enabling their deployment on devices with fewer computational resources.
Smart Images

Figure US2025035786_02012026_PF_FP_ABST
Abstract
Description
ARTIFICIAL INTELLIGENCE (Al) VALIDATION DATASET GENERATION SYSTEMS AND METHODS FOR GENERATING Al VALIDATION DATASETS FOR VALIDATING ALBASED PROTEIN SEQUENCE MODELSRELATED APPLICATIONS
[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 666,028 (filed on June 28, 2024); and U.S. Provisional Application No. 63 / 696,680 (filed on September 19, 2024). The entirety of each of the foregoing provisional applications is incorporated by reference herein.FIELD
[0002] The present disclosure generally relates to artificial intelligence (Al) validation dataset generation systems and methods, and more particularly to, Al validation dataset generation systems and methods for generating Al validation datasets for validating Al-based protein sequence models.BACKGROUND
[0003] Training large language models (LLMs) on protein sequences includes the creation of a validation set, which is a set of data for testing the LLM. This allows, for example, the calculation of metrics related to overfitting and model generalizability.
[0004] These metrics can demonstrate model performance and they are intrinsically linked to the validation set used; on an algorithmic level the validation set can be essential to computing the metric. The set itself serves as an essential tool for identifying which models have the best performance.
[0005] Two conventional methods for building validation sets include either to take a random split of the training data as a leave out set, or to cluster a given training set and take out a random set of these clusters. For example, the current state of the art family of protein LLMs, evolutionary scale modeling 2 (ESM2), was trained by Meta-AI using a validation set (e.g., UniRef50 or UniRef90), developed with these conventional methods based on a public dataset. Each of the validation datasets (e.g., UniRef50 or UniRef90) are preprocessed dataset having data removed or otherwise altered to reduce available information or data before training.
[0006] Using the public dataset as training data means that a conventional clustering step is part of a given experiment, wherein clustering is a manipulated variable. However, when validation sets are drawn from the same clustered sequences, this variable becomes part of the control and therefore corrupts or otherwise makes unreliable the given validation set, such that any validation metrics based on the conventional validation set are no longer able to effectively identify the effect in view of the conventional data manipulation techniques.
[0007] For the foregoing reasons, there is a need for Al validation dataset generation systems and methods for generating Al validation datasets for validating Al-based protein sequence models.SUMMARY
[0008] The Al validation dataset generation systems and methods described herein disclose techniques for generating Al validation datasets (e.g., a protein sequence validation dataset) for validating Al-based protein sequence models. The techniques here can create a protein sequence validation dataset that can be used to accurately validate or otherwise test for predictive accuracy of a trained Al-based protein sequence model that has been trained with protein sequences defined by or otherwise comprising the reference proteome dataset. The protein sequence validation dataset may enable a statistics test defining how to exclude data from the training set in a way that allows for accurate testing of the Al-based protein sequence model. Such accurate testing can eliminate errors (e.g., such as false positives and / or false negatives) as to the output (e.g., predictions or classifications) of the Al-based protein sequence model.
[0009] Further, the validation dataset generation systems and methods as described herein can maintain model quality and, at the same time, allow for generation and training of Al models that have fewer model parameters, and thus may be stored or installed on devices having lower computer memory. Such smaller memory models allow for distribution, deployment, and installation on deployable devices with fewer computational resources.
[0010] Still further, the protein sequence validation datasets as created or generated using the techniques herein allow for more accurate model assessment, where validation metrics (e.g., metrics for validation and assessing Al models) can be used to demonstrate Al model efficacy, including lower perplexity and loss, as well as higher accuracy, for a given Al-based protein sequence model. This demonstrates that training and validating using the proteinsequence validation dataset as described herein can be used to test Al models, update Al models, and thus improve performance, or otherwise identified accuracy, of the given AI- based protein sequence model. That is, in various aspects, the validation metrics provide predictors of performance for enabling updating and improvement of to improve real world performance of an Al model. The validation metrics enable measurement of perplexity and loss, which allows identification of when such metrics or low or otherwise unacceptable in the course of an experiment. Validation metrics provide a prediction of model quality. That is, a validation set, and related validation metrics for a given validation set, allows selection of models that have lower perplexity and loss and higher accuracy.
[0011] In addition, the present disclosure includes specific features other than what is well- understood, routine, conventional activity in the field, or adding unconventional steps that confine the claim to a particular useful application, e.g., Al validation dataset generation systems and methods for generating Al validation datasets for validating Al-based protein sequence models.
[0012] In some aspects, the techniques described herein relate to an artificial intelligence (Al) validation dataset generation system configured to generate Al validation datasets for validating Al-based protein sequence models, the Al validation dataset generation system including: one or more processors; a computer memory communicatively coupled to the one or more processors; a reference proteome dataset stored on the computer memory and including a plurality of reference proteomes; and computing instructions stored on the computer memory, and that when executed by the one or more processors, cause the one or more processors to: filter the reference proteome dataset to generate a high completeness proteome data subset including reference proteomes selected from the plurality of reference proteomes having a completeness value above a completeness threshold; filter the high completeness proteome data subset to generate a protein confidence-based proteome data subset including reference proteomes selected from the high completeness proteome data subset having a protein confidence score above a protein confidence threshold, wherein the protein confidence-based proteome data subset includes a plurality of protein sequences, cluster the plurality of protein sequences to generate a set of protein sequence clusters, each cluster of the plurality of protein sequences grouping one or more of the plurality of protein sequences of the protein confidence-based proteome data subset, randomly allocate a validation cluster subset selected from the set of protein sequence clusters, and create aprotein sequence validation dataset by selecting the protein sequences of the validation cluster subset.
[0013] In some aspects, the techniques described herein relate to an Al validation dataset generation system, wherein the computing instructions, when executed by the one or more processors, further cause the one or more processors to: train an Al-based protein sequence model with a protein sequence training dataset, wherein the protein sequence validation dataset includes a dataset independent of the protein sequence training dataset used to train the Al-based protein sequence model; and output, from the Al-based protein sequence model, one or more predicted protein sequences.
[0014] In some aspects, the techniques described herein relate to an Al validation dataset generation system, wherein the one or more predicted protein sequences include data for predicting or determining one or more of: viscosity, post-translational modification, isomerization, thermal stability, pH, expression titre, aggregation, immunogenicity, binding affinity, binding avidity, clipping, protein secondary structure, protein tertiary structure, protein quaternary structure, disorder, function, enzymatic activity, homology, cellular localization, phenotypic effect, and / or half-life.
[0015] In some aspects, the techniques described herein relate to an Al validation dataset generation system, wherein the protein sequence training dataset includes protein sequences selected from at least a portion of the protein confidence-based proteome data subset including the plurality of protein sequences.
[0016] In some aspects, the techniques described herein relate to an Al validation dataset generation system, wherein the protein sequence training dataset includes protein sequences different from those selected for the protein sequence validation dataset from the validation cluster subset.
[0017] In some aspects, the techniques described herein relate to an Al validation dataset generation system, wherein at least one of the plurality of reference proteomes is linked to one or more sub-fragments or one or more residues of a given protein sequence in the reference proteome dataset.
[0018] In some aspects, the techniques described herein relate to an Al validation dataset generation system, wherein at least one of the validation cluster subset or the protein sequence validation dataset includes the one or more sub-fragments or the one or more residues of the given protein sequence in the reference proteome dataset.
[0019] In some aspects, the techniques described herein relate to an Al validation dataset generation system, wherein the computing instructions when executed by the one or more processors, further cause the one or more processors to: validate the one or more predicted protein sequences as output by the Al-based protein sequence model with the protein sequence validation dataset.
[0020] In some aspects, the techniques described herein relate to an Al validation dataset generation system, wherein creation of the protein sequence validation dataset allocates a larger protein sequence training dataset than a randomly sampled validation dataset randomly sampled from the plurality of reference proteomes of the reference proteome dataset, and wherein the protein sequence validation dataset and the randomly sampled validation dataset have at least one of: (1) a same sequence identity threshold value; or (2) an equal number of protein sequences selected from the plurality of reference proteomes of the reference proteome dataset.
[0021] In some aspects, the techniques described herein relate to an Al validation dataset generation system, wherein the computing instructions when executed by the one or more processors, further cause the one or more processors to: remove, based on a sequence identity threshold value, one or more of the protein sequences of the protein sequence validation dataset to create an ablated protein sequence validation dataset.
[0022] In some aspects, the techniques described herein relate to an Al validation dataset generation system, wherein the removed one or more of the protein sequences includes no more than 5% to 10% of the protein sequence validation dataset.
[0023] In some aspects, the techniques described herein relate to an Al validation dataset generation system, wherein the computing instructions, when executed by the one or more processors, further cause the one or more processors to: update the protein sequence validation dataset to include protein sequences each having a maximum protein length that the Al-based protein sequence model was trained on.
[0024] In some aspects, the techniques described herein relate to an Al validation dataset generation system, wherein the reference proteome dataset includes proteins having different protein lengths, and wherein the computing instructions, when executed by the one or more processors, further cause the one or more processors to: truncate one or more of protein lengths of at least a subset of the proteins to have a truncated length, the truncated length including a maximum content length for training the Al-based protein sequence model, andtrain the Al-based protein sequence model with training data including at least a portion of the reference proteome dataset having the subset of the proteins having the truncated length.
[0025] In some aspects, the techniques described herein relate to an Al validation dataset generation system, wherein the completeness threshold includes a value of 2 / 3 completeness or greater.
[0026] In some aspects, the techniques described herein relate to an Al validation dataset generation system, wherein 5% to 15% of the set of protein sequence clusters are randomly selected for creation of the protein sequence validation dataset.
[0027] In some aspects, the techniques described herein relate to an Al validation dataset generation system, wherein 15% or more of the set of protein sequence clusters are randomly selected for creation of the protein sequence validation dataset.
[0028] In some aspects, the techniques described herein relate to an Al validation dataset generation system, wherein the computing instructions, when executed by the one or more processors, further cause the one or more processors to: generate validation metrics based on comparison between: (a) a set of predetermined values of the protein sequence validation dataset; and (b) the one or more predicted protein sequences as output by the Al-based protein sequence model .
[0029] In some aspects, the techniques described herein relate to an Al validation dataset generation system, wherein the validation metrics include at least one improved metric selected from: a perplexity metric, a loss metric, and / or an accuracy metric, when analyzed against validation metrics of a previously trained version of the Al-based protein sequence model.
[0030] In some aspects, the techniques described herein relate to an Al validation dataset generation system, wherein the at least one improved metric includes: a reduced perplexity value, a reduced loss value, and / or an increased accuracy value.
[0031] In some aspects, the techniques described herein relate to an Al validation dataset generation system, wherein the at least one improved metric includes: a reduced perplexity value when the Al-based protein sequence model is trained on a UniRefl 00 dataset, a reduced loss value when the Al-based protein sequence model is trained on the UniRefl 00 dataset, and / or an increased accuracy value when the Al-based protein sequence model is trained on the UniRefl 00 dataset.
[0032] In some aspects, the techniques described herein relate to an Al validation dataset generation system, wherein the at least one improved metric includes: a reduced perplexity value when the Al-based protein sequence model is trained on a UniRefl 00 dataset compared to a second Al-based protein sequence model trained on a UniRef50 dataset, a reduced loss value when the Al-based protein sequence model is trained on the UniRefl 00 dataset compared to a second Al-based protein sequence model trained on the UniRef50 dataset, and / or an increased accuracy value when the Al-based protein sequence model is trained on the UniRefl 00 dataset compared to a second Al-based protein sequence model trained on the UniRef50 dataset.
[0033] In some aspects, the techniques described herein relate to an Al validation dataset generation system, wherein the one or more predicted protein sequences as output by the AI- based protein sequence model is used to develop a protein-related product.
[0034] In some aspects, the techniques described herein relate to an Al validation dataset generation system, wherein the protein-related product is predicted to treat a disease.
[0035] In some aspects, the techniques described herein relate to an Al validation dataset generation system, wherein the protein-related product is predicted to provide a prophylaxis for a disease or medical condition.
[0036] In some aspects, the techniques described herein relate to an Al validation dataset generation system, wherein the output of the one or more predicted protein sequences is used to monitor or measure the protein-related product during manufacture of the protein-related product.
[0037] In some aspects, the techniques described herein relate to an artificial intelligence (Al) validation dataset generation method for generating Al validation datasets for validating AI- based protein sequence models, the Al validation dataset generation method including: filtering, by one or more processors, a reference proteome dataset to generate a high completeness proteome data subset including reference proteomes selected from a plurality of reference proteomes having a completeness value above a completeness threshold, wherein the reference proteome dataset is stored on a computer memory communicatively coupled to the one or more processors, and wherein the reference proteome dataset includes the plurality of reference proteomes; filtering, by the one or more processors, the high completeness proteome data subset to generate a protein confidence-based proteome data subset including reference proteomes selected from the high completeness proteome data subset having aprotein confidence score above a protein confidence threshold, wherein the protein confidence-based proteome data subset includes a plurality of protein sequences; clustering, by the one or more processors, the plurality of protein sequences to generate a set of protein sequence clusters, each cluster of the plurality of protein sequences grouping one or more of the plurality of protein sequences of the protein confidence-based proteome data subset; randomly allocating, by the one or more processors, a validation cluster subset selected from the set of protein sequence clusters; and creating, by the one or more processors, a protein sequence validation dataset by selecting the protein sequences of the validation cluster subset.
[0038] In some aspects, the techniques described herein relate to an Al validation dataset generation method further including: training an Al-based protein sequence model with a protein sequence training dataset, wherein the protein sequence validation dataset includes a dataset independent of the protein sequence training dataset used to train the Al-based protein sequence model, and outputting, from the Al-based protein sequence model, one or more predicted protein sequences.
[0039] In some aspects, the techniques described herein relate to an Al validation dataset generation method, wherein the one or more predicted protein sequences include data for predicting or determining one or more of viscosity, post-translational modification, isomerization, thermal stability, pH, expression titre, aggregation, immunogenicity, binding affinity, binding avidity, clipping, protein secondary structure, protein tertiary structure, protein quaternary structure, disorder, function, enzymatic activity, homology, cellular localization, phenotypic effect, and / or half-life.
[0040] In some aspects, the techniques described herein relate to an Al validation dataset generation method, wherein the protein sequence training dataset includes protein sequences selected from at least a portion of the protein confidence-based proteome data subset including the plurality of protein sequences.
[0041] In some aspects, the techniques described herein relate to an Al validation dataset generation method, wherein the protein sequence training dataset includes protein sequences different from those selected for the protein sequence validation dataset from the validation cluster subset.
[0042] In some aspects, the techniques described herein relate to an Al validation dataset generation method, wherein at least one of the plurality of reference proteomes is linked toone or more sub-fragments or one or more residues of a given protein sequence in the reference proteome dataset.
[0043] In some aspects, the techniques described herein relate to an Al validation dataset generation method, wherein at least one of the validation cluster subset or the protein sequence validation dataset includes the one or more sub-fragments or the one or more residues of the given protein sequence in the reference proteome dataset.
[0044] In some aspects, the techniques described herein relate to an Al validation dataset generation method further including validating the one or more predicted protein sequences as output by the Al-based protein sequence model with the protein sequence validation dataset.
[0045] In some aspects, the techniques described herein relate to an Al validation dataset generation method, wherein creation of the protein sequence validation dataset allocates a larger protein sequence training dataset than a randomly sampled validation dataset randomly sampled from the plurality of reference proteomes of the reference proteome dataset, and wherein the protein sequence validation dataset and the randomly sampled validation dataset have at least one of: (1) a same sequence identity threshold value; or (2) an equal number of protein sequences selected from the plurality of reference proteomes of the reference proteome dataset.
[0046] In some aspects, the techniques described herein relate to an Al validation dataset generation method further including: removing, based on a sequence identity threshold value, one or more of the protein sequences of the protein sequence validation dataset to create an ablated protein sequence validation dataset.
[0047] In some aspects, the techniques described herein relate to an Al validation dataset generation method, wherein the removed one or more of the protein sequences includes no more than 5% to 10% of the protein sequence validation dataset.
[0048] In some aspects, the techniques described herein relate to an Al validation dataset generation method further including: updating the protein sequence validation dataset to include protein sequences each having a maximum protein length that the Al-based protein sequence model was trained on.
[0049] In some aspects, the techniques described herein relate to an Al validation dataset generation method, wherein the reference proteome dataset includes proteins having different protein lengths, and wherein the validation dataset generation method further includes:truncating one or more of protein lengths of at least a subset of the proteins to have a truncated length, the truncated length including a maximum content length for training the Al-based protein sequence model, and training the Al-based protein sequence model with training data including at least a portion of the reference proteome dataset having the subset of the proteins having the truncated length.
[0050] In some aspects, the techniques described herein relate to an Al validation dataset generation method, wherein the completeness threshold includes a value of 2 / 3 completeness or greater.
[0051] In some aspects, the techniques described herein relate to an Al validation dataset generation method, wherein 5% to 15% of the set of protein sequence clusters are randomly selected for creation of the protein sequence validation dataset.
[0052] In some aspects, the techniques described herein relate to an Al validation dataset generation method, wherein 15% or more of the set of protein sequence clusters are randomly selected for creation of the protein sequence validation dataset.
[0053] In some aspects, the techniques described herein relate to an Al validation dataset generation method further including: generating, by the one or more processors, validation metrics based on comparison between: (a) a set of predetermined values of the protein sequence validation dataset; and (b) the one or more predicted protein sequences as output by the Al-based protein sequence model.
[0054] In some aspects, the techniques described herein relate to an Al validation dataset generation method, wherein the validation metrics include at least one improved metric selected from: a perplexity metric, a loss metric, and / or an accuracy metric, when analyzed against validation metrics of a previously trained version of the Al-based protein sequence model.
[0055] In some aspects, the techniques described herein relate to an Al validation dataset generation method, wherein the at least one improved metric includes: a reduced perplexity value, a reduced loss value, and / or an increased accuracy value.
[0056] In some aspects, the techniques described herein relate to an Al validation dataset generation method, wherein the at least one improved metric includes: a reduced perplexity value when the Al-based protein sequence model is trained on a UniRefl 00 dataset, a reduced loss value when the Al-based protein sequence model is trained on the UniRefl 00dataset, and / or an increased accuracy value when the Al-based protein sequence model is trained on the UniRefl 00 dataset.
[0057] In some aspects, the techniques described herein relate to an Al validation dataset generation method, wherein the at least one improved metric includes: a reduced perplexity value when the Al-based protein sequence model is trained on a UniRefl 00 dataset compared to a second Al-based protein sequence model trained on a UniRef50 dataset, a reduced loss value when the Al-based protein sequence model is trained on the UniRefl 00 dataset compared to a second Al-based protein sequence model trained on the UniRef50 dataset, and / or an increased accuracy value when the Al-based protein sequence model is trained on the UniRefl 00 dataset compared to a second Al-based protein sequence model trained on the UniRef50 dataset.
[0058] In some aspects, the techniques described herein relate to an Al validation dataset generation method, wherein the one or more predicted protein sequences as output by the AI- based protein sequence model is used to develop a protein-related product.
[0059] In some aspects, the techniques described herein relate to an Al validation dataset generation method, wherein the protein-related product is predicted to treat a disease.
[0060] In some aspects, the techniques described herein relate to an Al validation dataset generation method, wherein the protein-related product is predicted to provide a prophylaxis for a disease or medical condition.
[0061] In some aspects, the techniques described herein relate to an Al validation dataset generation method, wherein the output of the one or more predicted protein sequences is used to monitor or measure the protein-related product during manufacture of the protein-related product.
[0062] In some aspects, the techniques described herein relate to a tangible, non-transitory computer-readable medium storing instructions for generating artificial intelligence (Al) validation datasets for validating Al-based protein sequence models, that when executed by one or more processors cause the one or more processors to: filter a reference proteome dataset to generate a high completeness proteome data subset including reference proteomes selected from a plurality of reference proteomes having a completeness value above a completeness threshold, wherein the reference proteome dataset is stored on a computer memory communicatively coupled to the one or more processors, and wherein the reference proteome dataset includes the plurality of reference proteomes; filter the high completenessproteome data subset to generate a protein confidence-based proteome data subset including reference proteomes selected from the high completeness proteome data subset having a protein confidence score above a protein confidence threshold, wherein the protein confidence-based proteome data subset includes a plurality of protein sequences; cluster the plurality of protein sequences to generate a set of protein sequence clusters, each cluster of the plurality of protein sequences grouping one or more of the plurality of protein sequences of the protein confidence-based proteome data subset; randomly allocate a validation cluster subset selected from the set of protein sequence clusters; and create a protein sequence validation dataset by selecting the protein sequences of the validation cluster subset.
[0063] In some aspects, the techniques described herein relate to a tangible, non-transitory computer-readable medium, wherein the computing instructions, when executed by the one or more processors, further cause the one or more processors to: train an Al-based protein sequence model with a protein sequence training dataset, wherein the protein sequence validation dataset includes a dataset independent of the protein sequence training dataset used to train the Al-based protein sequence model; and output, from the Al-based protein sequence model, one or more predicted protein sequences.
[0064] In some aspects, the techniques described herein relate to a tangible, non-transitory computer-readable medium, wherein the one or more predicted protein sequences include data for predicting or determining one or more of viscosity, post-translational modification, isomerization, thermal stability, pH, expression titre, aggregation, immunogenicity, binding affinity, binding avidity, clipping, protein secondary structure, protein tertiary structure, protein quaternary structure, disorder, function, enzymatic activity, homology, cellular localization, phenotypic effect, and / or half-life.
[0065] In some aspects, the techniques described herein relate to a tangible, non-transitory computer-readable medium, wherein the protein sequence training dataset includes protein sequences selected from at least a portion of the protein confidence-based proteome data subset including the plurality of protein sequences.
[0066] In some aspects, the techniques described herein relate to a tangible, non-transitory computer-readable medium, wherein the protein sequence training dataset includes protein sequences different from those selected for the protein sequence validation dataset from the validation cluster subset.
[0067] In some aspects, the techniques described herein relate to a tangible, non-transitory computer-readable medium, wherein at least one of the plurality of reference proteomes is linked to one or more sub-fragments or one or more residues of a given protein sequence in the reference proteome dataset.
[0068] In some aspects, the techniques described herein relate to a tangible, non-transitory computer-readable medium, wherein at least one of the validation cluster subset or the protein sequence validation dataset includes the one or more sub-fragments or the one or more residues of the given protein sequence in the reference proteome dataset.
[0069] In some aspects, the techniques described herein relate to a tangible, non-transitory computer-readable medium, wherein the computing instructions when executed by the one or more processors, further cause the one or more processors to: validate the one or more predicted protein sequences as output by the Al-based protein sequence model with the protein sequence validation dataset.
[0070] In some aspects, the techniques described herein relate to a tangible, non-transitory computer-readable medium, wherein creation of the protein sequence validation dataset allocates a larger protein sequence training dataset than a randomly sampled validation dataset randomly sampled from the plurality of reference proteomes of the reference proteome dataset, and wherein the protein sequence validation dataset and the randomly sampled validation dataset have at least one of: (1) a same sequence identity threshold value; or (2) an equal number of protein sequences selected from the plurality of reference proteomes of the reference proteome dataset.
[0071] In some aspects, the techniques described herein relate to a tangible, non-transitory computer-readable medium, wherein the computing instructions when executed by the one or more processors, further cause the one or more processors to: remove, based on a sequence identity threshold value, one or more of the protein sequences of the protein sequence validation dataset to create an ablated protein sequence validation dataset.
[0072] In some aspects, the techniques described herein relate to a tangible, non-transitory computer-readable medium, wherein the removed one or more of the protein sequences includes no more than 5% to 10% of the protein sequence validation dataset.
[0073] In some aspects, the techniques described herein relate to a tangible, non-transitory computer-readable medium, wherein the computing instructions, when executed by the one or more processors, further cause the one or more processors to: update the protein sequencevalidation dataset to include protein sequences each having a maximum protein length that the Al-based protein sequence model was trained on.
[0074] In some aspects, the techniques described herein relate to a tangible, non-transitory computer-readable medium, wherein the reference proteome dataset includes proteins having different protein lengths, and wherein the computing instructions, when executed by the one or more processors, further cause the one or more processors to: truncate one or more of protein lengths of at least a subset of the proteins to have a truncated length, the truncated length including a maximum content length for training the Al-based protein sequence model, and train the Al-based protein sequence model with training data including at least a portion of the reference proteome dataset having the subset of the proteins having the truncated length.
[0075] In some aspects, the techniques described herein relate to a tangible, non-transitory computer-readable medium, wherein the completeness threshold includes a value of 2 / 3 completeness or greater.
[0076] In some aspects, the techniques described herein relate to a tangible, non-transitory computer-readable medium, wherein 5% to 15% of the set of protein sequence clusters are randomly selected for creation of the protein sequence validation dataset.
[0077] In some aspects, the techniques described herein relate to a tangible, non-transitory computer-readable medium, wherein 15% or more of the set of protein sequence clusters are randomly selected for creation of the protein sequence validation dataset.
[0078] In some aspects, the techniques described herein relate to a tangible, non-transitory computer-readable medium, wherein the computing instructions, when executed by the one or more processors, further cause the one or more processors to: generate validation metrics based on comparison between: (a) a set of predetermined values of the protein sequence validation dataset; and (b) the one or more predicted protein sequences as output by the AI- based protein sequence model .
[0079] In some aspects, the techniques described herein relate to a tangible, non-transitory computer-readable medium, wherein the validation metrics include at least one improved metric selected from: a perplexity metric, a loss metric, and / or an accuracy metric, when analyzed against validation metrics of a previously trained version of the Al-based protein sequence model.
[0080] In some aspects, the techniques described herein relate to a tangible, non-transitory computer-readable medium, wherein the at least one improved metric includes: a reduced perplexity value, a reduced loss value, and / or an increased accuracy value.
[0081] In some aspects, the techniques described herein relate to a tangible, non-transitory computer-readable medium, wherein the at least one improved metric includes: a reduced perplexity value when the Al-based protein sequence model is trained on a UniRefl 00 dataset, a reduced loss value when the Al-based protein sequence model is trained on the UniRefl 00 dataset, and / or an increased accuracy value when the Al-based protein sequence model is trained on the UniRefl 00 dataset.
[0082] In some aspects, the techniques described herein relate to a tangible, non-transitory computer-readable medium, wherein the at least one improved metric includes: a reduced perplexity value when the Al-based protein sequence model is trained on a UniRefl 00 dataset compared to a second Al-based protein sequence model trained on a UniRef50 dataset, a reduced loss value when the Al-based protein sequence model is trained on the UniRefl 00 dataset compared to a second Al-based protein sequence model trained on the UniRef50 dataset, and / or an increased accuracy value when the Al-based protein sequence model is trained on the UniRefl 00 dataset compared to a second Al-based protein sequence model trained on the UniRef50 dataset.
[0083] In some aspects, the techniques described herein relate to a tangible, non-transitory computer-readable medium, wherein the one or more predicted protein sequences as output by the Al-based protein sequence model is used to develop a protein-related product.
[0084] In some aspects, the techniques described herein relate to a tangible, non-transitory computer-readable medium, wherein the protein-related product is predicted to treat a disease.
[0085] In some aspects, the techniques described herein relate to a tangible, non-transitory computer-readable medium, wherein the protein-related product is predicted to provide a prophylaxis for a disease or medical condition.
[0086] In some aspects, the techniques described herein relate to a tangible, non-transitory computer-readable medium, wherein the output of the one or more predicted protein sequences is used to monitor or measure the protein-related product during manufacture of the protein-related product.
[0087] In some aspects, the techniques described herein relate to an artificial intelligence (Al) training dataset generation system configured to generate Al training datasets for training AI- based protein sequence models, the Al training dataset generation system including: one or more processors; a computer memory communicatively coupled to the one or more processors; a reference proteome dataset stored on the computer memory and including a plurality of reference proteomes; and computing instructions stored on the computer memory, and that when executed by the one or more processors, cause the one or more processors to: filter the reference proteome dataset to generate a high completeness proteome data subset including reference proteomes selected from the plurality of reference proteomes having a completeness value above a completeness threshold; filter the high completeness proteome data subset to generate a protein confidence-based proteome data subset including reference proteomes selected from the high completeness proteome data subset having a protein confidence score above a protein confidence threshold, wherein the protein confidence-based proteome data subset includes a plurality of protein sequences, cluster the plurality of protein sequences to generate a set of protein sequence clusters, each cluster of the plurality of protein sequences grouping one or more of the plurality of protein sequences of the protein confidence-based proteome data subset, randomly allocate a training cluster subset selected from the set of protein sequence clusters, create a protein sequence training dataset by selecting the protein sequences of the training cluster subset, train an Al-based protein sequence model with the protein sequence training dataset, and output, from the Al-based protein sequence model, one or more predicted protein sequences.
[0088] In some aspects, the techniques described herein relate to an Al training dataset generation system, wherein the computing instructions, when executed by the one or more processors, further cause the one or more processors to: generate validation metrics based on comparison between: (a) a set of predetermined values of a protein sequence validation dataset including data excluded from the protein sequence training dataset; and (b) a set of predicted values as output by the Al-based protein sequence model.
[0089] In some aspects, the techniques described herein relate to an artificial intelligence (Al) training dataset generation method for generating Al training datasets for training Al-based protein sequence models, the Al training dataset generation method including: filtering, by one or more processors, a reference proteome dataset to generate a high completeness proteome data subset including reference proteomes selected from a plurality of referenceproteomes having a completeness value above a completeness threshold, wherein the reference proteome dataset is stored on a computer memory communicatively coupled to the one or more processors, and wherein the reference proteome dataset includes the plurality of reference proteomes; filtering, by the one or more processors, the high completeness proteome data subset to generate a protein confidence-based proteome data subset including reference proteomes selected from the high completeness proteome data subset having a protein confidence score above a protein confidence threshold, wherein the protein confidence-based proteome data subset includes a plurality of protein sequences; clustering, by the one or more processors, the plurality of protein sequences to generate a set of protein sequence clusters, each cluster of the plurality of protein sequences grouping one or more of the plurality of protein sequences of the protein confidence-based proteome data subset; randomly allocating, by the one or more processors, a training cluster subset selected from the set of protein sequence clusters; creating, by the one or more processors, a protein sequence training dataset by selecting the protein sequences of the training cluster subset; training an Al-based protein sequence model with the protein sequence training dataset; and outputting, from the Al-based protein sequence model, one or more predicted protein sequences.
[0090] In some aspects, the techniques described herein relate to an Al training dataset generation method further including: generating validation metrics based on comparison between: (a) a set of predetermined values of a protein sequence validation dataset including data excluded from the protein sequence training dataset; and (b) a set of predicted values as output by the Al-based protein sequence model.
[0091] In some aspects, the techniques described herein relate to a tangible, non-transitory computer-readable medium storing computing instructions for generating artificial intelligence (Al) training datasets for training Al-based protein sequence models, that when executed by one or more processors cause the one or more processors to: filter a reference proteome dataset including a plurality of reference proteomes to generate a high completeness proteome data subset including reference proteomes selected from the plurality of reference proteomes having a completeness value above a completeness threshold; filter the high completeness proteome data subset to generate a protein confidence-based proteome data subset including reference proteomes selected from the high completeness proteome data subset having a protein confidence score above a protein confidence threshold, wherein theprotein confidence-based proteome data subset includes a plurality of protein sequences, cluster the plurality of protein sequences to generate a set of protein sequence clusters, each cluster of the plurality of protein sequences grouping one or more of the plurality of protein sequences of the protein confidence-based proteome data subset, randomly allocate a training cluster subset selected from the set of protein sequence clusters, create a protein sequence training dataset by selecting the protein sequences of the training cluster subset, train an AI- based protein sequence model with the protein sequence training dataset, and output, from the Al-based protein sequence model, one or more predicted protein sequences.
[0092] In some aspects, the techniques described herein relate to a tangible, non-transitory computer-readable medium, wherein the computing instructions, when executed by the one or more processors, further cause the one or more processors to: generate validation metrics based on comparison between: (a) a set of predetermined values of a protein sequence validation dataset including data excluded from the protein sequence training dataset; and (b) a set of predicted values as output by the Al-based protein sequence model.
[0093] In some aspects, the techniques described herein relate to an artificial intelligence (Al) model generation system for generating and validating Al-based protein sequence models, the Al model generation system including: one or more processors; a computer memory communicatively coupled to the one or more processors; a reference proteome dataset stored on the computer memory and including a plurality of reference proteomes; and computing instructions stored on the computer memory, and that when executed by the one or more processors, cause the one or more processors to: train an Al-based protein sequence model with a protein sequence training dataset, wherein the protein sequence training dataset includes a dataset independent of a protein sequence validation dataset used to validate the Al-based protein sequence model, wherein the Al-based protein sequence model is configured to output one or more predicted protein sequences, and wherein creation of the protein sequence validation dataset allocates the protein sequence training dataset as having a larger dataset size than a randomly sampled validation dataset randomly sampled from the plurality of reference proteomes of the reference proteome dataset, and wherein the protein sequence validation dataset and the randomly sampled validation dataset have at least one of: (1) a same sequence identity threshold value; or (2) an equal number of protein sequences selected from the plurality of reference proteomes of the reference proteome dataset.
[0094] In some aspects, the techniques described herein relate to an artificial intelligence (Al) model generation method for generating and validating Al-based protein sequence models, the Al model generation method including: training an Al-based protein sequence model with a protein sequence training dataset, wherein the protein sequence training dataset includes a dataset independent of a protein sequence validation dataset used to validate the Al-based protein sequence model, wherein the Al-based protein sequence model is configured to output one or more predicted protein sequences, and wherein creation of the protein sequence validation dataset allocates the protein sequence training dataset as having a larger dataset size than a randomly sampled validation dataset randomly sampled from a plurality of reference proteomes of a reference proteome dataset, and wherein the protein sequence validation dataset and the randomly sampled validation dataset have at least one of (1) a same sequence identity threshold value; or (2) an equal number of protein sequences selected from the plurality of reference proteomes of the reference proteome dataset.
[0095] In some aspects, the techniques described herein relate to a tangible, non-transitory computer-readable medium storing instructions for generating and validating Al-based protein sequence models, that when executed by one or more processors cause the one or more processors to: train an Al-based protein sequence model with a protein sequence training dataset, wherein the protein sequence training dataset includes a dataset independent of a protein sequence validation dataset used to validate the Al-based protein sequence model, wherein the Al-based protein sequence model is configured to output one or more predicted protein sequences, and wherein creation of the protein sequence validation dataset allocates the protein sequence training dataset as having a larger dataset size than a randomly sampled validation dataset randomly sampled from a plurality of reference proteomes of a reference proteome dataset, and wherein the protein sequence validation dataset and the randomly sampled validation dataset have at least one of: (1) a same sequence identity threshold value; or (2) an equal number of protein sequences selected from the plurality of reference proteomes of the reference proteome dataset.
[0096] In some aspects, the techniques described herein relate to a tangible, non-transitory computer-readable medium storing a protein sequence as output by an Al-based protein sequence model, the non-transitory computer-readable medium further storing instructions for validating the Al-based protein sequence model, that when executed by one or more processors cause the one or more processors to: filter, by one or more processors, a referenceproteome dataset to generate a high completeness proteome data subset including reference proteomes selected from a plurality of reference proteomes having a completeness value above a completeness threshold, wherein the reference proteome dataset is stored on a computer memory communicatively coupled to the one or more processors, and wherein the reference proteome dataset includes the plurality of reference proteomes; filter, by the one or more processors, the high completeness proteome data subset to generate a protein confidence-based proteome data subset including reference proteomes selected from the high completeness proteome data subset having a protein confidence score above a protein confidence threshold, wherein the protein confidence-based proteome data subset includes a plurality of protein sequences; cluster, by the one or more processors, the plurality of protein sequences to generate a set of protein sequence clusters, each cluster of the plurality of protein sequences grouping one or more of the plurality of protein sequences of the protein confidence-based proteome data subset; randomly allocate, by the one or more processors, a validation cluster subset selected from the set of protein sequence clusters; create, by the one or more processors, a protein sequence validation dataset by selecting the protein sequences of the validation cluster subset; and validate the Al-based protein sequence model with the protein sequence validation dataset.
[0097] In some aspects, the techniques described herein relate to an output protein sequence generated by an Al-based protein sequence model, wherein the Al-based protein sequence model is validated by an Al validation dataset generated by a method including: filtering, by one or more processors, a reference proteome dataset to generate a high completeness proteome data subset including reference proteomes selected from a plurality of reference proteomes having a completeness value above a completeness threshold, wherein the reference proteome dataset is stored on a computer memory communicatively coupled to the one or more processors, and wherein the reference proteome dataset includes the plurality of reference proteomes; filtering, by the one or more processors, the high completeness proteome data subset to generate a protein confidence-based proteome data subset including reference proteomes selected from the high completeness proteome data subset having a protein confidence score above a protein confidence threshold, wherein the protein confidence-based proteome data subset includes a plurality of protein sequences; clustering, by the one or more processors, the plurality of protein sequences to generate a set of protein sequence clusters, each cluster of the plurality of protein sequences grouping one or more ofthe plurality of protein sequences of the protein confidence-based proteome data subset; randomly allocating, by the one or more processors, a validation cluster subset selected from the set of protein sequence clusters; creating, by the one or more processors, a protein sequence validation dataset by selecting the protein sequences of the validation cluster subset; and validating the Al-based protein sequence model with the protein sequence validation dataset.
[0098] In some aspects, the techniques described herein relate to a method for evaluating generalization capability of one or more Al-based protein sequence models, the method including: generating, by one or more processors, a validation dataset including protein sequences obtained by: filtering a reference proteome dataset to generate a high completeness proteome data subset including reference proteomes selected from a plurality of reference proteomes having a completeness value above a completeness threshold, wherein the reference proteome dataset includes a plurality of reference proteomes; filtering the high completeness proteome data subset to generate a protein confidence-based proteome data subset including reference proteomes selected from the high completeness proteome data subset having a protein confidence score above a protein confidence threshold, wherein the protein confidence-based proteome data subset includes a plurality of protein sequences; clustering the plurality of protein sequences of the protein confidence-based proteome data subset to generate a set of protein sequence clusters, each cluster of the plurality of protein sequences grouping one or more of the plurality of protein sequences of the protein confidence-based proteome data subset; and creating the protein sequence validation dataset by selecting the protein sequences from one or more clusters in the set of protein sequence clusters; for each of two or more subsets of a training dataset, removing the subset from the training dataset to obtain a respective ablated training dataset, wherein each respective ablated training dataset includes at least one or more protein sequences excluded from the validation dataset, and wherein each respective ablated training dataset includes a unique set of protein sequences; implementing, for each respective ablated training dataset, an iterative model routine to evaluate an Al-based protein sequence model trained on the training dataset, the iterative model routine including: training an ablated Al-based protein sequence model on the ablated training dataset, generating an ablated performance score defining a predictive performance of an output of the ablated Al-based protein sequence model, and storing in memory the ablated performance score as part of a plurality of ablated performance scores;and determining, based on the plurality of ablated performance scores, a model generalization output for the Al-based protein sequence model.
[0099] In some aspects, the techniques described herein relate to a method, wherein the iterative model routine further includes determining, using the validation dataset as input to the ablated Al-based protein sequence model, an output of the ablated Al-based protein sequence model, and wherein generating the ablated performance score further includes comparing the output of the ablated Al-based protein sequence model to a ground truth value.
[0100] In some aspects, the techniques described herein relate to a method, wherein the iterative model routine further includes: determining, using the validation dataset as input to the ablated Al-based protein sequence model, an output of the ablated Al-based protein sequence model; and determining, using the validation dataset as input to the Al-based protein sequence model, an output of the Al-based protein sequence model; and wherein generating the ablated performance score further includes comparing the output of the ablated Al-based protein sequence model to the output of the Al-based protein sequence model.
[0101] In some aspects, the techniques described herein relate to a method, further including determining two or more sequence similarity thresholds by comparing protein sequences in the training dataset to protein sequences in the validation dataset, and wherein each of the two or more subsets is associated with a respective sequence similarity threshold selected from among the two or more sequence similarity thresholds, and removing the subset from the training dataset to obtain the respective ablated training dataset further includes applying the respective sequence similarity threshold as a filter to the training dataset to remove protein sequences having a sequence similarity above the respective sequence similarity threshold.
[0102] In some aspects, the techniques described herein relate to a method further including outputting, from the Al-based protein sequence model, one or more predicted protein sequences.
[0103] In some aspects, the techniques described herein relate to a method, further including outputting, from the Al-based protein sequence model, information specifying a protein sequence and manufacturing a protein-related product having the protein sequence.
[0104] In some aspects, the techniques described herein relate to a method, wherein the model generalization output defines one or more validation metrics including at least one of: a perplexity metric, a loss metric, and / or an accuracy metric.
[0105] In some aspects, the techniques described herein relate to a method, wherein the model generalization output defines one or more improved validation metrics when compared to validation metrics of a different Al-based protein sequence model trained on at least a portion of the reference proteome dataset, the one or more improved validation metrics including at least one of an improved perplexity metric, an improved loss metric, and / or an improved accuracy metric.
[0106] In some aspects, the techniques described herein relate to a method further including: generating each respective ablated training dataset by ablating no more than 5% to 10% of the training dataset.
[0107] In some aspects, the techniques described herein relate to a method, wherein creating the protein sequence validation dataset further includes selecting protein sequences in a cluster of the set of protein sequence clusters, and the model generalization output defines one or more validation metrics indicating performance of the Al-based protein sequence model for the selected protein sequences in the cluster.
[0108] In some aspects, the techniques described herein relate to a method, wherein the selected protein sequences in the cluster correspond to a protein class or protein family.
[0109] In some aspects, the techniques described herein relate to a method, wherein generating the validation dataset further includes allocating a validation cluster subset selected from the set of protein sequence clusters, and creating the protein sequence validation dataset further includes selecting protein sequences from the validation cluster subset as the protein sequence validation dataset.
[0110] In some aspects, the techniques described herein relate to a method, wherein generating the validation dataset further includes randomly allocating a validation cluster subset selected from the set of protein sequence clusters, and creating the protein sequence validation dataset further includes selecting the protein sequences of the validation cluster subset.
[0111] Advantages will become more apparent to those of ordinary skill in the art from the following description of the preferred embodiments which have been shown and described by way of illustration. As will be realized, the present embodiments may be capable of other and different embodiments, and their details are capable of modification in various respects. Accordingly, the drawings and description are to be regarded as illustrative in nature and not as restrictive.BRIEF DESCRIPTION OF DRAWINGS
[0112] The accompanying drawings are not intended to be drawn to scale. In the drawings, each identical or nearly identical component that is illustrated in various figures is represented by a like numeral. For purposes of clarity, not every component may be labeled in every drawing. In the drawings:
[0113] FIG. 1 A illustrates an example artificial intelligence (Al) validation dataset generation method for generating Al validation datasets for validating Al-based protein sequence models, in accordance with some embodiments.
[0114] FIG. IB illustrates an example artificial intelligence (Al) training dataset generation method for generating Al training datasets for training Al-based protein sequence models, in accordance with some embodiments.
[0115] FIG. 2A illustrates example validation metrics of Al-based protein sequence models in accordance with some embodiments.
[0116] FIG. 2B illustrates an example diagram depicting sequence identity threshold values (as a percentage) across a spectrum of training data removed (as a percentage) in accordance with some embodiments.
[0117] FIG. 3 A illustrates a further example artificial intelligence (Al) validation dataset generation method for generating Al validation datasets for validating Al-based protein sequence models, in accordance with some embodiments.
[0118] FIG. 3B illustrates a further example artificial intelligence (Al) training dataset generation method for generating Al training datasets for training Al-based protein sequence models, in accordance with some embodiments.
[0119] FIG. 3C illustrates an example method for evaluating generalization capability of one or more Al-based protein sequence models, in accordance with some embodiments.
[0120] FIG. 4 illustrates a diagram of illustrative implementation of a computer system that may be used in connection with any of the embodiments described herein.
[0121] The Figures depict preferred embodiments for purposes of illustration only.Alternative embodiments of the systems and methods illustrated herein may be employed without departing from the principles of the invention described herein.DETAILED DESCRIPTION
[0122] While principles of the present disclosure are described herein with reference to illustrative embodiments for particular applications, it should be understood that the disclosure is not limited thereto. Those having ordinary skill in the art and access to the teachings provided herein, will recognize that the features illustrated or described with respect to one embodiment, may be combined with the features of another embodiment. Therefore, additional modifications, applications, embodiments, and substitution of equivalents, all fall within the scope of the embodiments described herein. Accordingly, the invention is not to be considered as limited by the foregoing description. Various nonlimiting embodiments of the present disclosure will now be described to provide an overall understanding of the principles of the structure, function, and use of systems, methods, algorithms, software, and / or as otherwise described herein.
[0123] In general, as referred to herein, a proteome can refer to a set or otherwise a collection of proteins. For example, a proteome can refer to a set of proteins expressed by an organism and can also be used to describe the assortment of proteins produced at a specific time in a particular cell or tissue type. Generally, a proteome can comprise an expression of a given organism’s genome. Additionally, or alternatively, a proteome can refer to an artificial dataset that is created to simulate proteins or protein sequences.
[0124] In general, as referred to herein, a protein can refer to a class of molecule that contains polymeric chains of amino acids linked together in peptide bonds. In some aspects, a protein can comprise a natural proteins where polymeric chains are encoded in a gene and produced through the process of mRNA translation. However, a protein as referred to herein can refer to non-natural protein, such as an artificial protein or synthetically created protein, or other such molecule of polymeric chains of amino acids. In addition, or in the alternative, a protein can also refer to a sequence of amino acids.
[0125] FIG. 1A illustrates an example artificial intelligence (Al) validation dataset generation method 100 for generating Al validation datasets for validating Al-based protein sequence models, in accordance with some embodiments. Said another way, FIG. 1 A shows a validation set generation strategy. In the example of FIG. 1 A, reference proteomes are downloaded, or otherwise stored in a computer memory. Validation sets can be generated based on series of filtration, clustering, and selection phases. These phases are described for FIG. 1 A, and are further described for FIGs 2 and 3, and / or elsewhere herein.
[0126] In the example of FIG. 1A, creation and / or selection of protein sequences for validation sets creates exclusive validation dataset that is independent of the training data used for training an Al model. The exclusive validation dataset can allow assessment of the effects of training data manipulations, e.g., such as the use of the UniRef50 dataset compared to the UniRefl 00 dataset, e.g., as described for FIGs. 2A. The UniRef50 dataset can be a protein sequence database maintained by the UNIPROT consortium, updated on a regular schedule to provide a diverse source of non-redundant protein sequences. The UniRef50 dataset can be prepared by clustering predetermined sequences first to a 100% identity threshold (as taken from the UniRefl 00 dataset) and then further by clustering those down to a 50% identity threshold and selecting a single sequence from each cluster. The exclusive validation dataset can also allow assessment and testing of whether or not data augmentation (e.g., the addition of somatic antibody sequences) degrades performance on the germline sequences in or of the independent and exclusive validation dataset.
[0127] Generally, creation, implementation, and use of the exclusive validation dataset, as described herein, can demonstrate that using the UniRef50 dataset alone as the training set, in a conventional manner, degrades performance against confident proteins found in a distribution matching how often such proteins occur in high confidence proteomes.
[0128] Further, this has been shown in large language model (LLM) training using NVIDIA’s BIONEMO framework. As shown and described for FIG. 2A, creation and use of the protein sequence validation dataset as described herein provides consistently observed improvements in perplexity, loss, and accuracy when compared with conventional training on sequence datasets, e.g., conventional training applied to sequences taken from the UniRefl 00 dataset or the UniRef50 dataset.
[0129] With reference to FIG. 1 A, at block 101 (phase 1) method 100 may comprise obtaining and / or storing (e.g., in memory 420 of FIG. 4) reference proteomes. The reference proteome dataset may comprise a plurality of reference proteomes. In the example of FIG. 1 A, a full reference proteome dataset may be downloaded, or otherwise obtained, from uniprot.org that includes 22,925 reference proteomes. It should be understood, however, that while FIG. 1 A includes a specific number of 22,925 reference proteomes and later filtered protein sequences (as described further herein), additional and / or different numbers may be used. For example, the number of the plurality of reference proteomes may be at least 10, 100, 500, 1000, 5000, 10000, 20000, 30000, 50000 or greater. Additionally, or alternatively,the number of the plurality of reference proteomes may be at most 50000, 30000, 20000, 10000, 5000, 1000, 500, 100 or less. Additional and / or different numbers may be used.
[0130] At block 102 (phase 2) method 100 comprises filtering, by one or more processors (e.g., processor 410 of FIG. 4), the reference proteome dataset for generation of a high completeness proteome data subset. In the example of FIG. 1A, such filtering may comprise filtering via a Benchmarking Universal Single-Copy Orthologue (BUSCO) BUSCO completeness report, where proteomes are filtered down to those with a certain threshold (e.g., 2 / 3rdor more) completeness, which, in the example of FIG. 1 A comprises 10,662 filtered proteomes. BUSCO can provide an analytical tool for evaluation of genomic resources and can provide evolutionarily sound measures of completeness and redundancy in terms of expected gene content. It should be understood, however, that BUSCO is but one example for filtering and / or measuring completeness of a proteomes, and that additional and / or different filtering and / or measurements may be used. Such additional and / or different filtering and / or measurements can comprise Core Eukaryotic Genes Mapping Approach (CEGMA), Quality Assessment Tool for Genome Assemblies (QUAST), Recognition of Errors in Assemblies using Paired Reads (REAPR), and a Bioinformatics Application for Navigating De novo Assembly Graphs Easily (Bandage).
[0131] More generally, as used herein, completeness refers to how much of a real world or otherwise actual or full proteome a deposited proteome dataset has captured. In other words, when it is suspected that a proteome dataset is incomplete this means that all real world or otherwise actual proteomes have not yet been identified for a given proteome or its related dataset (e.g., dataset identifying proteins) in order to complete the proteome or its related dataset.
[0132] In this way, the completeness of a given proteome provides an estimate. High completeness (e.g., 66% or more) refers to a degree of completeness were most of the real world or actual proteome have been captured (e.g., captured in a given dataset), such that the dataset or otherwise data distribution in the dataset is more likely to reflect the data or distribution of protein(s) in the real world (or real space). High completeness can eliminate bias in how proteins are identified. For example, by contrast, if a proteome contains 10% of the expected number of proteins (low completeness), then it may not be a random 10% in an example that uses 10% test data, and, in such example the data may be very biased at 10%.
[0133] At block 103 (phase 3) method 100 comprises filtering, by the one or more processors, the high completeness proteome data subset to generate a protein confidencebased proteome data subset. In the example of FIG. 3 A, the high completeness proteome data subset of the reference proteome dataset is further filtered by confidence values (e.g., protein evidence). For example, in one example aspect, reference proteome data may be associated with an annotation in each sequence’s header for PE (protein evidence), where values of 1 (direct protein) and 2 (RNA) reflect direct evidence (from sequencing of protein or coding RNA / DNA), and all higher numbers reflect indirect evidence or algorithmic predictions. Protein evidence may refer to a proprietary value by UNIPROT. It should be understood that additional and / or different values may be used to determine confidence of a given protein for filtering reference proteins. Such values can be referenced to herein as confidence and / or a confidence score, the latter of which is a numerical value of the former. For example, the confidence score may be used to track the protein sequence back to the original protein data, e.g., for direct protein sequencing and / or coding RNA / DNA.
[0134] With reference to the example of FIG. 1 A, the reference proteomes of phase 2 are filtered for proteomes with greater than 20% of their reference proteins, which are annotated as protein evidence (PE)=1 or protein evidence (PE)=2. Reference proteomes are fixed sets of proteins. In the present example, 20% of the proteins in a set have PE=1 or PE=2, so they pass the filter. Such filtering leads to 22 reference proteomes, which may represent primarily well studied model organisms. In the example of FIG. 1A, the 22 proteomes consist of 412,531 protein sequences, with 106,836 protein sequences at PE=1 and 82,822 protein sequences at PE=2.
[0135] At block 104 (phase 4), method 100 comprises randomly allocating, by the one or more processors, a validation cluster subset selected from the set of protein sequence clusters. In one example, one software (e.g., MMseqs2 (Many-against-Many sequence searching)) can be used to cluster the 412,531 protein sequences using a command that yields a sequence identity threshold of 30%. A sequence alignment may output a sequence identity comprising a percentage of residues that are an exact match. The percentage can vary according to the sequence alignment algorithm used. Such software can search and cluster huge protein and nucleotide sequence sets. However, it should be understood that different and / or additional software libraries, packages, or otherwise computing instructions may be used to randomlyallocate, by the one or more processors, a validation cluster subset selected from the set of protein sequence clusters.
[0136] At block 105 (phase 5) method 100 comprises creating, by the one or more processors, a protein sequence validation dataset by selecting the protein sequences of the validation cluster subset. In the example of FIG. 1 A, approximately 5% of the clusters were randomly selected for the validation set, comprising 21,275 of the 412,531 protein sequences. It should be understood, however, that additional and / or different percentage of clusters may also be used. The random selection of proteins can be considered an exclusion set of data, which is excluded from training data. The randomly selected proteins can be used as validation set proteins for the purposes of training set exclusion for testing and validating a given Al model (e.g., an Al-based protein sequence model as described herein), and not for training the Al model.
[0137] At block 106 (phase 6) method 100 comprises calculating performance metrics (e.g., validation metrics), including, for example, generating the protein sequences of phase 5 to those protein sequences with protein evidence (PE)=1 or protein evidence (PE)=2 and / or a sequence length from 50 to 512. Said another way, execution, by one or more processors, at block 106 (phase 6) comprises filtering the protein sequence validation dataset (e.g., as described for block 105, phase 5) to generate a more confident set, which can include excluding data from the protein sequence validation dataset (e.g., as described for block 105, phase 5) to those to those protein sequences with protein evidence (PE)=1 or protein evidence (PE)=2 and / or a sequence length from 50 to 512, and then calculating validation metrics for a remaining subset of protein sequences at block 106 (phase 6). This selects 6,665 protein sequences from the set of 21,275 from phase 5, and determines a validation set for comparing or otherwise generating validation metrics for a given Al model (e.g., an Al-based protein sequence model), for validating and / or evaluating its output, e.g., the Al model’s outputted predicted values.
[0138] FIG. IB illustrates an example artificial intelligence (Al) training dataset generation method 150 for generating Al training datasets for training Al-based protein sequence models, in accordance with some embodiments. Said another way, FIG. IB shows a training set generation strategy. In the example of FIG. IB, the first three phases are the same as those described herein for Al validation dataset generation method 100 of FIG. 1 A, such that the description herein for Al validation dataset generation method 100 applies equally for Altraining dataset generation method 150. That is, each of block 101 (phase 1), block 102 (phase 2), and block 103 (phase 3) of Al validation dataset generation method 100 of FIG. 1 A applies to each of block 151 (phase 1), block 152 (phase 2), and block 153 (phase 3) of Al training dataset generation method 150 of FIG. IB, respectively.
[0139] At block 154 (phase 4), method 150 comprises randomly allocating, by the one or more processors, a training cluster subset selected from the set of protein sequence clusters. In one example, one software (e.g., MMseqs2) can be used to cluster the 412,531 protein sequences using a command that yields a sequence identity threshold of 30%. A sequence alignment may output a sequence identity comprising a percentage of residues that are an exact match. The percentage can vary according to the sequence alignment algorithm used. Such software can search and cluster huge protein and nucleotide sequence sets. However, it should be understood that different and / or additional software libraries, packages, or otherwise computing instructions may be used to randomly allocate, by the one or more processors, a training cluster subset selected from the set of protein sequence clusters.
[0140] In an alternative embodiment, the training cluster subset may be selected from protein sequence clusters that are different from the validation cluster subset selected from the set of protein sequence clusters as described herein for block 104 (phase 4). That is, at least a remaining portion of clusters may be designated as the training cluster subset after implementation of block 104 (phase 4), where the validation cluster subset is selected from the set of protein sequence clusters mutually exclusively from the training cluster subset.
[0141] At block 155 (phase 5) method 150 comprises creating, by the one or more processors, a protein sequence training dataset by selecting the protein sequences of the training cluster subset. In the example of FIG. IB, approximately 95% of the clusters were randomly selected for the training set, comprising 391,256 of the 412,531 protein sequences. It should be understood, however, that additional and / or different percentage of clusters may also be used. The random selection of proteins can be considered an exclusion set of data, which is excluded from validation data. The randomly selected proteins can be used as a protein sequence training dataset for training a given Al model (e.g., an Al-based protein sequence model as described herein), and not for validating the Al model. The randomly selected proteins can also be used to determine what data to exclude for validation purposes, where a protein sequence validation dataset can be selected from a remaining set (e.g., a percentage) of data not selected for the protein sequence training dataset.
[0142] At block 156 (phase 6) method 150 comprises calculating performance metrics (e.g., validation metrics). In such aspects, for example, the performance metrics (e.g., validation metrics) may be generated based on protein sequence validation dataset excluded from the protein sequence training dataset as described for block 155 (phase 5). In such aspects, the performance metrics (e.g., validation metrics) can determine how useful or otherwise accurate the excluded protein sequence validation dataset is testing a given Al model (e.g., an Al-based protein sequence model as described herein). Such Al model may be trained by the protein sequence training dataset as described for block 155 (phase 5) and validated by the protein sequence validation dataset as described for block 155 (phase 5). Additionally, or alternatively, such Al model may be trained by a filtered protein sequence training dataset, and validated by a filtered protein sequence validation dataset, where the further filtered protein sequence training dataset can include, for example, the protein sequence training dataset as described for block 155 (phase 5), but having selected those protein sequences with protein evidence (PE)=1 or protein evidence (PE)=2 and / or a sequence length from 50 to 512, and where the filtered protein sequence validation dataset can comprise a reminder (or some percentage thereof) of the protein sequences not selected for the further filtered protein sequence training. In some aspects, the validation metrics of an Al model trained and validated with protein sequences as described for block 155 (phase 5), can be compared to validation metrics of an Al model trained and validated with the filtered protein sequence training dataset and the filtered protein sequence validation dataset, to determine which dataset group indicates a more accurate model via such training and validation phases.
[0143] FIG. 3A illustrates a further example artificial intelligence (Al) validation dataset generation method 300 for generating Al validation datasets for validating Al-based protein sequence models, in accordance with some embodiments. At block 310, method 300 comprises filtering, by one or more processors (e.g., processor 410 of FIG. 4), a reference proteome dataset for generation of a high completeness proteome data subset comprising reference proteomes selected from a plurality of reference proteomes having a completeness value above a completeness threshold.
[0144] The reference proteome dataset may be stored on a computer memory (e.g., memory 420 of FIG. 4) communicatively coupled to the one or more processors (e.g., processor 410). The reference proteome dataset may comprise the plurality of reference proteomes, e.g., 22,925 reference proteomes as shown and described for phase 1 of FIG. 1 A.
[0145] In some aspects, the completeness threshold may comprise a value of 2 / 3 completeness or greater, indicating that a given reference proteome is at least 2 / 3 complete in order to be selected for inclusion as part of the high completeness proteome data subset, e.g., 10,662 proteomes as shown and described for phase 2 of FIG. 1 A.
[0146] Additionally, or alternatively, in some aspects, the completeness threshold may be adjusted. For example, in some aspects, the completeness threshold may be set at 50% or greater, e.g., where at least a majority of a given reference proteome is completed for the given reference proteome to be included as part of the high completeness proteome data subset. In some embodiments, the completeness threshold may comprise a value of at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90% completeness or greater. In some embodiments, the completeness threshold may comprise a value of at most 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, 10% completeness or less.
[0147] At block 320, method 300 comprises filtering, by the one or more processors, the high completeness proteome data subset to generate a protein confidence-based proteome data subset. The protein confidence-based proteome data subset may comprise reference proteomes selected from the high completeness proteome data subset having a protein confidence score above a protein confidence threshold, e.g., 22 proteomes as shown and described for phase 3 of FIG. 1 A. The protein confidence-based proteome data subset may comprise a plurality of protein sequences, e.g., 412,531 protein sequences as shown and described for phase 4 of FIG. 1 A.
[0148] At block 330, method 300 comprises clustering, by the one or more processors, the plurality of protein sequences to generate a set of protein sequence clusters. Each cluster of the plurality of protein sequences may group one or more of the plurality of protein sequences of the protein confidence-based proteome data subset. In various aspects, the set of protein sequence clusters enable or otherwise allow for a generation of a validation dataset sample having a smaller space or set of proteins that are related to one another. This allows for generation of a validation dataset (e.g., a protein sequence validation dataset) that when excluded from a given training data removes less of the training data, and allows for more effective Al-model training, and thus generation of an Al model (e.g., an Al-based protein sequence model) with less error and higher predictive output. Therefore, such validation dataset, exclusion, training, and otherwise improvements allow for the Al-based protein sequence model to be used for specific experiments related the Al-based protein sequencemodel specifically trained with the given validation dataset (e.g., protein sequence validation dataset).
[0149] In some aspects, a range of 5% to 15% of the set of protein sequence clusters may be randomly selected and / or grouped for creation of the protein sequence validation dataset.
[0150] Alternatively, a greater number of protein sequence clusters may also be selected including, for example, where 15% or more of the set of protein sequence clusters may be randomly selected for creation of the protein sequence validation dataset. In some embodiments, the percentage of the set of protein sequence clusters that may be randomly selected for creation of the protein sequence validation dataset may be at least 5%, 10%, 15%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90% or greater. In some embodiments, the percentage of the set of protein sequence clusters that may be randomly selected for creation of the protein sequence validation dataset may be at most 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, 15%, 10%, 5% or less.
[0151] At block 340, method 300 comprises randomly allocating, by the one or more processors, a validation cluster subset selected from the set of protein sequence clusters. For example, the validation cluster subset may comprise, for example, 5% (or approximately 5%) of the clusters, which may have randomly selected, e.g., a random selection of 21,275 of the 412,531 protein sequences as shown and described for phase 5 of FIG. 1A.
[0152] In some aspects, method 300 may further comprise removing, based on a sequence identity threshold value, one or more of the protein sequences of the protein sequence validation dataset to create an ablated protein sequence validation dataset. Ablation may be implemented to test how well a given Al model generalizes to data that it hasn’t been trained with. The sequence identity threshold value can be a percentage value (e.g., as shown and described for FIG. 2B, e.g., for sequence identity threshold values 236). Increasing the validation threshold’s stringency (by decreasing the sequence identity threshold value) removes protein training set sequences on a case-by-case basis depending on their maximum similarity to a given individual sequence defined, set, or otherwise found in the protein sequence validation dataset. In some aspects, the removed one or more of the protein sequences may comprise no more than 5% to 10% of the protein sequence validation dataset. This is superior to prior art methodologies, where more than 50% of the data would need to be removed in order to ablate at a same level. It is to be understood, however, that additional and / or different values of data may be removed (e.g., ablated). Such Al validation ablationallows for testing generalizability, which allows testing of the degree to which an Al model performance reflects generalized understanding compared to memorization (e.g., overfitting) of a training dataset, and provides a method to determine strength of an Al model. That is, a protein sequence validation dataset or an ablated protein sequence validation dataset, as the case may be, comprises an out-of-distribution sample with respect a related protein sequence training dataset, which allows testing and validation of an Al model with unseen data, otherwise data excluded from the protein sequence training dataset.
[0153] In an iterative implementation, a protein sequence validation dataset can be tested, one iteration after another, with each iteration removing additional protein sequences, to ultimately determine whether performance of the Al model degrades as measured by validation of a given ablated protein sequence validation dataset for a given ablation iteration.
[0154] In addition, a sequence identity threshold value as used to define exclusion or ablation determines how far out-of-distribution the related protein sequence training dataset needs to be. The lower the sequence identity threshold value, the more data that is excluded and the more out-of-distribution the given protein sequence validation dataset becomes. In some aspects, if such exclusion is performed iteratively (e.g., removing the validation set, while maintaining the general size of the training data), such implementation allows generation of a continuous validation metric for assessing how far out-of-distribution the validation dataset can be before the Al model being validated begins to lose predictive accuracy, has more loss, etc.
[0155] At block 350, method 300 comprises creating, by the one or more processors, a protein sequence validation dataset by selecting the protein sequences of the validation cluster subset. The protein sequence validation dataset can be determined as validation set proteins to be excluded from a training set for training an Al-based protein sequence model. In this way, the protein sequences in the protein sequence validation dataset can be considered as validation set proteins for the purposes of training set exclusion.
[0156] In such aspects, the protein sequence validation dataset comprises a dataset independent of a training dataset used to train the Al-based protein sequence model. For example, the protein sequence validation dataset may comprise 6,665 protein sequences as shown and described for phase 6 of FIG 1. The protein sequence validation dataset may be excluded from a training dataset used to train an Al-based protein sequence model. Thus, the protein sequence validation dataset can be different from the training dataset, where theprotein sequence validation dataset can be used for validating that Al-based protein sequence model after it is generated with other and / or different protein sequence data.
[0157] In various aspects, the Al-based protein sequence model may be trained on a training dataset that includes data not selected or otherwise excluded for the protein sequence validation dataset. For example, in various aspects, method 300 may further comprise training an Al-based protein sequence model with a protein sequence training dataset. The protein sequence validation dataset, as created and described by block 350 of method 300, may comprise a dataset independent of the protein sequence training dataset used to train the Al-based protein sequence model.
[0158] Still further, method 300 may further comprise outputting, from the Al-based protein sequence model, one or more predicted protein sequences. In one or more aspects, the one or more predicted protein sequences may comprise data for predicting or determining one or more protein attributes or related values, including, but not limited to at least one of: viscosity, post-translational modification, isomerization, thermal stability, pH, expression titre, aggregation, immunogenicity, binding affinity, binding avidity, clipping, protein secondary structure, protein tertiary structure, protein quaternary structure, disorder, function, enzymatic activity, homology, cellular localization, phenotypic effect, and / or half-life
[0159] In some aspects, the protein sequence training dataset may comprise protein sequences selected from at least a portion of the protein confidence-based proteome data subset comprising the plurality of protein sequences. Additionally, or alternatively, in some aspects the protein sequence training dataset may comprise protein sequences different from those selected for the protein sequence validation dataset from the validation cluster subset. In such aspects, the data is different between the protein sequence training dataset and the validation dataset. Further, in such aspects, the data may be different between the protein sequence training dataset and the validation dataset with respect to one or more data clusters.
[0160] In various aspects, method 300 may further comprise validating the one or more predicted protein sequences as output by the Al-based protein sequence model with the protein sequence validation dataset. In this way, the this the protein sequence validation dataset can be implemented to check the accuracy or otherwise validation metrics (e.g., as described for FIG. 2B) of the Al-based protein sequence model.
[0161] Still further, in various aspects, at least one of the plurality of reference proteomes may be linked to one or more sub-fragments or one or more residues of a given proteinsequence in the reference proteome dataset. For example, such reference proteome dataset may comprise the UniRefl 00 dataset. In such aspects, at least one of the validation cluster subset or the protein sequence validation dataset may comprise the one or more subfragments or the one or more residues of the given protein sequence in the reference proteome dataset. Accordingly, despite filtering the reference proteome dataset (e.g., as described herein for block 310 and / or block 320 of method 300), the sub-fragments and / or residues that are stored or linked in the reference proteome dataset (e.g., the UniRefl 00 dataset) can still remain; and do not have to be removed to implement clustering of such data as described herein, e.g., for blocks 330 and / or 340 of method 300.
[0162] In various aspects, one or more predicted protein sequences as output by the AI- based protein sequence model, may be used to develop a protein-related product. In such aspects, the protein-related product may be predicted to treat a disease. Additionally, or alternatively, the protein-related product may be predicted to provide a prophylaxis for a disease or medical condition. In various aspects, such diseases and / or medical conditions may comprise, by way of non-limiting example, one or more of cancer, cardiovascular disease, bone health, kidney and blood disorders, inflammatory and autoimmune diseases, neurology diseases, and / or rare genetic disorders. Still further, in some aspects, the output of the one or more predicted protein sequences may be used to monitor or measure the protein-related product during manufacture of the protein-related product. Such monitoring and / or measuring may be used to control a process or physical equipment during manufacture of the protein- related product, e.g., to determine safety and / or efficacy of the protein-related product.
[0163] At block 360, method 300 comprises generating, by the one or more processors, validation metrics based on comparison between: (a) a set of predetermined values of the protein sequence validation dataset, and (b) the one or more predicted protein sequences as output by an Al-based protein sequence model. Generation of the validation metrics allows for testing of the performance a given Al-based protein sequence model, where a validation set, and related validation metrics, can produce output showing improvements in perplexity, loss, and accuracy when compared with conventional training on conventional validation set protein sequences. The example validation metrics may comprise metrics as described herein for FIG. 2A. Additionally, or alternatively, the validation metrics may comprise a confusion matrix or data related thereto determining the accuracy of the Al-based protein sequence model.
[0164] In some aspects, a target protein length or maximum protein length may be applied, where, for example, the plurality of reference proteomes are filtered or preprocessed to have the target protein length and / or maximum protein length before or otherwise for creating the protein sequence validation dataset, e.g., as described herein. For example, in one aspect, the protein sequence validation dataset may be updated (e.g., filtered, truncated, or otherwise updated) to comprise protein sequences each having a maximum protein length, or otherwise target protein length (e.g., 2000 maximum allowable tokens or characters representing the target protein length), that the Al-based protein sequence model was trained on. For example, the Al-based protein sequence model, or the Al training algorithm used to train the Al-based protein sequence model (e.g., such as a deep learning training algorithm), may have a maximum allowable length of tokens or characters (e.g., 2000 maximum allowable tokens or characters) for a given LLM model. In such aspects, the Al-based protein sequence model may be trained on protein sequences that have been filtered, truncated, or otherwise updated to the maximum allowable length, or some target length that is at or less than the maximum allowable length.
[0165] For example, in some aspects, the reference proteome dataset may comprise proteins having different protein lengths. In such aspects, the validation dataset generation method may further comprise truncating one or more of protein lengths of at least a subset of the proteins to have a truncated length. In such aspects, the truncated length may comprise a maximum content length (e.g., 2000 maximum allowable tokens or characters) of data representing given protein(s) for training the Al-based protein sequence model. Still further, in such aspects, the validation dataset generation method may further comprise training the Al-based protein sequence model with training data comprising at least a portion of the reference proteome dataset having the subset of the proteins having the truncated length. In this way, one or more of the protein lengths of the proteins can have different protein lengths, but where such proteins are truncated to fit within a maximum context length that the AI- based protein sequence model can be trained on, e.g., as limited by a given Al algorithm or otherwise.
[0166] In some aspects, the validation metrics, as previously generated and described for phases 310-360 of method 300, comprise at least one improved metric selected from: a perplexity metric, a loss metric, and / or an accuracy metric. For example, an improved validation metric may comprise target value(s) of a perplexity metric, a loss metric, and / or anaccuracy metric for a given Al-based protein sequence model, where the target value(s) may be targets for increasing prediction accuracy of the given Al-based protein sequence model. Additionally, or alternatively, an improved metric may comprise any of an improved perplexity metric, an improved loss metric, and / or an improved accuracy metric when analyzed against validation metrics of a previously trained version of the Al-based protein sequence model. In some aspects, the previously trained version of the Al-based protein sequence model can comprise a previous iteration of the same Al-based protein sequence model, e.g., an Al-based protein sequence model updated with new, additional, and / or different training data. In other aspects, the previously trained version of the Al-based protein sequence model can comprise any conventional models (e.g., ESM2).
[0167] FIG. 3B illustrates a further example artificial intelligence (Al) training dataset generation method 370 for generating Al training datasets for training Al-based protein sequence models, in accordance with some embodiments. At block 372, method 370 comprises filtering, by one or more processors, a reference proteome dataset (e.g., the UniRefl 00 dataset) to generate a high completeness proteome data subset comprising reference proteomes selected from a plurality of reference proteomes having a completeness value above a completeness threshold. The reference proteome dataset may be stored on a computer memory communicatively coupled to the one or more processors. The reference proteome dataset may comprise the plurality of reference proteomes. In various aspects, block 372 of method 370 may be implemented in the same manner as block 310 of method 100, such that all description for block 310 of method 100 applies equally for block 372 of method 370.
[0168] At block 374, method 370 further comprises filtering, by the one or more processors, the high completeness proteome data subset to generate a protein confidence-based proteome data subset comprising reference proteomes selected from the high completeness proteome data subset having a protein confidence score above a protein confidence threshold. The protein confidence-based proteome data subset may comprise a plurality of protein sequences. In various aspects, block 374 of method 370 may be implemented in the same manner as block 320 of method 100, such that all description for block 320 of method 100 applies equally for block 374 of method 370.
[0169] At block 376, method 370 further comprises clustering, by the one or more processors, the plurality of protein sequences to generate a set of protein sequence clusters.Each cluster of the plurality of protein sequences may group or otherwise cluster one or more of the plurality of protein sequences of the protein confidence-based proteome data subset. In various aspects, block 376 of method 370 may be implemented in the same manner as block 330 of method 100, such that all description for block 330 of method 100 applies equally for block 376 of method 370.
[0170] At block 378, method 370 further comprises randomly allocating, by the one or more processors, a training cluster subset selected from the set of protein sequence clusters. In various aspects, block 378 of method 370 may be implemented in a similar manner as block 340 of method 100, but where instead of selecting the validation cluster subset from the set of protein sequence clusters, such clusters are selected or otherwise designated be the training cluster subset, and, additionally or alternatively, where the validation cluster subset is excluded from the training cluster subset. In various aspects, the clusters may comprise a greater number of clusters (or an remaining number of clusters or portion thereof) than the validation cluster subset as described for block 340 of method 100.
[0171] At block 380, method 370 further comprises creating, by the one or more processors, a protein sequence training dataset by selecting the protein sequences of the training cluster subset. In various aspects, block 380 of method 370 may be implemented in a similar manner as block 350 of method 100, but where instead of creating a protein sequence validation dataset from as selected from the validation cluster subset, such the protein sequence training dataset is selected from the training cluster subset, and, additionally or alternatively, where the protein sequence validation dataset is excluded from the protein sequence training dataset. In various aspects, the training cluster subset may comprise a greater number of data (or an remaining amount of data or portion thereof) than the protein sequence validation dataset as described for block 350 of method 100.
[0172] At block 382, method 370 further comprises training, by the one or more processors, an Al-based protein sequence model with the protein sequence training dataset.
[0173] At block 384, method 370 further comprises outputting, by the one or more processors from the Al-based protein sequence model, one or more predicted protein sequences.
[0174] In some aspects, at block 386, method 370 further comprises generating, by the one or more processors, validation metrics based on comparison between: (a) a set of predetermined values of a protein sequence validation dataset comprising data excluded fromthe protein sequence training dataset; and (b) a set of predicted values as output by the AI- based protein sequence model. In such aspects, the protein sequence validation dataset may be defined as data excluded from the protein sequence training dataset as implemented during any of blocks 378 and / or 380 as described above.
[0175] Example Results
[0176] FIG. 2A illustrates example validation metrics of Al-based protein sequence models in accordance with some embodiments. In particular, FIG. 2A shows validation metrics as calculated based on outputs of an Al-based protein sequence model when provided inputs for a same task. Al-based protein sequence model may comprise a multi-parameter LLM model, such as a 150 million parameter model trained on proteome data and proteins as described herein. Al-based protein sequence model can comprise any types of Al models related to prediction, analysis, and design of protein sequences, structures, functions, and interactions. Al-based protein sequence model can be trained to output predictions and / or classifications related to protein sequences and / or proteomes. In the example of FIG. 2A, two versions of an Al-based protein sequence model are illustrated, where a first example model was trained from a dataset derived from UniRef50 (validation metrics 207, 217, and 227), and a second example model was trained from a dataset derived from UniRefl 00 (validation metrics 209, 219, and 229).
[0177] The validation data or otherwise validation metrics can produce or otherwise provide diagnostic metrics for assessing the performance of data set and machine learning parameter changes for a given Al-based protein sequence model across various metrics.
[0178] As shown for FIG. 2A, validation set improvements in perplexity, loss, and accuracy are observed when compared with conventional training on conventional validation set sequences, or otherwise conventional validation set protein sequences, when using the UniRefl 00 and UniRef50 datasets, respectively. As shown, diagram 202 illustrates values of perplexity against number of steps (how many times) the model is updated (e.g., updated after processing batches of the training data), e.g., 0 to 500,000 steps (e.g., number of steps 206) in the example of FIG. 2A. As shown for diagram 202, the perplexity values of each of the two Al models decreased as the number of steps increased. Perplexity can measure how well an Al model’s output (e.g., predicted protein sequences) fits or otherwise compares to a test value, e.g., a test value in a validation data set (e.g., the protein sequence validation dataset as described herein). A perplexity value (e.g., any one of perplexity values 204)) canbe used to evaluate how well a given Al model (e.g., Al-based protein sequence model) has learned a distribution of the data (e.g., a protein sequence training dataset) it was trained on. For example, a perplexity value can measure how surprised a model is by a given input, e.g., indicating a degree or measure to whether the Al model was or was not previously trained on the given input, such as how distant or close the given input is to the training data the Al model experienced during training. A lower perplexity value indicates that the model is less surprised and, thus, more robust at predicting a value based on the given input. In some cases, a perplexity score of 1 can indicate perfect prediction, while higher scores can indicate less accurate performance or greater surprise, indicating that the model has not been previously trained to predict the given input with sufficient robustness or accuracy (e.g., values of 14 or greater may indicated insufficient perplexity). As shown for diagram 202, an Al-based protein sequence model trained with each of a UniRefl 00 based dataset (e.g., relating to validation metrics 209) and a UniRef50 dataset (e.g., relating to validation metrics 207) yields reduced, and therefore better, perplexity values over a number of steps in which the model is updated (e.g., number of steps 206). Further, this result demonstrates several improvements with respect to the systems and methods herein. First, for each of the UniRefl 00 based dataset and the UniRef50 dataset, an Al-based protein sequence model that is trained with a protein sequence training dataset, and validated with corresponding a protein sequence validation dataset as excluded from that protein sequence training dataset as described by the systems and methods herein, the Al-based based protein sequence model continuously improves its perplexity value as the Al-based based protein sequence model is updated over 0 to 500,000 steps, and beyond. Second, an Al-based protein sequence model trained with a UniRefl 00 based dataset yields a further improved perplexity value (i.e., a further reduced perplexity value) compared to an Al-based protein sequence model trained with UniRef50 based dataset, when being updated across the same 0 to 500,000 steps (and beyond). Such comparison illustrates that the systems and methods herein are able to improve an Al-based protein sequence model based on the raw data of the UniRefl 00 based dataset alone, without reliance on pre-processed data of the UniRef50 based dataset, which is commonly used by conventional models (e.g., ESM2). Further, training the Al-based protein sequence model with the raw data of the UniRefl 00 based dataset allows such Al-based protein sequence model to be validated with less validation data, and thus be trained withmore training data (e.g., as described for FIG. 2B), while also being more robust when analyzing or inputting new inputs as shown for 2A with respect to perplexity values.
[0179] Further, as shown, diagram 212 illustrates values of loss against number of steps (how many times) the model is updated (e.g., updated after processing batches of the training data), e.g., 0 to 500,000 steps (e.g., number of steps 216) in the example of FIG. 2A. As shown in diagram 212, the loss values of each of the two Al models decreased as the number of steps increased. Loss can measure how well an Al model’s output (e.g., a prediction of and / or related to a given protein or otherwise predicted protein sequences) compares to an actual value. This can include comparing a difference in value between a predicted value as output by an Al-based protein sequence model to an actual value (e.g., a ground truth value) in a validation data set (e.g., the protein sequence validation dataset as described herein). A loss value (e.g., any one of loss values 214) may represent a summation, average, or other statistical value indicating an error detected for each data sample in a given training or validation dataset. For example, loss values can be used in the training process to find the optimal parameter values for the Al model (e.g., weights in neural network). During the training process, reducing the loss value can result in an improved model. Loss values may be computed by executing a loss function, which can comprise any of log loss, cross-entropy loss, mean squared error of loss, likelihood of loss, or other statistical loss function, which can each take as input an error value (e.g., a difference in value between a predicted value as output by an Al-based protein sequence model to an actual value in a validation data set, e.g., the protein sequence validation dataset as described herein). As shown for diagram 212, an Al-based protein sequence model trained with each of a UniRefl 00 based dataset (e.g., relating to validation metrics 219) and a UniRef50 dataset (e.g., relating to validation metrics 217) yields reduced, and therefore better, loss values over a number of steps in which the model is updated (e.g., number of steps 216). Further, this result demonstrates several improvements with respect to the systems and methods herein. First, for each of the UniRefl 00 based dataset and the UniRef50 dataset, an Al-based protein sequence model that is trained with a protein sequence training dataset, and validated with corresponding a protein sequence validation dataset as excluded from that protein sequence training dataset as described by the systems and methods herein, the Al-based based protein sequence model continuously improves its loss value as the Al-based based protein sequence model is updated over 0 to 500,000 steps, and beyond. Second, an Al-based protein sequence modeltrained with a UniRefl 00 based dataset yields a further improved loss value (i.e., a further reduced loss value) compared to an Al-based protein sequence model trained with UniRef50 based dataset, when being updated across the same 0 to 500,000 steps (and beyond). Such comparison illustrates that the systems and methods herein are able to improve an Al-based protein sequence model based on the raw data of the UniRefl 00 based dataset alone, without reliance on pre-processed data of the UniRef50 based dataset, which is commonly used by conventional models (e.g., ESM2). Further, training the Al-based protein sequence model with the raw data of the UniRefl 00 based dataset allows such Al-based protein sequence model to be validated with less validation data, and thus be trained with more training data (e.g., as described for FIG. 2B), while also having less error (or otherwise experiencing reduced loss) as shown for 2A with respect to loss values.
[0180] Additionally, as shown, diagram 222 illustrates values of accuracy against a number of steps (how many times) the model is updated (e.g., updated after processing batches of the training data), e.g., 0 to 500,000 steps (e.g., number of steps 226) in the example of FIG. 2A. As shown in diagram 222, the accuracy values of each of the two Al models increased as the number of steps increased. Accuracy can measure a percentage of correct predictions or classifications (e.g., predicted protein sequences) that a given Al-based protein sequence model achieves, i.e., the number of correct predictions divided by the total number of predictions (e.g., 0.25 or 25%) across all classes, clusters, or other values of a dataset. The number of correct predictions can be determined by comparing a difference in value between an output predicted value of an Al-based protein sequence model to an actual value (e.g., a ground truth value) in a validation data set (e.g., the protein sequence validation dataset as described herein). That is, an accuracy value can measure an Al model’s performance with respect to predictive accuracy. It is typically expressed as a number, percentage, or other value or ratio. An accuracy value can comprise a number of predictions where the predicted value is equal to the true value (e.g., true value in a given validation dataset). Accuracy values can be graphed and monitored during the training phase of an Al model (e.g., as shown for diagram 222 of FIG 2A). Additionally, or alternatively, an accuracy value can define an overall accuracy of the model, when fully trained, e.g., in a last training step and / or when the model is deployed onto specific hardware or equipment. As shown for diagram 222, an Al-based protein sequence model trained with each of a UniRefl 00 based dataset (e.g., relating to validation metrics 229) and a UniRef50 dataset (e.g., relating to validation metrics227) yields increased, and therefore better, accuracy values over a number of steps in which the model is updated (e.g., number of steps 226). Further, this result demonstrates several improvements with respect to the systems and methods herein. First, for each of the UniRefl 00 based dataset and the UniRef50 dataset, an Al-based protein sequence model that is trained with a protein sequence training dataset, and validated with corresponding a protein sequence validation dataset as excluded from that protein sequence training dataset as described by the systems and methods herein, the Al-based based protein sequence model continuously improves its accuracy value as the Al-based based protein sequence model is updated over 0 to 500,000 steps, and beyond. Second, an Al-based protein sequence model trained with a UniRefl 00 based dataset yields a further improved accuracy value (i.e., a further increased accuracy value) compared to an Al-based protein sequence model trained with UniRef50 based dataset, when being updated across the same 0 to 500,000 steps (and beyond). Such comparison illustrates that the systems and methods herein are able to improve an Al-based protein sequence model based on the raw data of the UniRefl 00 based dataset alone, without reliance on pre-processed data of the UniRef50 based dataset, which is commonly used by conventional models (e.g., ESM2). Further, training the Al-based protein sequence model with the raw data of the UniRefl 00 based dataset allows such Al-based protein sequence model to be validated with less validation data, and thus be trained with more training data (e.g., as described for FIG. 2B), while also having increased accuracy error (or otherwise experiencing improved accuracy values) as shown for 2A with respect to accuracy values.
[0181] In this way, the perplexity, loss, and accuracy values, as illustrated for diagrams 202, 212, and 222, respectively, provide examples of large language model (LLM) training curves where the difference between the test runs is the training dataset origin (uniref50 vs unireflOO), where Al-based protein sequence model was trained on one of these datasets, and where each Al-based protein sequence model is improved by use of the systems and methods herein for generating Al validation datasets for validating Al-based protein sequence . In the example of FIG. 2A, lower perplexity and loss, as well as higher accuracy, show that training on unireflOO, and using the Al validation dataset generation systems and method described herein, improves performance on the task based on this validation set. However, it should be noted that lower perplexity and loss, as well as higher accuracy, are also achieved by using the UniRef50 dataset when compared to conventional techniques (e.g., training an ESM2model with conventional techniques). Thus, as shown for FIG. 2A, the Al validation dataset generation systems and method for generating Al validation datasets for validating Al-based protein sequence models can be used to improve validation metrics for, and thus assessment of, Al-based protein sequence model(s) when compared with conventional techniques.
[0182] For example, an Al validation dataset (e.g., a protein sequence validation dataset as described herein) can provide several improvements. Generally, a validation dataset is used to implement hypothesis testing, where a validation dataset seeks to produce an empirically valid test of a given hypothesis. As illustrated for FIG. 2 A, validation metrics of diagrams 202, 212, and 222 illustrate that respective protein sequence validation datasets (e.g., as provided for each of validation metrics 207, 209, 217, 219, 227, and 229) confirm the hypothesis that the training an Al model with the UniRefl 00 dataset results in an improvement (e.g., improved perplexity, loss, and accuracy) for the Al model compared to training an Al model with the UniRef50 dataset. Such improvement results primarily because generation of a protein sequence validation dataset using the UniRefl 00 dataset via the Al validation dataset generation systems and methods as described herein allows for creation of a validation dataset that is independent of the conventional clustering method used to produce standard UniRef50 data and any validation set derived from UniRef50.
[0183] Further, generation of a protein sequence validation dataset using the UniRefl 00 dataset via the Al validation dataset generation systems and methods as described herein provides a further improvement for testing hypothes(es) (e.g., with the protein sequence validation dataset against a given Al-based protein sequence model) involving out-of- distribution performance. For example, in various aspects, a protein sequence validation dataset as generated using the UniRefl 00 dataset via the Al validation dataset generation systems and methods as described herein allows for the protein sequence validation dataset to predict the outcomes of unrelated tests, e.g., and thereof provide predictive output for out-of- distribution performance.
[0184] Further, a protein sequence validation dataset, as generated via the Al validation dataset generation systems and methods as described herein, can result in an improvement in the field of validation testing Al models where such validation test results in one or both of: a) a validation test based on a given protein sequence validation dataset that can be used to prove a given hypothesis wrong, and / or b) a validation test based on a given protein sequencevalidation dataset that results in one or more conclusions (e.g., output) that, for multiple iterations of a given Al model, illustrates the Al model outputs accurate, predictive results.
[0185] Further, a protein sequence validation dataset, as generated via the Al validation dataset generation systems and methods as described herein, can be reused and / or shared. For example, a given protein sequence validation dataset may be used to validate or otherwise test different Al models on different computing devices allowing for comparisons of such different Al models using a common or single protein sequence validation dataset, which can be used to implement ranking of such different Al models against the common or single protein sequence validation dataset as benchmark protein sequence validation dataset. Such common or single protein sequence validation dataset can provides a common dataset for hypothesis testing as described herein. Further, the common or single protein sequence validation dataset can fit into computer memory, allowing for a reduced storage size when compared to storing various validation sets for each set of training data. This reduction results in an improvement to the underling computing device storing the common or single protein sequence validation dataset because less computational resources (e.g., memory) are used.
[0186] Further, an Al validation dataset can be used to identify which Al models contain lower perplexity and loss, as well as higher accuracy, even when such models have fewer model parameters. For example, the Al validation dataset can be used to generate or otherwise identify Al models that have fewer model parameters, and thus may be stored or installed on devices having lower computer memory. Such smaller memory models allow for distribution, deployment, and installation on deployable devices with fewer computational resources.
[0187] Still further, models having fewer parameters use less computing cycles (e.g., less floating point operations per second (FLOPs)) and power consumption, and therefore can be implemented or executed by a wider variety of systems, including smaller computing devices with low power consumption and / or less powerful processors.
[0188] Still further, the protein sequence validation datasets as created or generated using the techniques herein allow for more accurate model assessment, where validation metrics (e.g., metrics for validation and assessing Al models) can be used to demonstrate Al model efficacy, including lower perplexity and loss, as well as higher accuracy, for a given Al-based protein sequence model (e.g., as described herein for FIG. 2A). This demonstrates that training and validating using the protein sequence validation dataset as described herein canbe used to test Al models, update Al models, and thus improve performance, or otherwise identified accuracy, of the given Al-based protein sequence model. That is, in various aspects, the validation metrics provide predictors of performance for enabling updating and improvement of to improve real world performance of an Al model. The validation metrics enable measurement of perplexity and loss, which allows identification of when such metrics or low or otherwise unacceptable in the course of an experiment. Validation metrics provide a prediction of model quality. That is, a validation set, and related validation metrics for a given validation set, allows selection of models that have lower perplexity and loss and higher accuracy.
[0189] FIG. 2B illustrates an example diagram 232 depicting sequence identity threshold values 236 (as a percentage) across a spectrum of training data removed 234 (as a percentage) in accordance with some embodiments. In various aspects, sequence identity can refer to the occurrence of the same nucleotide or amino acid in the same position in aligned sequences, e.g., protein sequences. Further, in various aspects herein, sequence identity can refer to method to represent the similarity between two aligned biological sequences, e.g., protein sequences. In some aspects, sequence identity can be calculated by or otherwise determined by counting a number of aligned positions between two or more biological sequences, where the matching characters are identical. Still further, in some aspects, a sequence identity threshold can refer to a parameter used to define a cut-off or otherwise threshold for clustering biological sequences, e.g., protein sequences. In some aspects, the sequence identity threshold can be a value that ranges from 0 (complete mismatch) to 1 (identical sequences). Additionally, or alternatively, the sequence identity threshold can be represented as a percentage value (e.g., sequence identity threshold values 236), which range from 0% (complete mismatch) to 100% (identical sequences). A sequence identity threshold value can be chosen based on domain knowledge or a widely accepted value within that domain, e.g., for a given protein or otherwise biological sequence.
[0190] With reference to FIG. 2B, values 239 correspond to protein sequences as randomly sampled from a reference proteome dataset (e.g., the UniRefl 00 dataset). Further, values 237 correspond to protein sequences as clustered from the same reference proteome dataset (e.g., the UniRefl 00 dataset), but generated in accordance with the Al validation dataset generation systems and methods disclosed herein, which include generating Al validation datasets for validating Al-based protein sequence models. For example, in various aspects, values 237can correspond to protein sequence validation datasets as created by systems and methods as described herein for FIGs. 1 A and 3 A, or otherwise herein. As a further example, and in alternative various aspects, values 237 can correspond to protein sequence training datasets as created by systems and methods as described herein for FIGs. IB and 3B, or otherwise herein.
[0191] For example, in the embodiment of FIG. 2B, and a total of 15,1849 clusters were randomly allocated or otherwise defined to generate a validation cluster subset selected from a set of protein sequence clusters, e.g., as described and implemented by block 104 (phase 4) of method 100 of FIG. 1 A. Further, 7,593 clusters were selected or otherwise defined for validation exclusion as a protein sequence validation dataset, e.g., as described and implemented by block 105 (phase 5) of method 100 of FIG. 1 A. Still further, 2,806 clusters were selected for validation test set metrics (e.g., validation metrics), or otherwise performance metrics, for determining a validation set for comparing or otherwise generating validation metrics for a given Al model (e.g., an Al-based protein sequence model), for validating and / or evaluating its output, e.g., the Al model’s outputted predicted values, as described and implemented by block 106 (phase 6) of method 100 of FIG. 1A.
[0192] With further reference to FIG. 2B, diagram 232 illustrates that the protein sequence validation datasets, as generated by the systems and methods herein, and as shown for values 237, allow for training an Al-based protein sequence model with a complete set (or more complete set) of training data than values 239 of the randomly sampled protein sequences. This is because the protein sequence validation datasets (e.g., as shown for values 239), as generated by the systems and methods herein, are able to validate the Al-based protein sequence model (and with more accuracy and precision, e.g., as shown for FIGs. 2A and 2B) with less validation data compared to conventional random sampling validation techniques. That is, the validation cluster subset and / or related protein sequence validation datasets (as generated and created by the systems and methods herein (and as represented by values 239)), remove or otherwise exclude less training data (e.g., along spectrum of training data removed 234) for each sequence identity threshold value (e.g., across sequence identity threshold values 236) than an equally sized validation set chosen from random sampling (as represented by values 237). Still further, and said another way, for a given set of data (e.g., a reference proteome dataset comprising plurality of reference proteomes), generation, implementation, and / or use of validation cluster subset(s) and / or related protein sequencevalidation datasets, as described the systems and methods herein, allow for a greater percentage of that given set of data to be allocated as training data (e.g., a protein sequence training dataset) for training an Al-based protein sequence model because less data is used to be reserved or otherwise allocated from the same for given set of data for validation exclusion.
[0193] For example, in various aspects, a protein sequence validation dataset can be excluded or otherwise allocated to create or otherwise allow to remain a larger protein sequence training dataset than a randomly sampled validation dataset randomly sampled from the plurality of reference proteomes of the reference proteome dataset as described herein. In such aspects, the protein sequence validation dataset and the randomly sampled validation dataset have at least one of: (1) a same sequence identity threshold value, or (2) an equal number of protein sequences selected from the plurality of reference proteomes of the reference proteome dataset, but where the protein sequence validation dataset nonetheless is smaller in data size therefore allocating a larger portion of the protein sequence training dataset for the protein sequence training dataset.
[0194] Therefore, generation, implementation, and / or use of validation cluster subset(s) and / or related protein sequence validation datasets, as described the systems and methods herein, provide an improvement to underlying computing devices by providing an implementation to create, generate, or otherwise use a smaller sized validation dataset, which can free up training data for training a more accurate Al model, e.g., an Al-based protein sequence model as described herein. This is because, when less data is taken from a given dataset (e.g., a protein sequence training dataset) for validation purposes (e.g., validation exclusion) the Al model can have additional data (e.g., additional protein sequences, or subfragments or residues thereof) to be trained upon. Said another way, the Al model’s accuracy in predicting, identifying, or otherwise outputting one or more predicted protein sequences can improve, which can improve an underlying computing device upon which the Al model is deployed.
[0195] Still further, the validation cluster subset(s) and / or related protein sequence validation datasets can have a smaller data size compared to a randomly sampled validation dataset, e.g., as illustrated and described for FIG. 2B. These validation cluster subset(s) and / or related protein sequence validation datasets can also improve an underlying computing device because such validation cluster subset(s) and / or related protein sequence validationdatasets use less memory when stored in the underlying computing device. Similarly, a processor, implementing, executing, or otherwise using the validation cluster subset(s) and / or related protein sequence validation datasets use fewer computing resources (e.g., can be implemented with fewer FLOPs) given that the validation cluster subset(s) and / or related protein sequence validation datasets are comparatively smaller, with less data to process, when compared to conventional randomly sampled datasets.
[0196] Additional Aspects
[0197] The below additional aspects describe features consistent with the above related disclosure.
[0198] In some aspects, an artificial intelligence (Al) model generation method for generating and validating Al-based protein sequence models, the Al model generation method is disclosed. The method comprises training an Al-based protein sequence model with a protein sequence training dataset. The protein sequence training dataset may comprise a dataset independent of a protein sequence validation dataset used to validate the Al-based protein sequence model. Still further, the Al-based protein sequence model may be configured to output one or more predicted protein sequences. In addition, and as described herein for FIG. 2B, the creation of the protein sequence validation dataset may allocate allocates the protein sequence training dataset as having a larger dataset size than a randomly sampled validation dataset randomly sampled from a plurality of reference proteomes of a reference proteome dataset, and wherein the protein sequence validation dataset and the randomly sampled validation dataset have at least one of: (1) a same sequence identity threshold value; or (2) an equal number of protein sequences selected from the plurality of reference proteomes of the reference proteome dataset
[0199] In additional aspects, a tangible, non-transitory computer-readable medium may store a protein sequence as output by an Al-based protein sequence model. The non-transitory computer-readable medium may further store computing instructions for validating the AI- based protein sequence model, that when executed by one or more processors cause the one or more processors to filter, by one or more processors, a reference proteome dataset to generate a high completeness proteome data subset comprising reference proteomes selected from a plurality of reference proteomes having a completeness value above a completeness threshold. The reference proteome dataset may be stored on a computer memory communicatively coupled to the one or more processors. The reference proteome dataset maycomprise a plurality of reference proteomes. The computing instructions, when executed by the one or more processors, may cause the one or more processors to filter the high completeness proteome data subset to generate a protein confidence-based proteome data subset comprising reference proteomes selected from the high completeness proteome data subset having a protein confidence score above a protein confidence threshold, wherein the protein confidence-based proteome data subset comprises a plurality of protein sequences. The computing instructions, when executed by the one or more processors, may cause the one or more processors to cluster the plurality of protein sequences to generate a set of protein sequence clusters, each cluster of the plurality of protein sequences grouping one or more of the plurality of protein sequences of the protein confidence-based proteome data subset. The computing instructions, when executed by the one or more processors, may cause the one or more processors to randomly allocate a validation cluster subset selected from the set of protein sequence clusters. The computing instructions, when executed by the one or more processors, may cause the one or more processors to create a protein sequence validation dataset by selecting the protein sequences of the validation cluster subset. The computing instructions may cause the one or more processors to validate the Al-based protein sequence model with the protein sequence validation dataset.
[0200] In still further aspects, a protein sequence may be generated by an Al-based protein sequence model. The Al-based protein sequence model can produce a probability table indicating the probabilities or likelihoods of each amino acid residue at specific positions within a protein sequence. This table can be used to generate one or more protein sequences by selecting the amino acid residue with the highest probability or likelihood at a specific position. Still further, the Al-based protein sequence model may validated by an Al validation dataset generated by a method comprising filtering, by one or more processors, a reference proteome dataset to generate a high completeness proteome data subset comprising reference proteomes selected from a plurality of reference proteomes having a completeness value above a completeness threshold, wherein the reference proteome dataset is stored on a computer memory communicatively coupled to the one or more processors, and wherein the reference proteome dataset comprises the plurality of reference proteomes. The method may further comprise filtering, by the one or more processors, the high completeness proteome data subset to generate a protein confidence-based proteome data subset comprising reference proteomes selected from the high completeness proteome data subset having a proteinconfidence score above a protein confidence threshold. The protein confidence-based proteome data subset comprises a plurality of protein sequences. The method may further comprise clustering, by the one or more processors, the plurality of protein sequences to generate a set of protein sequence clusters. Each cluster of the plurality of protein sequences may group one or more of the plurality of protein sequences of the protein confidence-based proteome data subset. The method may further comprise randomly allocating, by the one or more processors, a validation cluster subset selected from the set of protein sequence clusters. The method may further comprise creating, by the one or more processors, a protein sequence validation dataset by selecting the protein sequences of the validation cluster subset. The method may further comprise validating the Al-based protein sequence model with the protein sequence validation dataset.
[0201] Example Method for Evaluating Generalization Capability of Al-based Protein Sequence Model(s)
[0202] Performing an ablation routine for a given Al model by ablating the training data it is trained on can be used to evaluate how well the Al model generalizes to data that it has been trained on. In the context of Al-based protein sequence models, an ablation routine may involve systematically removing subsets of training data (e.g., based on sequence similarity to a protein sequences in a validation dataset) to enable quantitative assessment of how the Al-based protein sequence model performs on sequences that differ to varying degrees from the distribution of sequences in the training dataset and provides insights into the model’s robustness and reliance on memorized patterns. This assessment can be used to evaluate bias and overfitting of the Al-based protein sequence model and quantify transferability of the model across protein classes and families.
[0203] FIG. 3C illustrates an example method 390 for evaluating generalization capability of one or more Al-based protein sequence models, in accordance with some embodiments. In various aspects, method 390, which may comprise an algorithm, may be implemented on a system (such as a system as illustrated by FIG. 4), including a system comprising one or more processors (e.g., processor 410 of FIG. 4). The one or more processors may implement computing instructions stored on a non-transitory computer-readable storage medium (e.g., (e.g., memory 420 of FIG. 4) to implement the algorithm shown and described for method 390.
[0204] For example, as shown for block 391, method 390 comprises generating, by one or more processors, a validation dataset comprising protein sequences. The validation dataset comprising protein sequences may be generated using the techniques described herein. Block 391 is implemented in subroutines 391a, 391b, 391c, and 391d. At block 391a, method 390 comprises obtaining the protein sequences of the validation dataset, in part, by filtering a reference proteome dataset to generate a high completeness proteome data subset comprising reference proteomes selected from a plurality of reference proteomes having a completeness value above a completeness threshold. The reference proteome dataset comprises a plurality of reference proteomes. Examples of reference proteome datasets are described herein, including in connection with block 101 of method 100 illustrated in FIG. 1 A. Additional techniques that may be used in filtering the reference proteome dataset to generate the high completeness proteome data subset are described, by way of non-limiting example, in connection with block 102 in method 100 illustrated in FIG. 1 A and block 310 in method 300 illustrated in FIG. 3 A.
[0205] At block 391b, method 390 comprises obtaining the protein sequences of the validation dataset, in part, by filtering the high completeness proteome data subset to generate a protein confidence-based proteome data subset comprising reference proteomes selected from the high completeness proteome data subset having a protein confidence score above a protein confidence threshold. The protein confidence-based proteome data subset comprises a plurality of protein sequences. Additional techniques that may be used in filtering the high completeness proteome data subset to generate the protein confidence-based proteome data subset are described in connection with block 103 in method 100 illustrated in FIG. 1 A and block 320 in method 300 illustrated in FIG. 3A.
[0206] At block 391c, method 390 comprises obtaining the protein sequences of the validation dataset, in part, by clustering the plurality of protein sequences of the protein confidence-based proteome data subset to generate a set of protein sequence clusters. Each cluster of the plurality of protein sequences groups one or more of the plurality of protein sequences of the protein confidence-based proteome data subset. Additional techniques that may be used in clustering the protein sequence validation dataset are described in connection with block 330 in method 300 illustrated in FIG. 3A.
[0207] At block 391 d, method 390 comprises obtaining the protein sequences of the validation dataset, in part, by creating the protein sequence validation dataset by selecting theprotein sequences from one or more clusters in the set of protein sequence clusters. The protein sequences of the protein sequence validation dataset may be considered as the set of sequences used for evaluating how well an Al-based protein sequence model performs. In some embodiments, for a given Al-based protein sequence model, the protein sequences in the validation dataset are excluded from the protein sequences included in the training dataset that the Al-based protein sequence model is trained on. In such embodiments, the Al-based protein sequence model would not have been trained on protein sequences in the one or more clusters selected for the protein sequence validation dataset. In such embodiments, performing the ablation routine allows for evaluation of the Al-based protein sequence model’s generalization capability with respect to protein sequences in the one or more clusters selected for the protein sequence validation dataset. In some embodiments, creating the protein sequence validation dataset comprises selecting all the protein sequences of the one or more clusters. In such embodiments, the one or more clusters may be considered to define the sequence space of the protein sequence validation dataset.
[0208] In some aspects, creating the protein sequence validation dataset may further comprise selecting protein sequences in a cluster of the set of protein sequence clusters. In such aspects, the model generalization output (as described for block 394) may define one or more validation metrics indicating performance of the Al-based protein sequence model for the selected protein sequences in the cluster. In such aspects, the selected protein sequences in the cluster may correspond to a protein class or protein family. In this way, the Al-based protein sequence model may be evaluated on its performance with respect to the protein class or protein family that is selected for the protein sequence validation dataset. In embodiments where the protein sequences in the validation dataset are excluded from the protein sequences included in the training dataset the Al-based protein sequence model is trained on, the AI- based protein sequence model would not have been trained on protein sequences associated with the protein class or protein family. In such embodiments, performing the ablation routine allows for evaluation of the Al-based protein sequence model’s generalization capability with respect to protein sequences it has not been trained on, e.g., the protein class or protein family.
[0209] Still further, in some aspects, generating the validation dataset further comprises allocating a validation cluster subset selected from the set of protein sequence clusters. In such aspects, creation of the protein sequence validation dataset further may further compriseselecting protein sequences from the validation cluster subset as the protein sequence validation dataset.
[0210] In additional aspects, generating the validation dataset may further comprise randomly allocating a validation cluster subset selected from the set of protein sequence clusters. In such aspects, creating the protein sequence validation dataset may further comprise selecting the protein sequences of the validation cluster subset. Additional techniques that may be used in allocating the validation cluster subset are described in connection with block 104 in method 100 illustrated in FIG. 1 A and block 340 in method 300 illustrated in FIG. 3 A.
[0211] At block 392, method 390 comprises for each of two or more subsets of a training dataset, removing the subset from the training dataset to obtain a respective ablated training dataset. Each respective ablated training dataset comprises at least one or more protein sequences excluded from the validation dataset. For example, in some aspects, each respective ablated training dataset may be generated by ablating no more than 5% to 10% of the training dataset. Further, each respective ablated training dataset comprises a unique set of protein sequences, e.g., comprising at least one different protein sequence between given ablated training datasets.
[0212] In some aspects, two or more sequence similarity thresholds may be determined by comparing protein sequences in the training dataset to protein sequences in the validation dataset. Each of the two or more subsets can be associated with a respective sequence similarity threshold selected from among the two or more sequence similarity thresholds. In such aspects, removing the subset from the training dataset to obtain the respective ablated training dataset may further comprise applying the respective sequence similarity threshold as a filter to the training dataset to remove protein sequences having a sequence similarity above the respective sequence similarity threshold. In this way, the ablated Al-based protein sequence model is trained and evaluated with a reduced training set, omitting the protein sequences that exceed the sequence similarity threshold. As discussed in connection with FIG. 2B, the sequence similarity thresholds can be a percentage value (e.g., 50%, 70%, 90%).
[0213] In some embodiments, determining the two or more sequence similarity thresholds involves determining maximum similarity based on comparing protein sequences in the validation dataset to protein sequences in the training dataset. In some embodiments, each protein sequence in the validation dataset is compared to each protein sequence in the trainingdataset to determine the maximum similarity. Similarity may be computed using any suitable sequence alignment algorithms (e.g., MMseqs2) or embedding-based metrics (e.g., cosine similarity between learned sequence embeddings obtained using a protein language model). The maximum similarity may represent the maximum identity or relatedness between a given protein sequence in the training dataset and any protein sequence in the validation dataset. Some embodiments involve using these maximum similarities for the protein sequences in the training dataset to define the two or more sequence similarity thresholds.
[0214] At block 393, method 390 comprises implementing, for each respective ablated training dataset, an iterative model routine to evaluate an Al-based protein sequence model trained on the training dataset. The iterative model routine comprises training an ablated AI- based protein sequence model on the ablated training dataset. As performing an ablation to evaluate an Al-based protein sequence model involves studying the effect of a variable, the Al-based protein sequence model being evaluated and each of the ablated Al-based protein sequence models may only differ by the training data used in training the models and other variables (e.g., hyperparameters, model architecture) remain constant or re-tuned for the ablated Al-based protein sequence models to ensure effective training. Examples of hyperparameters include model architecture hyperparameters (e.g., embedding dimension, number of transformer layers, position embedding type) and training hyperparameters (e.g., learning rate, batch size). In some aspects, each of the ablated Al-based protein sequence models are trained with a same set of hyperparameters as the Al-based protein sequence model. In some aspects, the Al-based protein sequence model is trained with a set of hyperparameters and each of the ablated Al-based protein sequence models are trained using the same set of hyperparameters as a baseline. This may allow for certain hyperparameters (e.g., batch size, learning rate) to be re-tuned for the ablated Al-based protein sequence models as they are trained on a reduced data set. Additionally, or alternatively, in some aspects each of the ablated Al-based protein sequence models has a same model architecture as the Al-based protein sequence model, which may be determined by the hyperparameter(s), including a number or type of hyperparam eter(s), selected or chosen for training the ablated Al-based protein sequence models. In some aspects, each of the ablated Al-based protein sequence models are trained with a same set of model architecture hyperparameters as the AI- based protein sequence model.
[0215] Still further, the iterative model routine may further comprise generating an ablated performance score defining a predictive performance of an output of the ablated Al-based protein sequence model. The iterative model routine may further comprise storing in memory the ablated performance score as part of a plurality of ablated performance scores.
[0216] In some aspects, the iterative model routine may further comprise determining, using the validation dataset as input to the ablated Al-based protein sequence model, an output of the ablated Al-based protein sequence model. Generating the ablated performance score may further comprise comparing the output of the ablated Al-based protein sequence model to a ground truth value. In such aspects, the performance score may define the proportion of correctly predicted amino acids relative to a ground truth protein sequence for a given ablated Al-based protein sequence model.
[0217] Still further, in some aspects, the iterative model routine may further comprise determining, using the validation dataset as input to the ablated Al-based protein sequence model, an output of the ablated Al-based protein sequence model. In such aspects, the iterative model routine may determine, using the validation dataset as input to the Al-based protein sequence model, an output of the Al-based protein sequence model. In such aspects, generating the ablated performance score may further comprise comparing the output of the ablated Al-based protein sequence model to the output of the Al-based protein sequence model. In such aspects, the performance score may define a delta value between the AI- based protein sequence model and the ablated Al-based protein sequence model (e.g., a delta value for a validation metric). Turning to FIG. 2A as reference, a delta value for validation metric between the Al-based protein sequence model and the ablated Al-based protein sequence model may be obtained by taking the difference between validation metric values at a particular number of steps (e.g., the final number of steps) or over a range of number of steps (e.g., 40,000-50,000) and averaged.
[0218] At block 394, method 390 comprises determining, based on the plurality of ablated performance scores, a model generalization output for the Al-based protein sequence model. In some aspects, the model generalization output can define one or more validation metrics comprising at least one of: a perplexity metric, a loss metric, and / or an accuracy metric, for example, as described herein.
[0219] Still further, in some aspects, the model generalization output can define one or more improved validation metrics when compared to validation metrics of a different Al-basedprotein sequence model trained on at least a portion of the reference proteome dataset. The one or more improved validation metrics comprising at least one of: an improved perplexity metric, an improved loss metric, and / or an improved accuracy metric.
[0220] The model generalization output may provide an indication of the Al-based protein sequence model’s dependency on the varying similarity of the protein sequences in the training dataset to the protein sequences in the validation dataset. In some embodiments, the plurality of performance scores may indicate that the Al-based protein sequence model has low dependency on the varying similarity of the protein sequences in the training dataset. In such embodiments, the model generalization output may indicate a strong generalization capability of the Al-based protein sequence model. In some embodiments, the plurality of performance scores may indicate that the Al-based protein sequence model has high dependency on the varying similarity of the protein sequences in the training dataset. In such embodiments, the model generalization output may indicate a low generalization capability of the Al-based protein sequence model.
[0221] At block 395, method 390 may optionally comprise outputting, from the Al-based protein sequence model, one or more predicted protein sequences, for example, as described herein.
[0222] At block 396, method 390 may optionally comprise outputting, from the Al-based protein sequence model, information specifying a protein sequence and manufacturing a protein-related product having the protein sequence, for example, as described herein.
[0223] Example Computing System
[0224] An illustrative implementation of computer system 400 that may be used in connection with any of the embodiments of the technology described herein is shown in FIG. 4. That is, FIG. 4 illustrates an Al validation dataset generation system configured to generate Al validation datasets for validating Al-based protein sequence models. Computer system 400 includes one or more processors 410 and one or more articles of manufacture that comprise non-transitory computer-readable storage media (e.g., memory 420 and one or more non-volatile storage media 430). The processor 410 may control writing data to and reading data from the memory 420 and the non-volatile storage media 430 in any suitable manner, as the aspects of the technology described herein are not limited to any particular techniques for writing or reading data. To perform any of the functionality described herein, the processor 410 may execute one or more processor-executable instructions stored in one or more non-transitory computer-readable storage media (e.g., the memory 420), which may serve as non- transitory computer-readable storage media storing processor-executable instructions for execution by the processor 410.
[0225] Computer system 400 may also include a network input / output (I / O) interface 440 via which the computing device may communicate with other computing devices (e.g., over a network), and may also include one or more user I / O interfaces 450, via which the computing device may provide output to and receive input from a user. The user I / O interfaces may include devices such as a keyboard, a mouse, a microphone, a display device (e.g., a monitor or touch screen), speakers, a camera, and / or various other types of I / O devices.
[0226] Example Protein-Related Product Manufacturing
[0227] Some embodiments of the technology described herein involve manufacturing a protein-related product having a protein sequence specified by an output of an Al-based protein sequence model disclosed herein. Manufacturing the protein-related product may involve producing a nucleic acid encoding the protein sequence and expressing the nucleic acid (e.g., using host cells, cell-free expression system) to produce a protein (e.g., a protein to include in the protein-related product). In some embodiments, producing the protein involves culturing a host cell so as to express the protein and harvesting the expressed protein from the cell culture. In some embodiments, producing the protein involves purifying the expressed protein from the cell culture.
[0228] The nucleic acid encoding the protein sequence may include either RNA or DNA, or analogs of RNA or DNA. The nucleic acid can be single-stranded or double-stranded. The nucleic acid can be made from nucleotides, nucleotide analogs, or modified nucleotides such as, though not limited, to methylated and / or capped nucleic acids. Non-natural and altered nucleotides are known in the art. It will be appreciated that a nucleic acid comprising altered or non-natural nucleotides will still be understood to be a nucleic acid.
[0229] In some embodiments, the protein-related product having a protein sequence specified by an output of the Al-based protein sequence model is an antibody. In such embodiments, one or more nucleic acids may encode the protein sequences of the antibody. In some embodiments, a single chain of nucleic acids encodes each polypeptide of the antibody (for example, the light chain and the heavy chain). In some aspects, two or more different nucleic acids chains encode the different polypeptides of the antibody (for example,a first nucleic acid encodes the light chain and a second nucleic acid encodes the heavy chain).
[0230] In some embodiments, the nucleic acids disclosed herein are produced through recombinant DNA techniques known in the art. The skilled artisan will appreciate that, due to the degeneracy of the genetic code, a protein sequence may be encoded by a number of different nucleic acid sequences. Nucleic acid sequences that allow protein-coding sequences to be cloned or expressed are often provided in form of a cloning or expression vector. Suitable cloning and expression vectors can be any of the pUC series (Fermentas Life Sciences), pBluescript series (Stratagene, LaJoIIa, CA), pET series (Novagen, Madison, WI), pGEX series (Pharmacia Biotech, Uppsala, Sweden), and pEX series (Clontech, Palo Alto, C A). Bacteriophage vectors, such as GTIO, GT1 1, ZapII (Stratagene), EMBL4, and NM1 149, also can be used. Examples of animal expression vectors include pEUK-Cl, pMAM and pMAMneo (Clontech).
[0231] The recombinant expression vectors encoding protein sequences described herein can comprise a nucleic acid in a form suitable for expression of the nucleic acid in a host cell. The recombinant expression vectors include one or more regulatory sequences, selected on the basis of the host cells to be used for expression, which is operably linked to the nucleic acid sequence to be expressed. Regulatory sequences may include promoter, enhancer, origin of replication, transcriptional termination sequence, and ribosome binding site. Examples of regulatory sequences include those that direct constitutive expression of a nucleotide sequence in many types of host cells (e.g., SV40 early gene enhancer, Rous sarcoma virus promoter and cytomegalovirus promoter), those that direct expression of the nucleotide sequence only in certain host cells (e.g., tissue-specific regulatory sequences, see Voss et al., 1986, Trends Biochem. Sci. 11 :287, Maniatis et al., 1987, Science 236: 1237), and those that direct inducible expression of a nucleotide sequence in response to particular treatment or condition (e.g., the metallothionin promoter in mammalian cells and the tet-responsive and / or streptomycin responsive promoter in both prokaryotic and eukaryotic systems). Commonly, expression vectors will contain selection markers, e.g., tetracycline, neomycin, and dihydrofolate reductase, to permit detection of those cells transformed with the desired DNA sequences. It will be appreciated by those skilled in the art that the design of the expression vector can depend on such factors as the choice of the host cell to be transformed, the level ofexpression of protein desired, etc. The vectors are typically replicable in the host organisms either as epi somes or as an integral part of the host chromosomal DNA.
[0232] Host cells that are capable of producing an expression product encoded by the nucleic acid (e.g., mRNA, protein) are important for recombinant production of the antibodies disclosed herein. The host cell in some aspects is an adherent cell or a suspended cell, i.e., a cell that grows in suspension. The host cell in exemplary aspects is a cultured cell or a primary cell, i.e., isolated directly from an organism, e.g., a human that is adapted for growth in culture conditions. The host cell can be of any cell type, can originate from any type of tissue, and can be of any developmental stage.
[0233] In exemplary aspects, the cell is a eukaryotic cell, including, but not limited to, a yeast cellor mammalian cell. Such host cells are described in the art. See, e.g., Kunert et al., Appl. Microbiol Biotechnol. 100: 3451-61 (2016). In exemplary aspects, the eukaryotic cells are mammalian cells. In exemplary aspects, the mammalian cells are non-human mammalian cells. In some aspects, the cells are Chinese Hamster Ovary (CHO) cells and derivatives thereof (e.g., CHO-K1, CHO pro-3, CS9), mouse myeloma cells (e.g., NS0, GS-NS0, Sp2 / 0), cells engineered to be deficient in dihydrofolatereductase (DHFR) activity (e.g., DUKX-X11, DG44), human embryonic kidney 293 (HEK293) cells or derivatives thereof (e.g., HEK293T, HEK293-EBNA), green African monkey kidney cells (e.g., COS cells, VERO cells), human cervical cancer cells (e.g., HeLa), human bone osteosarcoma epithelial cells U2-OS, adenocarcinomic human alveolar basal epithelial cells A549, human fibrosarcoma cells HT1080, mouse brain tumor cells CAD, embryonic carcinoma cells P19, mouse embryo fibroblast cells NTH 3T3, mouse fibroblast cells L929, mouse neuroblastoma cells N2a, human breast cancer cells MCF-7, retinoblastoma cells Y79, human retinoblastoma cells SO- Rb50, human liver cancer cells Hep G2, mouse B myeloma cells J558L, or baby hamster kidney (BHK) cells (Gaillet et al. 2007; Khan, Adv Pharm Bull 3(2): 257-263 (2013)).Commonly used host cells include CHO cells and HEK293 cells. In a particular embodiment, the host cell is CS9 (a CHO cell line). For purposes of amplifying or replicating the vector, the host cell in some aspects is a prokaryotic cell, e.g., a bacterial cell.
[0234] In exemplary aspects, the step of culturing a host cell comprises culturing the host cell in a growth medium to support the growth and expansion of the host cell, and, the growth medium increases cell density, culture viability and productivity in a timely manner. In further exemplary aspects, the growth medium comprises amino acids, vitamins, inorganicsalts, glucose, and serum as a source of growth factors, hormones, and attachment factors. In exemplary aspects, the growth medium is a fully chemically defined media consisting of amino acids, vitamins, trace elements, inorganic salts, lipids and insulin or insulin-like growth factors. In addition to nutrients, the growth medium also helps maintain pH and osmolality. Several growth media are commercially available and are described in the art. See, e.g., Arora, "Cell Culture Media: A Review" MATER METHODS 3: 175 (2013). Various methods of protein purification may be employed to purify the protein disclosed herein, and such methods are known in the art.
[0235] The protein-related products described herein can be biosynthesized, purified, and formulated for administration. For example, an appropriate host cell, such as HEK 293 or CHO, is either transiently or stably transfected with an expression system for secreting proteins using a suitable vector system. Vectors suitable for expression and secretion of antibodies from these commonly-used host cells are well-known. Following expression and secretion of the antibody, the medium is clarified to remove cells and the clarified medium is purified using any of many commonly-used techniques. Protein fractions are detected, such as by SDS-PAGE, and then are pooled. Further purification is optional, depending on the intended use. The protein may be concentrated and / or sterile filtered using common techniques. The protein may undergo purification by techniques such as size exclusion, hydrophobic interaction, cation exchange, anion exchange, affinity, or hydroxyapatite chromatography.
[0236] In some embodiments, the protein-related product is an antibody and manufacturing the protein-related product comprises producing the antibody. Producing the antibody may involve culturing a host cell so as to express the antibody and harvesting the expressed antibody from the cell culture. The host cell can be any of the host cells described herein (e.g., CHO cells, NS0 cells, COS cells, VERO cells, or BHK cells). The antibodies described herein can be biosynthesized, purified, and formulated for administration. For example, an appropriate host cell, such as HEK 293 or CHO, is either transiently or stably transfected with an expression system for secreting antibodies using a predetermined HC:LC vector ratio if two vectors are used, or a single vector system encoding both heavy chain and light chain. Vectors suitable for expression and secretion of antibodies from these commonly-used host cells are well-known. Following expression and secretion of the antibody, the medium is clarified to remove cells and the clarified medium is purified using any of many commonly-used techniques. For example, the medium may be applied to a Protein A or G column that has been equilibrated with a buffer, such as phosphate buffered saline (pH 7.4). The column is washed to remove nonspecific binding components. The bound antibody is eluted, for example, by a pH gradient (such as 0.1 M sodium phosphate buffer pH 6.8 to 0.1 M sodium citrate buffer pH 2.5). Antibody fractions are detected, such as by SDS-PAGE, and then are pooled. Further purification is optional, depending on the intended use. The antibody may be concentrated and / or sterile filtered using common techniques. The antibody may undergo purification by techniques such as size exclusion, hydrophobic interaction, cation exchange, anion exchange, affinity, or hydroxyapatite chromatography. The purity of the antibody after these chromatography steps is typically greater than 95%. The product may be formulated into a liquid formulation, held in storage by freezing at -70 °C or may be lyophilized.
[0237] ADDITIONAL CONSIDERATIONS
[0238] The above-described embodiments can be implemented in any of numerous ways. For example, the embodiments may be implemented using hardware, software, or a combination thereof. When implemented in software, the software code can be executed on any suitable processor (e.g., a microprocessor) or collection of processors, whether provided in a single computing device or distributed among multiple computing devices. It should be appreciated that any component or collection of components that perform the functions described above can be generically considered as one or more controllers that control the above-described functions. The one or more controllers can be implemented in numerous ways, such as with dedicated hardware, or with general purpose hardware (e.g., one or more processors) that is programmed using microcode or software to perform the functions recited above.
[0239] In this respect, it should be appreciated that one implementation of the embodiments described herein comprises at least one computer-readable storage medium (e.g., RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or other tangible, non-transitory computer-readable storage medium) encoded with a computer program (i.e., a plurality of executable instructions) that, when executed on one or more processors, performs the above-described functions of one or more embodiments. The computer-readable medium may be transportable such that the program stored thereon can be loaded onto any computing device to implementaspects of the techniques described herein. In addition, it should be appreciated that the reference to a computer program which, when executed, performs any of the above-described functions, is not limited to an application program running on a host computer. Rather, the terms computer program and software are used herein in a generic sense to reference any type of computer code (e.g., application software, firmware, microcode, or any other form of computer instruction) that can be employed to program one or more processors to implement aspects of the techniques described herein.
[0240] The foregoing description of implementations provides illustration and description but is not intended to be exhaustive or to limit the implementations to the precise form disclosed. Modifications and variations are possible in light of the above teachings or may be acquired from practice of the implementations. In other implementations the methods depicted in these figures may include fewer operations, different operations, differently ordered operations, and / or additional operations. Further, non-dependent blocks may be performed in parallel.
[0241] It will be apparent that example aspects, as described above, may be implemented in many different forms of software, firmware, and hardware in the implementations illustrated in the figures. Further, certain portions of the implementations may be implemented as a “module” that performs one or more functions. This module may include hardware, such as a processor, an application-specific integrated circuit (ASIC), or a field-programmable gate array (FPGA), or a combination of hardware and software.
[0242] Having thus described several aspects and embodiments of the technology set forth in the disclosure, it is to be appreciated that various alterations, modifications, and improvements will readily occur to those skilled in the art. Such alterations, modifications, and improvements are intended to be within the spirit and scope of the technology described herein. For example, those of ordinary skill in the art will readily envision a variety of other means and / or structures for performing the function and / or obtaining the results and / or one or more of the advantages described herein, and each of such variations and / or modifications is deemed to be within the scope of the embodiments described herein. Those skilled in the art will recognize or be able to ascertain using no more than routine experimentation many equivalents to the specific embodiments described herein. It is, therefore, to be understood that the foregoing embodiments are presented by way of example only and that, within the scope of the appended claims and equivalents thereto, inventive embodiments may bepracticed otherwise than as specifically described. In addition, any combination of two or more features, systems, articles, materials, kits, and / or methods described herein, if such features, systems, articles, materials, kits, and / or methods are not mutually inconsistent, is included within the scope of the present disclosure.
[0243] The above-described embodiments can be implemented in any of numerous ways. One or more aspects and embodiments of the present disclosure involving the performance of processes or methods may utilize program instructions executable by a device (e.g., a computer, a processor, or other device) to perform, or control performance of, the processes or methods. In this respect, various inventive concepts may be embodied as a computer readable storage medium (or multiple computer readable storage media) (e.g., a computer memory, one or more floppy discs, compact discs, optical discs, magnetic tapes, flash memories, circuit configurations in Field Programmable Gate Arrays or other semiconductor devices, or other tangible computer storage medium) encoded with one or more programs that, when executed on one or more computers or other processors, perform methods that implement one or more of the various embodiments described above. The computer readable medium or media can be transportable, such that the program or programs stored thereon can be loaded onto one or more different computers or other processors to implement various ones of the aspects described above. In some embodiments, computer readable media may be non-transitory media.
[0244] The terms “program” or “software” are used herein in a generic sense to refer to any type of computer code or set of computer-executable instructions that can be employed to program a computer or other processor to implement various aspects as described above. Additionally, it should be appreciated that according to one aspect, one or more computer programs that when executed perform methods of the present disclosure need not reside on a single computer or processor but may be distributed in a modular fashion among a number of different computers or processors to implement various aspects of the present disclosure.
[0245] Computer-executable instructions may be in many forms, such as program modules, executed by one or more computers or other devices. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. Typically, the functionality of the program modules may be combined or distributed as desired in various embodiments.
[0246] Also, data structures may be stored in computer-readable media in any suitable form. For simplicity of illustration, data structures may be shown to have fields that are related through location in the data structure. Such relationships may likewise be achieved by assigning storage for the fields with locations in a computer-readable medium that conveys a relationship between the fields. However, any suitable mechanism may be used to establish a relationship between information in fields of a data structure, including through the use of pointers, tags or other mechanisms that establish relationship between data elements.
[0247] When implemented in software, the software code can be executed on any suitable processor or collection of processors, whether provided on a single computer or distributed among multiple computers.
[0248] Also, a computer may have one or more input and output devices. These devices can be used, among other things, to present a user interface. Examples of output devices that can be used to provide a user interface include printers or display screens for visual presentation of output and speakers or other sound generating devices for audible presentation of output. Examples of input devices that can be used for a user interface include keyboards, and pointing devices, such as mice, touch pads, and digitizing tablets. As another example, a computer may receive input information through speech recognition or in other audible formats.
[0249] Such computers may be interconnected by one or more networks in any suitable form, including a local area network or a wide area network, such as an enterprise network, and intelligent network (IN) or the Internet. Such networks may be based on any suitable technology and may operate according to any suitable protocol and may include wireless networks, wired networks or fiber optic networks.
[0250] Also, as described, some aspects may be embodied as one or more methods. The acts performed as part of the method may be ordered in any suitable way. Accordingly, embodiments may be constructed in which acts are performed in an order different than illustrated, which may include performing some acts simultaneously, even though shown as sequential acts in illustrative embodiments.
[0251] All definitions, as defined and used herein, should be understood to control over dictionary definitions, definitions in documents incorporated by reference, and / or ordinary meanings of the defined terms.
[0252] The indefinite articles “a” and “an,” as used herein in the specification and in the claims, unless clearly indicated to the contrary, should be understood to mean “at least one.”
[0253] The phrase “and / or,” as used herein in the specification and in the claims, should be understood to mean “either or both” of the elements so conjoined, i.e., elements that are conjunctively present in some cases and disjunctively present in other cases. Multiple elements listed with “and / or” should be construed in the same fashion, i.e., “one or more” of the elements so conjoined. Other elements may optionally be present other than the elements specifically identified by the “and / or” clause, whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, a reference to “A and / or B,” when used in conjunction with open-ended language such as “comprising” can refer, in one embodiment, to A only (optionally including elements other than B); in another embodiment, to B only (optionally including elements other than A); in yet another embodiment, to both A and B (optionally including other elements); etc.
[0254] As used herein in the specification and in the claims, the phrase “at least one,” in reference to a list of one or more elements, should be understood to mean at least one element selected from any one or more of the elements in the list of elements, but not necessarily including at least one of each and every element specifically listed within the list of elements and not excluding any combinations of elements in the list of elements. This definition also allows that elements may optionally be present other than the elements specifically identified within the list of elements to which the phrase “at least one” refers, whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, “at least one of A and B” (or, equivalently, “at least one of A or B,” or, equivalently “at least one of A and / or B”) can refer, in one embodiment, to at least one, optionally including more than one, A, with no B present (and optionally including elements other than B); in another embodiment, to at least one, optionally including more than one, B, with no A present (and optionally including elements other than A); in yet another embodiment, to at least one, optionally including more than one, A, and at least one, optionally including more than one, B (and optionally including other elements); etc.
[0255] In the claims, as well as in the specification above, all transitional phrases such as “comprising,” “including,” “carrying,” “having,” “containing,” “involving,” “holding,”“composed of,” and the like are to be understood to be open-ended, i.e., to mean includingbut not limited to. Only the transitional phrases “consisting of’ and “consisting essentially of’ shall be closed or semi-closed transitional phrases, respectively.
[0256] The terms “approximately,” “substantially,” and “about” may be used to mean within ±20% of a target value in some embodiments, within ±10% of a target value in some embodiments, within ±5% of a target value in some embodiments, within ±2% of a target value in some embodiments. The terms “approximately,” “substantially,” and “about” may include the target value.
Claims
CLAIMSWhat is claimed is:
1. An artificial intelligence (Al) validation dataset generation system configured to generate Al validation datasets for validating Al-based protein sequence models, the Al validation dataset generation system comprising: one or more processors; a computer memory communicatively coupled to the one or more processors; a reference proteome dataset stored on the computer memory and comprising a plurality of reference proteomes; and computing instructions stored on the computer memory, and that when executed by the one or more processors, cause the one or more processors to: filter the reference proteome dataset to generate a high completeness proteome data subset comprising reference proteomes selected from the plurality of reference proteomes having a completeness value above a completeness threshold; filter the high completeness proteome data subset to generate a protein confidence-based proteome data subset comprising reference proteomes selected from the high completeness proteome data subset having a protein confidence score above a protein confidence threshold, wherein the protein confidence-based proteome data subset comprises a plurality of protein sequences, cluster the plurality of protein sequences to generate a set of protein sequence clusters, each cluster of the plurality of protein sequences grouping one or more of the plurality of protein sequences of the protein confidence-based proteome data subset, randomly allocate a validation cluster subset selected from the set of protein sequence clusters, and create a protein sequence validation dataset by selecting the protein sequences of the validation cluster subset.
2. The Al validation dataset generation system of claim 1, wherein the computing instructions, when executed by the one or more processors, further cause the one or more processors to:train an Al-based protein sequence model with a protein sequence training dataset, wherein the protein sequence validation dataset comprises a dataset independent of the protein sequence training dataset used to train the Al-based protein sequence model; and output, from the Al-based protein sequence model, one or more predicted protein sequences.
3. The Al validation dataset generation system of claim 2, wherein the one or more predicted protein sequences comprise data for predicting or determining one or more of: viscosity, post-translational modification, isomerization, thermal stability, pH, expression titre, aggregation, immunogenicity, binding affinity, binding avidity, clipping, protein secondary structure, protein tertiary structure, protein quaternary structure, disorder, function, enzymatic activity, homology, cellular localization, phenotypic effect, and / or half-life.
4. The Al validation dataset generation system of claim 2, wherein the protein sequence training dataset comprises protein sequences selected from at least a portion of the protein confidence-based proteome data subset comprising the plurality of protein sequences.
5. The Al validation dataset generation system of claim 4, wherein the protein sequence training dataset comprises protein sequences different from those selected for the protein sequence validation dataset from the validation cluster subset.
6. The Al validation dataset generation system of any one of claims 1-5, wherein at least one of the plurality of reference proteomes is linked to one or more sub-fragments or one or more residues of a given protein sequence in the reference proteome dataset.
7. The Al validation dataset generation system of claim 6, wherein at least one of the validation cluster subset or the protein sequence validation dataset comprises the one or more sub-fragments or the one or more residues of the given protein sequence in the reference proteome dataset.
8. The Al validation dataset generation system of claim 2, wherein the computing instructions when executed by the one or more processors, further cause the oneor more processors to: validate the one or more predicted protein sequences as output by the Al-based protein sequence model with the protein sequence validation dataset.
9. The Al validation dataset generation system of any one of claims 1-8, wherein creation of the protein sequence validation dataset allocates a larger protein sequence training dataset than a randomly sampled validation dataset randomly sampled from the plurality of reference proteomes of the reference proteome dataset, and wherein the protein sequence validation dataset and the randomly sampled validation dataset have at least one of: (1) a same sequence identity threshold value; or (2) an equal number of protein sequences selected from the plurality of reference proteomes of the reference proteome dataset.
10. The Al validation dataset generation system of any one of claims 1-9, wherein the computing instructions when executed by the one or more processors, further cause the one or more processors to: remove, based on a sequence identity threshold value, one or more of the protein sequences of the protein sequence validation dataset to create an ablated protein sequence validation dataset.
11. The Al validation dataset generation system of claim 10, wherein the removed one or more of the protein sequences comprises no more than 5% to 10% of the protein sequence validation dataset.
12. The Al validation dataset generation system of claim 2, wherein the computing instructions, when executed by the one or more processors, further cause the one or more processors to: update the protein sequence validation dataset to comprise protein sequences each having a maximum protein length that the Al-based protein sequence model was trained on.
13. The Al validation dataset generation system of claim 2, wherein the reference proteome dataset comprises proteins having different protein lengths, and wherein the computing instructions, when executed by the one or more processors, further cause the one or more processors to:truncate one or more of protein lengths of at least a subset of the proteins to have a truncated length, the truncated length comprising a maximum content length for training the Al-based protein sequence model, and train the Al-based protein sequence model with training data comprising at least a portion of the reference proteome dataset having the subset of the proteins having the truncated length.
14. The Al validation dataset generation system of any one of claims 1-13, wherein the completeness threshold comprises a value of 2 / 3 completeness or greater.
15. The Al validation dataset generation system of any one of claims 1-14, wherein 5% to 15% of the set of protein sequence clusters are randomly selected for creation of the protein sequence validation dataset.
16. The Al validation dataset generation system of any one of claims 1-15, wherein 15% or more of the set of protein sequence clusters are randomly selected for creation of the protein sequence validation dataset.
17. The Al validation dataset generation system of claim 2, wherein the computing instructions, when executed by the one or more processors, further cause the one or more processors to: generate validation metrics based on comparison between: (a) a set of predetermined values of the protein sequence validation dataset; and (b) the one or more predicted protein sequences as output by the Al-based protein sequence model .
18. The Al validation dataset generation system of claim 17, wherein the validation metrics comprise at least one improved metric selected from: a perplexity metric, a loss metric, and / or an accuracy metric, when analyzed against validation metrics of a previously trained version of the Al-based protein sequence model.
19. The Al validation dataset generation system of claim 18, wherein the at least one improved metric comprises: a reduced perplexity value, a reduced loss value, and / or an increased accuracy value.
20. The Al validation dataset generation system of claim 18, wherein the at least one improved metric comprises: a reduced perplexity value when the Al-based protein sequence model is trained on a UniRefl 00 dataset, a reduced loss value when the Al-based protein sequence model is trained on the UniRefl 00 dataset, and / or an increased accuracy value when the Al-based protein sequence model is trained on the UniRefl 00 dataset.
21. The Al validation dataset generation system of claim 20, wherein the at least one improved metric comprises: a reduced perplexity value when the Al-based protein sequence model is trained on a UniRefl 00 dataset compared to a second Al-based protein sequence model trained on a UniRef50 dataset, a reduced loss value when the Al-based protein sequence model is trained on the UniRefl 00 dataset compared to a second Al-based protein sequence model trained on the UniRef50 dataset, and / or an increased accuracy value when the Al-based protein sequence model is trained on the UniRefl 00 dataset compared to a second Al-based protein sequence model trained on the UniRef50 dataset.
22. The Al validation dataset generation system of claim 2, wherein the one or more predicted protein sequences as output by the Al-based protein sequence model is used to develop a protein-related product.
23. The Al validation dataset generation system of claim 22, wherein the protein- related product is predicted to treat a disease.
24. The Al validation dataset generation system of claim 22, wherein the protein- related product is predicted to provide a prophylaxis for a disease or medical condition.
25. The Al validation dataset generation system of claim 22, wherein the output of the one or more predicted protein sequences is used to monitor or measure the protein-related product during manufacture of the protein-related product.
26. An artificial intelligence (Al) validation dataset generation method for generating Al validation datasets for validating Al-based protein sequence models, the Al validation dataset generation method comprising:filtering, by one or more processors, a reference proteome dataset to generate a high completeness proteome data subset comprising reference proteomes selected from a plurality of reference proteomes having a completeness value above a completeness threshold, wherein the reference proteome dataset is stored on a computer memory communicatively coupled to the one or more processors, and wherein the reference proteome dataset comprises the plurality of reference proteomes; filtering, by the one or more processors, the high completeness proteome data subset to generate a protein confidence-based proteome data subset comprising reference proteomes selected from the high completeness proteome data subset having a protein confidence score above a protein confidence threshold, wherein the protein confidence-based proteome data subset comprises a plurality of protein sequences; clustering, by the one or more processors, the plurality of protein sequences to generate a set of protein sequence clusters, each cluster of the plurality of protein sequences grouping one or more of the plurality of protein sequences of the protein confidence-based proteome data subset; randomly allocating, by the one or more processors, a validation cluster subset selected from the set of protein sequence clusters; and creating, by the one or more processors, a protein sequence validation dataset by selecting the protein sequences of the validation cluster subset.
27. The Al validation dataset generation method of claim 26 further comprising: training an Al-based protein sequence model with a protein sequence training dataset, wherein the protein sequence validation dataset comprises a dataset independent of the protein sequence training dataset used to train the Al-based protein sequence model, and outputting, from the Al-based protein sequence model, one or more predicted protein sequences.
28. The Al validation dataset generation method of claim 27, wherein the one or more predicted protein sequences comprise data for predicting or determining one or more of viscosity, post-translational modification, isomerization, thermal stability, pH, expression titre, aggregation, immunogenicity, binding affinity, binding avidity, clipping, proteinsecondary structure, protein tertiary structure, protein quaternary structure, disorder, function, enzymatic activity, homology, cellular localization, phenotypic effect, and / or half-life.
29. The Al validation dataset generation method of claim 27, wherein the protein sequence training dataset comprises protein sequences selected from at least a portion of the protein confidence-based proteome data subset comprising the plurality of protein sequences.
30. The Al validation dataset generation method of claim 29, wherein the protein sequence training dataset comprises protein sequences different from those selected for the protein sequence validation dataset from the validation cluster subset.
31. The Al validation dataset generation method of any one of claims 26-30, wherein at least one of the plurality of reference proteomes is linked to one or more subfragments or one or more residues of a given protein sequence in the reference proteome dataset.
32. The Al validation dataset generation method of claim 31, wherein at least one of the validation cluster subset or the protein sequence validation dataset comprises the one or more sub-fragments or the one or more residues of the given protein sequence in the reference proteome dataset.
33. The Al validation dataset generation method of claim 27 further comprising validating the one or more predicted protein sequences as output by the Al-based protein sequence model with the protein sequence validation dataset.
34. The Al validation dataset generation method of any one of claims 26-33, wherein creation of the protein sequence validation dataset allocates a larger protein sequence training dataset than a randomly sampled validation dataset randomly sampled from the plurality of reference proteomes of the reference proteome dataset, and wherein the protein sequence validation dataset and the randomly sampled validation dataset have at least one of: (1) a same sequence identity threshold value; or (2) an equal number of protein sequences selected from the plurality of reference proteomes of the reference proteome dataset.
35. The Al validation dataset generation method of any one of claims 26-34 further comprising: removing, based on a sequence identity threshold value, one or more of the protein sequences of the protein sequence validation dataset to create an ablated protein sequence validation dataset.
36. The Al validation dataset generation method of claim 35, wherein the removed one or more of the protein sequences comprises no more than 5% to 10% of the protein sequence validation dataset.
37. The Al validation dataset generation method of any one of claims 26-36 further comprising: updating the protein sequence validation dataset to comprise protein sequences each having a maximum protein length that the Al-based protein sequence model was trained on.
38. The Al validation dataset generation method of claim 27, wherein the reference proteome dataset comprises proteins having different protein lengths, and wherein the validation dataset generation method further comprises: truncating one or more of protein lengths of at least a subset of the proteins to have a truncated length, the truncated length comprising a maximum content length for training the Al-based protein sequence model, and training the Al-based protein sequence model with training data comprising at least a portion of the reference proteome dataset having the subset of the proteins having the truncated length.
39. The Al validation dataset generation method of claim 27, wherein the completeness threshold comprises a value of 2 / 3 completeness or greater.
40. The Al validation dataset generation method of claim 27, wherein 5% to 15% of the set of protein sequence clusters are randomly selected for creation of the protein sequence validation dataset.
41. The Al validation dataset generation method of claim 27, wherein 15% or more of the set of protein sequence clusters are randomly selected for creation of the protein sequence validation dataset.
42. The Al validation dataset generation method of claim 27 further comprising: generating, by the one or more processors, validation metrics based on comparison between: (a) a set of predetermined values of the protein sequence validation dataset; and (b) the one or more predicted protein sequences as output by the Al-based protein sequence model.
43. The Al validation dataset generation method of claim 42, wherein the validation metrics comprise at least one improved metric selected from: a perplexity metric, a loss metric, and / or an accuracy metric, when analyzed against validation metrics of a previously trained version of the Al-based protein sequence model.
44. The Al validation dataset generation method of claim 43, wherein the at least one improved metric comprises: a reduced perplexity value, a reduced loss value, and / or an increased accuracy value.
45. The Al validation dataset generation method of claim 43, wherein the at least one improved metric comprises: a reduced perplexity value when the Al-based protein sequence model is trained on a UniRefl 00 dataset, a reduced loss value when the Al-based protein sequence model is trained on the UniRefl 00 dataset, and / or an increased accuracy value when the Al-based protein sequence model is trained on the UniRefl 00 dataset.
46. The Al validation dataset generation method of claim 45, wherein the at least one improved metric comprises: a reduced perplexity value when the Al-based protein sequence model is trained on a UniRefl 00 dataset compared to a second Al-based protein sequence model trained on a UniRef50 dataset, a reduced loss value when the Al-based protein sequence model is trained on the UniRefl 00 dataset compared to a second Al-based protein sequence model trained on the UniRef50 dataset, and / or an increased accuracy value when the Al-based protein sequence model is trained on the UniRefl 00 dataset compared to a second Al-based protein sequence model trained on the UniRef50 dataset.
47. The Al validation dataset generation method of claim 27, wherein the one or more predicted protein sequences as output by the Al-based protein sequence model is used to develop a protein-related product.
48. The Al validation dataset generation method of claim 47, wherein the protein- related product is predicted to treat a disease.
49. The Al validation dataset generation method of claim 47, wherein the protein- related product is predicted to provide a prophylaxis for a disease or medical condition.
50. The Al validation dataset generation method of claim 47, wherein the output of the one or more predicted protein sequences is used to monitor or measure the protein- related product during manufacture of the protein-related product.
51. A tangible, non-transitory computer-readable medium storing instructions for generating artificial intelligence (Al) validation datasets for validating Al-based protein sequence models, that when executed by one or more processors cause the one or more processors to: filter a reference proteome dataset to generate a high completeness proteome data subset comprising reference proteomes selected from a plurality of reference proteomes having a completeness value above a completeness threshold, wherein the reference proteome dataset is stored on a computer memory communicatively coupled to the one or more processors, and wherein the reference proteome dataset comprises the plurality of reference proteomes; filter the high completeness proteome data subset to generate a protein confidencebased proteome data subset comprising reference proteomes selected from the high completeness proteome data subset having a protein confidence score above a protein confidence threshold, wherein the protein confidence-based proteome data subset comprises a plurality of protein sequences; cluster the plurality of protein sequences to generate a set of protein sequence clusters, each cluster of the plurality of protein sequences grouping one or more of the plurality of protein sequences of the protein confidence-based proteome data subset;randomly allocate a validation cluster subset selected from the set of protein sequence clusters; and create a protein sequence validation dataset by selecting the protein sequences of the validation cluster subset.
52. The tangible, non-transitory computer-readable medium of claim 51, wherein the computing instructions, when executed by the one or more processors, further cause the one or more processors to: train an Al-based protein sequence model with a protein sequence training dataset, wherein the protein sequence validation dataset comprises a dataset independent of the protein sequence training dataset used to train the Al-based protein sequence model; and output, from the Al-based protein sequence model, one or more predicted protein sequences.
53. The tangible, non-transitory computer-readable medium of claim 52, wherein the one or more predicted protein sequences comprise data for predicting or determining one or more of: viscosity, post-translational modification, isomerization, thermal stability, pH, expression titre, aggregation, immunogenicity, binding affinity, binding avidity, clipping, protein secondary structure, protein tertiary structure, protein quaternary structure, disorder, function, enzymatic activity, homology, cellular localization, phenotypic effect, and / or halflife.
54. The tangible, non-transitory computer-readable medium of claim 52, wherein the protein sequence training dataset comprises protein sequences selected from at least a portion of the protein confidence-based proteome data subset comprising the plurality of protein sequences.
55. The tangible, non-transitory computer-readable medium of claim 54, wherein the protein sequence training dataset comprises protein sequences different from those selected for the protein sequence validation dataset from the validation cluster subset.
56. The tangible, non-transitory computer-readable medium of any one of claims 51-55, wherein at least one of the plurality of reference proteomes is linked to one or moresub-fragments or one or more residues of a given protein sequence in the reference proteome dataset.
57. The tangible, non-transitory computer-readable medium of claim 56, wherein at least one of the validation cluster subset or the protein sequence validation dataset comprises the one or more sub-fragments or the one or more residues of the given protein sequence in the reference proteome dataset.
58. The tangible, non-transitory computer-readable medium of claim 52, wherein the computing instructions when executed by the one or more processors, further cause the one or more processors to: validate the one or more predicted protein sequences as output by the Al-based protein sequence model with the protein sequence validation dataset.
59. The tangible, non-transitory computer-readable medium of any one of claims 51-58, wherein creation of the protein sequence validation dataset allocates a larger protein sequence training dataset than a randomly sampled validation dataset randomly sampled from the plurality of reference proteomes of the reference proteome dataset, and wherein the protein sequence validation dataset and the randomly sampled validation dataset have at least one of: (1) a same sequence identity threshold value; or (2) an equal number of protein sequences selected from the plurality of reference proteomes of the reference proteome dataset.
60. The tangible, non-transitory computer-readable medium of any one of claims 51-59, wherein the computing instructions when executed by the one or more processors, further cause the one or more processors to: remove, based on a sequence identity threshold value, one or more of the protein sequences of the protein sequence validation dataset to create an ablated protein sequence validation dataset.
61. The tangible, non-transitory computer-readable medium of claim 60, wherein the removed one or more of the protein sequences comprises no more than 5% to 10% of the protein sequence validation dataset.
62. The tangible, non-transitory computer-readable medium of claim 52, wherein the computing instructions, when executed by the one or more processors, further cause the one or more processors to: update the protein sequence validation dataset to comprise protein sequences each having a maximum protein length that the Al-based protein sequence model was trained on.
63. The tangible, non-transitory computer-readable medium of claim 52, wherein the reference proteome dataset comprises proteins having different protein lengths, and wherein the computing instructions, when executed by the one or more processors, further cause the one or more processors to: truncate one or more of protein lengths of at least a subset of the proteins to have a truncated length, the truncated length comprising a maximum content length for training the Al-based protein sequence model, and train the Al-based protein sequence model with training data comprising at least a portion of the reference proteome dataset having the subset of the proteins having the truncated length.
64. The tangible, non-transitory computer-readable medium of any one of claims 51-63, wherein the completeness threshold comprises a value of 2 / 3 completeness or greater.
65. The tangible, non-transitory computer-readable medium of any one of claims 51-64, wherein 5% to 15% of the set of protein sequence clusters are randomly selected for creation of the protein sequence validation dataset.
66. The tangible, non-transitory computer-readable medium of any one of claims 51-65, wherein 15% or more of the set of protein sequence clusters are randomly selected for creation of the protein sequence validation dataset.
67. The tangible, non-transitory computer-readable medium of claim 52, wherein the computing instructions, when executed by the one or more processors, further cause the one or more processors to:generate validation metrics based on comparison between: (a) a set of predetermined values of the protein sequence validation dataset; and (b) the one or more predicted protein sequences as output by the Al-based protein sequence model .
68. The tangible, non-transitory computer-readable medium of claim 67, wherein the validation metrics comprise at least one improved metric selected from: a perplexity metric, a loss metric, and / or an accuracy metric, when analyzed against validation metrics of a previously trained version of the Al-based protein sequence model.
69. The tangible, non-transitory computer-readable medium of claim 68, wherein the at least one improved metric comprises: a reduced perplexity value, a reduced loss value, and / or an increased accuracy value.
70. The tangible, non-transitory computer-readable medium of claim 68, wherein the at least one improved metric comprises: a reduced perplexity value when the Al-based protein sequence model is trained on a UniRefl 00 dataset, a reduced loss value when the AI- based protein sequence model is trained on the UniRefl 00 dataset, and / or an increased accuracy value when the Al-based protein sequence model is trained on the UniRefl 00 dataset.
71. The tangible, non-transitory computer-readable medium of claim 70, wherein the at least one improved metric comprises: a reduced perplexity value when the Al-based protein sequence model is trained on a UniRefl 00 dataset compared to a second Al-based protein sequence model trained on a UniRef50 dataset, a reduced loss value when the AI- based protein sequence model is trained on the UniRefl 00 dataset compared to a second AI- based protein sequence model trained on the UniRef50 dataset, and / or an increased accuracy value when the Al-based protein sequence model is trained on the UniRefl 00 dataset compared to a second Al-based protein sequence model trained on the UniRef50 dataset.
72. The tangible, non-transitory computer-readable medium of claim 52, wherein the one or more predicted protein sequences as output by the Al-based protein sequence model is used to develop a protein-related product.
73. The tangible, non-transitory computer-readable medium of claim 72, wherein the protein-related product is predicted to treat a disease.
74. The tangible, non-transitory computer-readable medium of claim 72, wherein the protein-related product is predicted to provide a prophylaxis for a disease or medical condition.
75. The tangible, non-transitory computer-readable medium of claim 72, wherein the output of the one or more predicted protein sequences is used to monitor or measure the protein-related product during manufacture of the protein-related product.
76. An artificial intelligence (Al) training dataset generation system configured to generate Al training datasets for training Al-based protein sequence models, the Al training dataset generation system comprising: one or more processors; a computer memory communicatively coupled to the one or more processors; a reference proteome dataset stored on the computer memory and comprising a plurality of reference proteomes; and computing instructions stored on the computer memory, and that when executed by the one or more processors, cause the one or more processors to: filter the reference proteome dataset to generate a high completeness proteome data subset comprising reference proteomes selected from the plurality of reference proteomes having a completeness value above a completeness threshold; filter the high completeness proteome data subset to generate a protein confidence-based proteome data subset comprising reference proteomes selected from the high completeness proteome data subset having a protein confidence score above a protein confidence threshold, wherein the protein confidence-based proteome data subset comprises a plurality of protein sequences, cluster the plurality of protein sequences to generate a set of protein sequence clusters, each cluster of the plurality of protein sequences grouping one or more of the plurality of protein sequences of the protein confidence-based proteome data subset,randomly allocate a training cluster subset selected from the set of protein sequence clusters, create a protein sequence training dataset by selecting the protein sequences of the training cluster subset, train an Al-based protein sequence model with the protein sequence training dataset, and output, from the Al-based protein sequence model, one or more predicted protein sequences.
77. The Al training dataset generation system of claim 76, wherein the computing instructions, when executed by the one or more processors, further cause the one or more processors to: generate validation metrics based on comparison between: (a) a set of predetermined values of a protein sequence validation dataset comprising data excluded from the protein sequence training dataset; and (b) a set of predicted values as output by the Al-based protein sequence model.
78. An artificial intelligence (Al) training dataset generation method for generating Al training datasets for training Al-based protein sequence models, the Al training dataset generation method comprising: filtering, by one or more processors, a reference proteome dataset to generate a high completeness proteome data subset comprising reference proteomes selected from a plurality of reference proteomes having a completeness value above a completeness threshold, wherein the reference proteome dataset is stored on a computer memory communicatively coupled to the one or more processors, and wherein the reference proteome dataset comprises the plurality of reference proteomes; filtering, by the one or more processors, the high completeness proteome data subset to generate a protein confidence-based proteome data subset comprising reference proteomes selected from the high completeness proteome data subset having a protein confidence score above a protein confidence threshold, wherein the protein confidence-based proteome data subset comprises a plurality of protein sequences;clustering, by the one or more processors, the plurality of protein sequences to generate a set of protein sequence clusters, each cluster of the plurality of protein sequences grouping one or more of the plurality of protein sequences of the protein confidence-based proteome data subset; randomly allocating, by the one or more processors, a training cluster subset selected from the set of protein sequence clusters; creating, by the one or more processors, a protein sequence training dataset by selecting the protein sequences of the training cluster subset; training an Al-based protein sequence model with the protein sequence training dataset; and outputting, from the Al-based protein sequence model, one or more predicted protein sequences.
79. The Al training dataset generation method of claim 78 further comprising: generating validation metrics based on comparison between: (a) a set of predetermined values of a protein sequence validation dataset comprising data excluded from the protein sequence training dataset; and (b) a set of predicted values as output by the Al-based protein sequence model.
80. A tangible, non-transitory computer-readable medium storing computing instructions for generating artificial intelligence (Al) training datasets for training Al-based protein sequence models, that when executed by one or more processors cause the one or more processors to: filter a reference proteome dataset comprising a plurality of reference proteomes to generate a high completeness proteome data subset comprising reference proteomes selected from the plurality of reference proteomes having a completeness value above a completeness threshold; filter the high completeness proteome data subset to generate a protein confidencebased proteome data subset comprising reference proteomes selected from the high completeness proteome data subset having a protein confidence score above a protein confidence threshold,wherein the protein confidence-based proteome data subset comprises a plurality of protein sequences, cluster the plurality of protein sequences to generate a set of protein sequence clusters, each cluster of the plurality of protein sequences grouping one or more of the plurality of protein sequences of the protein confidence-based proteome data subset, randomly allocate a training cluster subset selected from the set of protein sequence clusters, create a protein sequence training dataset by selecting the protein sequences of the training cluster subset, train an Al-based protein sequence model with the protein sequence training dataset, and output, from the Al-based protein sequence model, one or more predicted protein sequences.
81. The tangible, non-transitory computer-readable medium of claim 80, wherein the computing instructions, when executed by the one or more processors, further cause the one or more processors to: generate validation metrics based on comparison between: (a) a set of predetermined values of a protein sequence validation dataset comprising data excluded from the protein sequence training dataset; and (b) a set of predicted values as output by the Al-based protein sequence model.
82. An artificial intelligence (Al) model generation system for generating and validating Al-based protein sequence models, the Al model generation system comprising: one or more processors; a computer memory communicatively coupled to the one or more processors; a reference proteome dataset stored on the computer memory and comprising a plurality of reference proteomes; and computing instructions stored on the computer memory, and that when executed by the one or more processors, cause the one or more processors to: train an Al-based protein sequence model with a protein sequence training dataset,wherein the protein sequence training dataset comprises a dataset independent of a protein sequence validation dataset used to validate the Al-based protein sequence model, wherein the Al-based protein sequence model is configured to output one or more predicted protein sequences, and wherein creation of the protein sequence validation dataset allocates the protein sequence training dataset as having a larger dataset size than a randomly sampled validation dataset randomly sampled from the plurality of reference proteomes of the reference proteome dataset, and wherein the protein sequence validation dataset and the randomly sampled validation dataset have at least one of: (1) a same sequence identity threshold value; or (2) an equal number of protein sequences selected from the plurality of reference proteomes of the reference proteome dataset.
83. An artificial intelligence (Al) model generation method for generating and validating Al-based protein sequence models, the Al model generation method comprising: training an Al-based protein sequence model with a protein sequence training dataset, wherein the protein sequence training dataset comprises a dataset independent of a protein sequence validation dataset used to validate the Al-based protein sequence model, wherein the Al-based protein sequence model is configured to output one or more predicted protein sequences, and wherein creation of the protein sequence validation dataset allocates the protein sequence training dataset as having a larger dataset size than a randomly sampled validation dataset randomly sampled from a plurality of reference proteomes of a reference proteome dataset, and wherein the protein sequence validation dataset and the randomly sampled validation dataset have at least one of: (1) a same sequence identity threshold value; or (2) an equal number of protein sequences selected from the plurality of reference proteomes of the reference proteome dataset.
84. A tangible, non-transitory computer-readable medium storing instructions for generating and validating Al-based protein sequence models, that when executed by one or more processors cause the one or more processors to:train an Al-based protein sequence model with a protein sequence training dataset, wherein the protein sequence training dataset comprises a dataset independent of a protein sequence validation dataset used to validate the Al-based protein sequence model, wherein the Al-based protein sequence model is configured to output one or more predicted protein sequences, and wherein creation of the protein sequence validation dataset allocates the protein sequence training dataset as having a larger dataset size than a randomly sampled validation dataset randomly sampled from a plurality of reference proteomes of a reference proteome dataset, and wherein the protein sequence validation dataset and the randomly sampled validation dataset have at least one of: (1) a same sequence identity threshold value; or (2) an equal number of protein sequences selected from the plurality of reference proteomes of the reference proteome dataset.
85. A tangible, non-transitory computer-readable medium storing a protein sequence as output by an Al-based protein sequence model, the non-transitory computer- readable medium further storing instructions for validating the Al-based protein sequence model, that when executed by one or more processors cause the one or more processors to: filter, by one or more processors, a reference proteome dataset to generate a high completeness proteome data subset comprising reference proteomes selected from a plurality of reference proteomes having a completeness value above a completeness threshold, wherein the reference proteome dataset is stored on a computer memory communicatively coupled to the one or more processors, and wherein the reference proteome dataset comprises the plurality of reference proteomes; filter, by the one or more processors, the high completeness proteome data subset to generate a protein confidence-based proteome data subset comprising reference proteomes selected from the high completeness proteome data subset having a protein confidence score above a protein confidence threshold, wherein the protein confidence-based proteome data subset comprises a plurality of protein sequences; cluster, by the one or more processors, the plurality of protein sequences to generate a set of protein sequence clusters, each cluster of the plurality of protein sequences grouping one or more of the plurality of protein sequences of the protein confidence-based proteome data subset;randomly allocate, by the one or more processors, a validation cluster subset selected from the set of protein sequence clusters; create, by the one or more processors, a protein sequence validation dataset by selecting the protein sequences of the validation cluster subset; and validate the Al-based protein sequence model with the protein sequence validation dataset.
86. An output protein sequence generated by an Al-based protein sequence model, wherein the Al-based protein sequence model is validated by an Al validation dataset generated by a method comprising: filtering, by one or more processors, a reference proteome dataset to generate a high completeness proteome data subset comprising reference proteomes selected from a plurality of reference proteomes having a completeness value above a completeness threshold, wherein the reference proteome dataset is stored on a computer memory communicatively coupled to the one or more processors, and wherein the reference proteome dataset comprises the plurality of reference proteomes; filtering, by the one or more processors, the high completeness proteome data subset to generate a protein confidence-based proteome data subset comprising reference proteomes selected from the high completeness proteome data subset having a protein confidence score above a protein confidence threshold, wherein the protein confidence-based proteome data subset comprises a plurality of protein sequences; clustering, by the one or more processors, the plurality of protein sequences to generate a set of protein sequence clusters, each cluster of the plurality of protein sequences grouping one or more of the plurality of protein sequences of the protein confidence-based proteome data subset; randomly allocating, by the one or more processors, a validation cluster subset selected from the set of protein sequence clusters; creating, by the one or more processors, a protein sequence validation dataset by selecting the protein sequences of the validation cluster subset; and validating the Al-based protein sequence model with the protein sequence validation dataset.
87. A method for evaluating generalization capability of one or more Al-based protein sequence models, the method comprising: generating, by one or more processors, a validation dataset comprising protein sequences obtained by: filtering a reference proteome dataset to generate a high completeness proteome data subset comprising reference proteomes selected from a plurality of reference proteomes having a completeness value above a completeness threshold, wherein the reference proteome dataset comprises a plurality of reference proteomes; filtering the high completeness proteome data subset to generate a protein confidence-based proteome data subset comprising reference proteomes selected from the high completeness proteome data subset having a protein confidence score above a protein confidence threshold, wherein the protein confidence-based proteome data subset comprises a plurality of protein sequences; clustering the plurality of protein sequences of the protein confidence-based proteome data subset to generate a set of protein sequence clusters, each cluster of the plurality of protein sequences grouping one or more of the plurality of protein sequences of the protein confidence-based proteome data subset; and creating the protein sequence validation dataset by selecting the protein sequences from one or more clusters in the set of protein sequence clusters; for each of two or more subsets of a training dataset, removing the subset from the training dataset to obtain a respective ablated training dataset, wherein each respective ablated training dataset comprises at least one or more protein sequences excluded from the validation dataset, and wherein each respective ablated training dataset comprises a unique set of protein sequences; implementing, for each respective ablated training dataset, an iterative model routine to evaluate an Al-based protein sequence model trained on the training dataset, the iterative model routine comprising: training an ablated Al-based protein sequence model on the ablated training dataset, generating an ablated performance score defining a predictive performance of an output of the ablated Al-based protein sequence model, andstoring in memory the ablated performance score as part of a plurality of ablated performance scores; and determining, based on the plurality of ablated performance scores, a model generalization output for the Al-based protein sequence model.
88. The method of claim 87, wherein the iterative model routine further comprises determining, using the validation dataset as input to the ablated Al-based protein sequence model, an output of the ablated Al-based protein sequence model, and wherein generating the ablated performance score further comprises comparing the output of the ablated Al-based protein sequence model to a ground truth value.
89. The method of any one of claims 87 or 88, wherein the iterative model routine further comprises: determining, using the validation dataset as input to the ablated Al-based protein sequence model, an output of the ablated Al-based protein sequence model; and determining, using the validation dataset as input to the Al-based protein sequence model, an output of the Al-based protein sequence model; and wherein generating the ablated performance score further comprises comparing the output of the ablated Al-based protein sequence model to the output of the Al-based protein sequence model.
90. The method of any one of claims 87-89, further comprising determining two or more sequence similarity thresholds by comparing protein sequences in the training dataset to protein sequences in the validation dataset, and wherein each of the two or more subsets is associated with a respective sequence similarity threshold selected from among the two or more sequence similarity thresholds, and removing the subset from the training dataset to obtain the respective ablated training dataset further comprises applying the respective sequence similarity threshold as a filter to the training dataset to remove protein sequences having a sequence similarity above the respective sequence similarity threshold.
91. The method of any one of claims 87-90 further comprising outputting, from the Al-based protein sequence model, one or more predicted protein sequences.
92. The method of any one of claims 87-91, further comprising outputting, from the Al-based protein sequence model, information specifying a protein sequence and manufacturing a protein-related product having the protein sequence.
93. The method of any one of claims 87-92, wherein the model generalization output defines one or more validation metrics comprising at least one of: a perplexity metric, a loss metric, and / or an accuracy metric.
94. The method of any one of claims 87-93, wherein the model generalization output defines one or more improved validation metrics when compared to validation metrics of a different Al-based protein sequence model trained on at least a portion of the reference proteome dataset, the one or more improved validation metrics comprising at least one of: an improved perplexity metric, an improved loss metric, and / or an improved accuracy metric.
95. The method of any one of claims 87-94 further comprising: generating each respective ablated training dataset by ablating no more than 5% to 10% of the training dataset.
96. The method of any one of claims 87-95, wherein creating the protein sequence validation dataset further comprises selecting protein sequences in a cluster of the set of protein sequence clusters, and the model generalization output defines one or more validation metrics indicating performance of the Al-based protein sequence model for the selected protein sequences in the cluster.
97. The method of claim 96, wherein the selected protein sequences in the cluster correspond to a protein class or protein family.
98. The method of claim 87-97, wherein generating the validation dataset further comprises allocating a validation cluster subset selected from the set of protein sequence clusters, and creating the protein sequence validation dataset further comprises selecting protein sequences from the validation cluster subset as the protein sequence validation dataset.
99. The method of any one of claims 87-98, wherein generating the validation dataset further comprises randomly allocating a validation cluster subset selected from the set of protein sequence clusters, and creating the protein sequence validation dataset further comprises selecting the protein sequences of the validation cluster subset.
100. A system comprising: at least one processor; and at least one non-transitory computer-readable storage medium storing computing instructions that, when executed by the at least one processor, causes the at least one processor to perform any of claims 87-99.
101. At least one non-transitory computer-readable storage medium storing computing instructions that, when executed by the at least one processor, causes the at least one processor to perform any of claims 87-100.
Citation Information
Patent Citations
Method for generating functional protein sequences with generative adversarial networks
US20220367007A1