Protein Formulation Property Prediction With Two-Stage Machine Learning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The development of protein-based pharmaceuticals is hindered by the need for extensive empirical testing to optimize formulations for viscosity, stability, and manufacturability, which is time-consuming and resource-intensive, and there is a challenge in predicting protein formulation properties accurately.
Innovation Solution
A two-stage machine learning approach is employed to classify formulation descriptors into groups and then apply specific models to predict properties like viscosity, using trained regression models to enhance accuracy and speed up the formulation optimization process.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If extensive empirical testing is conducted to optimize formulations for viscosity, stability, and manufacturability, then formulation accuracy and reliability are improved, but development time and resource consumption increase significantly
Solution Approach 1:
The patent applies preliminary action by training machine learning models on historical formulation data before actual formulation development. The models are pre-trained to predict viscosity, stability, and other critical properties, allowing developers to screen candidate formulations computationally before conducting empirical testing. This preliminary computational screening reduces the number of formulations that require extensive empirical testing, thereby reducing development time while maintaining optimization accuracy.
Solution Approach 2:
The patent uses copying by creating virtual replicas of physical formulations through computational models. Instead of testing every formulation variant physically, the system creates digital twins of formulations using machine learning predictions. These virtual copies allow for rapid screening and comparison of multiple formulation options, reducing the need for time-consuming physical testing while maintaining the ability to identify optimal formulations.
2Reliability
If extensive empirical testing is conducted to optimize formulations, then formulation reliability is improved, but resource consumption increases significantly
Solution Approach 1:
The patent applies preliminary action by training machine learning models on historical formulation data before actual formulation development. The models are pre-trained to predict viscosity, stability, and other critical properties, allowing developers to screen candidate formulations computationally before conducting empirical testing. This preliminary computational screening reduces the number of formulations that require extensive empirical testing, thereby reducing resource consumption while maintaining optimization accuracy.
Solution Approach 2:
The patent uses copying by creating virtual replicas of physical formulations through computational models. Instead of testing every formulation variant physically, the system creates digital twins of formulations using machine learning predictions. These virtual copies allow for rapid screening and comparison of multiple formulation options, reducing the need for resource-intensive physical testing while maintaining the ability to identify optimal formulations.
3Quantity of substance
If protein concentration is increased to enable subcutaneous administration and reduce dosing frequency, then therapeutic efficacy is improved, but viscosity increases and device compatibility becomes more difficult
Solution Approach 1:
The patent applies parameter changes by using machine learning models to predict how different formulation parameters (excipient types and concentrations, pH, buffer composition) affect viscosity at high protein concentrations. The system optimizes these parameters to achieve the desired high protein concentration for subcutaneous administration while maintaining viscosity within acceptable limits for device compatibility. The models identify specific parameter combinations that balance concentration requirements with rheological constraints.
4Quantity of substance
If protein concentration is increased to enable subcutaneous administration, then therapeutic efficacy is improved, but formulation stability becomes more challenging to maintain
Solution Approach 1:
The patent applies parameter changes by using machine learning models to predict how different formulation parameters (excipient types and concentrations, pH, buffer composition) affect stability at high protein concentrations. The system optimizes these parameters to achieve the desired high protein concentration for subcutaneous administration while maintaining formulation stability. The models identify specific parameter combinations that balance concentration requirements with stability constraints, predicting potential degradation pathways and recommending stabilizing formulations.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
In a method for predicting a property of potential protein formulations, a set of formulation descriptors is classified as belonging to a specific one of a plurality of predetermined groups that each correspond to a different value range for a protein formulation property. Classifying the set of descriptors includes applying at least a first portion of the set of descriptors as inputs to a first machine learning model. The method also includes selecting, based on the classification, a second machine learning model from among multiple models corresponding to different groups. The method also includes predicting a value of the protein formulation property that corresponds to the set of descriptors, by applying at least a second portion of the set of formulation descriptors as inputs to the selected model. The method further includes causing the value of the protein formulation property to be displayed to a user and/or stored in a memory.