Protein Formulation Property Prediction With Two-Stage Machine Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The development of protein-based pharmaceuticals is hindered by the need for extensive empirical testing to optimize formulations for viscosity, stability, and manufacturability, which is time-consuming and resource-intensive, and there is a challenge in predicting protein formulation properties accurately.

Innovation Solution

A two-stage machine learning approach is employed to classify formulation descriptors into groups and then apply specific models to predict properties like viscosity, using trained regression models to enhance accuracy and speed up the formulation optimization process.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If extensive empirical testing is conducted to optimize formulations for viscosity, stability, and manufacturability, then formulation accuracy and reliability are improved, but development time and resource consumption increase significantly

Engineering Contradiction:
Improveformulation optimization accuracyVSAvoiddevelopment time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies preliminary action by training machine learning models on historical formulation data before actual formulation development. The models are pre-trained to predict viscosity, stability, and other critical properties, allowing developers to screen candidate formulations computationally before conducting empirical testing. This preliminary computational screening reduces the number of formulations that require extensive empirical testing, thereby reducing development time while maintaining optimization accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses copying by creating virtual replicas of physical formulations through computational models. Instead of testing every formulation variant physically, the system creates digital twins of formulations using machine learning predictions. These virtual copies allow for rapid screening and comparison of multiple formulation options, reducing the need for time-consuming physical testing while maintaining the ability to identify optimal formulations.

Inventive Principle:
Principle #26Copying

2Reliability

If extensive empirical testing is conducted to optimize formulations, then formulation reliability is improved, but resource consumption increases significantly

Engineering Contradiction:
Improveformulation optimization accuracyVSAvoidresource consumption
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The patent applies preliminary action by training machine learning models on historical formulation data before actual formulation development. The models are pre-trained to predict viscosity, stability, and other critical properties, allowing developers to screen candidate formulations computationally before conducting empirical testing. This preliminary computational screening reduces the number of formulations that require extensive empirical testing, thereby reducing resource consumption while maintaining optimization accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses copying by creating virtual replicas of physical formulations through computational models. Instead of testing every formulation variant physically, the system creates digital twins of formulations using machine learning predictions. These virtual copies allow for rapid screening and comparison of multiple formulation options, reducing the need for resource-intensive physical testing while maintaining the ability to identify optimal formulations.

Inventive Principle:
Principle #26Copying

3Quantity of substance

If protein concentration is increased to enable subcutaneous administration and reduce dosing frequency, then therapeutic efficacy is improved, but viscosity increases and device compatibility becomes more difficult

Engineering Contradiction:
Improveprotein concentrationVSAvoiddevice compatibility
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent applies parameter changes by using machine learning models to predict how different formulation parameters (excipient types and concentrations, pH, buffer composition) affect viscosity at high protein concentrations. The system optimizes these parameters to achieve the desired high protein concentration for subcutaneous administration while maintaining viscosity within acceptable limits for device compatibility. The models identify specific parameter combinations that balance concentration requirements with rheological constraints.

Inventive Principle:
Principle #35Parameter changes

4Quantity of substance

If protein concentration is increased to enable subcutaneous administration, then therapeutic efficacy is improved, but formulation stability becomes more challenging to maintain

Engineering Contradiction:
Improveprotein concentrationVSAvoidformulation stability
Core Design Contradiction:
Quantity of substanceVSStability of the object's composition

Solution Approach 1:

The patent applies parameter changes by using machine learning models to predict how different formulation parameters (excipient types and concentrations, pH, buffer composition) affect stability at high protein concentrations. The system optimizes these parameters to achieve the desired high protein concentration for subcutaneous administration while maintaining formulation stability. The models identify specific parameter combinations that balance concentration requirements with stability constraints, predicting potential degradation pathways and recommending stabilizing formulations.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentEP4022622B1Systems and methods for prediction of protein formulation properties
Publication Date: 2025.10.01 AMGEN INC
  • EP4022622B1 patent drawingFigure 1
  • EP4022622B1 patent drawingFigure 2
  • EP4022622B1 patent drawingFigure 3

AI summary

In a method for predicting a property of potential protein formulations, a set of formulation descriptors is classified as belonging to a specific one of a plurality of predetermined groups that each correspond to a different value range for a protein formulation property. Classifying the set of descriptors includes applying at least a first portion of the set of descriptors as inputs to a first machine learning model. The method also includes selecting, based on the classification, a second machine learning model from among multiple models corresponding to different groups. The method also includes predicting a value of the protein formulation property that corresponds to the set of descriptors, by applying at least a second portion of the set of formulation descriptors as inputs to the selected model. The method further includes causing the value of the protein formulation property to be displayed to a user and/or stored in a memory.