Pose-Sensitive vHTS Model for Polymer-Compound Interaction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional structure-based, virtual high throughput screening (vHTS) machine learning methods are pose insensitive, leading to inaccurate characterization of interactions between test compounds and target polymers, as they fail to distinguish between correct and incorrect poses of compounds and polymers.
Innovation Solution
Conditioning vHTS machine learning models to be pose sensitive by training them on both positive and negative poses of training compounds using an independent pose generation process, where positive and negative poses are defined based on interaction scores, allowing the models to differentiate between accurate and inaccurate poses.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If conventional vHTS machine learning methods represent compounds and polymers independently, then the model processing is simplified, but the model becomes pose insensitive and cannot distinguish between correct and incorrect poses
Solution Approach 1:
The patent combines the compound representation with pose information by integrating the compound's structural features with its spatial orientation data. This merging allows the model to simultaneously consider both the chemical structure and the pose configuration, enabling pose-sensitive predictions without requiring completely separate processing pipelines for structure and orientation.
Solution Approach 2:
The patent introduces pose as an additional dimensional aspect to the compound representation. By incorporating spatial orientation parameters (rotation angles, translation vectors) alongside the molecular structure, the model transitions from considering only chemical composition to considering both chemical structure and spatial configuration, thereby achieving pose sensitivity.
2Ease of manufacture
If the model is trained only on positive poses, then training data preparation is simpler, but the model cannot differentiate between accurate and inaccurate poses
Solution Approach 1:
The patent applies preliminary action by pre-generating both positive and negative pose examples during the training data preparation phase. By anticipating the need for pose differentiation, the system proactively creates diverse pose samples with known quality labels before training begins, enabling the model to learn pose sensitivity without requiring complex on-the-fly pose evaluation during training.
Solution Approach 2:
The patent utilizes parameter changes by systematically varying pose parameters (rotation angles, translation distances) to generate both positive and negative examples. By controlling and adjusting these spatial parameters during data generation, the system creates a balanced training set that covers the full range of pose quality variations, enabling the model to learn discriminative features.
3Productivity
If conventional methods provide categorical activity labels without pose information, then the screening process is faster, but a significant percentage of compounds are incorrectly labeled
Solution Approach 1:
The patent introduces pose information as an intermediary element between the compound structure and the activity label. Rather than directly mapping compound structure to activity, the model first evaluates the pose quality as an intermediate step, using pose-sensitive features to mediate the prediction process. This intermediary role of pose information resolves the contradiction by providing discriminative power while maintaining computational efficiency.
Data Source
AI summary
Systems and methods for characterizing an interaction between a test compound and a polymer use coordinates for the polymer and a training dataset of compounds. Each compound has a positive pose with respect to target polymer coordinates with a positive interaction score and a negative pose of the compound with respect to the target polymer coordinates and a negative interaction score. The model is trained by applying, for each compound, at least: (i) a positive score for the positive pose as input to the model, against the positive interaction score of the compound, and (ii) a negative score for the negative pose as input to the model, against the negative interaction score of the compound, thereby adjusting parameters of the model. In turn, an output of the model is used, at least in part, to characterize the interaction between the test compound and the polymer.


