Pharmacogenomic Protein Function Mapping With Contrastive Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies struggle to effectively link genetic variants, such as SNPs, with protein function to enhance drug discovery and personalized medicine, particularly in understanding disease resistance and drug resistance, and there is a lack of efficient methods for identifying novel therapeutic targets and drug repurposing.

Innovation Solution

A contrastive learning-based framework, PGxProt, is used to train a model that associates pharmacogenomic variants with protein sequences, enabling the discovery of multi-associations and generating embeddings for downstream analysis, including drug repurposing and personalized treatment.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional methods are used to link genetic variants with protein function, then the process is simple and straightforward, but the effectiveness and precision of linking are insufficient

Engineering Contradiction:
Improvelinking precision between variants and protein functionVSAvoidmethod complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent replaces traditional mechanical/biochemical methods of linking variants with protein function using a contrastive learning model that processes data through transformer architectures. The model substitutes manual or traditional computational approaches with an AI-based system that automatically learns associations between variants and protein functions from multi-omics data, significantly improving linking precision.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the parameter space by transforming biological data (variants, protein sequences, expressions) into embedded representations through transformer models. This parameter transformation enables the system to capture complex relationships between variants and protein functions that are not apparent in raw data, improving measurement precision without requiring overly complex biological experimentation.

Inventive Principle:
Principle #35Parameter changes

2Loss of information

If comprehensive multi-omics data is integrated to improve drug discovery, then the depth of understanding is enhanced, but the computational complexity and processing requirements increase

Engineering Contradiction:
Improveinformation completeness in drug discoveryVSAvoidcomputational system complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent segments the comprehensive multi-omics data into distinct components (genomic variants, protein sequences, protein expressions, drug information) and processes each through specialized transformer models. This segmentation allows the system to handle complex data systematically, maintaining information completeness while managing computational complexity through modular processing architecture.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The contrastive learning model acts as an intermediary that integrates multi-omics data without requiring direct complex interactions between all data types. The model processes different data modalities through separate transformers and combines them through contrastive learning objectives, reducing computational complexity while preserving comprehensive information for drug discovery.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If contrastive learning model is trained to associate variants with protein sequences, then the accuracy of association is improved, but the training time and computational resources increase

Engineering Contradiction:
Improveassociation accuracy between variants and proteinsVSAvoidmodel training time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary actions by pre-processing and embedding variant and protein sequence data using transformer models before the main contrastive learning training. This pre-processing creates ready-to-use embeddings that reduce the computational burden during actual training, improving association accuracy while minimizing training time through efficient data preparation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent changes parameters during training by using contrastive learning objectives that optimize embeddings in a transformed parameter space. This approach improves association accuracy by learning meaningful relationships between variants and proteins through contrastive optimization, while the parameter transformation enables more efficient training compared to traditional methods.

Inventive Principle:
Principle #35Parameter changes

4Adaptability or versatility

If existing databases and literature are used for drug target identification, then the available information is sufficient, but novel therapeutic targets cannot be effectively discovered

Engineering Contradiction:
Improvecapacity to identify novel therapeutic targetsVSAvoidutilization of available multi-omics data
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The patent creates a universal contrastive learning model that handles multiple data types (genomic variants, protein sequences, expressions, drug information) and performs multiple functions (association learning, embedding generation, target identification, repurposing). This multi-functional system maximizes the utilization of available multi-omics data to identify both known and novel therapeutic targets, enhancing adaptability beyond what specialized databases can provide.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20260004892A1Pharmacogenomics induced protein function of therapeutic targets
Publication Date: 2026.01.01 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US20260004892A1 patent drawing
  • US20260004892A1 patent drawing
  • US20260004892A1 patent drawing

AI summary

A set of candidate drugs is selected based on one or more outcomes related to one or more diseases and variants linked with the selected set of candidate drugs are obtained. One or more protein sequences related to the selected set of candidate drugs are collated. Pairs of variants and protein sequences are generated for each drug of the set of candidate drugs. A contrastive learning model is trained using the generated pairs of variants and protein sequences and a downstream task is performed using the contrastive learning model.