Single-cell RNA-seq Drug Response Prediction via Bulk Data Integration

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current computational tools for predicting drug response at the single-cell level face challenges due to differences between bulk and single-cell RNA-seq data, low detection rates, and stochastic drop-outs, leading to unreliable predictions.

Innovation Solution

The method integrates single-cell RNA-seq data with bulk RNA-seq data from cancer cell lines using canonical correlation analysis to capture shared gene expression patterns, allowing for the prediction of cellular drug response scores.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If bulk RNA-seq data is used to predict drug response, then average expression across cells is obtained, but cell type and composition information is obscured

Engineering Contradiction:
Improveaverage expression measurementVSAvoidcell type and composition information
Core Design Contradiction:
Quantity of substanceVSLoss of information

Solution Approach 1:

The patent segments bulk RNA-seq data into cell-type-specific expression profiles by integrating with scRNA-seq data. The method decomposes the bulk expression matrix into contributions from different cell types using reference scRNA-seq data, thereby recovering cell-type-specific information that was obscured in the bulk measurement.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent uses scRNA-seq data as an intermediary to transfer drug response information to bulk RNA-seq data. By establishing a mapping between bulk and single-cell expression spaces through canonical correlation analysis, the method enables bulk data to predict drug response at single-cell resolution without directly measuring single-cell drug responses.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If single-cell RNA-seq data is used for drug response prediction, then cell-level gene expression patterns are captured, but low detection rates and stochastic drop-outs reduce prediction reliability

Engineering Contradiction:
Improvecell-level gene expression measurementVSAvoidprediction reliability
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent merges bulk RNA-seq data with scRNA-seq data to create a hybrid prediction system. Bulk RNA-seq data provides robust average expression measurements that can compensate for drop-outs in scRNA-seq data, while scRNA-seq data provides cell-type-specific information. The integration through canonical correlation analysis combines the strengths of both data types to improve prediction reliability.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The method uses feedback from bulk RNA-seq data to correct and refine predictions made from scRNA-seq data. The bulk data serves as a reference that provides stable signal for genes with low detection rates in single-cell data, allowing the model to adjust predictions based on the more reliable bulk measurements.

Inventive Principle:
Principle #23Feedback

3Measurement precision

If computational tools integrate bulk and single-cell RNA-seq data, then drug response prediction accuracy is improved, but data integration complexity increases

Engineering Contradiction:
Improvedrug response prediction accuracyVSAvoiddata integration complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent transforms the data integration problem by changing the parameter space through canonical correlation analysis. Instead of directly integrating expression values from bulk and scRNA-seq data, the method finds linear combinations of genes (canonical variables) that maximize correlation between the two data types, thereby simplifying the integration process while preserving predictive information.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20250066840A1Single-cell RNA sequencing analysis pipeline for predicting cell sensitivities to various drugs
Publication Date: 2025.02.27 REGENTS OF THE UNIVERSITY OF MINNESOTA
  • US20250066840A1 patent drawing
  • US20250066840A1 patent drawing
  • US20250066840A1 patent drawing

AI summary

Methods for facilitating drug discovery in various models and proposing single cell-type-specific drug candidates, and predicting cellular drug sensitivities. Predictions comprise obtaining bulk RNA sequence data, wherein bulk RNA seq data is pan-cancer cell line transcriptomic data from one or more sources; obtaining single cell RNA sequence data (scRNA-seq) from one or more sources; integrating the bulk RNA sequence data and scRNA-seq data; capturing similarities between the bulk data and the scRNA-seq data based on a shared drug response relative gene and using canonical correlation analysis; extracting one or more drug-response relevant features and selecting a pharmacogenomic subspace for accurately predicting a single-cell drug response.