Clonotype Analysis via Consensus Contig Assembly

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for analyzing datasets from mRNA sequencing of single cells, particularly for T-cells and B-cells, face challenges in effectively utilizing the unique nucleotide sequences of CDR3 regions for clonotype identification and comparison, due to the imprecision of somatic rearrangement processes and the complexity of clonal grouping.

Innovation Solution

A system and method for analyzing datasets representing clonotypes, which includes obtaining and visualizing data on clonotypes from multiple cells, determining the frequency and proportion of each clonotype, and providing interactive visualizations and filtering options to compare clonotypes across different datasets, using barcodes and consensus sequences to identify and align contigs, and applying metrics like Morisita-Horn for clonotype commonality analysis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If somatic rearrangement process is used to generate CDR3 sequences, then unique nucleotide sequences are created for clonotype identification, but imprecision in the joining process creates variability that complicates clonotype grouping and analysis

Engineering Contradiction:
Improveclonotype identification accuracyVSAvoidclonotype grouping complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies parameter changes by transforming the analysis from raw nucleotide sequence level to amino acid sequence level, and further to clonotype level aggregation. This transformation changes the granularity and representation of the data, allowing imprecise nucleotide variations to be normalized while preserving biologically meaningful clonotype information through metrics like Morisita-Horn index for comparing clonotype distributions

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces computational intermediaries including contig assembly algorithms, consensus sequence generation, and clonotype grouping metrics that act as mediators between the raw sequencing data and the final clonotype analysis. These intermediaries process and standardize the imprecise rearrangement data into comparable clonotype representations

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If high throughput sequencing is used to obtain data from hundreds or thousands of cells, then large amounts of clonotype data are generated, but the complexity of analyzing and interpreting this data increases significantly

Engineering Contradiction:
Improvedata generation throughputVSAvoiddata analysis complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent applies segmentation by dividing the large-scale sequencing data into manageable units: individual cell barcodes, contigs per cell, consensus sequences, and finally clonotypes. This hierarchical segmentation allows the analysis pipeline to process high-throughput data in discrete, computationally tractable steps rather than as one overwhelming dataset

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent uses computational copying and consensus building where multiple sequencing reads from the same cell (identified by barcode) are copied and assembled into contigs, then into consensus sequences. This copying process creates standardized representations that simplify downstream analysis of the high-throughput data

Inventive Principle:
Principle #26Copying

3Reliability

If detailed contig sequences with barcodes are used for each cell, then accurate clonotype assignment is achieved, but the data structure becomes more complex requiring advanced computational methods

Engineering Contradiction:
Improveclonotype assignment accuracyVSAvoiddata structure complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies merging by combining multiple sequencing reads with the same barcode into a single consensus contig for each cell. This merging process consolidates the complex raw data structure into simplified consensus representations that maintain clonotype assignment accuracy while reducing structural complexity for analysis

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS20240218445A1Methods for clonotype screening
Publication Date: 2024.07.04 10X GENOMICS INC
  • US20240218445A1 patent drawing
  • US20240218445A1 patent drawing
  • US20240218445A1 patent drawing

AI summary

Methods for screening clonotypes are provided. Data representing a plurality of cells from a single subject is obtained. The data represents a plurality of clonotypes. The data includes a plurality of contigs for each respective clonotype in the plurality of clonotypes. Each respective contig in the plurality of contigs comprises (i) an indication of chain type for the respective contig and (ii) a contig sequence of an mRNA of the respective cell. There is determined, using the data, for each respective clonotype in the plurality of clonotypes, a number of the plurality of cells that represent the respective clonotype. In some instances, more than one cell in the plurality of cells have the same clonotype in the plurality of clonotypes. In some instances, the plurality of clonotypes comprises 25 clonotypes and where the plurality of cells includes at least one cell for each clonotype in the plurality of clonotypes.