Autoencoder Tissue Image Analysis via Sparse Dictionary Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional tissue image analysis methods face challenges due to tissue heterogeneity, large image sizes, and imprecise labeling, leading to misclassification and low throughput in diagnosing and predicting medical conditions like cancer.

Innovation Solution

The use of modified autoencoders for representation learning and dimensionality reduction, combined with convolutional networks to extract features that capture morphology and color, and enforce sparsity for improved accuracy and robustness, allowing for the generation of a dictionary of representative atoms that quantify tissue heterogeneity without prior cell classification.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional machine learning approaches are used for tissue image analysis, then the analysis can be performed with existing methods, but the classification accuracy deteriorates due to tissue heterogeneity and imprecise labeling

Engineering Contradiction:
Improveclassification accuracyVSAvoidanalysis system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the tissue image into multiple patches and further divides each patch into subpatches for processing. This segmentation approach allows the system to handle tissue heterogeneity by analyzing local regions independently, improving classification accuracy while managing computational complexity through distributed processing

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transforms the analysis from direct image classification to a two-stage process: first extracting representative atoms through dictionary learning, then classifying based on atom compositions. This dimensional transformation from pixel space to atom space improves reliability by capturing essential tissue characteristics while reducing the impact of labeling imprecision

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Productivity

If manual IHC image analysis is performed to emphasize certain characteristics, then staining intensity can be evaluated, but the throughput remains low and labor intensive

Engineering Contradiction:
Improveanalysis throughputVSAvoidstaining intensity measurement accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system performs automated dictionary learning and patch classification without manual intervention. The autoencoder architecture automatically learns representative atoms and their compositions from the training data, enabling high-throughput analysis while maintaining measurement precision through consistent algorithmic application across all images

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual visual analysis with an automated computational system based on dictionary learning and sparse coding. This substitution eliminates labor-intensive manual evaluation while maintaining or improving measurement precision through objective, repeatable algorithmic processing of staining intensity and morphological features

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Loss of information

If the full image is analyzed to capture global features, then comprehensive tissue characterization is achieved, but the computational burden increases due to large image sizes

Engineering Contradiction:
Improveglobal feature preservationVSAvoidcomputational energy consumption
Core Design Contradiction:
Loss of informationVSUse of energy by moving object

Solution Approach 1:

The patent divides large tissue images into multiple smaller patches for independent processing. This segmentation reduces the computational energy required for each processing unit while preserving global features by aggregating information from all patches. The system captures both local patch characteristics and global tissue composition through this divide-and-conquer approach

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent extracts only the most representative atoms and their compositions from the image patches, rather than processing all pixel data. This extraction approach significantly reduces computational energy consumption by focusing on essential features while preserving the information needed for accurate tissue characterization and classification

Inventive Principle:
Principle #2Taking out (Extraction)

4Measurement precision

If hard assignment clustering is used for patch characterization, then the classification is simple and fast, but the granularity of tissue heterogeneity representation is insufficient

Engineering Contradiction:
Improvetissue heterogeneity characterization precisionVSAvoidfeature representation complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent replaces static hard assignment clustering with dynamic soft assignment through sparse coding. Each patch is represented as a flexible combination of multiple atoms with varying weights, allowing the system to capture tissue heterogeneity with finer granularity. This dynamic representation adapts to the specific characteristics of each patch while maintaining a unified framework for the entire image

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS10339651B2Simultaneous feature extraction and dictionary learning using deep learning architectures for characterization of images of heterogeneous tissue samples
Publication Date: 2019.07.02 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US10339651B2 patent drawing
  • US10339651B2 patent drawing
  • US10339651B2 patent drawing

AI summary

Apparatus, methods, and computer-readable media are provided for simultaneous feature extraction and dictionary learning from heterogeneous tissue images, without the need of prior local labeling. A convolutional autoencoder is adapted and enhanced to jointly learn a feature extraction algorithm and a dictionary of representative atoms. While training the autoencoder an image patch is tiled in sub-patches and only the highest activation value per sub-patch is kept. Thus, only a subset of spatially constrained values per patch is used for reconstruction. The deconvolutional filters are the dictionary elements, and only a deconvolution layer is used for these elements. Embodiments described herein may be provided for use in models for representing local tissue heterogeneity for better disease progression understanding and thus treating, diagnosing, and/or predicting the occurrence (e.g., recurrence) of one or more medical conditions such as, for example, cancer or other types of disease.