Biomedical Acoustic Classification Using Image Representations

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for automated classification of biomedical acoustics, such as heart and lung sounds, rely heavily on segmentation, which is prone to errors due to reliance on a priori information and are not robust to variations in heart rate and additional sounds, limiting their clinical applicability.

Innovation Solution

Classifying biomedical acoustics based on image representations, such as spectrograms, recurrence plots, and Markov transition fields, using deep learning models that bypass the need for segmentation and incorporate time-frequency domain visual information for accurate classification.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If conventional segmentation methods are used for classifying biomedical acoustics, then the classification process can be structured, but the accuracy deteriorates due to reliance on a priori information and errors with complex sounds

Engineering Contradiction:
Improvestructured classification processVSAvoidclassification accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent replaces conventional mechanical segmentation methods with deep learning-based automated classification. The neural network directly processes raw acoustic signals without requiring manual segmentation steps, eliminating the need for a priori information about sound boundaries and structures. This substitution of mechanical processing with intelligent algorithms resolves the contradiction by maintaining operational simplicity while dramatically improving classification accuracy for complex and abnormal biomedical sounds.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The deep learning model performs self-service by automatically learning relevant features and patterns from the acoustic signals without human intervention or pre-programmed segmentation rules. The system adapts to different sound types and abnormalities autonomously, eliminating dependency on expert knowledge and a priori information while maintaining high accuracy across diverse clinical scenarios.

Inventive Principle:
Principle #25Self-service

2Measurement precision

If deep learning-based image representation is used, then classification accuracy improves, but device complexity increases

Engineering Contradiction:
Improveclassification accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary component that transforms complex acoustic signals into simplified image representations (spectrograms, scalograms, or other visual formats). This intermediary step converts the difficult audio classification problem into a more tractable image processing problem that can be solved by standard deep learning architectures, thereby improving accuracy while managing system complexity through effective problem transformation.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system transitions from one-dimensional temporal acoustic signals to two-dimensional or higher-dimensional image representations, adding visual dimensions (frequency, time, energy distribution) that make patterns more discernible to neural networks. This dimensional transformation enables more accurate classification by leveraging the strengths of computer vision algorithms while keeping the underlying system architecture relatively simple.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Reliability

If manual auscultation is performed by clinicians, then diagnostic insight can be obtained, but productivity deteriorates due to time consumption and human error

Engineering Contradiction:
Improvediagnostic insightVSAvoiddiagnosis efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The automated classification system performs self-service by independently analyzing acoustic signals and generating diagnostic classifications without requiring clinician time for manual auscultation. The system maintains high reliability by using deep learning models trained on extensive datasets that capture subtle diagnostic patterns, while simultaneously improving productivity by processing multiple signals rapidly and eliminating human fatigue and variability.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system implements feedback mechanisms where classification results are continuously refined based on performance metrics and clinical outcomes. This allows the system to maintain high diagnostic reliability while operating autonomously at scale, providing consistent, repeatable assessments that improve overall healthcare productivity without sacrificing the nuanced diagnostic insight previously requiring manual clinical evaluation.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12629100B2Classifying biomedical acoustics based on image representation
Publication Date: 2026.05.19 CORNELL UNIVERSITY
  • US12629100B2 patent drawing
  • US12629100B2 patent drawing
  • US12629100B2 patent drawing

AI summary

A method in an illustrative embodiment comprises obtaining an acoustic signal for a given individual, generating an image representation of at least a portion of the acoustic signal, processing the image representation in at least one neural network of an acoustics classifier to generate a classification for the acoustic signal, and executing at least one automated action based at least in part on the generated classification. The acoustic signal illustratively comprises, for example, at least one of a heart sound signal, a blood flow sound signal, a lung sound signal, a bowel sound signal, a cough sound signal, or other physiological sound signal of the given individual. Generating the image representation illustratively comprises generating at least one spectrogram. Additionally or alternatively, generating the image representation may comprise generating one or more recurrence plots, Markov transition field image representations and/or Gramian angular field image representations.