Multimodal Deep Memory Network for Diagnostic Inferencing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Clinicians are overwhelmed by the vast amount of data from historical medical records and databases, leading to increased diagnosis time and costs, and existing clinical guidelines often neglect non-textual medical data, potentially resulting in less accurate patient diagnoses.

Innovation Solution

A multimodal deep memory network is trained to generate embeddings from both medical images and documents, using attention data based on clinician eye movement to weight the importance of different data portions, allowing for more comprehensive patient diagnoses.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If clinicians review all available medical data manually, then diagnostic accuracy is improved, but diagnosis time and costs increase

Engineering Contradiction:
Improvediagnostic accuracyVSAvoiddiagnosis time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent introduces an attention mechanism as an intermediary component that automatically filters and weights medical data based on its relevance to the diagnosis. This mediator processes the overwhelming amount of medical data and presents only the most relevant information to clinicians, thereby maintaining diagnostic accuracy while significantly reducing the time required for data review.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces the manual mechanical process of clinician data review with an automated computational system using neural networks and attention mechanisms. This substitution allows the system to process and prioritize medical data much faster than manual review while preserving diagnostic quality through learned patterns from training data.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Ease of operation

If clinical guidelines focus only on textual data, then processing simplicity is improved, but diagnostic completeness deteriorates

Engineering Contradiction:
Improveprocessing simplicityVSAvoiddiagnostic completeness
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent merges multiple data modalities including textual data, image data, and tabular data into a unified diagnostic framework. The model processes these different data types simultaneously and integrates their information through the attention mechanism, ensuring that no important diagnostic information is omitted while maintaining operational simplicity through automated processing.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent creates a universal diagnostic model that can handle multiple types of medical data (text, images, tables) through a single integrated system. This multi-functional approach allows the model to process diverse data formats using the same architectural framework, improving diagnostic completeness without complicating the processing workflow for clinicians.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentEP3510505B1Systems, methods, and apparatus for diagnostic inferencing with a multimodal deep memory network
Publication Date: 2024.11.06 KONINKLIJKE PHILIPS NV
  • EP3510505B1 patent drawingFigure 1
  • EP3510505B1 patent drawingFigure 2
  • EP3510505B1 patent drawingFigure 3

AI summary

The described embodiments relate to systems, methods, and apparatus for providing a multimodal deep memory network (200) capable of generating patient diagnoses (222). The multimodal deep memory network can employ different neural networks, such as a recurrent neural network and a convolution neural network, for creating embeddings (204, 214, 216) from medical images (212) and electronic health records (206). Connections between the input embeddings (204) and diagnoses embeddings (222) can be based on an amount of attention that was given to the images and electronic health records when creating a particular diagnosis. For instance, the amount of attention can be characterized by data (110) that is generated based on sensors that monitor eye movements of clinicians observing the medical images and electronic health records. Resulting patient diagnoses can be provided according to a predetermined classification of weights, or a compilation of words that are generated over multiple iterations of the multimodal deep memory network.