Large Language Model Symptom Embeddings for Disease Similarity
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional symptom similarity measurement methods based on the human phenotype ontology (HPO) structure fail to reflect linguistic and medical contexts, leading to inaccurate symptom similarity calculations, especially for diseases with similar symptoms, and struggle with rare diseases due to limited data and complex symptom combinations.
Innovation Solution
A symptom similarity measurement system using a large language model and self-supervised learning, involving symptom embedding generation, synthetic data creation, and self-supervised learning with contrastive techniques to calculate symptom similarity accurately.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional symptom similarity measurement methods based on HPO structure are used, then standardized symptom representation is achieved, but linguistic and medical contexts are not reflected leading to inaccurate similarity calculations
Solution Approach 1:
The patent introduces a large language model as an intermediary between the HPO structured symptom data and the similarity calculation process. The LLM processes and enriches the structured symptom representations with contextual linguistic and medical knowledge, thereby resolving the contradiction between maintaining standardized representation and preserving contextual information.
Solution Approach 2:
The patent transforms the symptom representation from simple HPO structural parameters (distance, depth) to enriched representations that incorporate contextual parameters generated by the LLM. This parameter transformation enables the system to capture linguistic and medical contexts while maintaining the benefits of standardized HPO structure.
2Reliability
If HPO hierarchical depth is used to weight disease-specific symptoms, then standardized symptom classification is achieved, but discriminative power is reduced for diseases sharing similar symptoms
Solution Approach 1:
The large language model serves as an intermediary that re-evaluates symptom weights by considering the broader medical context and relationships between symptoms. Instead of relying solely on HPO hierarchical depth, the LLM analyzes the semantic and medical relevance of symptoms, thereby improving disease discrimination for conditions with overlapping symptom profiles.
Solution Approach 2:
The patent introduces dynamic symptom weighting through the LLM that can adapt to different disease contexts. The symptom weights are not fixed based on HPO depth but are dynamically adjusted based on the specific disease context and symptom relationships, enhancing the system's ability to discriminate between diseases with similar presentations.
3Quantity of substance
If conventional HPO-based methods are used for rare diseases, then existing symptom data is utilized, but accurate similarity calculation is difficult due to limited data and complex symptom combinations
Solution Approach 1:
The large language model acts as an intermediary that can infer and generate plausible symptom relationships for rare diseases based on patterns learned from more common diseases and general medical knowledge. This allows the system to work effectively even with limited rare disease data by leveraging the LLM's contextual understanding to fill data gaps and improve similarity measurements.
Data Source
AI summary
A symptom similarity measurement system using a large language model and self-supervised learning includes: a symptom embedding generation unit converting symptom text information of disease symptoms into embedding values for each symptom and generating disease symptom data using a large language model; a patient data generation unit generating synthetic data, a set of symptoms of a specific disease formed by randomly sampling all symptoms known for the specific disease; and a symptom model training unit performing self-supervised learning on a symptom model using the synthetic data.


