Synthetic Doctor-Patient Conversation Generation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The lack of readily available large datasets of doctor-patient conversations due to patient privacy and regulatory issues hinders the effective training of machine learning models for tasks like entity extraction and automatic speech recognition.
Innovation Solution
The generation of synthetic doctor-patient conversations using knowledge graph guided named entity recognition and existing conversation summaries, which allows for the creation of labeled data that can be used to train machine learning models without the constraints of actual conversation data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If actual doctor-patient conversation data is collected for training machine learning models, then model training quality is improved, but patient privacy and regulatory compliance deteriorate
Solution Approach 1:
The patent creates synthetic copies of doctor-patient conversations that replicate the structure, language patterns, and medical information flow of real conversations without containing actual patient data. These synthetic conversations are generated using template-based approaches where medical entities, symptoms, and dialogue structures are copied from real data while all personally identifiable information is replaced with fictional equivalents, thus maintaining training quality while eliminating privacy risks
Solution Approach 2:
The patent introduces an intermediary synthetic data generation system that sits between real patient data and machine learning model training. This intermediary process transforms real conversation patterns into anonymized synthetic versions through controlled template instantiation, allowing model training to proceed with realistic data characteristics while a layer of abstraction prevents direct access to actual patient information
2Quantity of substance
If manually curated doctor-patient conversation data is used for training, then data availability is improved, but representativeness of real-world conversations deteriorates
Solution Approach 1:
The patent implements dynamic template selection and randomization mechanisms that adapt synthetic conversation generation to match diverse real-world scenarios. Instead of static manual curation, the system dynamically instantiates conversation templates with varying medical conditions, patient backgrounds, and dialogue flows, ensuring the synthetic data reflects the variability and complexity of actual clinical interactions across different domains
Solution Approach 2:
The patent creates a universal template-based generation system that can produce synthetic conversations across multiple medical domains and conversation types. The same framework handles different specialties, patient demographics, and clinical scenarios by selecting and instantiating appropriate templates, making the synthetic data broadly representative without requiring domain-specific manual curation for each case
3Measurement precision
If human-generated doctor-patient conversation data is collected, then training data quality is improved, but data collection cost and time consumption deteriorate
Solution Approach 1:
The patent performs preliminary extraction and structuring of conversation patterns, medical entities, and dialogue templates from existing datasets before the actual training data generation is needed. By pre-processing and organizing the template structures, entity relationships, and conversation flows in advance, the system eliminates time-consuming manual data collection during model training phases while maintaining high data quality through carefully crafted templates
Solution Approach 2:
The patent implements self-service automated generation of synthetic training data using algorithmic template instantiation rather than human data collection. The system automatically generates realistic conversations by programmatically filling templates with medical entities, symptoms, and dialogue structures, eliminating the need for human annotators to manually create or curate training data while maintaining consistent quality through controlled generation parameters
Data Source
AI summary
Knowledge graph guide and entity controlled techniques for generating synthetic doctor-patient conversations. In one particular aspect, a method is provided that includes obtaining an original dataset containing textual dialogue associated with a plurality of individual doctor-patient conversations for training a machine learning model, constructing input data by using named entity recognition to capture and categorize named medical entities present in the dialogue, generating prepared input data by arranging the input data in an annotated turn-by-turn conversation format using an input data preparation algorithm having various control parameters, training the machine learning model using the prepared input data, utilizing a knowledge graph to identify a plurality of symptoms mapped to a randomly selected disease, and causing the trained machine learning model to generate a synthetic doctor-patient conversation by inputting the plurality of symptoms to the machine learning model as a first control parameter of a conversation generation control algorithm.


