Synthetic Doctor-Patient Conversation Generation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The lack of readily available large datasets of doctor-patient conversations due to patient privacy and regulatory issues hinders the effective training of machine learning models for tasks like entity extraction and automatic speech recognition.

Innovation Solution

The generation of synthetic doctor-patient conversations using knowledge graph guided named entity recognition and existing conversation summaries, which allows for the creation of labeled data that can be used to train machine learning models without the constraints of actual conversation data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If actual doctor-patient conversation data is collected for training machine learning models, then model training quality is improved, but patient privacy and regulatory compliance deteriorate

Engineering Contradiction:
Improvemodel training qualityVSAvoidpatient privacy risk
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent creates synthetic copies of doctor-patient conversations that replicate the structure, language patterns, and medical information flow of real conversations without containing actual patient data. These synthetic conversations are generated using template-based approaches where medical entities, symptoms, and dialogue structures are copied from real data while all personally identifiable information is replaced with fictional equivalents, thus maintaining training quality while eliminating privacy risks

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent introduces an intermediary synthetic data generation system that sits between real patient data and machine learning model training. This intermediary process transforms real conversation patterns into anonymized synthetic versions through controlled template instantiation, allowing model training to proceed with realistic data characteristics while a layer of abstraction prevents direct access to actual patient information

Inventive Principle:
Principle #24Intermediary (Mediator)

2Quantity of substance

If manually curated doctor-patient conversation data is used for training, then data availability is improved, but representativeness of real-world conversations deteriorates

Engineering Contradiction:
Improvedata availabilityVSAvoidconversation representativeness
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent implements dynamic template selection and randomization mechanisms that adapt synthetic conversation generation to match diverse real-world scenarios. Instead of static manual curation, the system dynamically instantiates conversation templates with varying medical conditions, patient backgrounds, and dialogue flows, ensuring the synthetic data reflects the variability and complexity of actual clinical interactions across different domains

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent creates a universal template-based generation system that can produce synthetic conversations across multiple medical domains and conversation types. The same framework handles different specialties, patient demographics, and clinical scenarios by selecting and instantiating appropriate templates, making the synthetic data broadly representative without requiring domain-specific manual curation for each case

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If human-generated doctor-patient conversation data is collected, then training data quality is improved, but data collection cost and time consumption deteriorate

Engineering Contradiction:
Improvetraining data qualityVSAvoiddata collection time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary extraction and structuring of conversation patterns, medical entities, and dialogue templates from existing datasets before the actual training data generation is needed. By pre-processing and organizing the template structures, entity relationships, and conversation flows in advance, the system eliminates time-consuming manual data collection during model training phases while maintaining high data quality through carefully crafted templates

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements self-service automated generation of synthetic training data using algorithmic template instantiation rather than human data collection. The system automatically generates realistic conversations by programmatically filling templates with medical entities, symptoms, and dialogue structures, eliminating the need for human annotators to manually create or curate training data while maintaining consistent quality through controlled generation parameters

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20250140404A1Generation of synthetic doctor-patient conversations
Publication Date: 2025.05.01 ORACLE INT CORP
  • US20250140404A1 patent drawing
  • US20250140404A1 patent drawing
  • US20250140404A1 patent drawing

AI summary

Knowledge graph guide and entity controlled techniques for generating synthetic doctor-patient conversations. In one particular aspect, a method is provided that includes obtaining an original dataset containing textual dialogue associated with a plurality of individual doctor-patient conversations for training a machine learning model, constructing input data by using named entity recognition to capture and categorize named medical entities present in the dialogue, generating prepared input data by arranging the input data in an annotated turn-by-turn conversation format using an input data preparation algorithm having various control parameters, training the machine learning model using the prepared input data, utilizing a knowledge graph to identify a plurality of symptoms mapped to a randomly selected disease, and causing the trained machine learning model to generate a synthetic doctor-patient conversation by inputting the plurality of symptoms to the machine learning model as a first control parameter of a conversation generation control algorithm.