AI Treatment Path Planning Using Reinforcement Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current medical AI technologies are limited in planning and exploring optimized treatment paths for patients, as they often rely on static treatment methods rather than dynamic, cumulative reward-based decision-making, which may not maximize patient improvement and can be biased towards specific patient groups or medical institutions.

Innovation Solution

An artificial intelligence apparatus and method that utilizes an episode conversion module, patient condition predictive intelligence deep learning, local policy intelligence reinforcement learning, and global policy intelligence management to explore and update optimized treatment paths based on electronic medical records, integrating data from multiple institutions to minimize bias and improve treatment planning.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If static treatment methods are used, then implementation simplicity is maintained, but treatment optimization and patient improvement are limited

Engineering Contradiction:
Improvetreatment optimizationVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent implements dynamic treatment path exploration by transitioning from static treatment methods to a reinforcement learning-based system that continuously adapts treatment paths based on patient responses. The policy intelligence dynamically adjusts treatment sequences by evaluating cumulative rewards from multiple episodes, enabling the system to optimize treatment paths in real-time according to individual patient conditions rather than following predetermined static protocols.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system incorporates feedback mechanisms through the reinforcement learning framework where treatment outcomes are continuously monitored and fed back into the policy intelligence. The episode conversion module processes patient responses and treatment results, which then update the policy intelligence to improve future treatment path selections. This closed-loop feedback system enables continuous optimization of treatment paths based on actual patient outcomes.

Inventive Principle:
Principle #23Feedback

2Reliability

If treatment paths are optimized for specific patient groups or institutions, then local effectiveness is improved, but generalizability and fairness across diverse populations deteriorates

Engineering Contradiction:
Improvelocal treatment effectivenessVSAvoidgeneralizability
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent implements a dual-level policy intelligence architecture where local policy intelligence is trained on institution-specific data to capture local characteristics, while global policy intelligence aggregates knowledge across multiple institutions to ensure generalizability. The global policy intelligence serves as a universal framework that can be applied across diverse populations and institutions, while local policy intelligence adapts this framework to specific local conditions, achieving both generalizability and local effectiveness.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system segments the policy intelligence into local and global components that operate at different levels. Local policy intelligence handles institution-specific optimization, while global policy intelligence manages cross-institutional knowledge sharing. This segmentation allows each component to specialize in its domain while collaborating to achieve overall optimization that is both locally effective and generally applicable.

Inventive Principle:
Principle #1Segmentation

3Productivity

If reinforcement learning with multiple episodes is implemented, then treatment path optimization is improved, but computational requirements and processing time increase

Engineering Contradiction:
Improvetreatment path optimization qualityVSAvoidprocessing time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system performs preliminary action by pre-converting electronic medical records into the episode format required by the reinforcement learning algorithm. The episode conversion module processes and structures patient data in advance, organizing it into standardized episodes that can be efficiently processed by the policy intelligence. This preliminary data preparation reduces computational overhead during the actual treatment path optimization process.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent extracts only the essential features and parameters from complex electronic medical records during the episode conversion process. By extracting and structuring only the relevant information needed for reinforcement learning, the system reduces the computational burden on the policy intelligence while maintaining the quality of treatment path optimization. This extraction approach filters out unnecessary data complexity.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS20230187069A1Artificial intelligence apparatus for planning and exploring optimized treatment path and operation method thereof
Publication Date: 2023.06.15 ELECTRONICS & TELECOMM RES INST
  • US20230187069A1 patent drawing
  • US20230187069A1 patent drawing
  • US20230187069A1 patent drawing

AI summary

Disclosed is an artificial intelligence apparatus, which includes an episode conversion module that receives an electronic medical record (EMR) of a patient and converts the received EMR into an episode including a condition of the patient, a treatment method, and a treatment history, a patient condition predictive intelligence deep learning module that trains a patient condition predictive intelligence for predicting a following condition of the patient after applying the treatment method, a local policy intelligence reinforcement learning module that performs reinforcement learning of a policy intelligence for planning an optimized treatment path for the patient based on the episode, an optimized treatment path exploration module that plans the optimized treatment path for the patient by using the policy intelligence, and a global policy intelligence management module that updates a global policy intelligence for planning and exploring the optimized treatment path based on the policy intelligence.