Reinforcement Learning for Data-Efficient Radiation Therapy Planning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional radiation therapy treatment planning (RTTP) is inefficient, time-consuming, and prone to errors due to reliance on subjective clinician interpretation and computationally intensive AI training methods that require large datasets.

Innovation Solution

A training system using reinforcement learning iteratively trains an AI model with a limited dataset, balancing exploration and exploitation to predict treatment attributes, combining reinforcement and supervised learning to enhance efficiency and accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional AI training methods are used for radiation therapy treatment planning, then the model can learn from historical data, but the training process is computationally intensive, costly, and requires a long training period with large datasets

Engineering Contradiction:
Improveprediction accuracyVSAvoidtraining time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary action by pre-defining multiple possible treatment attributes and their relationships before the actual training process. The treatment planning system pre-processes historical treatment data to create a structured knowledge base of treatment attributes, constraints, and relationships, which serves as a foundation for the AI model training and significantly reduces the training time required.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system segments the treatment planning process into distinct attribute categories (e.g., beam parameters, patient-specific parameters, treatment constraints) that can be independently analyzed and trained. This segmentation allows the AI model to learn from smaller, more manageable data subsets rather than requiring processing of entire large-scale datasets simultaneously.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If conventional AI training methods are used for radiation therapy treatment planning, then the model can predict treatment attributes, but the training process is computationally intensive and costly

Engineering Contradiction:
Improveprediction accuracyVSAvoidcomputational cost
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The system applies partial action by training the AI model on selected subsets of treatment attributes rather than requiring complete training on all possible attributes simultaneously. The system identifies and prioritizes the most critical treatment attributes for accurate prediction, training the model on these essential parameters first, which reduces computational resources required while maintaining prediction accuracy for the most important clinical decisions.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system changes parameters by transforming the training approach from conventional exhaustive training to a parameter-efficient training method. The predefined treatment attributes and their relationships serve as fixed parameters that guide the training process, allowing the model to learn with fewer computational iterations and reduced resource requirements while maintaining predictive accuracy.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If conventional AI training methods are used for radiation therapy treatment planning, then the model can learn from historical data, but it requires a large number of training data points

Engineering Contradiction:
Improveprediction accuracyVSAvoidtraining data volume
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The system uses copying by creating synthetic training data through the predefined treatment attribute relationships. Instead of requiring large volumes of real historical treatment data, the system generates additional training examples by copying and varying the relationships between treatment attributes from existing cases, effectively expanding the training dataset without requiring proportional increases in actual clinical data collection.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system achieves universality by creating a training framework based on universal treatment attribute relationships that can be applied across different treatment scenarios and patient cases. The predefined treatment attributes and their interrelationships form a universal knowledge structure that allows the AI model to learn from smaller datasets while maintaining the ability to generalize to new treatment situations.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12427339B2Training artificial intelligence models for radiation therapy
Publication Date: 2025.09.30 SIEMENS HEALTHINEERS INTERNATIONAL AG
  • US12427339B2 patent drawing
  • US12427339B2 patent drawing
  • US12427339B2 patent drawing

AI summary

Disclosed herein are systems and methods for iteratively training artificial intelligence models using reinforcement learning techniques. With each iteration, a training agent applies a random radiation therapy treatment attribute corresponding to the radiation therapy treatment attribute associated with previously performed radiation therapy treatments when an epsilon value indicative of a likelihood of exploration and exploitation training of the artificial intelligence model satisfies a threshold. When the epsilon value does not satisfy the threshold, the agent generates, using an existing policy, a first predicted radiation therapy treatment attribute, and generates, using a predefined model, a second predicted radiation therapy treatment attribute. The agent applies one of the first predicted radiation therapy treatment attribute or the second predicted radiation therapy treatment attribute that is associated with a higher reward. The agent iteratively repeats training the artificial intelligence model until the existing policy satisfies an accuracy threshold.