Reinforcement learning-based framework for adaptive decision support in radiotherapy

DE202025102741U1Active Publication Date: 2025-08-14AL-ADAILEH AHMED +3
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
DE202025102741
Authority / Receiving Office
DE · DE
Patent Type
Utility models
Current Assignee / Owner
Filing Date
2025-05-17
Publication Date
2025-08-14
Estimated Expiration
2035-05-31

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A computer-based adaptive decision support system for radiotherapy, including: • a patient data acquisition module configured to acquire patient-specific clinical data, including an anatomical image, a physiological signal, and genomic data; • a preprocessing and feature extraction modality compatible with normalizing, preprocessing and extracting statistical features from the acquired data; • a status estimator used to provide a dynamic representation of the patient's treatment evolution status based on radiation, biological, dosimetric characteristics; • customized action rescaler to define a series of clinically meaningful treatment adjustments depending on the patient's current condition; • a reward function engine used to calculate therapeutic outcome scores based on the probability of tumor control, the probability of complications in normal tissue, and other predetermined factors; • a reinforcement learning agent that can learn and update treatment adaptation policies based on deep reinforcement learning techniques; • a clinical decision dashboard that provides recommended treatment adjustments and personalized interaction with the physician; and • a clinical integration interface adapted for exporting the customized treatment plan to an external treatment planning or delivery system.
Need to check novelty before this filing date? Find Prior Art

Description

Field of the invention

[0001] The present invention relates to the field of medical technology and artificial intelligence, in particular to reinforcement learning systems for personalized and adaptive cancer treatment planning. In particular, it relates to an intelligent decision support system for optimizing radiotherapy plans using real-time patient data and deep reinforcement learning techniques. Background of the invention

[0002] Radiation therapy is one of the most common therapies for treating localized cancer. It relies on the targeted application of ionizing radiation to cancerous tissue while sparing healthy tissue. Although it is a popular treatment modality and has seen significant technological advances, it still faces significant limitations regarding personalization, adaptation, and real-time treatment optimization.

[0003] One of the major shortcomings of current radiotherapy practice is the reliance on static treatment plans. These are typically generated prior to treatment based on CT simulation data. However, the patient's anatomy can change significantly over the course of several weeks of treatment. These include tumor regression, changes due to respiration or peristalsis, weight loss or gain, and dynamic changes in other anatomical structures. Differences can adversely affect dose compliance and reduce treatment efficacy, for example, by underdosing the tumor or overdosing organs at risk (OARs).

[0004] To solve this problem, adaptive radiotherapy (ART) was developed. In ART, a patient is monitored during treatment, and acquired CT images are recorded to adjust treatment parameters based on these images. Despite their potential, current ART approaches are extremely manual, time-consuming, and rely on physician-driven heuristics that vary widely across physicians and centers. These steps often involve offline recalculations, iterative imaging, and subjective decisions, which delays ART adaptation and compromises ART's real-time potential.

[0005] Furthermore, conventional TPS lack intelligent feedback and learning capabilities. They are non-adaptive, meaning they do not learn from treatment history, observed outcomes, or patient-specific response profiles. This lack of learning and personalization capabilities limits the promise of ART to be truly individualized.

[0006] At the same time, (deep) artificial intelligence (AI), and in particular reinforcement learning (RL), has shown enormous potential for areas where decisions must be made in sequences under uncertainty. Reinforcement learning interacts with the environment like human learning, modifying strategies based on an observed outcome and a reward system. In complex, constantly changing medical environments, such as those found in radiotherapy, RL provides a natural framework for enabling automated and tailored decisions over time.

[0007] Although there are initial implementations of machine learning in treatment planning, none of the currently deployed clinical platforms is capable of directly integrating reinforcement learning to autonomously guide decisions regarding adaptive radiotherapy. One such gap exists in the use of patient-specific feedback, historical information, and predictive modeling to enable continuous learning and adaptation in radiotherapy.

[0008] Therefore, it is highly desirable for an intelligent decision support system to be able to automatically suggest and adjust radiotherapy in real time. Any system that proposes this must be able to: • Drawing on stories from previous treatments that have been developed over many treatments with many different patients, • Further expanding the coverage of policies for new patients and • Within the limitations of clinical time and • Improves clinical care outcomes by developing personalized and data-driven decisions.

[0009] The present invention addresses this need by proposing a reinforcement learning-based framework for adaptive radiotherapy planning and treatment formulation. It closes the loop between AI-assisted learning and clinical oncology practice, laying the foundation for a new generation of AI-assisted, responsive, and precise radiotherapy systems. Summary of the invention

[0010] The present invention proposes a novel, intelligent, and autonomous reinforcement learning-based framework for AR-DSS to revolutionize the planning, adjustment, and optimization of radiotherapy plans in cancer treatment. This framework leverages advanced reinforcement learning (RL) algorithms for personalized, real-time decision support, enabling the continuous adaptation of radiotherapy protocols based on clinical response, patient anatomical changes, and treatment impact.

[0011] At the core of the invention is the use of Markov decision processes (MDPs) to model the radiotherapy environment, where the patient's condition changes over time and sequential, context-dependent decisions must be made to achieve optimal clinical outcomes. In this model: • Conditions are the anatomical, radiological, biological and physiological conditions of the patient at a particular time. • Measures Countermeasures to possible treatment changes are taken (e.g. change of radiation dose, change of radiation configurations, adjustment of fractionation schemes). • Rewards are based on treatment effectiveness metrics (such as TCP / NTCP) or physician-defined clinical success criteria.

[0012] To effectively traverse this state-action-reward landscape, the invention employs a deep reinforcement learning (DRL) agent comprising a DQN, an actor-critic model, or PPO. This agent learns optimal strategies by executing simulated treatment scenarios, estimating predicted outcomes, and updating strategies in response to cumulative rewards. The agent is trained throughout therapy using a mix of retrospective treatment data and prospective patient feedback.

[0013] The proposed approach consists of several interconnected modules such as: • A data acquisition module that captures and consolidates multimodal inputs: QT / MRI / PET imaging, genomic profiles, organ movement data, and patient-specific biomarkers. • A preprocessing unit that normalizes / segments this data to create structured input vectors for feeding the model. • A state estimation module that creates a high-dimensional patient state representation that includes tumor dynamics, tissue deformation, and other clinical antecedents. • An action generator that indicates possible changes to the treatment plan with regard to clinical safety and regulatory compliance. • A reward engine that provides multi-objective feedback based on dosimetric predictors, predictive outcome models, and clinician implementation of reinforcement. • A decision support dashboard with interpretable visualizations of proposed interventions, predicted benefits, and risk warning signals, allowing the physician to override when appropriate. • A clinical integration interface that aligns the output with that of commercial radiotherapy treatment planning systems and delivery hardware, enabling easy integration into the clinical workflow.

[0014] Unlike conventional adaptive radiotherapy systems, which operate statically, reactively, and manually, the present invention offers a proactive, continuously learning system. It enables: • Tailored treatment adjustments according to the different response characteristics of the patients. • Evidence-based analysis This leads to greater clinical consistency and less variability between physicians. • Closed-loop feedback integration, where the system evolves with increasing clinical experience.

[0015] The system has two modes: it can operate in retrospective planning mode for historical case review and in adaptive online mode for live patient treatments. By combining real-time data interpretation with re-based, RL-driven guideline updates, it can help achieve the goals of precision oncology (e.g., better tumor control with less collateral damage, better utilization of clinical resources).

[0016] In short, the invention represents a quantum leap toward autonomous decision-making in radiotherapy, relieving clinical staff, minimizing planning errors, and maximizing treatment effectiveness. It is ideal for high-throughput cancer centers, mobile radiotherapy units, and cloud-connected oncology networks that require flexibility, speed, and personalization. Short description of the figure Fig. : Block diagram of the reinforcement learning-based framework for adaptive decision support in radiotherapy. Fig. shows a block diagram illustrating the main components of the system and the relationships between them, namely: data acquisition module (102), preprocessing and feature extraction unit (104), state estimator (106), action generator (108), reward function engine (110), reinforcement learning agent (112), decision support dashboard (114), and clinical integration interface (116). Detailed description of the invention

[0017] The present invention relates to a reinforcement learning-based framework for adaptive decision support in radiotherapy. It comprises several interconnected components that enable the intelligent collection, analysis, and immediate use of clinical data to guide radiotherapy plans. The invention, known as system 100, is configured to function in the clinical environment by communicating with external hospital systems such as treatment planning systems (TPS), electronic medical records (EMR), picture archiving and communication systems (PACS), and radiotherapy hardware. The system operates modularly and has units that run in backlog mode (for guideline training and simulation) and forward mode (for live patient treatment support).

[0018] System 100 is based on the data acquisition module (102), which is used to acquire high-dimensional clinical, multimodal data. This includes anatomical image data (e.g., CT, MRI, PET, 4D-CT), functional and molecular data (e.g., genome sequences and radiogenomic signatures), and physiological data (e.g., real-time respiratory signals, body weight fluctuations, sensor-based monitoring). Treatment data and previous radiotherapy results are also input to expand the learning model and adapt the response functions. This platform ensures the reliable, time-coordinated acquisition and processing of all required data for adaptive treatment decision support.

[0019] After this step, the data enters the preprocessing and feature extraction unit (104), which normalizes, standardizes, and adapts the heterogeneous input data. Modern algorithms, such as deep learning-based convolutional neural networks (CNNs), are used to segment tumors and organs at risk (OARs) in image data. Various radiomic features, including shape, intensity, texture, and spatial distribution, are calculated, and the genomic data is dimensionally reduced to achieve a concise, meaningful representation. All input data are temporally merged between treatment phases to obtain consistent snapshots of the patient's condition over time.

[0020] The output of the preprocessing module is fed to a state estimator (106), which provides a complete, time-dependent vector of patient states at each time point. This vector includes quantitative imaging metrics, DVH features, genome-based radiosensitivity profiles, estimates of tumor progression, and physiological trends. The state vector serves as input to the RL agent and is updated in real time after each data acquisition with the current patient anatomy or biology as it evolves.

[0021] The action generator 108 then specifies the clinically meaningful interventions or changes that the system can recommend at a given treatment timepoint. These interventions can range from dose adjustments (e.g., dose escalation or de-escalation), beam geometry changes, adaptive fractionation schemes, gating adjustments, to instructions for repositioning or re-imaging. The action space is defined within regulatory and clinically safe limits, with physically impossible or contraindicated actions automatically eliminated by internal constraint filters.

[0022] Following an action or therapeutic decision, the reward function (110) calculates a reward signal for the observed or predicted patient outcome. This utility function is a weighted combination of relevant objectives, including tumor control probability (TCP), normal tissue complication probability (NTCP), predicted survival, risk of toxicity from other limitations, and other clinically determined success measures. The reward function balances oncological efficacy against normal tissue sparing and can be personalized according to the patient's individual tolerance or risk profile.

[0023] At the core of the framework, the reinforcement learning agent (112) develops an optimal strategy π* by interacting with the patient through a sequence of state transitions, action selections, and reward feedback. The agent leverages state-of-the-art deep reinforcement learning architectures such as deep Q-networks (DQN), actor-critic models, and / or proximal policy optimization (PPO), depending on the complexity and dimensionality of the environment's state-action space. Experience repetition and occasional strategy updates through stochastic optimization are important components of learning. The agent can be pre-trained using historical patient data and subsequently fine-tuned using real-time feedback from real-world treatments, achieving continuous improvement.

[0024] For clinical benefit, the system includes a Decision Support Dashboard (114) that serves as a user interface for physicians and the radiation therapy team. This dashboard provides a comprehensive overview of the patient's condition, proposed treatment adjustments, anticipated outcomes, and confidence values. It includes explanatory features that highlight the condition parameters that significantly contributed to the decision, as well as a risk-benefit analysis for each intervention. The physician has the option to accept, modify, or reject the system's recommendations, enabling human-in-the-loop monitoring.

[0025] The Clinical Integration Interface (116) enables the System 100 to connect to commercial radiation therapy planning and delivery systems. The customized plans are then created in the standard DICOM RT format and transmitted to the treatment devices via secure APIs. The interface is also HIPAA and GDPR compliant and meets applicable healthcare data protection regulations. The system can be deployed both on-premises and in the cloud, making it suitable for large academic centers, mobile radiation therapy units, and clinical networks across multiple countries.

[0026] In summary, the reinforcement learning-based framework for adaptive decision support in radiotherapy represents a highly modular, intelligent, and user-friendly solution for personalized radiotherapy in clinical practice. By integrating reinforcement learning, predictive modeling, and clinical decision support into a monolithic structure, this invention enables patient-specific treatment plan optimization in real time, thus significantly improving treatment efficacy, safety, and efficiency.

Claims

[1] A computer-based adaptive decision support system for radiotherapy, including: • a patient data acquisition module configured to acquire patient-specific clinical data, including an anatomical image, a physiological signal, and genomic data; • a preprocessing and feature extraction modality compatible with normalizing, preprocessing and extracting statistical features from the acquired data; • a status estimator used to provide a dynamic representation of the patient's treatment evolution status based on radiation, biological, dosimetric characteristics; • customized action rescaler to define a series of clinically meaningful treatment adjustments depending on the patient's current condition; • a reward function engine used to calculate therapeutic outcome scores based on the probability of tumor control, the probability of complications in normal tissue, and other predetermined factors; • a reinforcement learning agent that can learn and update treatment adaptation policies based on deep reinforcement learning techniques; • a clinical decision dashboard that provides recommended treatment adjustments and personalized interaction with the physician; and • a clinical integration interface adapted for exporting the customized treatment plan to an external treatment planning or delivery system. [2] The system of claim 1, wherein the reinforcement learning agent uses a Deep Q-Network (DQN) to learn a policy for adapting to radiotherapy. [3] The system of claim 1, wherein the state estimator takes into account longitudinal changes in image data, genomic signatures, and physiology to evolve the patient state vector over time. [4] The system of claim 1, wherein the reward function is dynamically calculated based on a combination of predicted tumor control probability (TCP) and predicted normal tissue complication probability (NTCP). [5] The system of claim 1, further comprising: The treatment plan changes defined by the action generator include adjusting the dose, changing beams, selecting a gating technique, or adjusting a fractionation plan. [6] The system of claim 1, wherein the visual explainability is presented on the clinical decision dashboard for recommended interventions by feature association or model confidence. [7] The system of claim 1, wherein the experience replay mechanism stores historical state transitions of the patient and the corresponding results for optimizing the reinforcement learning policy. [8] The system of claim 1, wherein the clinical integration interface is DICOM-RT compatible and secure with commercial treatment planning systems. [9] The system of claim 1, wherein the reinforcement learning agent is pre-trained using historical patient data and fine-tuned in real time during clinical operation using online learning. [10] The system of claim 1, wherein the data includes patient-specific 4D CT scans, radiomic biomarkers, whole genome sequencing (WGS) data, and data generated by wearable sensors.

Citation Information

Cited By

  • Medical aid decision-making system based on multi-modal large model

    CN121302291A

  • Spinal cord injury state evaluation system fusing pathological data and clinical characteristics

    CN121354940A

  • Dynamic optimization method for individualized dosage of low-molecular heparin of cancer patient based on reinforcement learning

    CN121617545A

  • Intelligent intervention strategy generation method for full-period health management

    CN121905539A