Computer-implemented diagnostic image generation
Patent Information
- Application Number
- PCT/EP2026/055327
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-02-26
- Filing Date
- 2026-02-26
- Publication Date
- 2026-09-03
Smart Images

Figure EP2026055327_03092026_PF_FP_ABST
Abstract
Description
[0001] 26 / 02 / 2026
[0002] International patent application
[0003] Aileen Health
[0004] P383171WO
[0005] Computer-implemented diagnostic image generation
[0006] The present invention relates to the prediction of disease-related tissue changes that develop over a longer period and are detectable by longitudinal medical imaging. This includes, for example, malignant processes in oncological diseases, brain atrophy or lesions in neurological diseases, as well as structural and functional tissue changes in organic diseases that can ultimately lead to organ failure. The diseases affected include, among others, cardiovascular diseases such as heart failure, chronic respiratory diseases such as chronic obstructive pulmonary disease (COPD), endocrine and metabolic diseases such as diabetic retinopathy, autoimmune and inflammatory diseases such as endometriosis, and kidney diseases.
[0007] One application example of the present invention is the analysis of lesions in breast tissue with the aim of supporting the early detection and prognosis of the most likely disease course. In particular, two scenarios are considered: 1) the disease course without drug or therapeutic intervention, and 2) the disease course with such treatment measures. The invention enables a detailed analysis of the differences between these scenarios in order to evaluate the efficacy and effectiveness of the respective treatment methods.
[0008] Developing effective prognostic methods is a key challenge in medical imaging, as diagnoses for certain diseases are often only made after symptoms have appeared. Early assessment of a disease's aggressiveness, even at the first signs, is crucial for enabling timely and targeted interventions.
[0009] However, current diagnostic methods have several limitations. These include the challenge of differentiating between benign and malignant tissue changes, as well as the challenge of differentiating between normal aging processes and neurological diseases. Such problems can, among other things, lead to the development of interval cancers in breast tissue and delay the timely initiation of appropriate treatment measures.
[0010] This poses a significant challenge, particularly in the early stages of cancer or with small lesions, few of which progress to invasive cancer. A similar situation exists with early cognitive impairments, where there is an increased risk of dementia, but an Alzheimer's diagnosis is still unlikely. It is in physicians' interest to know which patients will soon be diagnosable with a rapidly progressing disease in order to begin treatment early. Furthermore, methods based on generative models are known in the art to predict the temporal evolution of medical image data. The architectural design of such generative models is typically highly specialized for processing a primary data modality, such as radiological image data driven by image-based biomarkers like hippocampal volume segmentation.
[0011] One technical challenge with these systems is that existing mechanisms for incorporating additional, non-image-based contextual information are often limited in their capacity and flexibility. Frequently, contextual information is fed into intermediate layers of the model as simple, low-dimensional vectors via separate input paths.
[0012] However, this architecture is only partially suitable for effectively processing rich, structured, and multimodal contextual biomarkers from non-image-based sources—such as a complete treatment history or complex biological tumor profiles—and for modeling their profound interactions with image morphology. The system's ability to generate a highly personalized prognosis based on complex chains of conditions may be limited by this architectural design, as the input paths of the generative model itself are not designed to process such complex, heterogeneous data.
[0013] Summary of the invention
[0014] The present disclosure aims to provide a method and an image processing system that enables the early detection of diseases, going beyond mere prognoses and allowing for the causal modeling of therapy effects. One aspect of this disclosure relates to a computer-implemented method for processing examination image data using a synthesis model. The examination image data captures at least one area of the subject in which tissue changes may be present, such as a lesion in cancer or the hippocampus in Alzheimer's disease. The method involves receiving examination image data from multiple examinations of the subject, which are separated in time. This examination image data includes temporal information reflecting the time of each examination.Furthermore, the procedure includes generating synthesis image data for one or more synthesis time points, based on the examination image data and the respective synthesis time point. The synthesis image data reflects the state of the tissue alteration area captured in the examination image data at the respective synthesis time point. The tissue alteration area is a region of the examination object in which potentially pathological tissue changes can occur.
[0015] In current technology, medical images, such as those from mammography screenings, are examined to identify and assess (potentially) pathological tissue changes. However, no images are generated or synthesized to predict the future development of these (potentially) pathological tissue changes. While current methods can detect (pathological) tissue changes in examination image data, the precise future timing of their development, or a prognosis regarding the timing of their development, remains uncertain. One reason why generative methods are not currently used to predict disease progression is that existing technologies are largely based on convolutional neural networks (CNNs).These networks are designed to extract relevant features from medical data and then make classifications that can serve as a basis for predictions.
[0016] An example of this is the assessment of the risk of developing breast cancer within 24 months based on mammograms, or the prediction of a drug's effectiveness using pathological images. However, classification, often in the form of numerical values, is frequently insufficient for physicians to make informed medical decisions. For this reason, medical guidelines have not yet been adapted accordingly. The present invention solves this problem by providing synthesis image data generated by the synthesis model, which contains comprehensive information and additionally takes into account the temporal data of the examination time points of previous examination image data. This integrates sequential information into the analysis. The training data for the synthesis model originates from longitudinal clinical datasets, but also from image data from clinical pharmaceutical studies.This means the image data includes clinical evidence previously unknown to physicians. Creating this synthesis image data thus enables a prognosis (with or without treatment), or even a diagnosis regarding the course of the disease and the determination of the precise time of its onset. Based on these prognoses, the surveillance interval can be adjusted to better detect interval cancers and reduce the risk of overlooking potential cancer diagnoses.
[0017] For the purposes of this disclosure, the term "tissue alteration" refers to a change in tissue that can be detected by medical imaging techniques such as X-ray, computed tomography (CT), magnetic resonance imaging (MRI), ultrasound or positron emission tomography (PET), microscopic imaging, optical coherence tomography (OCT), dopamine transporter imaging (DaTscan), or other functional or structural imaging techniques. A tissue alteration may be a functional or structural change in tissue caused by a disease in the tissue.
[0018] Here and in the following, the term "image data" is used to describe the information obtained through medical imaging procedures. This image data can be obtained from various imaging techniques, such as mammography, computed tomography (CT), magnetic resonance imaging (MRI), positron emission tomography (PET), and histopathological whole-slide images (WSI). The image data may be in digital form and acquired from different perspectives to obtain varying information about the subject. The image data can also be used in combination with other data, such as clinical or patient-specific contextual information or annotations, to enhance understanding of the subject and to support diagnosis and treatment planning.
[0019] The term "synthesis image data" refers to digital images generated by an algorithm based on input medical image data, specifically image data, and other information. This synthesis image data represents a predicted representation of the disease progression in the input medical image data and can be reviewed by a medical professional to support diagnosis and treatment planning. Synthesis image data can also be used to evaluate the effectiveness of a therapy and to adjust the monitoring interval for a specific disease.
[0020] For the purposes of the present invention, a "processing model" is a model configured to process input data and output data. The processing model can be a classical model created, for example, using classical optimization or analysis methods.
[0021] According to the present disclosure, a “machine learning modeH” is an algorithmic procedure, also called a processing model, which is trained on the basis of training data, comprising input data and annotations, to map the input data to the annotations, in order to then map input data not seen in the training in the same way.
[0022] For the purposes of the present invention, a "synthesis model" is an algorithmic method that is trained, based on training data comprising input data and annotations, to determine synthesis image data based on examination image data from multiple examinations. In particular, reference points, i.e., the examination times and a synthesis time point, are given for the examination image data and the synthesis image data, which place the images in a chronological order.
[0023] For the purposes of the present invention, "training data" comprises input data and target data, wherein each input data point corresponds to an annotation or label, called the target data point, of the target data. The target data is typically generated by complex processing or by manual labeling. A processing model is trained on the target data to execute a desired image. The synthesis time point refers to the point in time that the synthesis image data is intended to represent. The synthesis time point is often in the future to generate predictions, but can also be in the past, for example, to enable cardiological simulations. The synthesis time point can, in particular, be entered into the synthesis model as part of the input data. The synthesis model then generates the synthesis image data for the synthesis time point.
[0024] For the purposes of the present invention, an "input data" is data input into a processing model that is processed by the processing model. In the present invention, input data is, in particular, image data. The input data can, in particular, comprise a single image or multiple images, image stacks, or time series of images or image stacks. According to the present disclosure, input data is the data that is input into a machine learning model, such as the synthesis model, and is then processed according to the trained image, for example, to determine predictions or classifications.
[0025] According to the present disclosure, "annotations," also called target data, are data used in training a processing model, such as a machine learning model or a prediction model, to execute a processing mapping. The result data output by the processing model, based on the input data, is to be fitted to these annotations. This approximation is achieved using an objective function.
[0026] For the purposes of the present invention, an “objective function” is, in particular, a gain function or a loss function that specifies how differences between the output data of the processing model and the target data are evaluated. When training machine learning models, the training is carried out by optimizing the objective function, wherein the model parameters of the trained machine learning model are adjusted during training such that the objective function is optimized.
[0027] A preferred embodiment of the above method is characterized in that the image data includes object-specific context information and development-specific context information, wherein the development-specific context information includes the time of the examination and the time of the synthesis. The generation of synthesis image data comprises the following steps: First, generating an internal state representation, or simply representation, from the image data, the object-specific context information, and the development-specific context information using a coding model, in particular by encoding the image data together with the object-specific context information and the development-specific context information, especially by mapping them into a multi-scale feature space; and second, generating the synthesis image data from the internal representation using the synthesis model.The generation of synthesis image data from the internal representation includes controlling the synthesis model using a control model, and the control of the synthesis model includes conditioning the control model using development-specific context information. The control of the synthesis model is preferably based on patterns learned from training data, where the patterns arise in particular from biological and physical plausibility rules learned from the training data using a variety of investigation image data. The technical effect achieved by this specific architecture is the creation of a robust, controllable, and technically valid system for counterfactual simulation, which goes far beyond simple prediction. This effect results from the synergistic interaction of the individual models, each of which solves specific technical problems.
[0028] The function of the coding model is to solve the technical problem of processing extremely heterogeneous data (high-dimensional images, static patient data, dynamic progress data). It fuses all this information into a single, dense internal state representation, thus creating a unified, information-rich "digital fingerprint" of the patient's condition that is computationally efficient. The function of the synthesis model is the actual generation of the new, artificial image data from this internal representation.
[0029] The control model's function is the crucial innovation, producing several technical benefits. First, conditioning the control model with development-specific context information solves the technical problem of how to transform a generic predictive model into a specific, controllable simulation engine. This enables the targeted generation of "what-if" scenarios. Second, the control is based on learned biological and physical plausibility rules. The resulting technical benefit is an internal quality control mechanism that actively prevents the generation of medically nonsensical images or artifacts ("hallucinations")—a common technical problem with uncontrolled generative models. This significantly increases the technical fidelity and reliability of the generated synthesis image data.
[0030] According to this revelation, object-specific contextual information is static or slowly changing data that characterizes the object of study and its basic properties throughout the entire observation period. In contrast, developmental contextual information is dynamic, interval-related data that describes the events and changes that occur between two points in time or are defined for a future simulation. A coding model is a neural network trained to learn a compressed, information-carrying representation of input data called an internal state representation. This is a low-dimensional vector representation of the high-dimensional input data in a vector space learned by the model, the multiscale space.The synthesis model is a generative AI model designed to generate new, artificial data, specifically synthesis image data. A control model is a technical module, particularly a neural network, trained to actively direct and influence a generative process. In a preferred configuration, the control model is an integral part of the synthesis model or interacts with it. The technical process of providing the pre-trained control model with specific information as a control command during the application phase is called conditioning the control model. This control is based on learned patterns, understood as biological and physical plausibility rules, which represent regularities of disease development extracted from the training data.The architecture defined here, with its three interlocking models, achieves the effect of creating a robust, controllable, and technically valid simulation engine for counterfactual medical image time series. This system fundamentally surpasses the capabilities of conventional predictive models by not only forecasting a probable future but also generating a multitude of biologically plausible "what-if" scenarios. This overall effect results from the synergistic interplay of the individual technical features, each of which solves specific technical problems.
[0031] Conventional approaches to processing longitudinal data face the technical challenge of meaningfully combining extremely heterogeneous data sources—high-dimensional images, static patient data (object-specific), and dynamic trend data (development-specific). Often, this information is processed separately, or contextual information is merely appended as additional features to an image-based state representation ("concatenated"). This leads to a loss of information because the complex, nonlinear interactions between image morphology and clinical parameters are not fully captured.
[0032] The proposed solution, which uses a coding model to encode all this information into a single, shared internal state representation, solves this problem. The technical effect is not only efficient data compression, but also the generation of a holistic, information-dense state representation of the entire dynamic system "patient." In this multiscale space, the position of a vector encodes not only the appearance of an image, but also the entire clinical and historical context. This allows the subsequent synthesis process to operate on a significantly richer and contextually more valid data foundation, leading to more precise and personalized simulations.
[0033] Conventional, purely predictive models, especially sequential models, are designed to predict the most probable single next journal entry. They are "one-way streets" without a technical interface to actively guide the prediction process or explore alternative paths. The technical problem of how to translate an abstract instruction (e.g., "simulate therapy Y instead of X") into a targeted modification of the generation process remains unsolved.
[0034] The introduction of a control model explicitly conditioned by development-specific context information solves precisely this problem. The technical effect is the transformation of the system from a passive predictor into an active, controllable simulation engine. Conditioning is the technical mechanism that makes it possible to generate any number of different, but internally consistent, future scenarios for the same initial state (represented by the internal state representation). This ability to perform counterfactual simulation is the fundamental qualitative difference from pure forecasting models and constitutes the core of the invention.
[0035] A key technical problem with generative models is their tendency to "hallucinate" artifacts or biologically impossible results. For example, a model based purely on statistical correlations might simulate a tumor growing through bone tissue or changing at a physically impossible rate. Such results are technically worthless and clinically dangerous.
[0036] The proposed solution, which bases the control model on biological and physical plausibility rules learned from training data, implements an in-process technical validation. The technical effect is a dramatic increase in the robustness and technical fidelity of the generated data. The system does not simply produce any image, but rather an image that obeys the implicitly learned "laws of nature" governing disease progression. This solves the problem of the unreliability of generative models in the medical context and ensures that the generated counterfactual scenarios are not only statistically, but also biologically and physically meaningful. Synergistic overall effect: The combination of these three effects leads to a system capable of generating a completely new type of technical data: multiple, comparable, controllable, and technically valid visual future scenarios for a single subject under investigation.This solves the overarching technical problem that previously no direct, patient-specific and visual data basis existed for evaluating alternative treatment strategies.
[0037] For the purposes of the present invention, the term “feature space” is used to describe the high-dimensional vector space learned by the coding model (153). Within this feature space, the state of the object under investigation at a specific point in time is represented by an “internal state representation”. Since this representation reflects a temporal evolution, it can also be referred to as a spatiotemporal state representation.
[0038] This internal state representation is the concrete, dense, and numerical representation—for example, a vector, a tensor, or a set of vectors—that the coding model generates from the multimodal input data (image data and context information). The feature space, in turn, is not unstructured but acquires a specific semantic order through the training method according to the invention. Depending on the underlying architecture of the coding model, this feature space can exhibit different properties:
[0039] 1. In one embodiment, it can be conceived as a latent space, which typically represents a low-dimensional bottleneck in an autoencoder architecture. Here, the entire state of the object under investigation is encoded in a single, low-dimensional vector.
[0040] 2. In a second, preferred configuration, it can be designed as a multi-scale or hierarchical feature space, as generated by architectures such as UNets. These process information at different levels of abstraction in parallel and, through connections between these levels (so-called skip connections), enable a particularly detailed representation.
[0041] 3. In a third, also preferred, implementation, the multiscale feature space can be realized through a sequence-based architecture, such as a hierarchical vision transformer. Instead of compressing the input data into a single vector, the internal state representation is mapped here as a set of vectors (tokens) on different hierarchical levels. The multiscale space is created through a process that typically includes the following steps:
[0042] Tokenization: The input image is broken down into a sequence of small image patches (tokens) that form the finest scale of space. Hierarchical feature extraction: Across several stages, the relationships between the tokens are contextualized through a self-attention mechanism.
[0043] o Generation of coarser scales: Between the levels, neighboring tokens are combined to form new, semantically richer tokens on a coarser scale ("patch merging" or "token pooling"). This creates a pyramid of feature representations on different scales.
[0044] Regardless of the chosen architecture, the decisive technical effect of the invention lies in the fact that the position and trajectory of these internal state representations in the semantically ordered feature space reflect the clinical and biological relationships. This enables subsequent modules to derive properties such as the presence, type, and progression of a pathological tissue change directly from the structure of the space and the location of the representation, thus forming the basis for a precise, controllable simulation.
[0045] The crucial technical effect of the invention lies in the fact that the position and trajectory of these internal state representations in the semantically ordered feature space reflect the clinical and biological relationships. This enables subsequent modules to derive properties such as the presence, type, and progression of a pathological tissue change directly from the structure of the space and the location of the representation, thus forming the basis for a precise, controllable simulation.
[0046] Another preferred embodiment of the above procedure is characterized by the fact that the contextual information used is specifically defined in order to further increase the accuracy and clinical relevance of the simulation.
[0047] According to this design, the object-specific contextual information includes at least one or more pieces of information from the group consisting of patient-specific or subject-specific data, in particular demographic information such as age and sex, family history or hormonal status; TNM classification (tumor size, lymph node involvement, metastasis status), pathology-specific genetic mutations, genetic information, in particular DNA mutations such as BRCA1 / BRCA2 mutations; and molecular and histopathological markers, in particular information on receptor status, protein overexpression or proliferation markers.In addition, the development-specific contextual information includes at least one or more pieces of information from the group consisting of time information that defines the time interval to be simulated and the course of treatment within that interval, in particular drug information or active ingredient information that specifies a treatment type, dose or administration interval.
[0048] The technical effect achieved through this specific selection of contextual information is a significant increase in the specificity and precision of the simulation process. While the general architecture ensures basic controllability, this design solves the technical problem of how to supply the model with the most causally relevant biological and clinical parameters in order to make highly personalized and biologically sound predictions.
[0049] By explicitly encoding fundamental biological drivers (such as HER2 status, HR status, and BRCA mutations) into the internal state representation and using specific therapy information (such as treatment type and dose) for conditioning, the model is forced to apply the learned patterns to a specific, biologically accurate context. This reduces ambiguity in the multiscale space and leads to more precise and reliable control of the generation process. The result is synthesis image data whose predicted development is not only statistically plausible but also closely linked to the individual tumor biology and the patient's specific treatment scenario, significantly improving the technical quality and clinical utility of the simulation.
[0050] Advantageous forms of disclosure relate to the above procedure, and also comprehensively to:
[0051] Receiving examination image data of an examination of the object of study by an identification model, identifying the tissue change image data of the examination image data that capture the tissue change areas, and marking the identified tissue change image data so that the marked tissue change image data are prioritized when generating synthesis image data by the synthesis model, or can be prioritized.
[0052] Here and in the following, "examination image data" refers to digital image data obtained through various imaging techniques, such as X-rays, ultrasound (mammosonography), digital breast tomosynthesis (DBT), and magnetic resonance mammography (MRI). The examination image data comprises complete images of a single object under examination, for example, a breast.
[0053] The "object of examination" refers to the object captured during an examination. For example, the object of examination could be a breast, but other organs can also be captured. Organs that have been confirmed as particularly suitable for the procedure in oncological diseases include, in particular, the breast, lungs, prostate, brain, intestines, liver, skin, and cervix. The suitability of these organs is based on several key aspects: 1) Sufficient longitudinal imaging data from clinical settings is available, and / or the pharmaceutical industry has conducted extensive clinical trials that can provide corresponding longitudinal imaging data; 2) Imaging enables both the early detection and precise diagnosis of cancers as well as the continuous monitoring of disease progression. This allows for the evaluation of treatment success and, if necessary, adjustments to the treatment.Lung cancer is an example of a suitable organ and oncological disease, as extensive longitudinal imaging data from clinical trials and patient databases are available for this condition. This data allows for detailed tracking of tumor progression across different stages and provides a solid foundation for using AI-supported analysis to accurately model and predict disease progression. However, this method is also applicable to a wide variety of other cancers and organs. Organs that have been confirmed as particularly suitable for the method in neurological diseases include, in particular, the brain, spinal cord, and peripheral nerves. These organs are especially well-suited for the application of this method due to the availability of extensive imaging data, such as MRI and CT scans.Imaging enables precise detection and monitoring of neurological diseases, such as tumors, strokes, or neurodegenerative diseases, and provides a solid basis for AI-supported analysis of disease progression and response to therapeutic measures.
[0054] Organs that can ultimately lead to organ failure include the heart, lungs, eyes, kidneys, and liver. Diseases that can lead to failure of these organs include cardiovascular diseases such as heart failure, chronic respiratory diseases such as chronic obstructive pulmonary disease (COPD), endocrine and metabolic disorders such as diabetic retinopathy, autoimmune and inflammatory diseases such as endometriosis, and kidney diseases. These diseases are particularly well-suited for the application of the disease progression analysis method for the following reasons: 1) Longitudinal imaging data from clinical trials and patient observations are widely available for these diseases. 2) These diseases require continuous monitoring of treatment response.
[0055] According to the present disclosure, the term "change image area" refers to a specific region or section in a medical image that shows a tissue change, anomaly, or abnormality that has the potential to become pathological; therefore, it is also referred to as a potentially pathological tissue change, or alternatively, a lesion. In this disclosure, the term "tissue change image area" refers to the portion of an image that shows a potentially pathological tissue change. In particular, tissue change image areas are used by an algorithm for identifying signs of disease in medical images to automatically locate this area. The tissue change image area can have various shapes and sizes, depending on the type of disease and the medical imaging technique (e.g., MRI, PET) used to create the image.For a specific disease and a corresponding imaging procedure, the size and shape can be advantageously standardized. In particular, the tissue alteration image area is the area around the tissue alteration in the respective image where disease progression is evident, such as the atrophic tissue changes in the brain in Alzheimer's disease, the lesions in multiple sclerosis, or the areas around a lesion in which breast cancer develops or may develop.
[0056] According to the present invention, an examination is in particular an imaging method in which an object under investigation is captured. The examination can, for example, take place at regular intervals or when symptoms occur.
[0057] In current technology, the determination or classification of image data is often carried out manually by an operator; consequently, the quality of the assessment varies due to the inter- and intra-variability of the operators.
[0058] The inventors have recognized that for the creation of prognoses it is helpful to prioritize the processing of tissue change image data from the examination image data that captures the pathological tissue change, thereby significantly reducing the complexity and computing time.
[0059] Because the identification model is trained to automatically identify image data from the examination image data using a specially trained identification model, a consistent and precise preselection of tissue alteration image data can be ensured. This leads to improved predictive quality when synthesizing synthesis image data with the synthesis model, as the input data for the synthesis model is consistent and accurate. Automation reduces reliance on human intervention and thus minimizes sources of error that can arise from human carelessness or fatigue. This increases the accuracy and reliability of the generated synthesis image data, leading to better diagnosis and more effective monitoring of disease progression, such as the development of lesions or damage to brain tissue in Alzheimer's disease.
[0060] For the purposes of this disclosure, an “identification model” is a processing model that identifies tissue change image data in the examination image data; the marked tissue change image data can then be prioritized during processing using the synthesis model.
[0061] The term "tissue alteration image data" refers to one or more digital images showing the same tissue alteration, specifically multiple tissue alteration image areas of the same tissue alteration in the same subject. Tissue alteration image data may also contain additional information, known as contextual information, such as the time the image was acquired and specific disease-related biomarkers and risk factors. In particular, tissue alteration image data can be part of the examination image data, with the tissue alteration image data being marked within the examination image data so that it is prioritized for processing by the synthesis model.
[0062] Another preferred embodiment of the above method, wherein the coding model fulfills the function of the identification model for identifying tissue alteration image data and for labeling the identified tissue alteration image data by transferring the examination image data and the object-specific context information and the development-specific context information into the internal representation, wherein the internal representation includes or is associated with biological feature attributes of the tissue structure; and thus the control model extracts these biological feature attributes from the internal representation or uses them as a conditioning signal to localize tissue alteration areas and control their temporal development.
[0063] The technical effect achieved through this design is the creation of a semantically ordered and inherently interpretable multiscale space, leading to a drastic increase in the efficiency and accuracy of the entire simulation process.
[0064] Conventional coding model architectures often generate an unstructured, "chaotic" multiscale space where the position of a vector has no direct semantic meaning. Understanding what such an internal state representation represents would require complex, downstream analysis or classification steps. This is not only computationally intensive but also a potential source of errors, as the interpretation is decoupled from the actual coding process.
[0065] The proposed solution overcomes this technical problem by specifically training the coding model to create an intelligent "map" of the disease state. The crucial technical effect is the transformation of the identification task from a costly pattern recognition problem to a highly efficient position query problem. The control model no longer needs to search for complex features within the multiscale vector to detect a tissue change; instead, it can directly and reliably read this information from the vector's position in multiscale space.
[0066] This inherent encoding of meaning into the spatial structure offers several concrete technical advantages. First, identification by the control model becomes extremely fast and computationally efficient. Second, the controllability of the overall system improves dramatically. Because the control model "understands" multiscale space as an ordered map, it can plan the future trajectory much more precisely and robustly by navigating a meaningful, smooth path through the semantically ordered regions of space. Third, the overall system becomes more robust and reliable because the dependence on downstream, potentially error-prone classification steps is eliminated. In summary, this design solves the technical problem of efficiently and reliably interpreting internal state representations, thus creating a technically superior foundation for precise and fast simulation.
[0067] Another advantageous embodiment of the above procedure also includes:
[0068] Receiving the synthesis image data by an evaluation model, analyzing this synthesis image data using the evaluation model to create a synthesis evaluation, and outputting this synthesis evaluation, wherein the synthesis evaluation reflects a pathological escalation in the synthesis image data; the synthesis evaluation may, in particular, include one or more of the following evaluations:
[0069] • an assessment of the medical plausibility of a tissue alteration synthesized according to the synthesis image data,
[0070] • an assessment of the medical relevance of a tissue alteration synthesized according to the synthesis image data,
[0071] • an assessment of the progression of a tissue change synthesized according to the synthesis image data, in particular whether a health condition improves or deteriorates.
[0072] An example of this is the detection of malignant tissue changes in breast cancer, which typically occur in either the right or left breast, but very rarely in both breasts simultaneously.
[0073] For the purposes of the present invention, an "evaluation model" is a processing model, in particular a machine learning model, configured to evaluate the synthesis image data. The evaluation model can be configured as a classical processing model, in which case an evaluation is determined based on identified image features. Alternatively, the evaluation model can also be configured as a machine learning model.
[0074] A "synthesis score" refers to the result of a process in which an algorithm is used to generate a predicted output based on input data. In the context of the patent application, the synthesis score refers to the output of a scoring model trained to, for example, detect signs of malignant tissue changes in synthesis image data and output a score based on these signs. The synthesis score can be a number indicating whether, or with what probability, the synthesis image data exhibits a malignant tissue change, or it can be one or more segmentations around these signs. The synthesis score of a malignant tissue change can be based on a single image, but the reliability of the score is increased by analyzing multiple images.A typical synthesis assessment for brain atrophy includes segmentation of the area around the hippocampus and a numerical indication of brain volume reduction based on previous imaging data. This number should be within medically plausible limits—for example, the hippocampal atrophy rate in Alzheimer's disease is typically 4–8% per year.
[0075] In current technology, image data is manually analyzed by medical experts to identify pathological tissue changes. While this manual analysis is detailed, it is often insufficient for making a sound prognosis. For example, it is not possible for a breast cancer radiologist to predict who will develop interval cancer. Similarly, it is difficult for a neurologist to predict who with mild cognitive impairment will develop Alzheimer's disease within five years. Furthermore, generated synthesis images may be the result of hallucinations or contain artifacts that are difficult to detect, especially after long examination times.
[0076] The present disclosure addresses these shortcomings by employing a specially trained evaluation model designed to assess synthesis image data and generate a synthesis score. These scores indicate whether, and with what probability, the synthesis image data exhibits, for example, a malignant tissue alteration, based on the size and appearance of tumors. The evaluation model is trained to recognize both the normal image structures of malignant tissue alterations and potential artifacts in the generated images, enabling greater accuracy in the identification of malignant tissue alterations. By automating this analysis, the model provides a consistent and objective method for evaluating the synthesis image data.This leads to improved overall performance of diagnostic procedures and enables more precise and earlier detection of malignant tissue changes, which can ultimately contribute to better patient health.
[0077] A further advantageous embodiment of the above procedure also includes: receiving additional image data corresponding to further examinations at different time points, and comparing this additional examination image data with correspondingly generated synthesis image data for the respective examination time point. This is particularly important in decision-making regarding the treatment of diseases, as the generated synthesis image data represent scenarios both with and without treatment.
[0078] The examination time refers to the point in time at which a subject of examination—that is, a patient and, in particular, the tissue being examined, for example, breast tissue—is examined by a physician or other healthcare professional to determine whether a specific condition or abnormality persists or has changed. The examination may extend over a certain period, in which case the examination time is a representative point in time, such as the beginning or end of the examination, or a point in between. In particular, the time information with which the examination time is represented / corresponds does not need to be determined down to the second; a granularity of hours, days, weeks, or even months is sufficient. For image data that conforms to the DICOM standard, the examination time corresponds to the start date of the examination.
[0079] The term "time information" refers to information from which a point in time can be derived. Points in time can be specified with varying degrees of granularity. For example, down to the hour, day, week, or month; even smaller granularities such as seconds or minutes are conceivable. For the time information used here, a granularity down to the day, week, or even month is usually sufficient.
[0080] In the prior art, synthesis or predictive images are not generated; consequently, a prognosis cannot be verified using image data. Only estimated prognostics regarding the development of a lesion can be compared with the actual development of the lesion. Therefore, the accuracy of the predictions cannot be effectively verified. The present invention makes it possible to verify the accuracy of the synthesis model over repeated examinations by comparing further examination image data with the generated synthesis image data. This is of great benefit because it allows the performance of the model to be evaluated and, if necessary, adjustments to be made.
[0081] In particular, this allows researchers to determine the extent to which lifestyle habits and other external factors can influence the development of a disease, be it lesions, kidney disease, or multiple sclerosis. For example, if a person changes their diet or activity level after their last examination, this could lead to a positive outcome. There are also often medical preventative measures that can make it possible to avoid, for example, a kidney transplant if appropriate measures are taken in time.
[0082] Comparing recent imaging data with synthesis data predicted based on previous imaging data allows us to determine whether the actual progression deviates from the prediction and whether this deviation is positive. This not only supports monitoring disease progression but also the evaluation of the impact of interventions or lifestyle changes on the disease course. The ability to perform such comparisons thus provides valuable insights into the effectiveness of treatment and prevention strategies and enables adjustments to medical care.
[0083] According to a further advantageous embodiment, the examination image data capture at least one pathological tissue change, and the generation of synthesis image data comprises: creating a time series of synthesis image data, wherein the synthesis time points correspond to the examination time points of regularly or irregularly performed examinations. Based on the generated synthesis image data, the next examination time points are recalculated, depending on the synthesis time point in the time series and a temporal development of the pathological tissue changes in the synthesis image data, in particular a shortening of the intervals between examinations if the temporal development of the pathological tissue changes indicates a deterioration of a health condition.
[0084] Currently, no image-based prognostication methods exist for pathological tissue changes. As a result, preventive examinations and disease monitoring are usually performed at fixed intervals according to medical guidelines, without considering individual risk factors or the specific disease progression of patients, with the exception of special programs for high-risk patients, such as individuals with genetic mutations. These fixed intervals can lead to examinations being performed less frequently than would be optimal for a specific patient, resulting in an increased risk of developing so-called "interval cancers." It would be particularly beneficial for high-risk patients, who, for example, are examined every six months, to know whether they have an increased risk of interval cancers.Interval cancers are cancers that occur between scheduled screening examinations and are therefore potentially not detected early enough. The present invention makes it possible to individually adjust the next examination time points based on the generated synthesis image data. By predicting when a lesion might mutate malignantly, the examination time can be adjusted so that it occurs before the malignant tissue change, or, in particular, very shortly afterward, so that the malignant tissue change can already be identified in the image. Alternatively, a tissue sample can be taken to check whether a malignant tissue change has already occurred and, if so, the still small malignant tissue change can be removed directly. This enables earlier detection of potential malignant tissue changes and thus drastically reduces the risk of interval cancers.
[0085] Adjusting the screening interval based on individual risk factors not only leads to more efficient preventative care but also improves patient health by enabling early diagnosis and treatment of diseases. This can reduce the burden on the healthcare system from advanced stages of illness and allow patients to receive treatment more quickly. The screening interval refers to the time between scheduled examinations designed to regularly monitor a specific disease or area of the body. In the case of mammography, a specific type of routine examination, the screening interval is determined according to applicable national medical guidelines, for example, every two years for women between the ages of 40 and 75.This examination is primarily intended to reduce the risk of advanced breast cancer.
[0086] For conditions where medical guidelines do not mandate imaging, individualized measures are often available for high-risk patients, which may include imaging procedures, particularly if, for example, early cognitive impairment is already present. Sufficient imaging data may also exist for a patient that was originally acquired for other purposes but could reveal a condition incidentally. This is especially true for diseases where early diagnosis is rare, such as Alzheimer's disease.
[0087] Another preferred embodiment of the above method also includes controlling the synthesis model, by means of the control model, to generate multiple sets of synthesis image data. Here, a different set of development-specific context information is used for each set of synthesis image data. The different sets of development-specific context information preferably include different drug information or active ingredient information, or they may also vary different synthesis time points or other possible development-specific context information.
[0088] The technical effect achieved through this design is the enabling of a direct, visual comparative analysis of alternative future scenarios. Conventional prognostic methods are limited to predicting a single, most probable disease course. The technical problem of how to generate a technically consistent and directly comparable data basis for different hypothetical scenarios (e.g., "Therapy A" vs. "Therapy B") for a single patient remains unsolved.
[0089] The proposed solution utilizes the system's targeted controllability to iteratively perform the synthesis process with different conditionings (varying development-specific contextual information). The technical effect lies in the generation of a novel data product: a collection of multiple, alternative "disease films" based on the same initial state. This creates, for the first time, a technically valid and visual foundation that enables a direct "what-if" comparison and thus provides the technical basis for a more informed subsequent evaluation of treatment strategies.
[0090] Another advantageous embodiment of the above procedure also includes:
[0091] The image data contains not only the examination time but also additional contextual information, particularly risk-relevant factors such as disease-specific biomarkers. In current medical guidelines, only high-risk contextual information is typically considered. Specialized risk calculators such as the Gail score or Tyrer-Cuzick are used for risk assessment using this high-risk contextual information. In several European countries, the presence of genetic mutations, such as BRCA1 / 2, is used as a criterion for classifying breast cancer risk as "high," instead of a risk calculator. Approximately 0.2–0.5% of all women carry such mutations. National medical guidelines define measures for women at increased risk, including, for example, additional examinations such as MRI scans and semi-annual screening examinations.This means that preventive examinations for the majority of the population cannot be individually tailored to each patient. As a result, there is a risk that patients with an increased risk of cancer are not examined frequently enough, which can lead to the development of so-called "interval cancers." Interval cancers are cancers that occur between two screening examinations and therefore may not be detected early. Every disease has specific, risk-relevant contextual information that may be more relevant for diagnosis than imaging, such as clinical tests in Alzheimer's disease. In some cases, imaging can even be replaced by such information, as is the case with prostate-specific antigen (PSA) screening for prostate cancer. The goal, with regard to prognosis, is primarily to identify rapidly progressing diseases at an early stage.For example, in the case of prostate cancer, the diagnosis is generally less important than determining the aggressiveness of the cancer, especially in older men who may find cancer treatment more difficult to tolerate.
[0092] By integrating risk-relevant contextual information into the synthesis model, the method can more accurately predict when a tissue change might undergo pathological processes, particularly in rapidly progressing diseases. This allows for individualized adjustment of examination intervals, enabling more frequent screening of higher-risk patients. This, in turn, drastically reduces the incidence of interval cancers, as potential malignant tissue changes can be identified and treated earlier.
[0093] This approach not only improves patient health by reducing the risk of advanced cancer stages, but also increases efficiency within the healthcare system. By allocating resources to the patients who need them most, this method optimizes the costs and benefits of healthcare.
[0094] Another advantageous embodiment of the above procedure also includes:
[0095] • Determining a comprehensive measure of therapy success:
[0096] o Generating initial synthesis image data, whereby initial drug information is entered into the synthesis model as further contextual information along with the lesion image data. The drug information specifies a drug used, administered in the period between the last acquired lesion image data and the initial synthesis image data.
[0097] o Generating second synthesis images, wherein when generating the second synthesis image data, second drug information is entered into the synthesis model as further context information along with the lesion image data, and wherein the second drug information indicates that no drug or a different drug is administered during the corresponding period.
[0098] o Generating the first and second synthesis image data involves entering risk-relevant context information as input data into the synthesis model.
[0099] o Determining the therapy success measure based on a difference between the first and second synthesis image data.
[0100] Accordingly, the treatment success metric indicates how effective a particular medication was compared to an alternative medication or no treatment. Thus, the above method enables an individualized assessment of the success of a therapy. It analyzes how different medications influence the development of a disease and thereby supports the identification of the optimal treatment method for the patient.
[0101] No known method currently exists that reliably enables the individual prediction of a treatment outcome. Typically, treating physicians select drugs based on available clinical evidence derived from clinical trials and the associated confirmed clinical experience. The image data generated in clinical trials are not usually available to treating physicians; instead, treatment outcome data are typically available in aggregated numerical form for the entire treated patient group. Therefore, individual prediction of treatment outcomes is not feasible under the current state of the art without a synthesis model and the synthesis image data generated by this model. However, the invention described in this patent provides the basis for making this possible.
[0102] Nowadays, physicians utilize pathological diagnostic methods, where available, for example in oncological diseases where a biopsy can be performed, to identify specific drugs for individual conditions and to create personalized treatment plans. However, the selection of drugs is often limited by clinically proven protocols, and a precise method for predicting the effectiveness of different medications while considering all individual patient data is lacking.
[0103] The present invention provides an improved method that refines the selection of individual active ingredients in therapy and systematically evaluates the effectiveness of the treatment. This contributes significantly to improving patient health.
[0104] The present invention makes it possible to systematically determine a therapeutic success metric by incorporating risk-relevant contextual information into the processing by the synthesis model. The method generates synthesis image data taking into account drug information and risk-relevant contextual factors. A difference is then determined between image data with and without the application of a specific active ingredient in order to objectively evaluate the efficacy of this active ingredient.
[0105] In addition, the method enables improved analysis of the influence of various active substances on the temporal development of the disease, as well as a more precise prognosis of disease progression. The described method makes a significant contribution to optimizing medical diagnostics and therapy selection through image-based prognosis, which describes the need for action by enabling a precise and individualized assessment of drug efficacy. This leads to a significant improvement in patient health and treatment outcomes. One aspect of this disclosure relates to a method for selecting subjects for a clinical trial for a new active substance, comprising:
[0106] • Determining the therapeutic success measure according to the above procedure for the new active ingredient for several subjects.
[0107] • Selecting subjects with the greatest therapeutic success factors for the clinical trial, characterized by using a set of contextual information as input data for each subject, including at least gender information, age information and skin color.
[0108] This capability is not only crucial for individualized patient treatment but also for the development of new drugs. It allows for the use of so-called "virtual control arms," which represent an efficient and ethically advantageous alternative to traditional control groups. Drug development and clinical trials often define image-based endpoints, such as hippocampal atrophy, which is measured using brain MRI. The use of such virtual control arms eliminates the need to recruit additional patients for a placebo group. This allows clinical trials to be completed more quickly, increasing research efficiency and accelerating the availability of new drugs.
[0109] If, in the above-described procedure for determining a therapeutic success measure, no drug information is used when determining one of the synthesis image data points, or if the drug information is selected to indicate that no medication was administered, then this synthesis image data is also referred to as virtual control arms. The virtual control arms are based on real image data acquired before the start of the clinical trial. With this data, disease progression in patients is simulated as if they were not receiving treatment with the new drug. These simulated courses serve as an objective basis for comparison to evaluate the efficacy of the tested therapy. In contrast to conventional placebo groups, which are often subject to challenges such as selective patient selection or potential non-adherence (e.g.,Since treatment outcomes can be influenced by factors such as non-adherence to treatment plans, virtual control arms offer a consistent, reproducible, and unbiased basis for comparison. A crucial advantage of virtual control arms lies in their more precise representation of actual disease progression and treatment success. The ability to simulate individual disease courses allows for a more accurate isolation of the treatment's influence. This significantly increases the accuracy and relevance of the study results. Retrospectively, it would also be possible to compare the two arms of a clinical trial—treatment and placebo—to determine whether a similar rate of, for example, brain atrophy was to be expected. Such an analysis could retrospectively verify whether the treatment success criteria were correctly recorded and measured. This is particularly important for Phase I and Phase II trials, in which the number of participants is often smaller.
[0110] In addition, virtual control arms offer the possibility of incorporating synthesis data into study design. This allows for a detailed differentiation of patient groups and leads to a more precise evaluation of treatment outcomes. At the same time, they offer the flexibility to react more quickly to a lack of therapeutic success and to adapt study designs accordingly.
[0111] Another key advantage of virtual control arms lies in the ability to adjust or pause clinical trials early if an insufficient therapeutic response is observed. This is particularly relevant if it is determined early on that the treated subjects are not achieving the expected clinical progress. In such cases, early termination or adjustment of the study can help conserve resources and minimize potential risks for participants. Furthermore, the rapid identification of ineffective therapies enables the faster development of improved treatment approaches, as flawed hypotheses or ineffective methods can be eliminated early on.
[0112] In summary, virtual control arms offer an innovative solution to significantly reduce the duration, costs, and ethical challenges of clinical trials. At the same time, they contribute to improving the validity and reliability of the results, thereby sustainably benefiting both scientific research and clinical practice.
[0113] In summary, this method significantly contributes to improving drug development, thus allowing for a much better determination of the efficacy and safety of these drugs.
[0114] Another aspect of the present disclosure concerns a method for determining a differential prognostic index. This method comprises the previously described method for processing examination image data, wherein at least two sets of synthesis image data are generated for each of the different sets of development-specific context information.The procedure then comprises the following steps: performing a computer-implemented comparison between these two sets of synthesis image data to identify quantifiable differences in at least one or more of the following parameters derived from the image data: volume or growth rate of a pathological tissue change, shape or texture characteristics of a pathological tissue change, morphology or density of microcalcifications, or vascularization as determined by contrast enhancement; and determining the differential prognostic index based on the identified quantifiable differences. In a particularly advantageous embodiment, this index is determined as a technical surrogate marker for clinically relevant endpoints.For this purpose, the computer-implemented comparison includes not only the synthesis image data itself, but can also be based on internal state representations, object-specific context information, and / or development-specific context information used for generation. Through this multimodal analysis, the index is calibrated to show a high correlation with long-term clinical endpoints, such as radiological response (rCR / rPFS), pathological complete response (pCR), or overall survival (OS).
[0115] The technical effect achieved through this method goes far beyond simply providing a numerical value and addresses a fundamental problem of the state of the art: the "black box" nature of many AI-based predictive tools. Conventional systems often deliver only an abstract risk score without a rationale that is comprehensible to the user. This leads to a lack of transparency and trust, which hinders clinical acceptance.
[0116] The present invention solves this problem in an inventive way by generating the differential prognostic index not in isolation, but inextricably coupled with the underlying synthesis image data. The technical effect of this coupling is the creation of an inherently explainable and interpretable prognostic system. For a physician, the index is thus no longer an abstract number. Instead, the system enables direct visual verification: A high index, indicating a strong therapeutic effect, is visually supported by synthesis image data showing a significant reduction in tumor volume. A low index is explained by images showing only minor changes.This direct link between quantitative metrics and qualitative, visual evidence fosters trust in the algorithmic prediction and allows the user to understand the morphological reasons for the value calculated by the system, leading to more informed clinical decision-making. Beyond plausibility checks, the synthesized images also allow for a visualization of the uncertainty of the predictive index.
[0117] Another preferred embodiment of the above procedure includes an additional step performed before generating the internal state representation. This step involves automatically extracting additional parameters directly from the examination image data using an image analysis algorithm and subsequently adding these extracted parameters to the object-specific or development-specific contextual information. The extracted parameters include at least one or more pieces of information from the group consisting of general tissue or organ parameters, such as organ size, breast size, or breast tissue density, and parameters that characterize the pathological tissue changes, such as functional tumor volume (FTV), which quantifies the metabolically or vascularly active tumor burden, tissue morphology and architecture (e.g.,These include roundness, spiculation), texture features, or vascular parameters, especially perfusion, permeability, or the spatial distribution of vascularization within the pathological change.
[0118] The technical effect achieved through this design is the automatic enrichment and objectification of the input data, leading to a more robust and accurate internal state representation.
[0119] Manually recorded or report-derived contextual information can be incomplete or subject to subjective variations. This design solves this technical problem by introducing an automated image analysis step that extracts quantitative, reproducible, and objective parameters directly from the image data itself. Instead of relying solely on external input, the system thus generates some of its own high-quality contextual information.
[0120] This enriched and objectified dataset, fed into the coding model, enables the generation of a technically superior and more information-dense internal state representation. A more precise description of the initial state, in turn, improves the accuracy and reliability of the entire subsequent simulation process, as the control model can operate on a more robust and detailed data foundation.
[0121] Another preferred embodiment of the above method is characterized by hybrid control of the generation process, in which the final control instruction that guides the synthesis model is formed from both data-driven patterns and explicit boundary conditions based on mechanistic or physical rules of disease progression. This achieves a significant increase in the robustness, physical correctness, and controllability of the simulation. Purely data-driven models excel at learning complex patterns from training data but reach their limits when confronted with scenarios outside their training distribution. This can lead to implausible results or "hallucinations" (e.g., a tumor growing through bone tissue).This design solves this technical problem by restricting creative, data-driven prediction through a deterministic, rule-based framework.
[0122] This combination acts as a technical guardrail for the generation process. It ensures that the simulation always adheres to the fundamental laws of biology and physics. This prevents the generation of technically or biologically impossible synthesis patterns and makes the overall system significantly more reliable and trustworthy. Furthermore, it enables finer, rule-based control (e.g., "simulate with 10% less dose") that goes beyond mere pattern recognition.
[0123] Conceptually, the implementation of these rule-based boundary conditions is realized through an additional model. This technical module operates in parallel or integrated with the data-driven control model and serves to apply explicitly defined mechanistic or physical knowledge to the simulation process. It acts as a plausibility and correction engine. The way in which this hybrid control is implemented can vary depending on the underlying architecture of the synthesis model.
[0124] • Implementation for diffusion models (e.g., with ControlNet): In this implementation, the data-driven control model and the rule-based supplementary model can be implemented as two separate neural networks. The control model generates an initial, highly detailed, but potentially error-prone control instruction. In parallel, the supplementary model generates a second, rule-based control instruction (e.g., a correction vector). These two instructions are then combined (e.g., by weighted addition) to form the final control instruction, which then guides the denoising process of the diffusion model. • Implementation for transformer-based models: In architectures such as video transformers, a separate supplementary model is often not required. Instead, the rule-based boundary conditions are integrated directly into the architecture of the control model.This can be achieved, for example, by formulating the rules as conditioned embeddings or bias terms and feeding them into the attention blocks of the transformer. The final control instruction then arises implicitly from the internal calculations of the conditioned transformer.
[0125] • Implementation in ODE-based models (e.g., ImageFlowNet): In models based on ordinary differential equations (ODEs), the temporal evolution is simulated by integrating a vector field. Here, hybrid control is realized by directly restricting or modifying the derivative function (the vector field) of the ODE system through rule-based boundary conditions. The additional model acts as a component that corrects the vector field so that the resulting trajectory obeys the laws of physics.
[0126] Regardless of the specific architectural implementation, the rules represented by the additional model or the integrated boundary conditions can encompass various aspects of the disease progression, such as:
[0127] • Dose-response relationship rules: A mathematical function that describes how a change in drug dose affects the tumor growth rate.
[0128] • Rules regarding growth dynamics: Basic oncological growth models (e.g., Gompertz or exponential growth) as lower or upper limits. • Rules regarding anatomical boundaries: Knowledge of anatomy that prevents tumor growth beyond organ boundaries (e.g., bones, major blood vessels).
[0129] • Rules for therapy logic: Simple logical conditions, such as "Set the effect of anti-HER2 therapy to zero if the object-specific context information indicates a HER2-negative status."
[0130] One aspect of the present disclosure relates to a synthesis model for processing examination image data, wherein the examination image data capture at least one tissue alteration area of an examination subject encompassing a potentially pathological tissue change, in particular for use in the above procedure, comprising:
[0131] • an input layer configured to receive examination image data of an examination object with time information and optional context information, wherein the examination image data includes, in particular, tissue change image data corresponding to the tissue change areas, • processing and output layers, wherein the processing and output layers are configured to process the examination image data, in particular to process the marked tissue change image data in a prioritized manner, and to output synthesis image data.
[0132] In the context of the present invention, an "output layer" of a machine learning model, particularly a neural network, is a final layer among several layers of the machine learning model. The output layer outputs the synthesized image data along with contextual information, including the time information (output date). In the current state of the art, no image data is synthesized or generated for analysis or diagnosis. Therefore, it is not possible to make image-based predictions about future tissue changes and the development of lesions such as breast cancer. Instead, only numerical risk assessments for the occurrence of lesions can be provided. Without the generation of synthesized image data, however, it remains difficult to predict when or if interval cancers will occur.These tumors can develop between two screening or follow-up examinations and are often diagnosed late, which can impair the effectiveness and, in particular, the success of therapies.
[0133] The present invention enables the prediction of the further development of a disease, e.g., the lesion, by generating synthesis image data using a synthesis model. The model receives examination image data with temporal information and optionally contextual information and processes this data to generate synthesis image data for future time points. This synthesis image data can then be used to optimize the examination interval and the choice of therapy.
[0134] By considering the generated synthesis image data, physicians can adjust the timing of subsequent examinations accordingly, enabling the earliest possible diagnosis. This significantly reduces the risk of overlooking malignant tissue changes and allows physicians to detect interval cancers earlier and more effectively. Furthermore, the predictions not only support diagnosis but also the further course of a disease, thus facilitating decisions regarding treatment selection and success.
[0135] In summary, the synthesis model represents a significant advancement in the diagnosis and treatment of diseases. Precise, image-based predictions can have a positive impact on patient health, particularly for patients with rapidly progressing diseases.
[0136] One aspect of the present disclosure relates to a computer-implemented method for training a synthesis model to process examination image data, wherein the examination image data capture at least one tissue alteration area of an examination subject encompassing a potentially pathological tissue change, in particular the synthesis model used in the method described above, comprising:
[0137] • Receiving training data.
[0138] • Processing image data included in the training data; and • Adjusting the model parameters of the synthesis model to optimize an objective function that captures a difference between the synthesis image data output by the synthesis model and the annotations contained in the training data, wherein the training data includes image data of a study object for at least three different examination time points, as well as time information associated with the image data for each examination time point. The synthesis model processes the image data of at least one first examination time point, preferably from a first and a second examination time point, and generates corresponding synthesis image data for a third examination time point, while the annotations correspond to the image data of the third examination time point.
[0139] In current technology, trained models are primarily used for classifying or segmenting existing images, but not for generating new images for a specific subject or patient. This leads to a deficit in precise predictions about the potential future course of diseases, particularly in the early detection of malignant tissue changes such as interval cancers. Since image synthesis is not performed, generating precise prognoses is difficult, which significantly hinders the early identification of potential malignant tissue changes.In the case of incurable diseases such as Alzheimer's, for which only disease-modifying treatments exist, predicting the course of the disease – especially the speed of disease progression – is extremely valuable information not only for the patients themselves, but also for their relatives: A more accurate prediction of the course has a positive impact on the quality of life of the patients and their relatives, as it enables well-founded planning of support measures and care needs.
[0140] For the purposes of the present invention, "model parameters" are parameters of a machine learning model that determine the calculation of output data from input data of the machine learning model. During the training of the machine learning model, the model parameters are adjusted so that the output data of the machine learning model matches the desired outputs, the so-called target data or annotations, as closely as possible; that is, the output data matches the annotations as closely as possible. The machine learning model learns a desired mapping by appropriately adjusting the model parameters.
[0141] The present disclosure enables the training of a generative model incorporating temporal information. This makes it possible to generate synthesis image data for future time points, allowing for precise predictions about disease progression. This not only helps to determine the optimal time for diagnosis earlier, but also to improve the effectiveness of therapies.
[0142] In summary, this method makes a significant contribution to improving the early detection of tissue changes. By training a generative model that incorporates temporal information, reliable predictions about potential disease-related tissue changes can be made, ultimately leading to a positive impact on patients' health.
[0143] One aspect of the present disclosure relates to a computer-implemented method for generating a training dataset for training a synthesis model, in particular the synthesis model according to the model described above, comprising:
[0144] • Filtering a database of examination image data of examination objects, filtering out examination image data of examination objects for which at least examination image data for at least three different examination times are available.
[0145] • Determining lesion image data from the examination image data.
[0146] • Saving the lesion image data as a training dataset.
[0147] Extensive medical databases already exist, containing, for example, images from mammography screenings. These databases often comprise hundreds of thousands or even millions of image data points with detailed information at various examination time points. However, the number of image data points that include disease progression is often limited, as, for example, only about 0.6% of mammography screenings lead to a cancer diagnosis. The training data includes not only the image files themselves but also annotated contextual information, such as patient information, diagnostic results, and clinical assessments.
[0148] These databases are traditionally used by medical professionals to analyze and assess diseases and can be used to develop image processing and recognition software. However, the data are often so extensive that it is difficult to efficiently create complete digital twins with generative models, such as diffusion models. The method described above, however, can generate a suitable training dataset for a synthesis model based on these databases.
[0149] By filtering and selecting relevant image data, the training data can be restricted to the contextual material important for malignant tissue changes. This allows for the selection of specific image files that are significant for predicting the development of tissue changes over different time periods. In contrast to a classification model that focuses on the presence or absence of a disease, the focus here is on how quickly the disease progresses. Therefore, the training data must have sufficient temporal variability to enable imprecise predictions about disease progression and its associated dynamics.
[0150] This makes it possible to train generative models efficiently and precisely, which contributes to improved progression prediction. Reducing the data to relevant contextual information simplifies training and improves predictive quality by considering only the necessary data.
[0151] Another aspect of the present disclosure relates to a method for training an identification model, in particular an identification model implemented as a coding model. The method includes receiving training data, which for a multitude of investigation objects comprises a time series of examination image data, associated contextual information, and annotations, wherein the annotations classify the investigation objects based on clinical or biological characteristics. This is followed by the processing of the training data by the identification model to create an internal representation for each examination image and to generate a reconstructed version of the examination image from this representation. Finally, the model parameters of the identification model are adjusted by optimizing an objective function that includes at least one of the following components: an overlap component,which measures spatial agreement; a classification component that evaluates pixel-wise assignment accuracy; optionally, a weighting component to account for class imbalances; a reconstruction loss component that minimizes pixel-wise differences; a perceptual loss component that minimizes differences between high-level feature representations; a KL divergence loss component that regularizes the latent space distribution; an adversarial loss component that improves image quality using a discriminator; a diffusion loss component that minimizes the noise prediction error of a conditioned diffusion model; another reconstruction loss component that minimizes the difference between original and reconstructed images; a trajectory loss component,which minimizes the distance between internal state representations of temporally adjacent investigations; and a similarity loss component that, based on the annotations, approximates the internal state representations of objects classified as similar and removes those of dissimilar objects to semantically structure the multiscale feature space. The technical effect achieved by this specific training procedure is the generation of an intelligent, highly structured, and semantically ordered feature space. In the prior art, coding models are often trained with a single objective function, resulting in an unstructured, "chaotic" feature space that is unsuitable for precise control. The present invention solves this fundamental technical problem by optimizing an innovative, multi-component objective function. This "toolkit" of specialized loss components makes it possible toThe training objective must be precisely tailored to the desired architecture and application. In particular, the combination of trajectory and similarity loss components forces the model to learn an explicit temporal and semantic order. The result is not merely a data-compressed representation, but an intelligent "map" of disease states, in which the position and path of a representation have direct clinical and biological significance. This learned structure is the crucial technical prerequisite that enables the downstream control model to perform its function of precisely controlling the synthesis process based on position.
[0152] Another preferred embodiment involves a computer-implemented method for generating a training dataset for training an identification model, in particular an identification model implemented as a coding model. The method includes filtering a dataset to obtain image data of subjects for which image data is available for at least two different examination time points. Furthermore, the method includes assigning annotations to the filtered image data, wherein the annotations classify the subjects based on clinical or biological characteristics relevant for the semantic structuring of a multiscale space. Finally, the method includes storing the filtered image data together with the associated contextual information and the assigned annotations as a training dataset.
[0153] The technical effect achieved through this design is the provision of a structurally and semantically enriched training dataset, which serves as a technical prerequisite for the innovative training procedure of the coding model.
[0154] Standard datasets for generative models are often not designed for longitudinal analyses or for learning clinically relevant similarity relationships. This method solves the technical problem of how to generate a highly specialized training dataset from a general database, enabling the optimization of a multi-component objective function.
[0155] The technical effect manifests itself in two steps:
[0156] 1. Filtering for longitudinal data with at least two time points ensures that the dataset has the necessary temporal depth to optimize the trajectory loss component during training. 2. Explicitly assigning annotations enriches the dataset with the crucial labels required for optimizing the similarity loss component.
[0157] Thus, this method creates the specific, technically necessary data basis that enables the coding model to generate the intelligent and semantically ordered multiscale space in the first place.
[0158] To semantically structure the multiscale space, clinical or biological characteristics that significantly influence the behavior or prognosis of the disease are preferentially used when assigning annotations. These annotations serve as "labels" for the training process, helping users learn which objects of study should be grouped as "similar" within the multiscale space. Relevant annotations include, in particular:
[0159] • Tumor biological characteristics: These are the fundamental drivers of disease behavior.
[0160] The HER2 status (positive / negative) and hormone receptor status (HR) are crucial, as they fundamentally determine the response to specific targeted or hormonal therapies. The coding model learns to represent these biologically distinct tumor types in separate regions of the multiscale space.
[0161] The Ki-67 proliferation marker serves as a measure of the cell division rate. Annotation based on high vs. low Ki-67 values allows the model to distinguish aggressive, fast-growing tumors from indolent, slow-growing ones and to group their trajectories accordingly.
[0162] • Genetic and genomic traits:
[0163] The presence of specific DNA mutations such as BRCA1 / 2. This annotation helps to distinguish high-risk patients with a hereditary predisposition from the general population.
[0164] Results from genomic signatures such as the MammaPrint test, which provide a comprehensive risk profile, are used. The model learns to position patients with high and low recurrence risk in multiscale space according to their genomic signature.
[0165] • Clinical endpoints and classifications:
[0166] The pathological complete response (pCR) achieved from past treatments. By annotating patients as "responders" (pCR=1) and "non-responders" (pCR=0), the model learns to identify the subtle features that correlate with a successful treatment response. The histological grade or TNM stage of the disease. These annotations allow for structuring the multiscale space according to the aggressiveness and progression of the disease at the time of diagnosis. By using these annotations in the training process, the multiscale space is transformed from a simple compression result into an intelligent, clinically relevant "map" where the position of an internal state representation directly reflects its biological and clinical significance.
[0167] Another aspect of the present disclosure relates to a data processing device comprising a processor configured to perform one or more of the above methods.
[0168] Another aspect of the present disclosure relates to a computer program product comprising instructions which, when the program is executed by a computer, cause it to perform one or more of the above methods.
[0169] Another aspect of the present disclosure relates to a computer-readable storage medium comprising instructions which, when the program is executed by a computer, cause it to perform one or more of the above methods.
[0170] Here and in the following, various types of contextual information are discussed. The respective contextual information can include the object of investigation, the time of investigation, but also very different information about the object of investigation, the examination image data, the tissue change image data, or the synthesis image data. For the purposes of the present invention, "contextual information" is information that influences processing by means of a machine learning model or is processed as part of the input data by a machine learning model and taken into account during processing. This can include one or more of the following:
[0171] • For each data point, also temporal information, e.g.:
[0172] o Time of image capture
[0173] o Time of diagnosis
[0174] o Course of symptoms or measurements over time
[0175] • Patient-specific or subject-specific data:
[0176] o Demographic information (age, gender, ethnicity) o Family history
[0177] o In women: Pre- or post-menopause in women, number of pregnancies.
[0178] Lifestyle factors (e.g., diet, physical activity, smoking, alcohol consumption)
[0179] Environmental factors, e.g., geographical location with environmental pollution (e.g.,
[0180] Air quality, pollution exposure). Socioeconomic factors, e.g., level of education, occupation and working conditions, social environment and support system
[0181] • Clinical data, each relating to the subjects:
[0182] • Pre-existing conditions and comorbidities: Recording pre-existing conditions and comorbidities is crucial for a comprehensive assessment of a patient's health status. They provide valuable information for risk analysis and treatment planning, as certain conditions can increase the risk of other diseases. In clinical trials, it is of great importance to consider the history of pre-existing conditions such as cardiovascular diseases, diabetes, cancer, neurological disorders, and metabolic disorders, as these can influence the choice of treatment and the response to therapy. Identifying comorbidities also helps to monitor potential drug interactions and to individualize treatment. • Symptoms and their duration: Documenting symptoms, their intensity, and their duration is crucial for the diagnosis and monitoring of the disease.In clinical trials, this area is recorded in detail to assess the impact of therapy on symptoms and to evaluate its effectiveness. For example, in neurodegenerative diseases such as Alzheimer's or Parkinson's, the duration and severity of symptoms like memory loss, motor impairments, and cognitive deficits can influence the assessment of treatment success. In cancer treatments, symptoms such as pain, fatigue, or loss of appetite can also be important parameters. A widely used screening tool for the early detection of cognitive impairment, particularly in patients suspected of having neurodegenerative diseases, is MoCA. The MoCA test assesses a variety of cognitive abilities, including memory, attention, language, abstract reasoning, and spatial skills.MoCA is frequently used to assess the severity of cognitive deficits and as a starting point for monitoring disease progression. Combined with neuroimaging data such as MRI and PET scans, it is particularly useful in the diagnosis and monitoring of diseases like Alzheimer's, Parkinson's, and other dementias. The combination of MoCA and imaging techniques allows for a more precise assessment of disease progression, as changes in brain structure and function become detectable over time. This helps clinicians closely monitor disease progression and optimize treatment.
[0183] • Medical imaging: These procedures offer the possibility of visualizing the body's internal structures, identifying pathological changes, and assessing their extent and impact on adjacent tissues and organs. The most important imaging techniques include X-rays, computed tomography (CT), magnetic resonance imaging (MRI), and ultrasound. These techniques provide detailed information about organs, tissues, and the distribution of tumors. Imaging techniques offer insights not only into anatomical structure but also into the functionality of tissues, as with positron emission tomography (PET) or functional MRI scans. In addition to these techniques, microscopic imaging is of crucial importance, particularly for examining cells and tissues on a much smaller scale, which is described under the heading "Laboratory Values" here.Accurate interpretation of the imaging data is only possible by considering relevant contextual information. This includes:
[0184] RECIST: The Response Evaluation Criteria in Solid Tumors (RECIST) is a standardized and internationally recognized system for the objective measurement of tumor size and the assessment of treatment response. The application of RECIST in clinical practice is of central importance for monitoring the course of cancer, evaluating the response to therapy, and adjusting treatment accordingly. Accurate measurement and analysis of tumor size make it possible to detect disease progression early and optimize individualized therapy.
[0185] BI-RADS (Breast Imaging-Reporting and Data System), PI-RADS (Prostate Imaging-Reporting and Data System), and KI-RADS (Kidney Imaging Reporting and Data System): These systems serve the standardized classification of imaging findings in the examination of breast and prostate tissue. BI-RADS is used particularly in mammography, ultrasound, and MRI examinations to assess the risk of breast cancer based on imaging results. The BI-RADS classification ranges from 0 to 6, with higher scores indicating a greater likelihood of cancer. PI-RADS is used to assess the likelihood of prostate cancer based on MRI data. This classification helps to differentiate between benign and malignant changes in the prostate. Both systems are crucial for refining diagnoses and making targeted treatment decisions.AI-RADS is used to classify kidney lesions in imaging procedures, particularly CT and MRI, and helps assess the risk of kidney cancer. All three systems are crucial for refining diagnosis and making targeted treatment decisions.
[0186] Fazekas scale (for white matter lesions): Used in neuroimaging, particularly in MRI scans, to assess the severity of white matter lesions associated with diseases such as multiple sclerosis (MS), small vessel disease, and the aging process. The scale categorizes lesions into mild, moderate, and severe stages and helps clinicians assess disease progression.
[0187] Unified Parkinson's Disease Rating Scale (UPDRS): Frequently used to assess the severity of Parkinson's disease. Imaging techniques such as DAT (dopamine transporter) scans and SPECT (single-photon emission computed tomography) are used to assess dopamine activity and degeneration in the brain.
[0188] Other imaging findings: Other imaging findings include changes such as nipple retraction, asymmetry, tumor classifications, and calcifications, which can be identified by mammography or ultrasound. These changes play an important role in the diagnosis, monitoring, and risk assessment of breast cancer, as well as in the evaluation of other breast conditions. Nipple retraction and asymmetry can indicate pathological processes, while calcifications are often associated with benign or malignant tumors. Similar changes, such as asymmetry or tumor masses, can also be detected in other oncological diseases, such as lung cancer, colorectal cancer, or prostate cancer, using imaging techniques such as X-ray, CT, MRI, or ultrasound. In lung cancer, for example, nodules and lesions in the lungs are detected, while in colorectal cancer, changes in the lymph nodes or metastases are often visible.In neurological diagnostics, imaging findings such as changes in the brain, including white matter lesions, tumors, or vascular alterations, are crucial for the diagnosis and monitoring of diseases such as multiple sclerosis (MS), Alzheimer's disease, Parkinson's disease, and other neurodegenerative disorders. Regarding kidney diseases, imaging techniques such as ultrasound, CT, and MRI are used to identify changes such as kidney cysts, tumors, or the development of fibrosis.
[0189] Perspective: The precise perspective and imaging modality used during the examination are crucial for accurate diagnosis. This includes both selecting the appropriate imaging modality (e.g., MRI, CT, ultrasound) and considering the perspective when interpreting the images to enable the best possible assessment of the disease and treatment outcomes. Some examples of perspectives and their significance are: Axial perspective (horizontal section): Particularly common in CT and MRI scans to obtain a comprehensive view of organ structures and to visualize tumors or lesions in various areas of the body. Sagittal perspective (vertical section): Helpful in assessing changes in spinal cord structure, such as in neurological disorders, or in depicting tumors along the spine.Coronal perspective (frontal view): Frequently used to visualize organs such as the heart and lungs, and useful in assessing tumors in the breast, lungs, or abdomen. Craniocaudal perspective (CC): This perspective is also used in mammography and provides a top-down view of the breast. It helps assess the shape and volume of breast tissue and detect changes such as tumors or calcifications.
[0190] • Laboratory values: Medical imaging provides high-resolution anatomical and functional information and is therefore a central tool for the diagnosis, monitoring of disease progression, and treatment response. However, imaging alone is often insufficient to ensure a comprehensive assessment of disease progression and treatment response. Laboratory values and biomarkers, such as conventional tissue biopsies or modern liquid biopsies, play a crucial role in addition. Liquid biopsies enable non-invasive tumor monitoring and are particularly valuable for real-time tracking of tumor progression and treatment adjustment.They are based on the analysis of circulating tumor DNA (ctDNA), that is, DNA fragments released into the bloodstream by tumor cells. These fragments originate from apoptotic or necrotic tumor cells and circulate freely in the blood plasma. Diagnostic accuracy can be further increased by combining imaging and biopsies, either conventional or liquid biopsies.
[0191] Genetic information concerns DNA mutations, deletions, or amplifications that are detectable in the genome of a cell and directly affect disease biology.Examples of genetic information include BRCA1 / BRCA2 mutations, which increase the risk of breast and ovarian cancer; EGFR mutations, which are relevant for targeted therapies in lung cancer; BRAF V600E mutations, which are important for treatment decisions in melanoma and colorectal cancer; HTT (CAG repeat expansion) gene mutations, which are responsible for the development of Huntington's disease; APP, PSEN1, and PSEN2 mutations, which are inherited genetic alterations associated with familial Alzheimer's disease; GRN and C9orf72 mutations, which are associated with frontotemporal dementia (FTD) and amyotrophic lateral sclerosis (ALS); PKD1 mutations, which are associated with polycystic kidney disease; APOL1 mutations, which are associated with an increased risk of kidney disease in certain ethnic groups; and LMNA mutations, which are associated with dilated cardiomyopathy and are associated with heart failure.
[0192] Molecular markers encompass biochemical or genetic signals that are not solely based on DNA alterations but can also include protein expression, RNA levels, or epigenetic modifications. Examples of molecular markers include HER2 overexpression (a protein marker) as a decision factor for anti-HER2 therapies in breast cancer; the absence of estrogen, progesterone, and HER2 receptors in triple-negative breast cancer (TNBC); PD-L1 expression (an immune marker) as an indicator of response to immune checkpoint inhibitors; Ki-67 (a proliferation marker) for determining the cell division rate of a tumor; p53 protein status for assessing tumor suppressor function; tumor mutational burden (TMB), which can indicate treatment success; PSA (prostate-specific antigen) for the diagnosis and monitoring of prostate cancer; and microsatellite instability (MSI-H) as a biomarker for immunotherapy sensitivity in colorectal cancer.Tau protein as a biomarker for neurodegenerative processes in Alzheimer's disease, alpha-synuclein in Parkinson's disease, neurofilament proteins as markers for axonal degeneration in various neurodegenerative diseases, BNP (B-type natriuretic peptide) as a diagnostic indicator for heart failure, galectin-3, which is associated with heart failure and cardiac fibrosis, KIM-1 (kidney injury molecule-1) as a marker for acute kidney injury, and neutrophil gelatinase-associated lipocalin (NGAL) as an early marker for kidney damage.
[0193] • Treatment history: The treatment history includes a detailed record of the therapeutic measures carried out so far and their effects on the patient.
[0194] Previous therapies and their effectiveness: Documenting previously administered treatments is crucial for assessing their effectiveness in influencing disease progression. This includes both pharmacological therapies such as chemotherapy and immunotherapy (e.g., checkpoint inhibitors like Keytruda) and non-pharmacological measures such as surgery, radiation therapy, and targeted therapies. Treatment effectiveness is assessed and evaluated through clinical outcomes, imaging (e.g., MRI, CT, PET scans), and biomarkers (e.g., PD-L1 expression, TMB). Clinical outcomes include objective measurements such as tumor size, survival rates, disease progression, and response to therapy. These can be determined through routine examinations and tests such as blood tests, physical examinations, and clinical assessments.In addition to diagnostic imaging and biomarkers, clinical parameters such as patient well-being, functional status, quality of life, and the presence of symptoms contribute significantly to the overall assessment of treatment success. A complete list of medications used, including dosage, treatment duration, administration period, dosing interval, active ingredient, and dosage, provides a comprehensive overview of the patient's pharmacological treatment history and its potential impact on disease progression. This includes both approved medications used in clinical practice and experimental medications used in clinical trials.
[0195] o Currently used medications: It is documented which medications are currently being used as part of the ongoing treatment in order to monitor effectiveness and potential interactions compared to previous medications.
[0196] Contextual information can be categorized into different classes; accordingly, a class value can be entered when inputting contextual information. Information particularly helpful for therapy or its success is preferably entered as contextual information. Contextual information can be divided into two main categories: object-specific contextual information and development-specific contextual information.
[0197] The object-specific context information encompasses all static or very slowly changing data that characterize the object of study—that is, the patient and their fundamental disease biology—over the entire observation and simulation period. This information is typically compiled in an object-specific context file and provided to the system as structured input. This file can, for example, be in CSV (Comma-Separated Values) format, where each column or row corresponds to specific, predefined information. However, differently encoded files can also be used, as long as they exhibit the corresponding recurring structure. This structured format ensures that the processing model, particularly the coding model, can interpret and process the data unambiguously and correctly.
[0198] The essential object-specific contextual information includes, in particular:
[0199] • Patient-specific or subject-specific data: This includes basic demographic information such as age and gender, family history, and systemic physiological status.
[0200] • Pathology-specific genetic information: This includes information about known DNA mutations relevant to the respective disease.
[0201] • Pathology-specific molecular markers and laboratory values: These data describe the intrinsic properties of the tissue change.
[0202] This includes, for example, the molecular receptor status, the expression of certain proteins, or proliferation markers.
[0203] By separately recording and structuring these static, fundamental properties of the object under investigation, a robust and reproducible basis is created for the subsequent simulation of dynamic disease development.
[0204] For a particularly accurate and clinically relevant synthesis of prognostic image data, it is advantageous to use certain object-specific contextual information that provides a deeper biological and anatomical basis to the synthesis process.
[0205] In a preferred implementation, the initial magnetic resonance imaging (MRI) images, particularly T1- or T2-weighted scans, are used as object-specific contextual information. These serve as an essential visual anchor point, defining the initial morphology, size, and position of the tissue change from which future development is simulated.
[0206] In another preferred configuration, the object-specific contextual information includes fundamental, pathology-specific biological drivers of the disease. For example, in the case of breast cancer, this includes HER2 status and hormone receptor (HR) status. This information is crucial for the model to correctly interpret the effect of therapy-specific conditioning. For instance, the simulation of targeted anti-HER2 therapy can only be technically accurate if the model receives the information that a HER2-positive tumor is present.
[0207] In another preferred configuration, the object-specific contextual information includes the patient's sex, particularly in the case of sex-specific diseases such as breast cancer. This defines the basic anatomical and hormonal framework within which the trained model operates.
[0208] In a particularly preferred configuration, this information is combined. For example, the initial MRI images are used together with the HER2 and HR status as object-specific contextual information. This combination enables the system to perform a visually anchored simulation that is simultaneously biologically tailored to the specific behavior of the tumor.
[0209] In a particularly preferred configuration, all the aforementioned information is combined: The initial MRI images, the biological drivers (HER2 / HR status), and the sex are provided together in the object-specific context file. This comprehensive combination of visual, biological, and demographic data enables the synthesis model to generate highly personalized, accurate, and clinically relevant prognostic image data, as the baseline state of the subject is defined with maximum precision.
[0210] For practical implementation, the object-specific context file could be structured as a CSV file, in which each row corresponds to a unique study object (patient) and the columns define the specific characteristics. Such a structure ensures a structured and machine-readable input.
[0211] An example setup could look like this:
[0212] Patient_ID,Age,Gender,HER2_Status,HR_Status,MammaPrint_Score,MRT_Da-tei_PfadP001, 55, 1,1,0, 0.8, , / data / P001 / mri_t0.nii.gz P002, 62, 1,0,1, 0.2, . / data / P002 / mri_t0.nii. gz
[0213] Here, categorical data such as sex (e.g., 1 for female) or HER2 status (e.g., 1 for positive, 0 for negative) are numerically coded to make them processable by neural networks. The image data itself is not stored directly in the file; instead, a column such as MRI_File_Path points to the location of the corresponding image file. This machine-readable format ensures that the model can uniquely assign each piece of information to the correct patient and combine all data points as a unified input vector for processing in the coding model.
[0214] In contrast to static, object-specific contextual information, developmental contextual information describes the dynamic events and changes that occur in the time interval between two successive examination time points or between an examination time point and a synthesis time point to be simulated. It defines the trajectory of disease development and is crucial for teaching the model how specific interventions or time courses affect visual morphology. Similar to object-specific data, this information is typically captured in a structured format, such as a CSV file, to ensure machine processing.
[0215] In particular, all input data to be passed to the various models can be entered via a CSV file. This CSV file then contains the respective context information and a link to the image data to be read, so that the context information and the image data are, in effect, passed together via a single file. The essential development-specific context information includes, in particular:
[0216] • The timing information: This includes the examination time and the synthesis time, thus defining the exact time interval between the images. This is fundamental information for scaling the speed of the simulated development.
[0217] • The course of treatment within this interval: This explicitly includes drug information or active ingredient information. This includes the specific treatment type (e.g., chemotherapy such as AC-THP), the administered dosage, and the timing of the treatment. This information is crucial for conditioning the model to a specific treatment scenario and enabling counterfactual simulation.
[0218] • Measurable changes in the subject of study: These can be quantitative data such as the percentage change in tumor size or tumor volume within the interval, as recorded, for example, by RECIST criteria. This information serves as important validation and as additional input to describe the dynamics of the disease's progression to date.
[0219] The combination of this interval-related, dynamic information enables the system to precisely understand the previous course of the disease and, based on this, to create an accurate, conditional prognosis for the future.
[0220] For a technically meaningful and accurate simulation of disease development, it is advantageous to use certain development-specific contextual information that defines the dynamic framework of the synthesis.
[0221] In a preferred configuration, the development-specific contextual information includes the treatment type that was applied or is to be applied during the respective time interval. This information, which specifies, for example, a specific chemotherapy such as AC-THP, is crucial for conditioning the model to a concrete treatment scenario. Without this information, a treatment-specific prediction would be technically indeterminate, as the model would not know what kind of influence it is supposed to simulate.
[0222] In another preferred embodiment, the development-specific contextual information includes the pathological complete response (pCR) status, which quantifies the outcome of a previous treatment phase. This information serves the model as an important learnable endpoint and as an indicator of the tumor's previous response to therapies. In a particularly preferred embodiment, these two pieces of information are combined: The treatment type is provided together with the pCR status as development-specific contextual information. This combination enables the system to perform a simulation that is not only based on a specific treatment scenario but also takes into account the previously observed, individual response of the tumor. This significantly increases the accuracy and clinical relevance of the generated synthesis image data, as the model incorporates both the type of intervention and its previous effect into its calculations.
[0223] To process the dynamic, interval-based data automatically, a corresponding set of development-specific contextual information is provided for each set of examination image data at a specific examination time. This information does not describe a closed past, but rather defines the development up to the next, subsequent time point. The system thus models the disease progression as a sequential chain of states and transitions.
[0224] For practical implementation, this can be represented in a single CSV file for one study subject (patient), where each row corresponds to a time point in the examination. The columns then describe the development parameters that apply to the interval until the next time point.
[0225] An example structure for such a sequential file could look like this:
[0226] Patient JD, Time, Treatment type_for_next_interval, Interval_to_next_ZP_days, Tumor_detected P001, 2023-01-15, AC-THP, 186.0 P001 ,2023-07-20, None, 179.0 P001, 2024-01-15, **Alternative_Therapy_Y**, **365**, 1
[0227] Explanation of this sequential model:
[0228] • The first line (time 2023-01-15) contains the development-specific context information (treatment with AC-THP, interval of 186 days) that led to the creation of the actual examination image data of the next time point (2023-07-20).
[0229] • This continues for all real, past examinations. As you correctly point out, this information can accumulate over time. For example, if a tumor is discovered on January 15, 2024 (Tumor_discovered=1), additional tumor-specific data (e.g., RECIST scores) can be added for the next row.
[0230] The transition to synthesis:
[0231] The crucial point is the last line, which corresponds to the last actual point in time for the investigation. This line no longer contains a description of a past development, but rather the instruction for the future synthesis. In the example above, the last line describes the instruction to generate a forecast from the state of 2024-01-15.
[0232] to generate for 365 days under Alternative_Therapy_Y.
[0233] Flexibility for simulation:
[0234] As you correctly analyzed, this "last line" is completely flexible and interchangeable. To simulate different scenarios, it is simply modified or replaced at runtime:
[0235] • Scenario 1 (as above): Simulation with Alternative_Therapy_Y for 365 days.
[0236] • Scenario 2: The last line is changed to ...,**None**,**180**,1, to generate a simulation without treatment for 6 months.
[0237] • Scenario 3: The last line is changed to ...,**Existing_Therapy_Z**,**90**,1, to predict the course under the current therapy for the next 3 months.
[0238] This sequential model is technically superior because it depicts the course of the disease as a dynamic, adaptive process and makes it possible to generate an unlimited number of flexible, counterfactual future scenarios from any real endpoint.
[0239] Brief summary of the characters
[0240] The invention is explained in more detail below with reference to the examples shown in the drawings. The drawings show:
[0241] FIG. 1 schematically shows a processing device for use in a method according to an aspect according to at least one embodiment,
[0242] FIG. 2 schematically shows a synthesis model for use with a process
[0243] according to one embodiment;
[0244] FIG. 3 schematically shows one aspect according to at least one embodiment;
[0245] FIG. 4 schematically shows one aspect according to at least one embodiment;
[0246] FIG. 5 schematically shows one aspect according to at least one embodiment;
[0247] FIG. 6 schematically shows one aspect according to at least one embodiment;
[0248] FIG. 7 schematically shows an aspect according to at least one embodiment;
[0249] FIG. 8 schematically shows one aspect according to at least one embodiment;
[0250] FIG. 9 schematically shows one aspect according to at least one embodiment. Detailed description of the embodiments
[0251] One embodiment of an image processing system 100 comprises an image acquisition device 110 and a processing device 120, also called a control and / or evaluation device (see FIG. 1). The processing device 120 is communicatively connected to the image acquisition device 110, for example, via a wired or wireless communication link. The processing device 120 can evaluate image data 130 acquired by the image acquisition device 110 and control the image acquisition device 110, for example, based on the acquired and, if applicable, evaluated image data 130. In some embodiments of the present invention, an image acquisition device is, in particular, a mammography X-ray device, a magnetic resonance imaging (MRI) scanner, a PET scanner, or an ultrasound imaging device. The image acquisition device 110 can, for example, be an X-ray mammography device used to perform mammograms.Mammography is an imaging procedure that takes images of the breast, used for the early detection of breast cancer. In many countries, mammography screenings are offered or recommended for women from a certain age. These screenings take place at regular intervals, also known as examination intervals. The interval between two examinations is chosen based on various criteria.
[0252] Mammography screenings typically involve taking two standard images of each breast:
[0253] • Mediolateral oblique view (MLO projection):
[0254] o This oblique view captures the chest from the side and slightly from above, o It allows for good visualization of the axillary extension and the pectoralis muscle.
[0255] • Craniocaudal projection (CC projection):
[0256] This image is taken from top to bottom and compresses the breast from above.
[0257] Together, both projections provide a comprehensive view of the breast tissue and allow radiologists to assess potential abnormalities from different angles. In some cases, additional images may be required.
[0258] • In cases of large breasts (macromastia), several overlapping images may be necessary to depict all relevant areas.
[0259] • In special cases, a 90° lateral view (LM / ML view) can be taken to better assess certain findings, such as calcifications.
[0260] A distinction is then made between images of the right breast and images of the left breast; these are accordingly referred to as LMLO and RMLO projections, or as LCC and RCC projections.
[0261] As an alternative to x-ray-based imaging, other imaging techniques can also be used in mammography to enable more comprehensive diagnostics:
[0262] • Ultrasound (mammosonography):
[0263] It is often used as a supplement to mammography.
[0264] o particularly useful in cases of dense breast tissue and younger patients o can visualize small tissue changes better than mammography o increases cancer detection by about 40% compared to mammography alone
[0265] • Digital breast tomosynthesis (DBT):
[0266] a further development of classic 2D mammography
[0267] • Creates layered images of the breast and reduces superimposition artifacts • Can improve diagnostic accuracy and reduce unnecessary recalls • Magnetic resonance mammography (MRI):
[0268] o is used as an additional imaging technique
[0269] especially useful in certain diagnostic situations
[0270] • Breast computed tomography (mammography CT scan):
[0271] or another alternative imaging method
[0272] o is not routinely done due to the higher radiation exposure
[0273] Used in screening.
[0274] The choice of imaging technique depends on various factors, such as the patient's age, breast density, and specific diagnostic questions. Often, several methods are combined to obtain the most accurate diagnosis possible.
[0275] The processing unit 120 comprises an evaluation module 122, a storage module 124, and a control module 126. The modules can be designed as either software or hardware modules and are interconnected via channels 128 (see FIG. 1). The channels 128 are logical data connections between the individual modules. Alternatively, the modules can also be interconnected via a bus system.
[0276] The storage module 124 stores the image data 130 acquired by the image acquisition device 110. The image data 130 includes, in particular, examination image data 131, tissue alteration image data 132, and synthesis image data 133. The image data 130 can also include training data 135 for training processing models 150. The training data 135 includes training inputs 136 and annotations 137. Depending on the processing model 150 to be trained, the annotations 137 and the training inputs 136 can include examination image data 131, tissue alteration image data 132, and also synthesis image data 133. Furthermore, during the development of the processing models, model parameters 151 for processing models 150 can be stored in the storage module 124.
[0277] The evaluation module 122 is configured to process the image data 130 acquired by the image acquisition device 110. In particular, one or more processing models 150 for processing the image data 130 can be implemented in the evaluation module 122, for example a synthesis model 154, an identification model 152 or an evaluation model 156.
[0278] In accordance with the present invention, the control module 126 is configured to control the processing unit 120. The control module 126 reads data from the storage module 124 and forwards the data to the evaluation module 122. According to one embodiment, the processing unit 120 can also be configured to control the image acquisition unit 110. This includes, for example, the transfer of image data 130 to the storage module 124. In particular, the data can also be transferred back to the image acquisition unit. The return transmission of the synthesis image data 133 to the image acquisition unit 110 can, for example, be used to display the image data 130, view it, and use it for diagnostic purposes. According to another embodiment, however, the processing unit 120 can also be used independently of the image acquisition unit 110 solely for evaluating the image data 130.The image data 130 can, for example, be transmitted to the processing unit 120 via a suitable communication interface. In particular, the control module 126 is configured to instantiate suitable processing models 150 with suitable model parameters 151 in the evaluation module 122.
[0279] In particular, the processing unit 120 can also be configured with multiple computers. Specifically, the analysis and processing can take place on a cloud platform, while the diagnosis is then performed on the computer system in the hospital. Specifically, the processing unit 120 can be used to train machine learning models 200. During training, the training data 135 is processed, and the model parameters 151 of the respective trained machine learning model are adjusted. The training data includes training inputs 136 and annotations 137 corresponding to the image data. Once the training of a machine learning model 200 is complete, the processing unit 120 can be used to infer the machine learning model 200.
[0280] The various machine learning models 200 can be, in particular, one or more synthesis models 154, one or more identification models 152, or one or more evaluation models 156. The machine learning model 200 (see FIG. 2) can, in particular, be a multi-layered neural network, such as a convolutional neural network (CNN), a transformer, a generative adversarial network (GAN), a diffusion model, or flow matching. However, partially non-neural approaches, such as ordinary differential equations (ODEs), can also be used. A machine learning model must, however, be an image-based generative model in order to generate the synthesis image data.
[0281] The machine learning model 200 comprises an input layer 202, one or more intermediate layers 204, and an output layer 206. The input layer 202 receives input data 208, here the image data 130, processes the input data 208 by means of the input layer 202, the intermediate layers 204, and the output layer 206, and outputs result data 210. Depending on the type or implementation of processing models 150 used, the form and scope of the input data 208 and the result data 210 vary. For some machine learning models 200, intermediate data 212 can also be output. For the purposes of the present invention, an "intermediate data," also called an intermediate output or intermediate layer output, is an output of a layer of a multilayer processing model 150 that is not the last layer of the processing model, i.e., an intermediate layer.For the purposes of the present invention, an "intermediate layer" 204 is a layer that receives input data from a previous layer in a machine learning model 200, in particular a network or a neural network, and passes output data to a subsequent layer of the network. For the purposes of the present invention, an "input layer" 202 is a first layer of a multi-layered machine learning model 200, in particular a first layer of a neural network.
[0282] According to this embodiment, the machine learning models 200 can be implemented in particular as picture-to-picture models, as video transformers, as classifiers, as detectors or as regressors.
[0283] For the purposes of the present invention, a "picture-to-picture model" is a processing model 150 that is configured to perform a picture-to-picture mapping. The picture-to-picture model assigns a value in the result data 210 to each entry of an input data 208. The synthesis model 154 can, for example, be understood as a picture-to-picture model that assigns a value in the generated synthesis image 145 to each of the input data 208, for example, several images registered with each other or change image data.
[0284] For the purposes of the present invention, a video transformer is a processing model 150 that analyzes not only information about multiple frames but also temporal relationships between them. The video transformer assigns a value in the result data 210 to each entry of the input data 208. The synthesis model 154 can, for example, be supported by a video transformer that calculates pixels and associated values in the generated synthesis image 145 for each input data 208, such as multiple images registered relative to each other or tissue change image data.
[0285] For the purposes of the present invention, a “result date” 210 is a date output by a processing model 150, which is calculated and output by the processing model 150 by processing the input date 208. In a multi-layered model, a final layer of the model, the so-called output layer 106, also called the result layer, outputs the result date 210.
[0286] For the purposes of the present invention, the term "classifier" (also called classification model) refers to a processing model 200 that assigns a class to an input data 208 or can be trained to assign a class to the input data 208. The classifier can, in particular, be a machine learning model 200.
[0287] For the purposes of the present invention, a "detector," also called a detection model, is a machine learning model 200 trained to identify predetermined detection patterns in input data 208 and to output a list. In particular, the list is a list of localizations, for example, a localization in the input data 208. The input data 208 can, in particular, be an image, a stack of images, or an input tensor. The exact format of the localization depends, in particular, on the format of the input data 208 and the detection patterns to be identified.
[0288] For the purposes of the present invention, a "regressor," also called a regression model, is a processing model 150, in particular a machine learning model 200, that performs a regression. If the regressor is a machine learning model 200, it is trained to perform the regression, in particular by supervised or unsupervised learning.
[0289] In particular, the machine learning model 200 is configured to process or generate the image data 130 acquired by the image acquisition device 110. Depending on which of the processing models 150 is currently instantiated, different parts of the image data 130 are processed.
[0290] For the purposes of the present invention, "supervised learning" is a learning process in which a machine learning model 200 is trained to perform a desired mapping using an annotated data set.
[0291] For the purposes of the present invention, "unsupervised learning" of a machine learning model 200 is training or learning in which training takes place without specifying a desired goal, solely on the basis of a non-annotated data set, wherein the machine learning model 200 automatically finds, or should find, certain clustering points in the data set. Regardless of the implementation, the machine learning models 200 must be trained to execute a processing mapping. During the training of the machine learning model 200, the evaluation module 122, controlled by the control module 126, reads a portion of the training data 135 from the memory module 124 and inputs the read training data 135 into the respective machine learning model 200. The evaluation module 122 determines, based on the output data, theThe result data 210 of the machine learning model 200 and based on target data contained in the annotated data set, an objective function is defined and the objective function is optimized by adjusting the model parameters 151 of the machine learning model 200 based on the optimization of the objective function.
[0292] In particular, the objective function is optimized using a stochastic gradient descent method. In this method, only a small subset of the training data from the annotated dataset, called a batch, is used at any given time. The control module 126 determines the objective function for each input data 208 of the batch, based on the output data 210 generated by the machine learning model 200 and the annotation 137 corresponding to the input data 208. This objective function is a loss function that captures the difference between the output data 210 and the annotation 137. The control module 126 then calculates a gradient for each of the calculated objective functions with respect to the model parameters 151 of the machine learning model 200, sums the calculated gradients across the batch, and determines the mean.From the mean value, the control module 126 determines updated model parameters 151 for the machine learning model 200 by means of so-called backpropagation. The control module 126 then reinitializes the machine learning model 200 with the updated model parameters 151 in the evaluation module 122, and a next step of the stochastic gradient descent method is performed. For the purposes of the present invention, a "loss function" is a function that detects differences between the result data 210 and the specified target data. If the result data 210 and the target data are, for example, images, the comparison can be performed pixel by pixel. If the result data 210 and the target data are, for example, vectors or tensors, the difference can be determined by entry. The differences can be added in absolute terms (as absolute values) in an L1 loss function. In an L2 loss function, the sum of the squares of the differences is calculated.To minimize the loss function, the values of model parameters 151 of the processing model 150 are changed, which can be calculated, for example, by gradient descent and backpropagation.
[0293] The training of the machine learning model 200 terminates as soon as the optimization of the objective function results in the objective function reaching a predetermined limit. Once the training is complete, the control module 126 stores the last used model parameters 151 of the machine learning model 200 in the memory module 124, in particular together with contextual information, so that the newly trained machine learning model 200 can later be identified again and, for example, initialized for further training or inference.
[0294] As an alternative to the stochastic gradient descent method, other methods can also be used. In particular, any other training method can be used.
[0295] Once the training of a machine learning model 200 is complete, the corresponding model parameters 151 are stored in the memory module 124 and can later be read out in the inference to execute the learned processing mapping.
[0296] According to the present invention, machine learning models 200 are used, which are implemented as video transformers, classifiers, detectors, segmentation models, and image-based generative models. Furthermore, three different machine learning models 200 are used, each of which is trained with different training datasets to perform a processing operation, in particular an identification model 152, a synthesis model 154, and an evaluation model 156.
[0297] A method for processing examination image data 131 according to an embodiment of the present disclosure with reference to FIG. 3 to FIG. 6 is described below.
[0298] The method comprises steps 302 to 330. The method begins with step 302, receiving examination image data 131. According to the first embodiment, receiving the examination image data 131 is receiving examination image data 131 acquired with the image acquisition device 110.
[0299] The examination image data 131 are processed by the identification model 152; that is, the input data 208 of the identification model 152 are the examination image data 131. The processing of the examination image data 131 with the identification model 152 is shown schematically in FIG. 4. The processing comprises step 304, identifying, using the identification model 152, areas of tissue alteration in the input data, the so-called tissue alteration image data 132 in the examination image data 131. The identification model 152 is preferably implemented as a convolutional neural network (CNN) that has been trained to recognize certain patterns that correspond precisely to image patterns, for example, those caused by tissue alterations 141 such as lesions and by the tissue alteration areas surrounding the lesions in the examination image data 131.For this purpose, the folding layers of the CNN include certain feature maps by means of which the different types of tissue changes 141 can be recognized.
[0300] The identification of tissue change image data is illustrated by way of example in FIG. 4. The examination image data 131 comprise one or more examination images 140. The examination image data 131 can, as in the examination image 140 shown in FIG. 4 as input data 208, exhibit several tissue changes 141a, 141b, again lesions. For better understanding, tissue changes 141 and tissue change image areas 142 are marked as examples in the examination image 140 in FIG. 4. The tissue change image areas 142a, 142b arranged around the lesions 141a, 141b overlap in this example. However, the tissue alteration image areas 142 can also appear individually in the examination images 140, several tissue alteration image areas 142 can appear without overlapping, or an examination image 140 may show no lesions 141 at all, and accordingly, no tissue alteration image areas 142 are found. As shown in FIG.As shown in Figure 4, the identification model 152 would output the two tissue change image areas 142 as tissue change image data 132a, 132b, each together with the examination time, or mark them in the examination image data 131, such that the synthesis model preferentially processes the tissue change image data. The term "examination image" 140 refers to a single image captured during an examination with an image acquisition device 110, in particular the entire captured image.
[0301] According to the present disclosure, the investigation image data 131 can, in particular, each comprise pairs of MLO and CC projections. These are jointly entered into and processed by the identification model 152. Alternatively, the identification model 152 can process the different projection images each as a single investigation image 140, as shown in FIG. 4.
[0302] The identification model 152 can, for example, be implemented as a detector that then outputs a list of localizations of the tissue change image areas 142 in the respective examination image data 131 for the processed examination image data 131.
[0303] The processing further includes step 306, outputting the tissue alteration image data 132 as input data 208 for the synthesis model 154. If the identification model 152 is implemented as a detector, the list output by the identification model 152, together with the respective examination image data 131, can be further processed such that, based on the locations in the list, the tissue alteration image areas 142 in the respective examination image data 131 can be prioritized and processed by the synthesis model, for example, using an attention mechanism. The tissue alteration image data 132 also include examination time points as temporal information of the respective corresponding examination image data 131.
[0304] Alternatively, the identification model 152 can also be implemented as a transformer or without machine learning. For example, a transformer could process the examination image data 131 from several intervals sequentially, whereby the transformer then determines an output tissue change image data 132 based on several previously input examination image data 131.
[0305] Alternatively, the identification model 152 can also be implemented as a classifier that classifies image areas of the examination image data 131 and assigns each image section a class indicating whether or not a lesion 141 is detected in the respective image section. The image areas that detect lesions 141 can then be further processed as tissue change image data 132.
[0306] The procedure further comprises step 310, receiving the examination image data 131 from several examinations of a subject performed at different times, wherein the tissue change image data are appropriately marked in the examination image data 131 as described above. The received examination image data 131 also include time information about the examination times.
[0307] Figure 5 shows an example of tissue change image data 132. Although the synthesis model receives the examination image data 131 as input data 208, with the tissue change image data 132 marked in the input examination image data 131 as described above, only tissue change image areas 142a to 142f are shown here to illustrate the temporal progression of the tissue changes recorded in these image areas. The designations LCC-1 to LCC-3 and LMLO-1 to LMLO-3, which refer to the different projections of the left breast, are also shown in tissue change image areas 142.The pairs with the same number, here 1, 2, or 3, each belong to the same examination, with the tissue change image areas numbered 1 corresponding to the earliest examination and the tissue change image areas numbered 3 corresponding to the most recent examination. The tissue change image areas 142 of the same projections are registered relative to each other as closely as possible, such that at least the lesion 141 detected in each case is located at the same point in the different tissue change image areas 142 of the different examinations. Since the breast is repositioned in the image acquisition device 110 for each new examination, registration is only possible to the extent that the breast tissue is also located in the same position.
[0308] In Germany, all women between the ages of 50 and 75 are advised to participate in mammography screening every two years; that is, the typical monitoring interval for mammography screening is two years. Accordingly, the tissue change image data 132, for example, show a time interval of two years. If individual risk factors are present, a shorter monitoring interval can be set, for example, six months, one year, or even every three months, or any other suitable period. According to step 312, the synthesis model is configured to generate synthesis image data 133 for a synthesis time point from the input examination image data 131. According to the illustrated embodiment, a synthesis image 145a, 145b is generated for LCC and LMLO, respectively.The synthesis images 145 shown again only depict the tissue alteration areas in order to better visualize the development of the tissue change, even though the synthesis model synthesizes the complete images as in the examination image data. The synthesis time point can, for example, again have exactly the same nominal time interval as the previous examination interval, i.e., again 2 years. Alternatively, synthesis image data 133 with shorter intervals can also be generated, for example, to prevent the occurrence of an interval cancer from being missed. For example, one or more synthesis time points can also be entered into the synthesis model 154; synthesis images 145 are then generated for each synthesis time point according to the entered projections.
[0309] According to the illustrated embodiment, images of different projections, here LCC and LMLO images, are jointly input into the synthesis model 154, and the synthesis model 154 accordingly outputs synthesis image data 133 for different projections, again LCC and LMLO images. Alternatively, the different projections can also be processed independently of each other.
[0310] The image data 130 shown here are only examples; in addition to mammography images, other objects can also be used as subjects for examination. Figure 6 shows steps 320 to 324. Step 320 involves receiving the synthesis image data 133 using the evaluation model 156. The evaluation model 156 is configured to evaluate the synthesis image data 133. The input data 208 of the evaluation model 156 includes at least the synthesis image data 133 that the synthesis model 154 output. This is followed by step 322, which analyzes, also called processes, the synthesis image data 133. The evaluation model 156 is configured or trained to, for example, recognize signs of malignant tissue changes in the synthesis image data 133 and to create a synthesis evaluation 160, whereby the synthesis evaluation reflects a pathological escalation in the synthesis image data 133.For example, the pre-diagnosis score might be a number between 0 and 10, with a 10 indicating a high probability of a malignant, rapidly progressing tissue change. After diagnosis, this score is replaced by conventional methods such as RECIST (Response Evaluation Criteria in Solid Tumors) to accurately monitor and evaluate tumor size and response to treatment. RECIST is a system used to measure changes in tumor size, typically as part of assessing the effectiveness of treatment. It focuses on changes in tumor size over time, which can indicate whether a tumor is shrinking (response), growing (progression), or remaining the same (stable disease).For diseases where RECIST is not applicable, the number is used continuously to estimate the rate of disease progression, even after diagnosis. This applies, for example, to neurological diseases such as Alzheimer's. In step 324, the synthesis score 160 is output by the scoring model 156. The scoring model 156 can be implemented as a classical regression model, a classifier, or a spatial pattern recognition model using segmentation. It can output both continuous and discrete values. The scoring model 156 may have been trained using manually created annotations 137 from synthesis image data 133. It may also have been trained using annotations 137 from publicly available databases.
[0311] If pattern recognition in Synthesis Evaluation 160 indicates, for example, that a malignant tissue change is present, the examination interval can be shortened accordingly, for example by halving it, or, if the malignant tissue change is particularly aggressive, reduced to a few months. Synthesis Evaluation 160 also uses plausibility rules to check whether the Synthesis Image Data 133 have generated hallucinations. In this context, for example, it is checked whether the Synthesis Image Data of the left breast in the LCC and LMLO perspectives correctly reflect the tumor position.
[0312] According to one embodiment of the evaluation model 156, the evaluation model 156 can also be designed to process both the synthesis image data 133 and the examination image data 131, wherein the tissue change image data 132 are marked in the examination image data 131 as for synthesis using the identification model. For this embodiment, the evaluation model includes a transformer, in particular a video transformer such as a Temporal Recurrent Video Transformer, TRecViT, which processes the various image data 130 from the different examination time points, including the time information. For example, the transformer can include a segmenter. The evaluation model 156 outputs a segmentation mask in which different tissue types and tissue changes 131 are each assigned an instance of the segmentation mask. Each instance is also assigned a class.The different classes encompass, for example, various diseases, such as different types of cancer. FIG. 8 shows, as an example, the examination image data 131 and the synthesis image data 133, as well as one instance of the segmentation mask, in this case a breast carcinoma identified by the evaluation model 156 as synthesis evaluation 160, or as part of the synthesis evaluation 160. In FIG. 8, the same instance of the segmentation mask is also shown for the examination image data, in order to mark, as an example, the image area in which the breast carcinoma developed. This can also be displayed in this way for a therapist, who can then identify corresponding signs in the image area.
[0313] For each identified instance, its size can also be determined. From the determined size of the various instances, a rate of change can then be easily calculated in post-processing. For example, if it is a 3D image, the volume change of the tissue alteration is determined; otherwise, the area of the alteration can be determined, and if examination image data 131 from several examinations are available in addition to the synthesis image data 133, a rate of change can also be determined.If the evaluation model 156 detects during post-processing that a tumor volume is increasing at a higher rate according to the synthesis image data 133 than in the previously acquired examination image data 131, this growth is identified as critical. For example, the evaluation model 156 can also output the growth rate before and at the time of synthesis and mark it accordingly as a deterioration of the patient's health status. Thus, a synthesis assessment is issued indicating that a pathological escalation has occurred, and the monitoring interval can be shortened accordingly.
[0314] Furthermore, the transformer can also detect hallucinations or changes in tissue alteration 141 that are not medically relevant or medically implausible. For example, the synthesis data may include both LMLO and LCC images. In the clinical data used for training, the tissue alterations 141 appear on corresponding instances of the synthesis evaluation 160, as exemplified in Fig. 9 for LMLO and LCC images. If different developments of tissue alteration 141 are synthesized at corresponding locations in the LMLO and LCC images, the evaluation model recognizes this, since the clinical training data does not include such mutually implausible images, and can assign the synthesis image data to a corresponding class for implausible synthesis image data in this case.A similar principle applies, for example, if tissue changes 141 occur in locations within the subject where they cannot occur and where they cannot occur in the training data. Such synthesis image data 133 can also be recognized by the evaluation model 156 and labeled or classified accordingly. For such implausible or medically irrelevant tissue changes in the synthesis image data 133, corresponding catch-all classes can be provided within the possible classes to be assigned by the segmenter. Analogous exclusion or evaluation criteria also exist for other diseases or tissue changes. For example, in the case of Alzheimer's disease, which is diagnosed based on brain imaging data, if the hippocampal atrophy rate is determined by the evaluation model 156 and lies between 4 and 8%, the evaluation model 156 would assess this as a deterioration of the patient's health.
[0315] As described above for the first embodiment, image data 130 from mammography examinations were used. However, the described method can also be applied to other objects under investigation, provided that imaging examinations of the object under investigation are carried out in each case.
[0316] According to one embodiment of the first model, steps 302 to 306 and steps 320 to 324 are not performed using the identification model 152 and the evaluation model 156 as described above, but rather the identification is carried out manually or with the support of one of the machine learning models 200 known from the prior art. For example, a machine learning model 200 can identify a lesion 141, and a therapy renderer can then, for example, select the appropriate tissue change image data 132 and accordingly start the input into the synthesis model 154. The evaluation can also be performed by the therapy renderer.
[0317] According to a second embodiment, the present disclosure provides a method for determining a therapeutic success parameter, as shown in FIG. 7. The method begins with step 702, generating initial synthesis image data 133. To generate the initial synthesis image data 133, steps 302 to 312 of the first embodiment are performed, wherein, in addition, initial drug information is entered into the synthesis model 154 as further contextual information along with the tissue change image data 132 during the generation of the initial synthesis image data 133. The initial drug information specifies an active ingredient to be administered between a last examination time and the synthesis time. The contextual information can relate to…The active ingredient may, for example, specify a particular class of active ingredients; various classes of active ingredients include, in particular, disease-modifying drugs such as Herceptin, Tysabri, Opdivo, Lecanemab, or other active ingredients known in the prior art or combinations thereof.
[0318] The procedure further includes step 704, "Generating second synthesis image data 133," in which, during the generation of the second synthesis image data 133, second drug information, different from the first, is entered into the synthesis model 154 as context information. In addition to the drug information, further risk-relevant context information can also be entered into the synthesis model 154, whereby the same risk-relevant context information is entered in each case.
[0319] In step 706, "Determining the therapeutic success measure," the success achievable through therapy with the drug is quantified based on the difference between the first and second synthesis image data. Here, either the first and second synthesis image data can be compared, for example, pixel by pixel, to identify where changes occur due to therapy with the drug. Alternatively, steps 320 to 324 can be performed again for the first and second synthesis image data 133, and the therapeutic success measure can be determined based on the difference between the resulting first and second synthesis assessments 160.
[0320] The third embodiment relates to a method for selecting subjects for a clinical trial for a new active ingredient. For each set of possible or potential subjects, the method comprises the procedure according to the second embodiment, wherein, in addition to the active ingredient-specific context information, subject-specific context information is also entered into the synthesis model 154 together with the examination image data 131 and the tissue change image data 132 in order to predict the image-based endpoint for a clinical trial.
[0321] The tissue change image data 132 can be identified as in the first embodiment.
[0322] The contextual information about the new drug can include, for example, information about the drug class—tumor therapy encompasses a wide variety of drug classes—as well as information about its mechanism of action or similar details. The new drug could, for instance, be a chemotherapeutic agent specifically developed for the treatment of breast cancer; the drug may have been specifically developed for certain pre-existing conditions or for specific tumor types. This information is processed accordingly in the contextual information. Synthesis model 154 then determines the therapeutic success metric for different subjects and subject-specific contextual information, enabling the quantification of the therapeutic effect of the new drug on each potential subject, for example, using established criteria such as RECIST (“Response Evaluation Criteria in Solid Tumors”).For example, the lesions 141 of the potential subjects may already show malignant tissue changes, and tissue of the malignant tissue changes has been classified, for example, by means of a puncture of the breast; corresponding contextual information is also entered into the synthesis model 154.
[0323] By processing information from many potential participants, researchers can ultimately select those with the best therapeutic prospects for a clinical trial and conduct specific studies for particular patient groups. This increases the overall likelihood of bringing new, effective therapies to market. Furthermore, these predictions can serve as virtual control arms in clinical trials.
[0324] The fourth embodiment relates to a computer-implemented method for training a synthesis model 154 for processing image data 130.
[0325] The procedure includes the step of receiving training data 135. The training data 135 comprise examination image data 131 from at least two, preferably three examinations and the respective examination times.
[0326] The training further includes processing tissue change image data 132 contained in the training data and marked accordingly as described in the embodiments above, as well as adjusting the model parameters 151 of the synthesis model 154 to optimize an objective function that detects a difference between synthesis image data 133 output by the synthesis model and annotations 137 contained in the training data. The adjustment of the model parameters 151 during training can, for example, be performed using a stochastic gradient descent method with backpropagation.
[0327] Any other training method is also possible. The training data comprises image data 130 of an object under investigation for at least two, preferably three, different examination time points, as well as the corresponding time information for each examination time point. During training, the synthesis model 154 is always fed, for example, the image data 131 from two of the three examinations, each together with the time information, and the third examination time point is used as the synthesis time point for the synthesis model 154. Typically, the image data 131 from the two earlier examinations are entered as input data 208, and the image data 131 from the last examination is used as annotations 137. The difference between annotation 137 and result date 208 is recorded, for example, pixel by pixel.
[0328] Alternatively, not only predictions but also interpolations can be generated using the synthesis model 154. For this purpose, for example, the examination image data 131 from the first and last examinations are used as input data 208, and the synthesis time point is the mean of the examination time points. In this way, the synthesis model 154 can also be trained to accurately determine the time point of a detected malignant tissue change. The training data 135 are obtained from state-of-the-art databases. There are various databases worldwide in which temporally longitudinal image data 130 from mammography screenings are collected, for example, Cohort of Screen-age Women - Case control (CSAW-CC) and EMoryBrEast imaging Dataset (EMBED). These often also include annotations, sometimes even pixel-level, so these datasets can be used accordingly in the training.According to a further modification, the context information, as used in the second and third embodiments, is also input during training, so that the synthesis model 154 is also trained according to the respective context.
[0329] According to a further embodiment, different synthesis models 154 can be trained for different contexts. For example, specific synthesis models 154 can be trained for certain genetic predispositions. Alternatively, individual synthesis models 154 can be trained for specific active ingredients or classes of active ingredients. The appropriate synthesis model 154 is then selected for processing upon input of the corresponding context information.
[0330] A further feature concerns the training of the identification model 152 and the evaluation model 156. The databases described above also provide corresponding annotations 137 for these machine learning models 200. If, for example, the lesions 141 and the corresponding malignant tissue changes are identified in the databases, an image area of a predetermined size around the marked lesion 141 can be selected for training and then used to train the identification model 152.
[0331] Similar annotations 137 to the malignant tissue changes with corresponding ratings such as Ki67 ratings, a grading and similar rating criteria can be used to determine the synthesis rating 160 from the annotations and then the rating model 156 is trained on the synthesis ratings 160 derived in this way to output the synthesis rating 160.
[0332] A fifth embodiment according to the present disclosure comprises the method according to the first embodiment and additionally a step for receiving further examination image data 131, wherein the further examination image data 131 correspond to one or more further examinations. The further examinations have correspondingly different examination times; these further examination times can be entered into the synthesis model 154 as the synthesis times, which then generates synthesis image data 133 with corresponding synthesis times. This is followed by a step comparing the further examination image data 132 with the synthesis image data 133 generated for the respective examination times. From the comparison, it can be deduced, for example, whether the synthesis model 154 still reliably generates synthesis image data 133 or not.If the comparison reveals a lack of agreement, a new training session can be initiated, for example. A fifth embodiment relates to a method for training the evaluation model 156. Training data is again required for this purpose. Ultimately, the same databases used for training the synthesis model can be searched to determine the training data. The training data differs only in which image data is used as input data during training. According to a first embodiment, image data from a single examination is used as input data during training, and corresponding diagnoses or segmentations representing the diagnosis are used as annotations and thus as target data. According to an alternative embodiment, all image data from an examination subject is used as input data during training.The annotations are selected according to the desired output of assessment model 156. Assessment model 156 can, for example, generate segmentation masks for each examination time point, in which the tissue change is marked. Furthermore, the area, or in the case of three-dimensional image data, the volume, of the tissue change can be determined, as well as a rate of change. For many diseases, the rate of change provides information about the pathological progression of the respective disease and can directly serve as an indicator of the disease's course.
[0333] The variants and configurations described for the various figures can be combined with one another. The configurations shown and described are purely illustrative, and modifications thereof are possible within the scope of the attached claims. According to a further embodiment, the training data 135 are provided.
[0334] According to another embodiment, a processing device 120 is provided with means for carrying out the procedures according to the above embodiments.
[0335] According to another embodiment, a computer program product is provided which includes instructions which, when the program is executed by a computer, cause it to execute the method according to the first three embodiments described above.
[0336] According to another embodiment, an image processing system 100 is provided, comprising the processing device 120 described above.
[0337] The variants and configurations described for the various figures can be combined with one another. The configurations shown and described are purely illustrative, and modifications thereof are possible within the scope of the appended claims. FIG. 10 shows a more detailed, preferred configuration of the evaluation module 122 shown in FIG. 1. While FIG. 1 describes the processing models 150 at a functional level, FIG. 10 shows specific technical implementations and additional optional modules that further increase the performance and precision of the system. In this configuration, the processing models 150 can comprise the following modules, specifically implemented as neural networks:
[0338] As already described with reference to FIG. 1, the identification model 152 is responsible for the analysis and processing of the input data. In the preferred embodiment shown here, this function is implemented by a coding model 153.
[0339] Synthesis model 154 is responsible for the actual generation of the synthesis image data. Its functional task is to generate new, future image data based on an internal state representation and a control instruction.
[0340] In a preferred embodiment, it is implemented as a conditioned generative model 155. The term "generative" means that the model is capable of generating new, previously unseen data. The crucial term "conditioned" means that the generation process is not random, but is deliberately guided by external information—here, the control instruction provided by the control model 157.
[0341] The technical advantage of modern generative architectures, such as diffusion models, generative adversarial networks (GANs) or flow-based models, lies in their ability to generate particularly detailed, high-resolution and artifact-free images, which significantly improves the visual quality and thus the clinical utility of the generated synthesis image data.
[0342] The evaluation model 156, whose basic function is to check plausibility, is retained as a conceptual module. In a preferred embodiment, its function can be integrated into the control model 157, as described below.
[0343] The control model 157 is a newly introduced, central module that assumes active, conditioned control of the generation process of the generative model 155. Its main technical function is the transformation of abstract, development-specific contextual information (e.g., the therapy scenario to be simulated) into a precise, low-threshold control instruction that guides the generation process. This control is based on patterns learned from training data, which in particular reflect biological and physical plausibility rules, thus allowing the function of an evaluation model 156 to be integrated here.
[0344] The combination of coding model 153, conditioned generative model 155, and control model 157 is also referred to as a 3-model framework or 3-model architecture. The three models of the 3-model architecture are interlinked and coordinated to further improve the quality of the synthesis image data 133.
[0345] The Preprocessor 158 is another optional module that introduces new functionality. Its technical task is the automatic extraction of additional parameters directly from the examination image data. This solves the technical problem of enriching the input data by obtaining objectively measurable parameters such as general tissue or organ parameters (e.g., breast density, organ size) or specific parameters of pathological tissue changes (e.g., tumor volume, shape features such as spiculation, texture features) and adding them to the contextual information.
[0346] The additional model 159 is also an optional module that further enhances the robustness of the simulation. Unlike the data-driven control model 157, the function of the additional model 159 is based on explicitly defined mechanistic or physical rules of disease progression. Its technical task is to generate a second, rule-based control instruction. This is combined with the data-driven instruction of control model 157 to form a final control instruction. This acts as a technical "guardrail" that ensures the simulation always adheres to physically and biologically plausible limits.
[0347] Coding Model 153 is a processing model, typically implemented as an artificial neural network, whose fundamental task is to learn an efficient and meaningful encoding of data. This is usually done for the purpose of dimensionality reduction or feature extraction.
[0348] Within the scope of the present invention, however, the function of the coding model 153 goes far beyond the conventional method of pure reconstruction training. While a classically trained model often generates an unstructured or "chaotic" feature space in which a state representation has no inherent, interpretable meaning, the goal of the training method according to the invention is to generate a semantically ordered and structured feature space. In this space, the arrangement of the internal state representations reflects the clinical and biological relationships of the objects under investigation.
[0349] To create this structured feature space, the coding model 153 is trained by optimizing an innovative, multi-component objective function. This objective function can consist of several of the following components, which act synergistically:
[0350] • A reconstruction loss component that minimizes the difference between the original and reconstructed examination images, thus ensuring low-loss coding.
[0351] • A trajectory loss component that minimizes the distance between the state representations of temporally adjacent images of the same object under investigation in order to map the dynamic evolution as a smooth, continuous path in feature space.
[0352] • A similarity loss component that, based on clinical annotations (e.g., tumor subtype), actively approximates the state representations of objects classified as similar and removes those of dissimilar objects. Through this combined optimization, the coding model is trained to create an intelligent "map" of disease states. This learned structure is the technical prerequisite that enables the downstream control model 157 to perform its function of precise, position-based control of the synthesis process. Within the scope of the present invention, this coding model 153 can be implemented by various technologically different architectures:
[0353] In a first embodiment, the coding model is implemented as an autoencoder. An autoencoder typically consists of two symmetrical main components: an encoder, which compresses the high-dimensional input data into a low-dimensional, dense internal state representation ("code"), and a decoder, which attempts to reconstruct the original data from this compressed form as accurately as possible. Within the three-model framework of this invention, the encoder part is primarily used as the identification model, while the decoder part can be conceptually integrated into the synthesis model 154 to enable advanced, controllable reconstruction.
[0354] In a second alternative implementation, the coding model can be realized as a sequence-based architecture, such as a hierarchical vision transformer. Instead of compressing the input data into a single latent vector, the internal state representation is mapped here as a set of vectors (tokens) at different hierarchical levels.
[0355] The generation of the multi-scale feature space works as follows in this configuration:
[0356] 1. Tokenization at the finest scale: The input image is divided into a sequence of small, non-overlapping image patches (tokens). Each of these tokens is transformed into a high-dimensional vector, a so-called feature embedding. This sequence of vectors forms the first and finest scale of the feature space.
[0357] 2. Hierarchical Feature Extraction: This token sequence is processed through several stages of the transformer encoder. At each stage, the relationships between all tokens are calculated using a self-attention mechanism, thereby enriching the feature embeddings with global context.
[0358] 3. Generating the coarser scales: A crucial feature of hierarchical vision transformers is the "patch merging" or "token pooling" step between stages. Here, adjacent groups of tokens (e.g., 2x2 tokens) are combined into a single, new token at a coarser scale. This reduces the number of tokens (the "spatial resolution") while increasing the dimensional depth of the feature embeddings (the "semantic information"). The result of this process is a pyramid of feature representations: a collection of token sequences at different scales, ranging from a high-resolution representation with many tokens in the early layers to a low-resolution but semantically rich representation with few tokens in the deep layers. This totality of representations at all levels forms the multiscale feature space.The final internal state representation passed to the synthesis model can then consist of the tokens from the deepest level or a combination of information from multiple scales.
[0359] In a third alternative embodiment, the coding model can be based on an architecture that generates a multiscale or hierarchical feature space, as is the case, for example, with the encoder path of a UNet model. Such architectures process information at different resolution levels simultaneously, resulting in particularly robust feature extraction and being relevant for models like ImageFlowNet. Regardless of the chosen architecture, the technical task of the coding model 153 is to transform the high-dimensional examination image data 131 and the heterogeneous contextual information into a single, dense, and information-carrying internal state representation. Through the training according to the invention, the feature space is structured such that the properties of the tissue change (presence, location, progression) are encoded in the state representation.This enables the downstream control model 157 to identify the relevant tissue alteration areas based on this structured state representation and to selectively control the synthesis process. While architectures in which control is primarily external via the control model 157 (e.g., in the ControlNet-based design) can, in principle, be initialized with the image data from a single examination time point, a transformer-based architecture preferably requires the image data from at least two previous, sequential examination time points.
[0360] The technical reason for this lies in the fact that sequence-based models like Transformer derive their strength from the analysis of temporal dependencies and trajectories. By providing two starting points, the model can directly derive the previous developmental speed and direction—that is, the initial dynamic vectors—from the data and use this information for a more precise and biologically plausible extrapolation into the future.
[0361] This design underlines the flexibility of the invention, which allows for different initialization strategies depending on the available data and the chosen model architecture, in order to always achieve optimal forecast quality.
[0362] A conditioned generative model, or simply generative model, is an advanced class of AI models trained to generate new, synthetic data that resembles the data from a training dataset. Their functionality is based on learning a complex generation process.
[0363] In contrast to conventional generative models, which often learn to generate data from a random starting point (e.g., pure noise), the generative model 155 according to this revelation is trained to predict a conditioned state transition from one time to the next.
[0364] The generation process is initialized by the internal state representation of the previous image, generated by coding model 153. This representation serves as the starting point for the synthesis process. Simultaneously, the process is guided by the final control instruction provided by control model 157, which contains the context information to be simulated (e.g., the applied therapy and the time interval).
[0365] During training, the generative model 155 learns to modify its synthesis process so that, guided by the control instruction, it does not reconstruct the original state but generates a new image synthesis that corresponds to the actual, subsequent image from the training sequence. The objective function for the training is therefore to minimize the difference between the generated synthesis image and the actual, subsequent image.
[0366] In this way, the generative model learns not only the general ability to generate images, but also the specific, conditional ability to simulate a plausible and accurate temporal evolution from a given initial state to a future state under a specific condition. Thus, in our invention, the generative model transforms from a mere image generator into a precisely controllable simulation engine capable of generating high-quality visual data for specific counterfactual scenarios.
[0367] The way in which the control model 157 translates the abstract clinical instruction into a spatial control instruction is closely linked to the architecture of the generative model 155 used. The invention is not limited to a single implementation but can be realized through various architectures:
[0368] 1. Implementation in Diffusion Models: In an implementation where the generative model 155 is implemented as a diffusion model, the control model 157 can be realized as a ControlNet architecture. Here, the architecture of the encoder part of the diffusion model is duplicated and created as a trainable copy. During the generation process, the outputs of these trainable ControlNet blocks are added to the outputs of the corresponding "frozen" blocks of the main model to control the denoising process.
[0369] 2. Design in Transformer-based models: When using a Transformer architecture, the control model 157 can have the function of converting the development-specific context information into specific conditioning tokens. These tokens are then fed into the self-attention or cross-attention layers of the Transformer to control the generation process by modifying the attention weights.
[0370] 3. Implementation in ODE-based models: In models based on ordinary differential equations (ODEs), image generation is modeled by integrating a vector field over time. Here, the control model 157 can dynamically adjust the parameters of the differential equation defining the vector field. It learns how the development-specific context information influences the "flow direction" of image development at the pixel level.
[0371] Regardless of the implementation, the trained control model solves the technical problem of how to translate abstract clinical scenarios into concrete, spatial instructions for a generative process.
[0372] According to the present disclosure, the control model 157 is trained on a significantly more abstract technical problem, which differs fundamentally from the standard conditioning of generative models in the prior art.
[0373] In the current state of the art, a conditioned model is typically trained to translate a visual cue (e.g., edge detection, a segmentation mask, or a human pose) into a photorealistic output. Training is performed using pairs of (conditioning image, target image).
[0374] In contrast, the control model 157 of the present invention does not learn to copy a visual instruction, but to translate an abstract, non-visual clinical instruction into a spatial control instruction.
[0375] The training data therefore does not consist of simple image pairs, but of inventive tuples composed of three components:
[0376] 1. Initial state: The internal state representation of the object under investigation at a first point in time.
[0377] 2. Developmental Information: The development-specific contextual information that describes the transition to the next point in time (e.g., the vector representing "Therapy A, 365 days").
[0378] 3. Target state: The internal state representation of the object of investigation at the subsequent, second time point.
[0379] The control model 157 thus learns the complex, non-linear function that transforms the abstract, development-specific contextual information into a high-dimensional control instruction (e.g., a "control map"). This instruction specifies to the generative model 155, at a granular level, such as pixel level, how the morphology and texture of the tissue should change to achieve the target state. This specific training procedure solves the technical problem of how to translate abstract clinical scenarios into concrete, spatial instructions for a generative process. It forms the basis for the system's ability to function as a precise "scenario switch" during the inference phase.
[0380] The way in which the trained control model 157 is used in the inference phase constitutes the core of the counterfactual simulation capability and also differs fundamentally from the state of the art.
[0381] In the state of the art, a conditioned model is typically given a new conditioning to produce a single, corresponding output (a "one-to-one" mapping).
[0382] In the present disclosure, the system is used as an iterative simulation engine during the inference phase. Based on a single internal state representation of the patient's condition, the control model 157 is conditioned multiple times with different development-specific context information.
[0383] In the first pass, it is conditioned with the information "simulate therapy A" to generate the first set of synthesis image data. In the second pass, based on the same initial state, it is conditioned with the information "simulate therapy B" to generate the second set of synthesis image data.
[0384] This process stands in direct contrast to purely predictive models, which can only predict a single, most probable future path. The use of the control model 157 as a "scenario switch" according to the invention is the technical mechanism that makes the generation of directly comparable, counterfactual visual data possible in the first place.
[0385] The preprocessor 158 is an optional, but in many configurations advantageous, module whose technical task is the preparation, standardization, and enrichment of the raw examination image data before it is fed into the identification model 152 (the coding model 153). This ensures higher data quality and consistency, which significantly improves the efficiency and accuracy of the subsequent training and inference process.
[0386] In a particularly preferred embodiment, the preprocessor 158 performs an automatic extraction of additional parameters directly from the image data. Using an image analysis algorithm, it obtains objectively measurable parameters such as general tissue or organ parameters (e.g., tumor volume, breast density, shape characteristics) and adds these to the contextual information. This automatic enrichment of the input data for the coding model 153 represents a significant technical advantage.
[0387] Furthermore, the preprocessor 158 can include one or more of the following preprocessing steps, which are known in the prior art but are important for the overall function of the system and can be performed alternatively or in combination with the automatic feature extraction:
[0388] Standardizing image size (resizing / resampling): The image data from different sources or time points may have different resolutions and dimensions. The preprocessor can rescale the images to a uniform size (e.g., using trilinear interpolation) to ensure that all images have the same input dimension for the neural network.
[0389] Intensity normalization: To compensate for differences arising from various scanner settings or protocols, the preprocessor can normalize the pixel or voxel intensity values. Common methods for this include Z-score normalization (subtracting the mean and dividing by the standard deviation) or scaling the values to a fixed range, for example [0, 1].
[0390] Noise reduction (denoising): Medical images often contain noise that can make it difficult to detect fine structures. The preprocessor can apply filter algorithms (e.g., Gaussian filters, median filters, or more advanced non-local mean filters) to reduce noise and improve image quality.
[0391] Image registration: This is a crucial step, especially for longitudinal data. The preprocessor can spatially align the images in a time series to compensate for motion artifacts between acquisitions caused by the patient's varying position in the scanner. This ensures that the development of a lesion is tracked at the correct anatomical location and is not distorted by spatial shifts.
[0392] By combining these different preprocessing functions, the preprocessor 158 ensures that the identification model 152 always works with high-quality, consistent and information-rich data, which is the basis for a precise and reliable simulation.
[0393] The additional model 159 is an optional module implemented in a particularly advanced configuration. Its technical function is to supplement and correct the purely data-driven simulation of the control model 157 with explicit mechanistic or physical knowledge. It acts as a kind of "plausibility engine" or "reality check" that ensures that the generated synthesis image data is not only statistically probable but also biologically and physically sound.
[0394] While the control model 157 implicitly learns its knowledge from correlations in vast datasets, the supplementary model 159 is based on explicitly defined, often mathematically formulated rules that describe causal relationships in disease progression. It solves the technical problem that data-driven models can tend to produce implausible results in scenarios that lie outside their training data distribution ("out-of-distribution"). The supplementary model implements fundamental "laws of nature" of the simulated system as technical boundary conditions.
[0395] Technically, the additional model is not an image generator, but a specialized computational model that processes numerical inputs and outputs a numerical "rule-based control instruction" in the form of a tensor (e.g., a correction vector or a gradient field). It receives the same information as the control model—the internal state representation of the current state and the prediction-conditioning contextual information. Based on this, it calculates a correction for the development proposed by the control model. The final control instruction, which guides the generation process of the synthesis model 154, then results from the mathematical combination (e.g., a weighted addition) of the data-driven instruction of the control model and the rule-based instruction of the additional model.
[0396] Examples of the technical implementation of mechanistic rules:
[0397] • Implementation of growth dynamics: The additional model can implement a simplified but established oncological growth model, such as a Gompertz function or an exponential growth model. Based on the current tumor volume extracted from the internal state representation and the time interval, it calculates an expected minimum or maximum growth. If the growth rate proposed by the control model deviates significantly from this, the additional model generates a scalar correction factor that adjusts the final control instruction so that the simulated growth remains within a physically plausible range.
[0398] • Implementation of dose-response relationships: The additional model can include an explicit mathematical function (e.g., a sigmoidal or logarithmic model) that describes the relationship between a drug dose and the expected growth inhibition. When it receives a specific dose as conditioning, it calculates a corresponding inhibition factor. This is converted into a correction vector that modifies the control map generated by the control model so that the simulated tumor regression or progression corresponds to the expected dose-response. This enables fine-grained control, for example, the simulation of a 10% dose reduction.
[0399] • Implementation of anatomical boundaries: The additional model can receive a digital 3D map of critical anatomical structures (e.g., bones, major blood vessels) as input. If the development proposed by the control model implies a collision with these boundaries, the additional model calculates a repulsive gradient field. This vector field, directed away from the boundaries, is combined with the control map of the control model and "pushes" the simulated development in an anatomically possible direction, thus preventing physically impossible infiltration.
[0400] By integrating the additional model, the overall system transforms from a pure pattern recognition and updating system into a hybrid system that combines the flexibility of data-driven learning with the robustness and reliability of mechanistic models.
[0401] FIG. 11 schematically shows the data flow and interaction of the processing models 150 described in FIG. 10 during the inference phase to generate one or more sets of counterfactual synthesis image data 133a, 133b, 133c.
[0402] The process begins with the processing of the input data. In an optional configuration, the preprocessor 158 first receives the raw examination image data 131 and associated contextual information for processing and enrichment.
[0403] The coding model 153, serving as the identification model 152, then receives all data describing the current state of the subject under investigation. This includes the examination image data 131, as well as static, object-specific contextual information (e.g., genetics, age) and development-specific contextual information describing past development (e.g., previous therapies and time intervals). The technical task of the coding model 153 is to encode this entire, multimodal wealth of information into a single, holistic internal state representation. This vector, which represents the output of the coding model 153, represents the complete current state and the subject's previous course in a multiscale space.
[0404] This internal state representation is now passed on via two parallel paths to enable the controlled generation process:
[0405] 1. It is directly transferred to the diffusion model 155, which serves as the synthesis model 154. Here, it serves as the technical starting point or "seed" for the generative process and ensures that the synthesis is based on the correct, patient-specific condition.
[0406] 2. It is simultaneously passed to the control model 157. Here, it serves as context for the control model 157 to "understand" the current state from which future development is to be controlled.
[0407] Control model 157 is now conditioned separately with the development-specific contextual information that defines the future scenario to be simulated. This is the actual "what-if" instruction (e.g., "simulate therapy A for 12 months"). In an optional, but particularly robust, configuration, control model 157 interacts with the additional model 159. Control model 157 generates a data-driven control instruction, while additional model 159 generates a rule-based control instruction based on mechanistic rules. These two instructions are combined to form a final control instruction that is both data-informed and physically plausible.
[0408] This final control instruction is passed from the control model 157 to the diffusion model 155. The diffusion model 155 now uses its two inputs – the internal state representation as a starting point and the final control instruction as a "roadmap" – to carry out the iterative generation process and produce a first set of synthesis image data, shown here as an example in 133a.
[0409] The system's particular strength, illustrated in FIG. 11 by the branching from diffusion model 155 to outputs 133a, 133b, and 133c, lies in its ability to perform counterfactual simulation. The described process, beginning with the conditioning of the control model 157, can be repeated multiple times based on the same internal state representation, using a different set of future-defining, development-specific contextual information each time. For example, the system can generate a first set of synthesis image data 133a for "Therapy A," a second set 133b for "Therapy B," and a third set 133c for "no treatment." This creates a collection of directly comparable, alternative visual future scenarios for one and the same subject under investigation.
[0410] For technical implementation, the object-specific and development-specific context information can be provided in various structured and machine-readable formats to ensure consistent and error-free processing by the processing models 150.
[0411] The object-specific contextual information, which describes the static properties of a study subject, is typically provided as a single digital file or dataset per patient. Common formats for this include, for example, a CSV (Comma-Separated Values) file, a JSON (JavaScript Object Notation) file, or an XML (Extensible Markup Language) file. In these formats, the information is usually organized as key-value pairs, where each key corresponds to a specific characteristic (e.g., "HER2_Status") and the value to its level (e.g., "positive"). To make this data processable for a neural network, non-numeric data is typically encoded numerically. Categorical characteristics such as sex or HER2 status, for example, can be converted into a binary vector using one-hot coding, while numeric values such as age can be normalized.
[0412] The development-specific contextual information describing the dynamic, interval-related events is preferably organized in a sequential or time-series-based structure. A practical implementation for this is also a CSV file, in which each row corresponds to a specific examination point in a patient's time series. The columns of this file then define the parameters that describe the development at each subsequent time point, such as the time interval in days, the type of treatment applied, or a measured change in tumor volume. Here, too, categorical information such as the type of treatment is numerically coded. This sequential format has the technical advantage of being able to flexibly handle patient histories of varying lengths.
[0413] Regardless of the original file format, these two types of contextual information are combined into one or more numerical tensors before being fed into the coding model 153. These tensors are then combined with the image data tensor to form the holistic, multimodal input for generating the internal state representation. This structured approach ensures consistent, reproducible, and technically valid processing of the heterogeneous input data.
[0414] In the preceding description, the coding model 153, the generative model 155, and the control model 157 were described as separate, functional modules for the sake of simplicity and clarity. However, it is evident to those skilled in the art that this functional separation does not necessarily require a separate architectural implementation. The invention also includes embodiments in which these functions are integrated to varying degrees into one or more model architectures.
[0415] 1. Design as a partially integrated two-model architecture: In another preferred embodiment, the control model 157 and the generative model 155 can be merged into a single, conditioned generative model. In this case, the control function is no longer a separate component, but an integral part of the generative model's architecture. The model then receives two main inputs: the internal state representation generated by the coding model 153 as a starting point, and the development-specific context information as direct conditioning. The model then learns end-to-end to adapt the synthesis process based on this conditioning. In this embodiment, the overall system consists of two main models: a coding model and a conditioned generative model.
[0416] 2. Design as a monolithic single-model architecture: In another, also preferred, design, all three functions—encoding, control, and synthesis—can be integrated into a single, monolithic end-to-end model. Such a model would receive the image data and all contextual information as direct input and generate the final synthesis image data as direct output.
[0417] In this case, the "internal state representation" and the "control instruction" are no longer explicit outputs of separate submodules. Instead, they are to be understood as intermediate states, activations, or implicit information flows within the hidden layers of the monolithic model. The training method according to the invention, with its multi-component objective function, ensures that the model internally learns to form a structured representation and conditionally controls the synthesis process, even if these processes are not explicitly separated architecturally.
[0418] Regardless of the chosen architectural layout—whether fully modular (three-model architecture), partially integrated (two-model architecture), or monolithic (one-model architecture)—the inventive core remains: the generation of a structured internal state representation and the conditional and controllable synthesis of future image data based on this representation to enable counterfactual simulations. The claims are therefore to be understood as encompassing all these architectural variations.
[0419] In a particularly preferred embodiment of the invention, the generative model 155 can be implemented as a sequence-aware model, such as a sequence-aware diffusion model (SADM). Such models are specialized in learning and processing temporal dependencies in a sequence of input data.
[0420] In contrast to a standard SADM, whose goal is to continue a sequence purely autoregressively by predicting the most probable next state, the generative model in the present invention is used in an inventively extended manner. Here, the generation process is guided not only by the past sequence but, crucially, by the external control instruction provided by the control model 157.
[0421] This enables the system to generate, from a given initial state, not only the one most probable path, but a multitude of different, specifically conditioned, and counterfactual future scenarios. The invention thus distinguishes itself from the prior art by transforming a sequence-aware model from a purely predictive tool into a controllable simulation engine.
[0422] 100 Image processing system 35 150 Processing model
[0423] 110 Image acquisition device 151 Model parameters
[0424] 120 Processing facility 152 Identification model
[0425] 122 Evaluation module 153 Coding model
[0426] 124 Memory module 154 Synthesis model
[0427] 126 Control module 40 155 Conditioned generative model 128 Channels Diffusion model
[0428] 130 image data 156 rating model
[0429] 131 Examination image data 157 Control model
[0430] 132 Tissue change image data 158 Preprocessor
[0431] 132a Tissue change imaging data 45 159 Additional model
[0432] 132b Tissue change imaging data 160 Synthesis evaluation
[0433] 133 synthesis image data, 200 machine learning model
[0434] 133a Synthesis image data 202 Input layer
[0435] 133b Synthesis image data 204 Interlayer
[0436] 133c Synthesis image data 50 206 Output layer
[0437] 135 training data points, 208 input data points
[0438] 136 training entries, 210 results data
[0439] 137 annotations, 212 intermediate data
[0440] 140 Examination image 302 Step
[0441] 141 Tissue change 55 304 Step
[0442] 141a Tissue change 306 Step
[0443] 141b Tissue change 310 Step
[0444] 142 Tissue Change Image Area 312 Step
[0445] 142a Tissue Change Image Area 320 Step
[0446] 142b Tissue Change Image Area 60 322 Step
[0447] 142c Tissue Change Image Area 324 Step
[0448] 142d Tissue Change Image Area 330 Step
[0449] 142e Tissue Change Image Area 332 Step
[0450] 142f Tissue Change Image Area 702 Step
[0451] 145 Synthesis image 65 704 Step
[0452] 145a Synthesis image 706 Step
[0453] 145b Synthesis image
Claims
Claims 1. Computer-implemented method for processing examination image data using a synthesis model, wherein the examination image data capture at least one area of tissue alteration of an examination object encompassing a potentially pathological tissue change, comprising: • Receiving the examination image data from one or more examinations of the object of investigation carried out at different times, wherein the examination image data includes time information that represents an examination time of the respective examination, and • Generating, using the synthesis model, synthesis image data for one or more synthesis time points, based on the examination image data and the respective synthesis time point, wherein the synthesis image data represent a state of the tissue alteration area captured in the examination image data at the respective synthesis time point and the tissue alteration area is an area of the examination object captured in the examinations in which the potentially pathological tissue changes can occur.
2. Method according to claim 1, characterized in that the examination image data comprise object-specific context information and development-specific context information, wherein the development-specific context information comprises the examination time and the synthesis time, and the generation of synthesis image data comprises the following steps: • Generating, by means of a coding model, an internal representation from the investigation image data, the object-specific context information and the development-specific context information, in particular by coding the investigation image data together with the object-specific context information and the development-specific context information, in particular by comprehensively mapping it into a multi-scale feature space; and • Generating the synthesis image data from the internal representation using the synthesis model, wherein generating the synthesis image data from the internal representation includes controlling the synthesis model using a control model, and controlling the synthesis model includes conditioning the control model using development-specific context information, wherein controlling the synthesis model is preferably based on patterns learned from training data, wherein the patterns in particular result from biological and physical plausibility rules learned from the training data with a variety of investigation image data.
3. Method according to claim 2, characterized in that: • the object-specific context information includes at least one or more pieces of information from the group consisting of: • patient-specific or subject-specific data, in particular demographic information such as age and gender, family history or hormonal status; • TNM classification (tumor size, lymph node involvement, metastasis status), • pathology-specific genetic mutations, • genetic information, especially DNA mutations such as BRCA1 / BRCA2 mutations; and • molecular and histopathological markers, in particular information on receptor status, protein overexpression or proliferation markers; include; and the developmental contextual information, including at least one or more pieces of information from the group consisting of: • Time information, in particular the time of investigation and the time of synthesis, which define a time interval; and • the course of treatment within the time interval, in particular drug information or active ingredient information that specifies a treatment type, dose or administration interval, include.
4. Method according to any one of claims 1 to 3, further comprising: • Received, from an identification model, from investigation image data of an investigation of the object under investigation, encompassing: • Identifying tissue change image data from the examination image data that captures the tissue change areas, and • Marking the identified tissue alteration image data such that the marked tissue alteration image data are prioritized for processing by the synthesis model when generating synthesis image data.
5. Method according to any one of claims 2 to 4, wherein the coding model fulfills the function of the identification model for identifying tissue alteration image data and for marking the identified tissue alteration image data by transferring the examination image data and the object-specific context information and the development-specific context information into the internal representation, wherein the internal representation includes or is associated with biological feature attributes of the tissue structure; and thus the control model extracts these biological feature attributes from the internal representation or uses them as a conditioning signal to localize tissue alteration areas and control their temporal development.
6. A method according to any one of claims 1 to 5, wherein the examination image data capture at least one pathological tissue change and the generation of synthesis image data comprises generating a time series of synthesis image data, wherein the synthesis time points correspond to examination time points of examinations taking place at regular or irregular intervals, and, based on the generated synthesis image data, the next examination time points are re-determined, depending on a synthesis time of synthesis image data in the time series and a temporal development of the pathological tissue changes in the synthesis image data, in particular a shortening of the intervals of the examinations if the temporal development of the pathological tissue changes indicates a deterioration of a health condition.
7. Method according to any one of claims 2 to 6, further comprising: Controlling the synthesis model, by means of the control model, to generate multiple sets of synthesis image data, wherein for each set of synthesis image data a different set of development-specific context information is used, wherein the different sets of development-specific context information preferably include different drug information or active ingredient information, or also different synthesis time points or also vary other of the possible development-specific context information.
8. Method for determining a differential prognostic index, comprising the method for determining and processing examination image data according to claim 7, further comprising the steps of: • Performing a computer-implemented comparison between two sets of synthesis image data to determine quantifiable differences in at least one or more of the following parameters derived from the image data: o Volume or growth rate of a pathological tissue change, o Shape or texture characteristics of a pathological tissue change, o Morphology or density of microcalcifications, or o Vascularization, as determined by contrast agent uptake; and • Determining the differential prognostic index based on the identified quantifiable differences.
9. The method of claim 8, characterized in that the determination of the differential prediction index comprises a computer-implemented comparison based on one or more of the following elements: • the internal state representations derived from the two sets of synthesis image data; • the synthesis image data itself; • the development-specific contextual information used to generate the two sentences; or • the object-specific contextual information of the respective subject of investigation, and wherein the identified quantifiable differences are chosen such that the resulting index serves as a technical surrogate marker for clinically relevant endpoints, in particular for endpoints selected from the group comprising: radiological response (rCR / rPFS), pathological complete response (pCR), major pathological remission (MPR), progression-free survival (PFS) and overall survival (OS).
10. A method according to any of the preceding claims, further comprising, prior to the step of generating the internal state representation: • Automatic extraction of additional parameters directly from the examination image data using an image analysis algorithm; and • Adding the extracted parameters to the object-specific or development-specific context information, • wherein the extracted parameters include at least one or more pieces of information from the group consisting of: general tissue or organ parameters, in particular organ size, breast size, breast tissue density, host tissue parameters such as density and volume, or the composition of adipose and glandular tissue; and These parameters include those that characterize the pathological tissue change, in particular the functional tumor volume (FTV), which quantifies the metabolically or vascularly active tumor burden, tissue morphology and architecture (e.g., roundness, spiculation), the texture features of the lesion, the location of the lesion within the organ, the presence and morphology of microcalcifications, or vascular parameters, in particular perfusion, permeability, or the spatial distribution of vascularization within the pathological change.
11. Method according to any one of the preceding claims 2 to 10, wherein the control by means of a control model comprises generating a control instruction, wherein the control instruction controls the generation of the synthesis image data and the control instruction is formed based on one or more of the following: • from data-driven patterns learned by the tax model; and • from boundary conditions based on mechanistic or physical rules of disease progression.
12. A synthesis model for processing examination image data, wherein the examination image data captures at least one tissue alteration area of an examination object encompassing a potentially pathological tissue change, in particular for use in a method according to any one of claims 1 to 11, comprising: • an input layer configured to receive examination image data of an examination object with time information, wherein the examination image data specifically includes tissue change image data corresponding to the tissue change areas, • Processing and output layers, wherein the processing and output layers are configured to process the examination image data, in particular to prioritize the processing of the marked tissue change image data, and to output synthesis image data.
13. Computer-implemented method for training a synthesis model to process examination image data, wherein the examination image data captures at least one tissue alteration area of an examination object encompassing a potentially pathological tissue change, in particular the synthesis model according to claim 10, comprising: • Receiving training data; • Processing examination image data included in the training data; and • Adjusting the model parameters of the synthesis model to optimize an objective function that captures a difference between synthesis image data output by the synthesis model and annotations contained in the training data, wherein the training data includes investigation image data of an investigation object for at least three different investigation time points and time information associated with the investigation image data for each investigation time point, wherein the synthesis model processes the investigation image data of a first and a second investigation time point and, for a synthesis time point for which the synthesis image data are generated, corresponds to a third investigation time point and the annotations correspond exactly to the investigation image data of the third investigation time point.
14. Computer-implemented method for generating a training dataset for training a synthesis model for processing examination image data, wherein the examination image data captures at least one tissue alteration area of an examination object encompassing a potentially pathological tissue change, in particular the synthesis model according to claim 10, comprising: • Filtering a database of examination image data of examination objects, filtering out examination image data of examination objects for which at least examination image data for at least three different examination times are available, • in particular, determining tissue change image data capturing the tissue change areas in the examination image data, • Storing the examination image data as a training dataset, in particular in such a way that the specific tissue change image data in the examination image data can be prioritized during the training of the synthesis model.
15. A training dataset for use in a method for training a synthesis model according to claim 11, in particular the synthesis model according to claim 10, comprising: • Examination image data from at least two examination time points.
16. Method for training an identification model, in particular an identification model implemented as a coding model, for use in a method according to any one of the preceding claims 1 to 9, comprising: • Receiving training data that includes, for a large number of study subjects, a time series of study image data, associated context information and annotations, where the annotations classify the study subjects based on clinical or biological characteristics; • Processing the training data by the identification model to create an internal representation for each examination image and to generate a reconstructed version of the examination image from this; and • Adjusting the model parameters of the identification model by optimizing an objective function that includes at least one of the following components: o an overlap component that measures the spatial agreement between predicted and actual segmentation (e.g., Dice, Jaccard); o a classification component that evaluates the pixel-wise assignment accuracy (e.g., cross-entropy); o optionally a weighting component to take class imbalances into account; o a reconstruction loss component (L1 loss) that minimizes a pixel-wise difference between the original and the reconstructed examination images; o a perceptual loss component that minimizes a difference between high-level feature representations of the original and reconstructed images in a pre-trained neural network to preserve visually meaningful structures; o a KL divergence loss component that regularizes a deviation of the learned latent distribution from a standard normal distribution to ensure a structured latent space; o an adversarial loss component that minimizes the distinguishability between real and reconstructed images using a discriminator network to promote photorealistic image quality; and a diffusion loss component (MSE loss) that minimizes the difference between the actual noise and the noise predicted by the diffusion model, where the diffusion model is conditioned by clinical context variables; a reconstruction loss component that minimizes the difference between the original and reconstructed images; a trajectory loss component that minimizes the distance between the internal state representations of temporally adjacent images of the same subject; and o a similarity loss component that, based on the annotations, brings the internal state representations of objects classified as similar closer together and moves those of objects classified as dissimilar away from each other in order to semantically structure the multiscale feature space.
17. Computer-implemented method for generating a training data set for training an identification model, in particular an identification model implemented as a coding model, for use in a method according to any one of the preceding claims 1 to 9, comprising: o Filtering a data collection containing examination image data of examination objects, whereby examination image data of examination objects are filtered out for which at least examination image data for at least two different examination times are available; o Assigning annotations to the filtered examination image data, wherein the annotations classify the examination objects based on clinical or biological characteristics that are relevant for the semantic structuring of a multi-scale feature space; and o Saving the filtered examination image data together with the associated context information and the assigned annotations as a training dataset.
18. Processing device comprising a processor configured to execute the method according to any one of claims 1 to 9 or 11 to 12 and 16.
19. Image processing system comprising a processing device according to claim 13.
20. Computer program product comprising instructions which, when the program is executed by a computer, cause the computer to execute the method according to any one of claims 1 to 9 or 11, 12 and 16.
21. Computer-readable storage medium comprising instructions which, when the program is executed by a computer, cause the computer to execute the method according to any one of claims 1 to 9, 11, 12 or 16.