Computer-implemented method for generating a classification of a dose and / or a contrast agent.

A computer-implemented method using a language model processes multimodal data to automate and enhance the classification of dose and contrast agent thresholds, addressing the limitations of manual methods by improving accuracy and reliability.

DE102024209385A1Pending Publication Date: 2026-04-02SIEMENS HEALTHINEERS AG
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-27
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

The manual classification of dose and contrast agent thresholds in medical imaging is time-consuming, lacks automation, and suffers from low accuracy and reliability due to varying data formats and expert availability, complicating the justification of exceedances.

Method used

A computer-implemented method using a language model that processes multimodal data, including radiology reports and medical images, to generate classifications and justifications for dose and contrast agent thresholds, supported by a knowledge base and user interaction.

Benefits of technology

Automates the classification process, improving accuracy and reliability while providing interpretable justifications, enabling timely corrective actions and reducing manual effort.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The invention relates to a computer-implemented method for generating a classification (C) of whether a dose and / or contrast agent administered to a patient during an examination is justified or unjustified. Multimodal query data (QD) containing information about the radiological conditions of the examination are collected. A knowledge base (KB) is queried to obtain multimodal knowledge data (KD), which is at least partially based on the multimodal query data (QD). The obtained query data (QD) and the obtained knowledge data (KD) are provided to a language model (LM), and the language model (LM) is used to generate a classification (C) of whether the dose and / or contrast agent administered to the patient during the examination is justified or unjustified.Based on the query data (QD) and the knowledge data (KD), a natural language justification (R) for the classification (C) is created, and the classification (C) and the justification (R) are presented on an interaction interface (10) for approval or modification by a user (U).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The invention relates to a computer-implemented method for generating a classification of a dose and / or contrast agent administered to a patient during an examination. The invention further relates to a corresponding data processing device, a computer program product, and a computer-readable storage medium.

[0002] Regardless of the grammatical term used, the term encompasses people with male, female, or other gender identities.

[0003] In many countries, healthcare providers are legally obligated to justify and document any exceedance of dose limits, such as X-ray dose, and / or the limits for contrast agents administered to patients during prescribed imaging procedures. This assessment of dose and / or contrast agent threshold exceedances, also known as outliers, in medical facilities is crucial for protecting patients from the effects of high radiation doses while simultaneously ensuring high image quality.

[0004] Evaluation is a repetitive and time-consuming task requiring manual assessment by a trained medical physics expert (MPE) who considers various case-specific factors describing the radiological conditions of the examination, such as patient characteristics, examination conditions, medical indication, and required image quality. The MPE reviews dose and / or contrast agent alerts generated by an alarm processing system. The alarm processing system can detect dose and / or contrast agent alerts based on threshold values, such as national reference levels or facility-specific reference levels. However, the manual classification process is limited by the availability and expertise of the MPE, the lack of automation and consistency, and the low accuracy and reliability of the classification.

[0005] Automating dose and / or contrast agent classification is difficult because the input data varies considerably and a rationale for the classification must be provided. Automation is further complicated by the fact that the format of the available data and the information contained within the data differ from manufacturer to manufacturer and from scanner model to scanner model (even within the same manufacturer).

[0006] Therefore, there is a need for an improved solution that can automate the classification of dose warnings and / or contrast agent warnings and increase the accuracy and reliability of the classification. The improved solution should also provide an interpretable and well-founded rationale for the classification.

[0007] This problem is solved by the features of the independent claims. Further advantageous examples are contained in the dependent claims.

[0008] In the following specification, the explanations are mostly limited to the classification of dose threshold exceedances in order to keep this specification concise and clear. Of course, the explanations and disclosed features are also applicable to the classification of limit exceedances with respect to media.

[0009] According to a first aspect of the invention, a computer-implemented method is provided for generating a classification of whether a dose and / or contrast agent administered to a patient during an examination is justified or unjustified. The method comprises the following steps: - Receiving multimodal query data including information about the radiological conditions of a performed examination; - Querying a knowledge base to obtain multimodal knowledge data, which is at least partially based on the multimodal query data, wherein the knowledge data includes information about radiological conditions of an examination, a corresponding classification of whether a dose and / or contrast agent applied to a patient during the examination was justified or unjustified, and / or a corresponding natural language justification for the corresponding classification; - Providing the obtained query data and knowledge data for a language model; - Generating a classification using the language model as to whether the dose and / or contrast agent applied to the patient during the examination is justified or unjustified, based on the query data and the knowledge data; - Generating, using the language model, a natural language justification for the classification based on the query data and the knowledge data; - Presenting the classification and justification on an interaction interface for approval or modification by a user.

[0010] In general, the invention provides a method for classifying a radiation dose and / or a contrast agent applied to a patient into at least two categories. For example, a dose and / or a contrast agent can be assigned to a first class, which may be designated as "justified", "OK" or "1", or to a second class, which may be designated as "not justified", "not OK" or "0".

[0011] The execution of the claimed method, i.e., the generation of a dose and / or contrast agent classification, can be triggered based on a comparison of the dose and / or contrast agent with threshold values, also called dose limits. The threshold value can specify the maximum (or minimum) permissible radiation dose and / or contrast agent for an examination, e.g., for a CT scan. The threshold value can, for example, be a reference value belonging to the standard protocol to which an examination-related scan is mapped. If an applied dose and / or administered contrast agent deviates from a standard range, this can be classified as an outlier. This can be the case, for example, if the amount of dose and / or contrast agent exceeds the threshold value. Optionally, a warning message can be generated, as explained later.

[0012] For example, a dose value can be classified as "justified" if there are good reasons for the applied dose exceeding the dose limit, such as when an image with above-average contrast is to be acquired or when there are additional objects in the field of view that result in a higher dose. Other possible reasons include the patient's body type requiring a higher dose, for example, due to a high body mass index, or the examination being misclassified, with the actual examination typically requiring a higher dose than incorrectly indicated. If the reasons for the higher dose value cannot be demonstrated or are inadmissible, the dose can be classified as "unjustified."

[0013] An examination can be one in which the patient is exposed to electromagnetic radiation, such as an X-ray, a CT scan, or the like. In an X-ray or CT scan, radiation in the form of X-rays is absorbed by the body. The amount of X-rays absorbed by the patient contributes to the patient's radiation dose. There are various ways to measure the dose applied in an examination; for example, the "effective dose," measured in millisieverts (mSv), can be considered. Other units of measurement for radiation dose include rad, rem, roentgen, sievert, and gray. A contrast agent (or contrast medium) is a substance used to increase the contrast of structures or fluids in the body when a medical image is taken. Contrast agents absorb or alter external electromagnetic radiation, such as...to make certain types of tissue, such as blood vessels, visible on an X-ray image.

[0014] Although the present explanations focus on examinations in which X-rays are used for image acquisition, the invention is not limited to examinations in which an ionizing radiation dose is applied to a patient. For example, the invention is also applicable in principle to MRI (magnetic resonance imaging) examinations in which images are acquired using magnetic resonance. In this case, the development of a classification of the contrast agents used may be of particular interest.

[0015] Multimodal data encompasses data from multiple data sources. In other words, multimodal data utilizes several data modalities or data types, such as text, audio, tables, sensor data, and / or image data. Combining and processing information from diverse data sources enables the language model to generate more accurate classifications and reasoning.

[0016] According to this disclosure, the multimodal data preferably includes at least written or textual data, such as radiology reports (RIS), electronic health records (EHR), and / or electronic medical records (EMR). Additionally or alternatively, the query data may include medical images, which may be stored, for example, in DICOM format. Data in DICOM format may also allow the extraction of information from the DICOM header. Additionally or alternatively, information may be extracted from a structured radiation dose report (RDSR) and / or a structured contrast report (Contrast SR). Additionally or alternatively, sensor data may be part of the multimodal data, such as data acquired by the sensors of an imaging modality used to perform the examination.Another example: The multimodal data can be provided by an image and data management infrastructure (IDM) where data, especially patient data, of an organization can be stored, managed and retrieved.

[0017] In general, "obtaining" data can encompass "collecting" data, "determining" data, "receiving" data, "retrieving" data, and / or "querying a database to obtain" data, regardless of whether the data is embodied as query data or knowledge data. Furthermore, "querying a database to obtain data" can include the step of "obtaining data."

[0018] The query data contains information suitable for describing the specific radiological conditions of an examination. In other words, the information contained in the query data is related to and / or specific to the radiological conditions of an examination. When manually classifying dose and / or contrast agent outliers, a medical physics expert (MPE) can investigate the case based on the information contained in the query data to determine whether exceeding a prescribed dose and / or contrast agent limit is justified. Thus, the information in the query data can provide insights into whether the dose and / or contrast agent administered to a patient during an examination is justified.The dose and / or amount of contrast agent that is justified varies from case to case and depends on the specific radiological conditions of an examination, such as the patient's physique or the desired image quality.

[0019] The query data may contain information about the actual dose administered to a patient during an examination and / or information about the amount of contrast agent administered to a patient. Information about the dose used or administered during an examination may be found, for example, in the header of DICOM image data and / or in textual patient reports, such as EMR or EHR. Whether an examination, such as a CT scan, was performed with contrast agent administration can be determined, for example, by checking whether the metadata (DICOM tag) of the corresponding series indicates the presence of a contrast agent during the procedure. Furthermore, contrast agent administration can be inferred if a "surveillance" or "preliminary" examination was performed immediately prior to the corresponding series or examination.Furthermore, contrast agent-related sensor data, i.e., injection data, can be measured and provided by a contrast agent injection device and linked to the examination data, so that it can be determined whether and at what time or in which scan of the examination a contrast agent was administered.

[0020] When querying the knowledge base to obtain multimodal knowledge data, the existing specific query data can be taken into account so that the knowledge data retrieved from the knowledge base is related to and / or similar to the query data to some extent. For example, knowledge data may be present in the knowledge base relating to a case with similar patient characteristics and / or a similar type of examination. This knowledge data may include, for example, query data from previous examinations, a corresponding approved classification, and a corresponding approved natural language rationale. Furthermore, the knowledge data may include case studies of radiological examinations, medical journals, and research papers on radiological examinations, etc.The knowledge data may contain annotations that specify the reasons / justification for certain classifications or special considerations that were taken into account. The knowledge data may be at least partially automatically labeled, and / or a human reviewer, e.g., an MPE, may label the knowledge data at least partially.

[0021] At least parts of the knowledge base can be stored on an organization's premises, e.g., in a hospital's IT infrastructure. Additionally or alternatively, at least parts of the knowledge base can be stored in a cloud.

[0022] After retrieving the knowledge data, the query data and the knowledge data can be provided to a language model to generate a classification of the dose and / or contrast agent outlier and a corresponding natural language explanation. Since the knowledge data includes information about the radiological conditions of an examination combined with a corresponding classification and / or natural language explanation, it is suitable for providing decision support to the language model. This improves the quality of the generated classification and / or natural language explanation.

[0023] The language model can comprise a deep neural network, for example, one that uses a transformer architecture. The language model can be a large model with a large number of neurons, such as one billion neurons or more, preferably seven billion neurons or more, even more preferably 13 billion neurons or more, and most preferably 70 billion neurons or more. Specifically, the language model can be configured to take a sequence of tokens, such as words, syllables, or lexemes, as input and output a token considered the "most likely" next token. In particular, the language model can operate in a generative or autoregressive manner.An entry, which is a sequence of tokens, can be provided to the language model as input; the first most likely next token output by the language model can be appended to the entry to obtain a new entry, which is provided to the language model as input in the next iteration; and by repeating this sequence of iterations, the language model can generate a meaningful most likely response to the entry, comprising several sentences or even paragraphs.

[0024] The language model can be trained to create a classification of dose and / or contrast agent outliers. Furthermore, the language model can be trained to generate a natural language rationale for the classification. As previously explained, the language model can classify outliers into at least two categories, which could be labeled, for example, "OK" or "not OK." The corresponding natural language rationale can explain the classification based on the provided query data and / or knowledge data. This means that the language model can determine and state the reasons for the classification.

[0025] In particular, the language model can be pre-trained to understand natural language and provide meaningful answers on any given topic. Such pre-training is possible by using a text database, the result of crawling a portion, preferably a large portion, of the internet, as language training data. Therefore, no labeling of the language data used for pre-training is required. However, pre-training the large language model can optionally include supervised training steps using labeled, human-generated language training data to further improve the model's ability to provide meaningful answers to the user.

[0026] Examples of pre-trained generative language models include OpenAL's ChatGPT 3 and 4, Google AI's PaLM, BERT, and Gemini, DeepMind's Chinchilla, and Meta's Llama 1, 2, and 3. Meta's latter two large pre-trained language models can be downloaded from the internet, while the former two can be accessed and licensed online, including the option for further customization.

[0027] The language model can be trained from scratch or by fine-tuning an existing model, for example, using supervised, unsupervised, or semi-supervised learning methods. In particular, the language model can be fine-tuned for the tasks to which it is to be applied, according to the proposed solution. That is, fine-tuning can involve further unsupervised or supervised training of the language model using language training data related to the proposed system.

[0028] The language model can be trained, for example, using a labeled training dataset. This labeled training dataset can contain query data with information about the radiological conditions of an examination. The sentences or paragraphs in the query data can include, for example, values ​​for one or more patient attributes and specific keywords corresponding to relevant radiological information (e.g., body mass, indication, patient history, age, radiation dose, external beam radiation, etc.). Other relevant information can be included in the query data in DICOM format, such as information about the type of scans acquired.

[0029] The language model can be trained using backpropagation techniques. For example, the language model can be provided with various entries (e.g., prompts, training patterns of query data and / or knowledge data) in a training dataset, and a classification and / or a justification for the classification can be generated using the language model.

[0030] The difference between the output of the language model and a label for the entry can be determined using a loss function. Based on this difference, backpropagation techniques can be used to adjust the parameters and / or weights of the language model. The language model can thus be trained over time with various labeled entries, for example, until it is accurate to a certain threshold. At this point, the language model can be used to generate classifications and / or corresponding justifications.

[0031] For example, the training dataset can be divided into three groups (i.e., training, validation, and testing). The language model can be trained based on the first group (training). The model can then be run (at least partially) to predict the results for the second data group (validation). Using the method described above, it can be assessed whether the language model has been trained correctly. The accuracy of the model can be measured using the remaining data points from the training dataset (testing) (e.g., area under the curve, precision, and recognition).

[0032] The language model takes multimodal query data and multimodal knowledge data as input. Due to the multimodal structure of the data, the language model has an expanded knowledge base for creating a classification and / or a corresponding natural language justification. For example, the language model can directly derive insights about the image quality of the examination from the images provided as input data. Insights about the patient's habitus and / or condition can be directly derived from patient reports. In this way, multiple data sources are considered, and the resulting predictions are more reliable.

[0033] The language model can be stored, at least partially, on the premises of an organization, e.g., in the IT infrastructure of a hospital. Additionally or alternatively, the language model can be stored in a cloud.

[0034] In addition to transmitting query data and acquired knowledge data to the language model, a text input, such as a prompt, can be provided to the language model. The text input can be generated before being fed to the language model. This text input can describe the task to be performed by the language model, i.e., a classification and a corresponding natural language justification based on the provided query and knowledge data. Preferably, the text input is based on natural language and can be a so-called prompt. For the purposes of this disclosure, a prompt is a command signal or input provided to the language model to trigger a specific response or action, i.e., the creation of a classification and / or a natural language justification.The text input could be, for example: "Create a classification as to whether the dose (and / or the contrast agent) is justified or not justified based on the data provided. Provide a rationale for the classification based on the data provided."

[0035] The generated classification and / or justification is presented so that a user can approve or modify it. To improve interaction with the language model, particularly to facilitate user approval of the classification and / or justifications, a user interface can be provided. The user interface can be designed to support conversational interaction with the language model, thus facilitating user feedback. For example, the user interface could include a data area that displays relevant data (e.g., in natural language and / or non-natural language), such as CT images, relevant patient information, and / or radiological data. Specifically, the data—i.e.,The query data and / or the knowledge data on which the classification and / or justifications are based are at least partially displayed in the data area. This provides the user / medical staff with visual information that presents the data, relevant parameters, and information necessary for assessing the justification of the administered dose and / or contrast agent in an easily understandable manner, enabling rapid analysis of the decision proposals generated by the language model.

[0036] In one example, the data window can contain a chat window that allows the user to interact with the language model. The chat window can enable iterative querying of the language model. The user can request a regeneration and / or update of the classification and / or justification generated by the language model if they are not satisfied with the generated classification and / or justification. Preferably, the user can specify how the classification and / or justification should be adjusted to be acceptable. The newly generated or updated classification and / or justification can then be presented to the user again for approval.

[0037] In one embodiment, the method further includes the step of initiating dose and / or contrast agent correction measures, particularly if the applied dose and / or contrast agent is classified as unjustified, i.e., "not OK". Preferably, the classification of the dose and / or contrast agent as unjustified can be approved by a user before corrective measures are automatically initiated. In one example, a suggestion for suitable corrective measures can be generated, e.g., by the language model, and presented to a user. The generated suggestion can then be confirmed by the user.Similar to the iterative updating of the classification and / or justifications, a dialogue can take place between the language model and the user, in which the language model suggests a corrective action, the user provides feedback on the suggested action, and the corrective action can be iteratively adjusted based on the user's feedback. This accelerates the process of identifying appropriate corrective actions.

[0038] In general, initiating dose and / or contrast correction measures can include controlling at least one device for performing a subsequent examination of the patient. In one example, initiating dose and / or contrast correction measures includes providing adapted control signals to control at least one device for performing a subsequent examination of the patient. The controlled device can be a medical imaging device, such as a CT or MRI scanner. Preferably, the adapted control signals provided to the device result in a reduction of the dose and / or contrast agent administered to the patient.

[0039] Various corrective measures can be initiated, e.g., a scan or CT protocol used on the modality can be adapted so that the dose and / or contrast agent applied to a patient is adjusted, in particular optimized and / or reduced.

[0040] The following are some non-restrictive examples of corrective actions related to CT scans. CT scans of patients typically include one or more CT scans and one or more topograms. A topogram might be, for example, a two-dimensional (2D) fluoroscopic image or a cross-sectional image (similar to a conventional X-ray) of a region of the patient's body to provide an overview of that region. A CT scan provides a three-dimensional (3D) image of a region of the patient's body. Each CT scan comprises one or more series, with each series containing multiple (slice) images. One or more series of the CT scan extend along a scan direction within the specified body region. For example, a CT scan of a patient's abdomen might contain four series. Each series contains multiple 2D (slice) images, each representing a cross-sectional view in a slice of the specified body region.The (layer) images follow each other along the scan direction of the corresponding CT scan and have a predefined layer thickness and a predefined distance from each other.

[0041] A CT protocol refers to a set of guidelines and procedures used to perform a CT scan on a patient. CT protocols are developed by radiologists and other medical professionals specializing in CT imaging and are based on evidence-based best practices and safety guidelines. The protocol defines the technical parameters of the scan, including the type of CT scanner to be used, the radiation dose, and / or scan parameters such as slice thickness, scan speed, and / or contrast agent injection protocols. The CT protocol may also include instructions for patient preparation, positioning, and breath-holding techniques to ensure the best possible image quality. The CT protocol has a significant impact on the diagnostic quality of the images produced by the CT scanner and can substantially influence patient outcomes.Standardized CT protocols help ensure the consistency and reproducibility of imaging, which can improve diagnostic accuracy and reduce the need for repeat examinations. However, once the radiological conditions of an examination deviate from the standard conditions, standardized CT protocols may no longer be suitable, and an adaptation of the CT protocols may be necessary.

[0042] In certain examples, corrective measures may include adjusting at least one of the following parameters to achieve, for example, a dose reduction: - mA (milliampere-hours): This parameter controls the current flowing through the X-ray tube during the scan. A lower mA setting can reduce the radiation dose but can also lead to increased image noise. - kVp (kilovoltage peak): This parameter controls the energy of the X-ray beam. A lower kVp setting can reduce the radiation dose, but can also lead to increased image noise and reduced contrast. - Pitch: The pitch is a measure of how fast the CT table moves relative to the speed of the X-ray beam. Increasing the pitch value can reduce the radiation dose, but can also lead to lower image resolution. - Scan length: Reducing the scan length to the necessary anatomical area can reduce the radiation dose. - Reconstruction algorithms: The use of iterative reconstruction algorithms can help reduce image noise and improve image quality, allowing for a lower radiation dose. - Shielding: The use of lead shielding around sensitive areas of the body (e.g. thyroid, breasts and gonads) can reduce the radiation dose. - Automatic Exposure Control (AEC): AEC systems can automatically adjust the mA and kVp settings to the size and shape of the patient's body, which can help optimize the radiation dose while maintaining image quality.

[0043] The specific parameters for dose reduction can vary depending on other radiological conditions of the examination, such as the patient's age, size, and clinical indication. Reducing the dose directly affects image quality and therefore impacts the radiologist's ability to interpret image findings and reduces comfort during image review. Therefore, unnecessary and / or inefficient dose reductions should be avoided. It is thus advantageous to use multimodal data as input for the language model to improve the proposed classification and / or justification and to enable the model to make an accurate suggestion for an appropriate corrective action.

[0044] The set parameters can be made available to a device, such as an imaging modality, and an examination can be performed with the device based on these parameters. In other words, control signals containing the set parameters are provided to the device.

[0045] As explained above, in a preferred embodiment the multimodal query data and the multimodal knowledge data each comprise at least two different data types, in particular text data, medical image data and / or sensor data.

[0046] The query data and / or knowledge data preferably include radiology reports, DICOM data, and / or sensor data. In particular, the query data and / or knowledge data may include topographic information, information about acquired scans, information about other examination conditions, information about the requested procedure, patient information, information about indications, information about the patient history, information about the examination performed, the device usage time, and / or the device processing time.

[0047] Patient information can include any type of patient characteristics, such as height, sex, weight, treatment options, treatment characteristics (e.g., gantry movements, gantry positions, X-ray tube movement, X-ray detector movement, etc.), treatment goals, tumor characteristics (e.g., size or shape), images of the patient or tumor, tumor stage, primary treatment site, endpoints, tumor enlargement, body mass index, blood pressure, medical history (e.g., previous medical treatments), etc. For example, patient information may be contained in text reports and / or DICOM image data.

[0048] Image data, especially medical image data, can be acquired by various imaging modalities or devices. For example, image data can be acquired using a CT scanner, an MRI scanner, an angiography system, an ultrasound machine, and the like. The image data can be stored and / or retrieved in DICOM format.

[0049] Sensor data can be in various forms. For example, it can be obtained from logs of the imaging modality used. This data can provide information about device utilization and / or processing times. It can also contain information about the functionality of the imaging modality, such as malfunctions or downtime. Alternatively or additionally, sensor data can be obtained from other relevant devices, such as a contrast agent injection device.

[0050] The generated classification and its corresponding justifications are thus based on multiple data sources, particularly data sources that were previously inaccessible / intelligible, such as images or natural language information from EMR. This "multimodal data integration" leads to higher accuracy, better quality, and consistency when the language model generates a classification and justification.

[0051] In a preferred embodiment, the language model comprises a plurality of branches, each branch being suitable for processing a specific type of query or knowledge data. For example, one part of the model could be specifically adapted for processing text data, such as by applying natural language processing techniques. Another branch of the model could be specifically adapted for processing image data, such as using convolutional neural networks. The model branches can be merged before the classification of whether a dose and / or amount of contrast agent is justified is generated. This architecture facilitates data processing and improves the results generated by the language model.

[0052] Preferably, the query data and / or knowledge data can be in embedded form. This "embedding" can include data encryption. In other words, the multimodal data can be translated into a numerical array—that is, a vector, matrix, tensor, or the like—to enable / facilitate the processing of the multimodal data by the language model, with vector embedding being preferred. Even when multimodal data from different data sources is embedded, the embedding still expresses the original meaning of the data, so any observed similarity between embeddings indicates a similarity in the underlying data. Representing data points in embedded form, especially as vectors, enables interoperability of data from different sources, such as text and / or image data.In particular, embedding the data in the same embedding space enables the processing of multimodal data by a single language model.

[0053] For example, the multimodal knowledge data in the knowledge base can be stored as a multitude of embedded knowledge vectors and / or the query data can be embodied as at least one embedded query vector.

[0054] Alternatively or additionally, the embedded data can serve as a reference to the location where the multimodal data—that is, the underlying multimodal query data and / or knowledge data—is stored. The storage capacity required for the embedded data can be less than the storage capacity required for the (raw / non-embedded) query data and / or knowledge data. Embedding the data thus enables efficient use of available storage capacity. For example, the embedded data can be stored in a location directly accessible to a computing unit executing the claimed procedure, requiring only minimal storage capacity. The non-embedded data can be stored in an external database with greater storage capacity, such as in a cloud.

[0055] In a preferred embodiment, the method also includes the following step: - Obtaining at least one embedded multimodal query vector based on the multimodal query data, in particular by embedding the multimodal query data using at least one feature extraction technique and combining the extracted features to form the multimodal query vector.

[0056] Similarly, at least one feature extraction technique can be applied to embed the knowledge data, wherein the embedding of the knowledge data is preferably carried out before the execution of the claimed method. Embedding the query data and / or the knowledge data facilitates data processing and reduces the required computing power and storage space for data processing.

[0057] Various feature extraction techniques can be used, such as natural language processing and / or convolutional neural networks. An embedding model can be used to embed the data. Embedding models can be trained on a large dataset specifically tailored to the task at hand. Alternatively, pre-trained embedding models can be used, which can optionally be fine-tuned. Feature extraction techniques can be suitable for dimensionality reduction. Preferred dimensionality reduction methods include autocoders, convolutions, principal component analysis, and T-distributed stochastic neighbor embedding (t-SNE). The selection of a specific dimensionality reduction method depends on several factors, such as the nature and / or information of the data to be embedded and / or compressed.

[0058] For text data, embedding models such as Google's Word2Vec or Stanford University's Global Vectors (GloVe) can be trained from scratch. However, pre-trained versions are also available, often based on public text data like Wikipedia and Common Crawl. Similarly, large language model encoders / decoders (LLMs), such as BERT and its numerous variants, can be used for embedding. For image embedding, pre-trained image classification models like ImageNet, ResNet, or VGG can be adapted to embed output by removing their final, fully connected prediction layer.

[0059] Especially in the medical field, highly specialized vocabulary is used, which is unlikely to be present in the training data of generalist models. Supplementing the basic knowledge of pre-trained models through further training, i.e., fine-tuning, on domain-specific examples, such as medical images or patient reports, can help the embedding model produce more effective embeddings.

[0060] Preferably, the multimodal knowledge data to be obtained from the knowledge base are selected on the basis of comparing the embedded query vector with the embedded knowledge vectors stored in the knowledge base and determining a similarity between the embedded query vector and the embedded knowledge vectors.

[0061] To determine the relative similarity of different vector embeddings, various methods can be used.

[0062] In one embodiment, the Euclidean distance can be used. According to this approach, the difference between two n-dimensional vectors a and b is calculated by first adding the squares of the differences between their respective components and then taking the square root of this sum. Alternatively, the cosine distance, also called cosine similarity, can be used. The cosine distance ranges from -1 to 1, where 1 represents identical vectors, 0 represents orthogonal (or unrelated) vectors, and -1 represents completely opposite vectors. Cosine similarity is well-suited for embeddings used in natural language processing because it naturally normalizes the sizes of the vectors and is less sensitive to the relative frequency of words in the training data than the Euclidean distance. As a third example, the dot product can be used.

[0063] Preferably, the step of querying the knowledge base further comprises obtaining n knowledge vectors, where n is a natural number greater than 1. These n knowledge vectors have a higher similarity to the query vector than the other knowledge vectors stored in the knowledge base. In other words, the n knowledge vectors with the greatest similarity to the query vector are determined from the knowledge base. The knowledge vectors thus obtained are considered particularly relevant as background information for the language model used to classify the applied dose and / or contrast agent, and for the natural language justifications to be generated.

[0064] The acquired knowledge vectors and / or the query vector can be made available to the language model. For example, the acquired knowledge vectors and / or the query vector can be loaded into the language model's cache.

[0065] In the classification generation step, the language model can derive insights into the radiological conditions of the performed examination, represented by the query data / vector, by comparing the specifics of the examination with the examples represented by the knowledge data / vector. Based on the classification contained in the knowledge data / vector, a similar classification can be generated for the dose and / or the contrast agents used in the examination. The natural language rationale generated by the language model can also be based on the rationale contained in the knowledge data / vector. Thus, the generated natural language rationale can include a reference to the rationales contained in the knowledge data / vector upon which it is based.If necessary, the generated natural language justification may contain arguments as to why it follows and / or deviates from the classifications and / or the justifications contained in the knowledge data / knowledge vectors.

[0066] Using the methods described above, a Retrieval-Augmented Generation (RAG) framework can be designed that provides additional context to the language model based on the query data for a given study. Thus, a RAG framework can improve the quality of the responses generated by the language model. By supplying the language model with external knowledge, the model gains access to the most current and reliable knowledge data relevant to the responses it generates. Furthermore, the sources on which the model's generated response is based can be presented to users, allowing them to verify the accuracy and correctness of the model's responses. Using a RAG framework also reduces the risk of hallucinations.

[0067] In a preferred embodiment, the RAG framework comprises two phases: a retrieval phase and a content creation phase.

[0068] In the retrieval phase, the information relevant for evaluating the multimodal query data—that is, the classification and / or the justifications generated by the language model—is retrieved from the knowledge base. The query data can be embedded as query vectors, and the knowledge data in the knowledge base can be embedded as knowledge vectors, as explained above. The most relevant information in the knowledge base that correlates with the query data is identified. A vector search can be performed in the knowledge base for this purpose. Before performing the vector search, the multimodal query data can be converted into a vector representation. The knowledge base can then be searched for embedded knowledge vectors that match the vector representation of the query data. Various search algorithms, such as k-Nearest Neighbors (KNN) or Hierarchical Navigable Small World (HNSW), can be used for this.

[0069] During the content generation phase, the language model is then prompted to create a classification and explanation for the dose data outlier. The prompt is augmented with the query data and the retrieved knowledge data. This augmentation can occur after the prompt, query data, and / or knowledge data have been embedded. This allows the language model to generate an accurate response, which can be based on explainable artificial intelligence, making the response generation process traceable and / or understandable to a human operator. The credibility of the response can be enhanced by providing links to its sources.

[0070] In a preferred embodiment, the method comprises the step: - Anonymizing the query data and / or the multimodal knowledge data, or - Pseudonymizing the query data and / or the multimodal knowledge data.

[0071] In this disclosure, pseudonymization refers to a process in which all personal data are replaced by unique identifiers, so that the data can be traced back to its origin and a link between patient and data is possible. Anonymization, on the other hand, renders all personal data that could allow for tracing unrecognizable. Therefore, anonymized data must not contain any information about the corresponding patient, which may be required, for example, by legal data protection regulations.

[0072] Preferably, query data is pseudonymized before the knowledge base is queried to obtain similar knowledge data. This ensures the confidentiality of the medical query data. Using a pseudonymized identifier assigned to a specific patient to pseudonymize the query data and / or the knowledge base still allows the retrieval of that specific patient's medical data from the knowledge base. This is particularly important when the knowledge base is hosted in a cloud or other system located outside of an organization's premises, such as remote from a hospital's local network.

[0073] In this way, the knowledge base can be effectively searched for data belonging to the patient examined in the study to be classified, and the corresponding knowledge data can be made available to the language model after retrieval. The ability to securely store patient data in cloud storage significantly reduces the load on the hospital's local storage, such as a PACS. The privacy of the patient data is not compromised by pseudonymization.

[0074] In a preferred embodiment, the language model is an explainable AI language model or an interpretable AI model. The language model may, in particular, include interpretable components such as attention mechanisms and salience maps. Especially when using a large model, for example, with multiple hidden layers, the decisions of such a model can be difficult for a human user to predict. Providing an explainable AI model improves the comprehensibility of the results generated by the language model. This enhances the transparency of the classification and justification process. It thus makes it easier for a user to decide whether the generated classification and justification can be accepted or require modification.Especially in the medical field, the reliability of decisions is very important, as persistent misclassification can have serious consequences for the patient's health.

[0075] Attention mechanisms are a method for processing high-dimensional input data in which different parts of the data or the information they contain are assigned different levels of importance. Salience maps can highlight regions in medical images, such as images included in the query data or knowledge data, that have the greatest impact on the generated classification and / or justifications. Salience maps can also be used to indicate the sections of input text, such as patient reports, that are most relevant to the generated classification and / or justification. Attention mechanisms and salience maps can be combined to further enhance understanding of the model. This can be done, for example, to...This is achieved by providing the user with a salience map that highlights the parts of the input data that have received more attention due to the different importance assigned by an attention mechanism in use.

[0076] Alternatively or additionally, the comprehensibility of the generated classification and the generated justification can be improved by carrying out the following further procedural step: - Conducting and presenting an analysis of the significance of specific features of the data for classification and / or justification, e.g. by conducting and presenting a Shapley Additive Explanations (SHAP) analysis and / or a Local Interpretable Model-Agnostic Explanation (LIME) analysis.

[0077] SHAP is based on Shapley scores, which determine the contribution of features to a resulting prediction by the language model. In other words, it assesses how much a feature contributes to the created classification and the justification for that classification. Shapley scores can be visualized in various ways, such as using different charts. Using SHAP and its corresponding visualization can help the user better understand the response generated by the language model.

[0078] In a LIME analysis, the language model, or parts of it, are treated as a black box and locally approximated with an interpretable model. This interpretable model modifies the feature values ​​of an input sample—that is, the available query data—and observes how these changes affect the predictions generated by the model. In this way, a range of explanations can be generated that shed light on the extent to which a feature of the query data and / or the knowledge data contributes to the generated classification and / or justification.

[0079] Providing the results of an analysis of the significance of individual features enables the user to better assess the reliability of the created classification and / or the generated justifications. Furthermore, based on the provided analysis, the user can deduce what modifications need to be made to the generated classifications and / or the generated justification.

[0080] In summary, the use of explainable AI components for the design of models and / or for the analysis of results facilitates decision-making for users, who can better understand the results of the model and better decide what measures are needed based on the results.

[0081] Preferably, the procedure also includes the following steps: - Receiving feedback from a user about the classification and / or the natural language justifications provided for the classification; - Fine-tuning the parameters of the language model to minimize a cost function, and / or updating the knowledge data stored in the knowledge base based on the feedback provided.

[0082] In particular, the user can provide feedback on the correctness of the classification and / or the provided justification. Furthermore, the user can provide feedback on how the classification and / or the natural language justification can be modified to make it more usable.

[0083] In this way, even if the type of query data and / or the information contained in the query data changes, retraining the entire language model from scratch is avoided. If negative user feedback is received, the knowledge base itself can be updated with the corrected data samples, and / or the most recent data samples can be used to fine-tune the model. This keeps the knowledge base and / or the model structure up to date. By incorporating user feedback, the language model can learn based on previously made suggestions. This improves the quality of the classifications and the justifications provided. Furthermore, the computational overhead for executing the claimed procedure and for hosting a computer system that executes the procedure can be reduced.

[0084] In a preferred embodiment, the method also includes the following step: - Receiving an alert indicating that a dose and / or contrast agent administered to a patient during an examination is classified as an outlier, particularly exceeding a threshold.

[0085] Receiving an alarm can also include generating an alarm. Receiving the alarm can trigger subsequent process steps. In other words, each process step following "receiving the alarm" can include the phrase "after receiving the alarm." The claimed method is particularly suitable for the automatic or semi-automatic processing of alerts, which is advantageous because processing alerts is a time-consuming and difficult task when performed manually. By implementing alert systems that notify the responsible personnel when an outlier in the data is detected, a timely response to anomalies in radiological performance is ensured.

[0086] The threshold for determining that a dose and / or contrast agent is an outlier can be predefined. Alternatively, the threshold for identifying an outlier can be dynamically or periodically adjusted based on knowledge data stored in the knowledge base to account for evolving data characteristics. This reduces the probability of false alarms, also known as false positives and false negatives. This adjustment of the alarm threshold based on knowledge data also results in the thresholds of individual institutions being "learned." This means the method can learn what to consider an outlier for a particular institution. This institution-specific threshold might depend, for example, on the imaging modalities used and / or the imaging procedures performed.Therefore, threshold values ​​for specific devices / device types can even be learned based on the knowledge data.

[0087] The warning threshold can be calculated, for example, using a statistical method such as z-scores or percentiles, based on the data stored in the knowledge base. For this purpose, existing classifications in the knowledge base can be analyzed, and it can be verified whether a particular dosage and / or amount of contrast agent was justified.

[0088] User feedback can be taken into account when adjusting the threshold. For example, users such as radiologists and / or medical physicists can determine whether an adjusted threshold meets clinical expectations. This makes the alarm processing system even more tailored to the needs of the user, such as the medical staff of a hospital.

[0089] Optionally, the threshold can be automatically recalculated at regular intervals or when the knowledge data in the knowledge database is updated.

[0090] The knowledge data in the knowledge base can be updated at regular intervals or on request. In particular, the latest data, i.e., the "delta" between the last update of the knowledge base and the current data set, can be added to the knowledge base to increase computational efficiency and ensure that the knowledge base is up to date.

[0091] According to another aspect, a data processing device is provided with means for carrying out the steps of the method according to one of the claimed embodiments.

[0092] According to another aspect, a computer program product is provided, wherein the computer program product includes instructions which, when the program is executed by a computer, cause the computer to execute one of the above-mentioned embodiments.

[0093] The computer program product, e.g., a computer program, can be stored on a memory card, a USB flash drive, a CD-ROM, a DVD, or as a file that can be downloaded from a server on a network. Such a file can be provided, for example, by transferring the file along with the computer program product from a wireless communication network.

[0094] According to another aspect, a computer-readable storage medium is provided which contains instructions which, when executed by a computer, cause the computer to perform the steps of the procedure according to one of the embodiments mentioned above.

[0095] The invention also relates to an alarm processing system for generating a classification as to whether a dose and / or contrast agent administered to a patient during an examination is justified or not. The alarm processing system is configured and / or adapted to execute the computer-implemented method according to the invention. The alarm processing system may consist of a data storage device, an interface, and / or a processor. The processor is configured to at least partially incorporate, implement, and / or execute the features of claim 1. The processor may be configured to receive multimodal query data, including information about the radiological conditions of the examination performed.The processor can further be configured to query a knowledge database to obtain multimodal knowledge data, which is at least partially based on the multimodal query data. This knowledge data includes information about the radiological conditions of an examination, a corresponding classification, whether a dose and / or contrast agent administered to a patient during the examination was justified or unjustified, and / or a corresponding natural language explanation for the corresponding classification. The processor can further be configured to provide the obtained query data and the obtained knowledge data to a language model. The processor can further be configured to use the language model to generate a classification as to whether the dose and / or contrast agent administered to a patient during the examination was justified or unjustified.The processor determines whether the classification applied to the patient during the examination is justified or not, based on the query data and the knowledge data. Furthermore, the processor can be configured to use the language model to generate a natural language rationale for the classification based on the query data and the knowledge data. The interface and / or the processor can be configured to display the classification and rationale on an interaction interface for user approval or modification.

[0096] The embodiments and features described with reference to the method of the present invention apply mutatis mutandis to the device and / or system of the present invention.

[0097] Other possible embodiments or alternative solutions of the invention also include combinations—not explicitly mentioned here—of features described above or below in relation to the embodiments. A person skilled in the art can also extend the basic form of the invention by adding individual or isolated aspects and features.

[0098] Further embodiments, features and advantages of the present invention will become apparent from the following description and the dependent claims in conjunction with the accompanying drawings, in which: Fig. 1 A schematic representation of an embodiment of a system for creating a classification shows whether a dose and / or contrast agent applied to a patient was justified; Fig. 2 shows a schematic representation of another embodiment of a system that generates a classification as to whether a dose and / or contrast agent applied to a patient was justified; Fig. 3. A schematic representation of a procedure for creating a classification shows whether a dose and / or contrast agent applied to a patient was justified; Fig. 4 A schematic representation of an exemplary data flow in a procedure for creating a classification shows whether a dose and / or contrast agent applied to a patient was justified; Fig. 5. A schematic representation of a further exemplary data flow in a procedure for generating a classification shows whether a dose and / or contrast agent applied to a patient was justified; and Fig. 6 A schematic representation of another exemplary data flow in a procedure for creating a classification shows whether a dose and / or contrast agent applied to a patient was justified.

[0099] In the figures, identical reference numbers denote identical or functionally equivalent elements, unless otherwise specified. To improve the clarity of this specification, some of the described examples are explained only in the context of dose classification. However, the principles can also be applied to the classification of the contrast agents used.

[0100] Fig. Figure 1 shows a system for generating a classification C indicating whether a dose and / or contrast agent administered to a patient during an examination is justified, i.e., "OK," or unjustified, i.e., "not OK." Thus, system 1 is suitable for distinguishing between at least two different states of the administered dose and / or contrast agent, namely justified and unjustified. However, system 1 can also be suitable for distinguishing between more than two different states, e.g., three or four states.

[0101] System 1 is capable of executing a procedure for generating a classification C, determining whether a dose and / or contrast agent applied or administered to a patient during an examination is justified or unjustified. System 1 comprises an interaction interface 10, in particular a user interface, and a computing unit 20. System 1 may also include an imaging modality 30. System 1 may contain a knowledge base KB. Additionally or alternatively, System 1 may be connected to an external knowledge base KB. The interaction interface 10, the computing unit 20, the imaging modality 30, and the knowledge base KB, i.e., the components of System 1, may be part of a health information system of an organization ORG. Fig. 1 and Fig. 2 The dashed border indicates the premises of the organization ORG.

[0102] The components mentioned above can be interconnected via a network N. Examples of network N include private or public local area networks (LANs), wireless local area networks (WLANs), metropolitan area networks (MANs), wide area networks (WANs), the internet, and / or combinations thereof. The network N can encompass wired and / or wireless communication according to one or more standards and / or over one or more transport media. Communication over the network N can be conducted in accordance with various communication protocols such as Transmission Control Protocol and Internet Protocol (TCP / IP), User Datagram Protocol (UDP), and IEEE communication protocols. For example, the network N can include wireless communication according to Bluetooth specifications or another standardized or proprietary wireless communication protocol.In another example, the network N can also include communication over a cellular network, e.g. a GSM (Global System for Mobile Communications), CDMA (Code Division Multiple Access) or EDGE (Enhanced Data for Global Evolution) network.

[0103] The knowledge base (KB) can generally be configured for the collection, storage, and / or forwarding of multimodal knowledge data (KD). The knowledge base (KB) can be embodied by one or more storage locations. In particular, the knowledge base (KB) can be implemented as local storage, cloud storage, or distributed storage, e.g.,...

[0104] The multimodal knowledge data stored in the knowledge base (KB) can encompass various data types, such as written data (e.g., patient reports) and image data (e.g., medical images). Knowledge data can include information about the radiological conditions of an examination, such as medical information about one or more patients, data on the amount of dose and / or contrast agent administered during the examination, classifications of dose and / or contrast agent as outliers, and / or a corresponding explanation for why a dose and / or contrast agent was classified as an outlier.

[0105] The knowledge data (KD) in the knowledge database (KB) can be structured as multiple data samples, with each data sample preferably containing information about a test. A data sample can, in particular, contain information about the radiological conditions of a performed examination. The data sample can also contain a finding, e.g., a classification of a dose administered during the examination and / or a corresponding natural language explanation of why the administered dose was classified as "OK" or "not OK".

[0106] Knowledge data (KD) can contain medical data. This medical data can encompass various types of data, such as a patient's medical history, electronic health records, medical images, or similar information. Knowledge data (KD) stored in the knowledge database (KB) can include, for example, image and non-image data.

[0107] The image data can consist of three-dimensional image datasets acquired, for example, with an X-ray system, a CT system, an MRI system, or other systems. Additionally or alternatively, the image data can also include two-dimensional medical image data. According to some examples, this two-dimensional image data may have been extracted from three-dimensional medical image datasets.

[0108] Additionally or alternatively, two-dimensional medical image data may have been generated by special imaging modalities such as slide scanners used in digital pathology.

[0109] For example, information about the radiological conditions of an examination—that is, information relevant to the radiological evaluation of an examination—can be included in the DICOM attributes of a DICOM object. The DICOM attributes can contain information about the acquisition type and / or the acquisition parameters of the acquired scans. They can also include information about other examination conditions, such as the patient's position and / or the administration of a contrast agent. Furthermore, the DICOM attributes can contain information about the requested procedure and / or the examination itself, such as which scans were performed or whether the image is a topogram.

[0110] Non-image data can include, for example, procedural or diagnostic data that provide information about the patient. Electronic health records can be considered non-image data. Non-image data can relate to non-imaging examination results, such as laboratory data, vital sign records (e.g., ECG data, blood pressure readings, ventilation parameters, oxygen saturation levels), etc. Furthermore, non-image data can include structured and unstructured medical reports relating to, for example, previous examinations of the patient. Non-image data can also include the patient's personal information, such as gender, age, weight, insurance information, etc.

[0111] To give a concrete example: Information from the EMR, such as patient information, indications and / or the patient history, can be relevant for the radiological assessment of an examination.

[0112] The interaction interface 10 can include an output unit, such as a display, and an input unit. The interaction interface 10 can be a smartphone or a tablet computer, i.e., a mobile device. The interaction interface 10 can also be a workstation in the form of a desktop PC or laptop. The input unit can be integrated into the display, for example, in the form of a touchscreen. Alternatively or additionally, the input unit can include a keyboard, a mouse or digital pen, a microphone, a combination thereof, or the like.

[0113] The processing unit 20 can determine which elements are displayed on the output unit of the user interface 10. Alternatively or additionally, the interaction interface 10 can include an interface processing unit configured to execute at least one software component to force the output unit to provide a graphical user interface (GUI). The graphical user interface can display a generated classification (C) and / or a justification (R) so that the user (U) can review it.

[0114] The interface processing unit can be a general-purpose processor, a central processing unit, a control processor, a graphics processing unit, a digital signal processor, a three-dimensional rendering processor, an image processor, an application-specific integrated circuit, a field-programmable gate array, a digital circuit, an analog circuit, combinations thereof, or other known devices for processing image data. User interface 10 can also be run as a client.

[0115] The computing unit 20 can be configured to generate a classification C and a justification R for an applied dose and / or contrast agent. Furthermore, the computing unit 20 can be configured to control communication over the network N and / or process or preprocess data. The computing unit 20 can be a processor. The processor can be a general-purpose processor, a central processing unit, a control processor, a graphics processing unit, a digital signal processor, a three-dimensional rendering processor, an image processor, an application-specific integrated circuit, a field-programmable gate array, a digital circuit, an analog circuit, combinations thereof, or any other known device for processing image data. The processor can be a single device or multiple devices operating serially, in parallel, or separately.The processor can be the main processor of a computer, such as a laptop or desktop computer, or a processor that performs some tasks in a larger system, such as a medical information system or a server. The processor is configured by instructions, design, hardware, and / or software to perform the steps described here. The processing unit 20 can be contained within the interaction interface 10. Alternatively, the processing unit 20 can be separate from the interaction interface 10. The processing unit 20 can comprise a real or virtual group of computers, such as a so-called "cluster" or a "cloud." The processing unit 20 can be implemented, at least partially, as a server system or be part of a server system. The server system can be a central server, such as a cloud server, or a local server, such as a desktop computer.in a hospital or radiology department, or the server may contain some parts as a cloud server and other parts as a local server.

[0116] Furthermore, the computing system 20 can include storage, e.g., main memory, to temporarily store query data QD and / or knowledge data KD from the knowledge database KB. According to some examples, such storage can also be included in the interaction interface 10.

[0117] Imaging modality 30 can be any device capable of generating medical images, such as an X-ray scanner, a CT scanner, or an MRI scanner. Imaging modality 30 can be equipped with measuring devices capable of providing sensor signals. This allows, for example, the monitoring of the imaging modality's operating time and / or measurements of the applied dose.

[0118] After an examination has been performed using imaging modality 30, query data QD is received. This query data QD can be provided, for example, by imaging modality 30, an EHR / EMR system of the organization ORG, and / or a PACS image archiving and communication system. The query data QD can be sent to computer system 20 (push procedure). Alternatively, computer system 20 can initiate the sending of the query data QD (pull procedure).

[0119] Then the computing system 20 can pass the query data QD to a language model LM, as will be explained later.

[0120] The language model LM can be configured or trained to automatically generate text responses based on received input, such as text, image, video, and / or audio. The language model LM can be trained to create a classification C, based on the provided input data, particularly the query data QD and / or the knowledge data KD, indicating whether the dose and / or contrast agent administered to a patient during a procedure is justified or not. The language model LM can be further trained to generate a natural language rationale R for the classification C based on the provided data. The language model LM can conduct or simulate a conversation with a user U, in which the results, i.e.,The classification C and / or the justification R will be discussed, whereby the user U can confirm the result, request further explanations and / or demand an adjustment / correction of the generated classification C or the justification R.

[0121] In the Fig. In the embodiment shown in Figure 1, the language model LM is hosted outside the premises of the organization ORG. Thus, the query data QD and / or the knowledge data KD leave the organization ORG when they are entered into the language model LM. To protect the confidential information contained in the data, such as a patient's personal data, the data leaving the organization ORG can be encrypted, anonymized, and / or pseudonymized.

[0122] The in Fig. The system shown in section 2 essentially corresponds to the one in Fig. 1 system shown.

[0123] The language model LM is hosted on the premises of the organization ORG, specifically on computing unit 20, thus ensuring the confidentiality of the information exchanged with the language model. Furthermore, the knowledge base KB of System 1 comprises several data silos. A first data silo, KBo, belongs to the organization. The organizational data silo KBo can be part of an IT system of the organization ORG and can be used by the organization ORG to store, for example, patient-related medical information. The knowledge data KD stored in the organizational knowledge base KBo must therefore only be accessible from within the premises of the organization ORG and must not be shared with other organizations. A second data silo, KBp, belongs to a public sector. In an exemplary embodiment, the public data silo KBp contains information about public law requirements, such as...Dose limits, and / or a knowledge base shared by several organizations, which may be based on publicly available knowledge data (KD).

[0124] In Fig. Figure 3 illustrates an embodiment of a method for generating a classification C as to whether a dose and / or contrast agent administered to a patient was justified. The steps of the method need not be performed in the sequence shown. Rather, the steps can be performed in any order or in parallel. In particular, the steps shown within a dashed frame are considered optional.

[0125] The process can be executed by a computing unit, e.g., by the one associated with Fig. 1 or Fig. 2. Computing unit 20. However, one or more steps of the procedure can be performed by any number of computing devices operating as servers, in a distributed computing system, and / or on the premises of an organization (ORG). For example, one or more computing devices can locally perform some or all of the steps described in Fig. Perform the 3 described steps.

[0126] Optionally, the procedure can be triggered by a step S0 in which a warning message A is received indicating that a dose and / or contrast agent administered to a patient is classified as an outlier, i.e., that the dose and / or contrast agent is below or above a threshold.

[0127] In step S10, multimodal query data QD is determined. The query data QD contains information about the radiological conditions of a performed examination. The query data QD can be obtained as raw data. The query data QD can be preprocessed, in particular, it can be embedded. Thus, in an optional step S12, an embedded multimodal query vector can be obtained.

[0128] The embedded multimodal query vector can be generated from the raw data by the computation unit 20. In an optional step S14, the multimodal query data QD can be embedded using at least one feature extraction method. In an optional step S16, the extracted features can be combined to form the multimodal query vector.

[0129] In step S20, a knowledge base KB is queried to obtain multimodal knowledge data KD, which is at least partially based on the multimodal query data QD. The knowledge data KD contains information about the radiological conditions of an examination, a corresponding classification C, whether a dose and / or contrast agent administered to a patient during the examination was justified or unjustified, and / or a corresponding natural language justification R.

[0130] In step S30, the received query data QD and the received knowledge data KD are made available to a language model LM. For this purpose, a corresponding input prompt P can be entered into the language model LM, which forces the language model LM to process the input data and, in step S40, to generate a classification C and, in step S50, a justification R.

[0131] In step S40, the language model LM is used to create a classification C based on the query data QD and the knowledge data KD, determining whether the dose and / or contrast agent administered to the patient during the examination is justified or not justified.

[0132] In step S50, the language model LM is used to generate a natural language justification R for the classification C based on the query data QD and the knowledge data KD.

[0133] In step S60, the classification C and the justification R are submitted to a user U for approval or modification. Step S60 may also include conducting and presenting an analysis of the significance of certain data characteristics for the classification C and / or the justification R.

[0134] Optionally, in a step S70, dose and / or contrast agent correction measures are initiated if the applied dose and / or contrast agent is deemed unjustified and classification C is approved by user U.

[0135] Optionally, in step S80, feedback from the user U can be obtained regarding the correctness of the classification C and / or the provided natural language justification R.

[0136] In an optional step S90, the parameters of the language model LM can be fine-tuned based on the provided user feedback, thus minimizing a cost function. Additionally or alternatively, the knowledge data KD stored in the knowledge base KB can be updated based on the given feedback, which is the optional step S92.

[0137] Optionally, in step S100 the query data QD and / or the knowledge data KD can be anonymized or pseudonymized.

[0138] The Fig. Figures 4 to 6 show schematic representations of various exemplary data flows in a process for generating a classification of a dose and / or contrast agent as justified, i.e., "OK", or as unjustified, i.e., "not OK". The process can be orchestrated by a computing unit 20, the specific embodiment of which is not critical for the functionality of the process.

[0139] An alarm A can be received by the processing unit 20. Alarm A can be generated by an alarm processing system and sent to the processing unit 20. Alternatively, the processing unit 20 can be configured to generate the alarm A. For this purpose, the processing unit 20 can be capable of monitoring the dose values ​​applied during an examination. Based on the determined dose values, the processing unit 20 can determine whether the applied dose is an outlier. This means that the dose is above or below a certain threshold. The threshold can be defined, for example, based on national regulations.

[0140] The query data QD can be retrieved by Computing Unit 20. The query data QD 20 can be sent to Computing Unit 20 from various components of System 1. This process can be initiated by Computing Unit 20. This means that Computing Unit 20 can send a request to specific components of System 1, for example, when Computing Unit 20 receives an alarm A. The request can then trigger the sending of query data QD to Computing Unit 20. As explained above, the query data QD can originate from various sources, such as a PACS, EMR / EHR, or directly from Imaging Modality 30.

[0141] Alternatively or additionally, the components of system 1 that supply query data QD to the computing unit 20 can also send the query data QD without a request from the computing unit 20. This can be the case, for example, if an alarm A is generated directly in one of the components of system 1. Upon receiving the query data QD, the computing unit 20 can begin to execute the procedure according to one of the described embodiments.

[0142] Multimodal knowledge data (KD) can be retrieved from the knowledge base (KB). Advantageously, the knowledge data (KD) is linked as closely as possible to the query data (QD), since the knowledge data (KD) is provided to the language model (LM) to provide context and improve the generated response of the language model (LM), i.e., the classification (C) and the natural language justification (R) for the classification (C).

[0143] The embedding of query data QD and knowledge data KD into the knowledge base KB can be achieved by generating a vector representation of the data. Since the query data QD and the knowledge data KD originate from different sources, embedding the data and translating it into a numerical representation allows the embedded vectors to be directly provided to the language model LM. Furthermore, this facilitates the processing of the multimodal data, as the features are translated into values ​​regardless of their origin.

[0144] Furthermore, embedding the query data QD and the knowledge data KD facilitates the identification of similarities between the query data QD and the knowledge data KD stored in the knowledge base KB, since the degree of similarity between vectors can be determined mathematically, e.g. using the k-nearest neighbor algorithm.

[0145] The knowledge data KD that are most relevant for assessing the dose value based on the query data QD are retrieved from the knowledge base KB. For example, the n most relevant data samples from the knowledge data KD can be determined, where n is a natural number > 1, preferably > 3, and more preferably > 5.

[0146] All in all, retrieving relevant knowledge data (KD) based on query data (QD) can be part of a Retrieval Augmented Generation (RAG) framework that improves the output generated by the language model (LM).

[0147] Both the query data QD and the knowledge data KD are made available to the language model LM.

[0148] The language model LM can be instructed to generate a classification C of the dose value for which the alarm was generated, based on the query data QD and the knowledge data KD. The applied dose can be classified as "OK" if the dose value is justified considering the specific radiological conditions of the examination. For example, a high dose may be necessary when examining a patient with a high body mass index. The language model LM is then instructed to generate a corresponding natural language justification R for the classification C.

[0149] There are various ways to instruct the LM language model, as shown in the figures. As in Fig. As shown in Figure 4, an input prompt P can be entered into the language model LM, which causes the language model LM to create a classification C and a corresponding justification R.

[0150] As in Fig. As shown in Figure 5, an initial prompt P can alternatively be entered into the language model LM, which causes the language model LM to generate the classification C. The classification C and a second prompt can be provided sequentially as input to the language model LM, which causes the language model LM to create a natural language justification R for the classification C.

[0151] Another alternative will be presented in Fig.Figure 6 illustrates this. The language model LM can contain several different models LM1, LM2 (and optionally LM3, LM4, etc.). A first model, specifically a first language model LM1, can be specifically trained to create a classification C of the applied dose based on the provided query data QD and knowledge data KD. A second model, specifically a second language model LM2, can be trained to generate a natural language justification R for the classification C. The creation of the classification C and the justification R can be completely independent of each other. For example, the first model can be prompted with a first prompt P1 to create the classification C, and the second model can be prompted with a second prompt P2 to create a justification R.Preferably, however, the classification C is provided as a further input for the second model in order to improve the quality of the generated natural language justifications R.

[0152] Regardless of how they were generated, the classification C and the justification R are displayed on interface 10 and can be viewed by a user U. User U can provide corresponding feedback to system 1, specifically to processing unit 20, as visualized by the bidirectional arrows. Alternatively, user U can consider the classification C and / or the natural language justification R to be incorrect and inform system 1, specifically processing unit 20, accordingly.

[0153] The user's response, i.e., user feedback, can be provided to the language model LM to generate an updated classification C and / or natural language justifications R. For example, a corresponding prompt P can be entered by the user U or generated automatically based on the user feedback. This process can be repeated iteratively until the user U is satisfied with the classification C and the natural language justifications R.

[0154] If negative user feedback is received and an update to the classification C and / or the justification R leads to user U's agreement, the language model LM can be fine-tuned with the corresponding data sample and / or the knowledge base KB can be updated by adding the corresponding data sample and / or by replacing an outdated data sample with the new, correct data sample.

[0155] In this way, System 1, especially the knowledge base KB, remains up to date and changes in legal requirements and / or technical developments can be taken into account without having to retrain the language model LM from scratch.

[0156] The various logic blocks, modules, circuits, and algorithm steps described in connection with the embodiments described herein can be implemented as electronic hardware, computer software, or a combination of both. To illustrate this interchangeability of hardware and software, various components, blocks, modules, circuits, and steps have been described above in general terms with respect to their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design requirements for the overall system. Skilled professionals may implement the described functionality in different ways for each specific application, but such implementation decisions should not be interpreted as a departure from the scope of this disclosure or the claims.

[0157] Implemented forms of computer software can be implemented in software, firmware, middleware, microcode, hardware description languages, or any combination thereof. A code segment, or machine-executable instruction, can represent a procedure, function, subroutine, program, routine, module, software package, class, or any combination of instructions, data structures, or program commands. A code segment can be coupled to another code segment or hardware circuitry by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc., can be passed, relayed, or transmitted in any suitable manner, such as memory sharing, message passing, token passing, network transmission, etc.

[0158] The actual software code or specific control hardware used to implement these systems and methods does not constitute a limitation of the claimed features or this disclosure. Therefore, the operation and behavior of the systems and methods have been described without reference to the specific software code, assuming that software and control hardware can be designed such that the systems and methods can be implemented based on this description.

[0159] When implemented in software, the functions can be stored as one or more instructions or code on a non-volatile, computer-readable or processor-readable storage medium. The steps of a method or algorithm disclosed herein can be embodied in a processor-executable software module, which may reside on a computer-readable or processor-readable storage medium. A non-volatile, computer-readable or processor-readable medium includes both computer storage media and physical storage media that facilitate the transfer of a computer program from one location to another. A non-volatile, processor-readable storage medium can be any available medium accessible to a computer.Such non-volatile, processor-readable media can include, for example, RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other tangible storage medium that can be used to store the desired program code in the form of instructions or data structures and that a computer or processor can access. The terms disk and disc used here include Compact Disc (CD), Laserdisc, Optical Disc, Digital Versatile Disc (DVD), Floppy Disc, and Blu-ray Disc, whereby disks typically reproduce data magnetically, while discs reproduce data optically using lasers. Combinations of the media mentioned above should also fall within the scope of computer-readable media.Furthermore, the operations of a procedure or algorithm may be contained as one or any combination or set of codes and / or instructions on a non-volatile processor-readable medium and / or a computer-readable medium that can be included in a computer program product.

[0160] The foregoing description of the disclosed embodiments is intended to enable any person skilled in the art to manufacture or use the embodiments and variants thereof described herein. Various modifications of these embodiments are readily apparent to a person skilled in the art, and the principles defined herein can be applied to other embodiments without departing from the spirit or scope of the subject matter disclosed herein. Therefore, the present disclosure is not intended to be limited to the embodiments shown herein, but rather to have the broadest possible scope of application compatible with the following claims and the principles and new features disclosed herein.

[0161] While various aspects and embodiments have been disclosed, other aspects and embodiments are conceivable. The various aspects and embodiments disclosed serve for illustration and are not to be understood as limiting, the true scope and spirit being specified by the following claims.

Claims

[1] Computer-implemented method for generating a classification (C) as to whether a dose and / or contrast medium administered to a patient during an examination is justified or unjustified, the method comprising the following steps: - Receiving (S10) multimodal query data (QD) containing information about the radiological conditions of a performed examination; - Queries (S20) of a knowledge base (KB) to obtain multimodal knowledge data (KD) at least partially based on the multimodal query data (QD), wherein the knowledge data (KD) includes information about radiological conditions of an examination, a corresponding classification of whether a dose and / or contrast agent applied to a patient during the examination was justified or unjustified, and / or a corresponding natural language explanation for the corresponding classification; - Providing (S30) the received query data (QD) and the received knowledge data (KD) to a language model (LM); - Generating (S40), using the language model (LM), a classification (C) of whether the dose and / or contrast agent applied to the patient during the examination performed is justified or unjustified, based on the query data (QD) and the knowledge data (KD); - Generating (S50), using the language model (LM), a natural language justification (R) for the classification (C) based on the query data (QD) and the knowledge data (KD); - Presenting (S60) the classification (C) and the justification (R) at an interaction interface (10) for approval or modification by a user (U). [2] The method of claim 1, wherein the method further comprises the following step: - Initiate (S70) dose and / or contrast media correction measures if the dose and / or contrast media used is deemed unjustified and the classification (C) is approved by a user (U). [3] Method according to any of the preceding claims, wherein the multimodal query data (QD) and the multimodal knowledge data (KD) each comprise at least two different data types, in particular text data, medical image data and / or sensor data. [4] Method according to any of the preceding claims, wherein the multimodal knowledge data (KD) are stored in the knowledge base (KB) as a plurality of embedded knowledge vectors, and the method further comprises the following step: - Obtaining (S12) at least one embedded multimodal query vector based on the multimodal query data (QD); wherein the multimodal knowledge data (KD) to be obtained from the knowledge base (KB) are selected based on a comparison of the embedded query vector with the embedded knowledge vectors and based on a determination of a similarity between the embedded query vector and the embedded knowledge vectors. [5] Method according to claim 4, wherein the acquisition of the embedded multimodal query vector comprises the following: - Embedding (S14) the multimodal query data (QD) using at least one feature extraction technique; - Combining (S16) the extracted features to form the multimodal query vector. [6] Method according to one of claims 4 or 5, wherein the step of querying the knowledge base (KB) comprises obtaining n knowledge vectors, where n is a natural number greater than 1, and the n knowledge vectors have a higher similarity to the query vector than the other knowledge vectors stored in the knowledge base (KB). [7] A method according to any of the preceding claims, wherein the method comprises the following step: - Anonymizing the query data (QD) and / or the multimodal knowledge data (KD), or - Pseudonymizing (S100) the query data (QD) and / or the multimodal knowledge data (KD). [8] Method according to any of the preceding claims, wherein the language model (LM) is an explainable AI language model, in particular the language model includes interpretable components such as attention mechanisms and saliency maps. [9] A method according to any of the preceding claims, wherein the method further comprises the following step: - Conducting and presenting (S60) an analysis of the significance of certain features of the data for classification (C) and / or justification (R), e.g. by conducting and presenting a Shapley Additive Explanations (SHAP) analysis and / or a Local Interpretable Model-Agnostic Explanation (LIME) analysis. [10] A method according to any of the preceding claims, the method further comprising the following steps: - Receiving (S80) user feedback on the classification (C) and / or on the natural language justification (R) provided for the classification (C); - Fine-tuning (S90) the parameters of the language model (LM) so that a cost function is minimized, and / or updating (S92) the knowledge data (KD) stored in the knowledge base (KB) based on the feedback provided. [11] A method according to any of the preceding claims, wherein the method further comprises the following step: - Receiving (S0) an alarm (A) indicating that a dose and / or contrast agent applied to a patient during an examination is an outlier, wherein a threshold on the basis of which the alarm (A) is received is determined by a statistical procedure, such as Z-scores or percentiles, based on the knowledge data (KD) stored in the knowledge database (KB), wherein the threshold is automatically updated at regular intervals or when the knowledge data (KD) in the knowledge database (KB) is updated. [12] Method according to any of the preceding claims, wherein the knowledge data (KD) in the knowledge base (KB) is updated at regular intervals or on request. [13] Data processing device with means for carrying out the steps of the method according to any one of claims 1 to 12. [14] Computer program product comprising instructions which, when the program is executed by a computer, cause the computer to perform the steps of the method according to any one of claims 1 to 12. [15] Computer-readable storage medium containing instructions which, when executed by a computer, cause the computer to perform the steps of the method according to any one of claims 1 to 12.