Method and infrastructure system for providing machine learning models, anonymizing medical data, and processing and / or analyzing medical data

Through unsupervised or self-supervised training methods, basic machine learning models are trained based on medical data, which solves the problems of limited medical data availability and high training costs, achieves efficient feature extraction and anonymization, simplifies cross-hospital data utilization, and reduces training costs.

CN120671760APending Publication Date: 2025-09-19SIEMENS HEALTHINEERS AG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510313222.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-03-18
Filing Date
2025-03-17
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing technologies for training machine learning models (MLMs) in the medical field face problems such as limited data availability, high training costs, and insufficient customization for specific downstream tasks.

Method used

Using unsupervised or self-supervised training methods, basic machine learning models (fMLMs) are trained based on medical data, features are extracted and stored for anonymization and further training, reducing dependence on raw data. This is particularly suitable for multimodal medical data.

Benefits of technology

It improves the efficiency and quality of medical data feature extraction, reduces training costs, enables data utilization across hospitals and clinical environments under data protection compliance, and simplifies the training process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120671760A_ABST
    Figure CN120671760A_ABST
Patent Text Reader

Abstract

In order to provide a trained machine learning model xMLM (7) for feature extraction from medical data (10, 10 ', 10' '), a base machine learning model fMLM (7, 8) in an untrained or partially trained state is provided, where the fMLM (7, 8) has an architecture that can be trained by means of unsupervised or self-supervised training. The fMLM (7, 8) has an xMLM (7) and at least one downstream machine model nMLM (8) for executing at least one corresponding downstream task. First medical data (10, 10 ', 10' ') is acquired and fMLM (7, 8) is trained in an unsupervised or self-supervised manner based on the first medical data (10, 10', 10 ''). XMLM (7) is extracted from the trained fMLM (7, 8) and stored.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for providing a trained machine learning model for extracting features from medical data. The present invention also relates to a method for anonymizing medical data containing personal data, a method for providing a trained machine learning model for processing and / or analyzing medical data, and a method for processing and / or analyzing medical data. Furthermore, the present invention relates to a corresponding infrastructure system and a computer program product. Background Art

[0002] Trained machine learning models (MLMs), particularly artificial neural networks (ANNs), are widely used in medical settings, particularly in the field of medical imaging. In particular, such MLMs are used for anatomical image classification, identification of medical abnormalities, medical image segmentation, and image processing.

[0003] The publication by O. Ronneberger et al.: "U-Net: Convolutional Networks for Biomedical Image Segmentation" (arXiv:1505.04597) describes the U-Net architecture, a widely used CNN architecture for segmenting images, which can however also be used for other tasks, in particular for image-to-image translation or artifact reduction.

[0004] ResNet is a well-known architecture for artificial neural networks (ANNs), in particular convolutional neural networks (CNNs), described in the publication by K. He et al.: "Deep Residual Learning for Image Recognition", Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 770-778. The ResNet model is widely used in deep learning tasks due to its ability to train very deep neural networks.

[0005] The development and training of MLMs in medicine present numerous challenges. For example, training deep neural networks with millions of parameters requires a large amount of training data to find a set of parameters that performs well. The generalizability of the parameters found depends, among other things, on the amount of training data used, the augmentation of the training data, and the distribution on which the training data is based—that is, the specific choice of training data.

[0006] Some approaches attempt to address this problem by using pre-trained MLMs, such as ImageNet for computer vision, or training large language models using large amounts of book text and strong augmentation or text publicly available on the Internet.

[0007] In the medical environment, many MLM architectures are known that can be trained unsupervised or self-supervised. Sometimes a pre-trained model is assumed here, which can also be pre-trained in a supervised manner, for example.

[0008] For example, in the publication "Self-supervised Learning from 100 Million Medical Images" (arXiv:2201.01283), FC Ghesu et al. propose a method for self-supervised learning of large-scale image features based on contrastive learning and online feature clustering. For this purpose, a large training dataset with over 100 million medical images from different imaging modalities, including X-ray recordings, computed tomography data, magnetic resonance tomography data, and ultrasound data, is used.

[0009] In the publication by A. Dosovitskiy et al.: "An Image is Worth 16x16Words: Transformers for Image Recognition at Scale" (arXiv:2010.11929), the Vision-Transformer architecture is proposed, which makes the concept of neural transformer networks applicable to image data.

[0010] In the publication by A. Kirillov et al.: “Segment Anything” (arXiv:2304.02643), the Segment Anything (SA) project for image segmentation is introduced.

[0011] In the publication "Meta-Transformer: A Unified Framework for Multimodal Learning" (arXiv:2307.10802), Y. Zhang et al. describe a network called Meta-Transformer for processing multimodal data such as natural language, 2D images, 3D point clouds, audio data, video data, time series, and tabular data. Meta-Transformer uses a frozen encoder to perform multimodal perception tasks without the need for paired multimodal training data.

[0012] In the publication "Hierarchical Text-Conditional Image Generation with CLIP Latents" by A. Ramesh et al., a two-stage model is described, in which a prior is used to generate CLIP image embeddings with the help of text labels and a decoder is used, which generates an image related to the image embedding.

[0013] Autoencoders, such as those described in G. Hinton and R. Salakhutdinov, "Reducing the Dimensionality of Data with Neural Networks," Science 313, 504-507 (2006), are multi-layer artificial neural networks with small intermediate layers that can convert high-dimensional data and reconstruct it into a low-dimensional encoding. The architecture of an autoencoder consists of an encoder module followed by a decoder module. The encoder module maps the input to a hidden representation, also called a latent representation, or a representation in a latent space, for example, via multiple fully connected layers. The decoder module maps the hidden representation back to the original input space, for example, via fully connected layers. Autoencoders are particularly unsupervised in their training.

[0014] Pre-trained MLMs, for example from the field of computer vision, can also be helpful for applications in medicine, but they are often insufficient. This is due in particular to the fact that the data typically used for applications in computer vision differ from the image data in medicine, both in terms of image content and in terms of the type of image data.

[0015] For many tasks in medicine, the availability of medical data is significantly limited for various reasons. For example, data protection regulations based on government regulations and laws can make the use of large amounts of data outside of hospitals difficult, cumbersome, and costly. When using medical data from multiple hospitals, the applicable local data protection regulations must be implemented.

[0016] Medical data preparation for training MLMs is not integrated into clinical workflows, requiring additional effort to extract and prepare the data for use in deep learning training. This includes, for example, anonymizing the data, generating annotations in the case of supervised learning, and normalizing data across different hospitals. Data anonymization is particularly application-dependent, as different downstream tasks require varying amounts of additional patient information, such as gender, age, and medical history.

[0017] Furthermore, training a complete MLM for each individual task is inefficient in terms of development effort and data usage, and will result in a trained MLM that is only tailored to a specific ad hoc downstream task. Summary of the Invention

[0018] The object of the present invention is to at least partially overcome the above-mentioned disadvantages.

[0019] This object is achieved by the subject matter of exemplary embodiments according to the present invention. Advantageous developments and preferred embodiments are the subject matter of exemplary embodiments according to the present invention.

[0020] The present invention is based on the following concept: training a basic machine learning model fMLM based on medical data by unsupervised or self-supervised training, wherein the basic machine learning model has an MLM, xMLM, for feature extraction, and extracting the trained xMLM from the trained fMLM and storing it for subsequent use, such as for anonymizing medical data and / or for training an MLM, sMLM for processing and / or analyzing medical data.

[0021] According to a first aspect of the present invention, a method, in particular a computer-implemented method, for providing a trained MLM, xMLM for extracting features from medical data is proposed. In this case, a basic machine learning model fMLM is acquired in an untrained or partially trained state, wherein the fMLM has an architecture that can be trained by means of unsupervised or self-supervised training. The fMLM comprises an xMLM and at least one downstream MLM (nMLM) for performing at least one corresponding downstream task, with the downstream task being in particular a medical task. First medical data are acquired and, based on the first medical data, an fMLM comprising an xMLM and in particular an xMLM and at least one nMLM is trained in an unsupervised and / or self-supervised manner. The trained xMLM is stored, for example extracted from the trained fMLM and stored, in particular on a computer-readable storage medium and / or as a hardware implementation.

[0022] The at least one trained nMLM can also be stored, in particular on a computer-readable storage medium and / or as a hardware implementation, for example together with or separately from the trained xMLM.

[0023] Unless otherwise specified, all steps of the computer-implemented method for providing a trained xMLM can be performed by a first data processing system comprising at least one first data processing device. In particular, the at least one first data processing device is configured or adapted to perform the steps of the computer-implemented method. For this purpose, the at least one first data processing device can, for example, store a computer program comprising instructions that, when executed by the at least one first data processing device, cause the at least one first data processing device to perform the computer-implemented method. The expressions "data processing system" and "at least one data processing device" are used interchangeably here and hereinafter. This also applies to corresponding expressions derived therefrom.

[0024] In the case where the at least one first data processing device includes two or more first data processing devices, a specific step performed by the at least one first data processing device can also be understood as different first data processing devices performing different steps or performing different parts of a step. In particular, it is not necessary for every first data processing device to perform the step. In other words, the execution of the step can be distributed among the two or more first data processing devices.

[0025] From each embodiment of the computer-implemented method for providing a trained xMLM, a corresponding embodiment of the non-purely computer-implemented method follows, in that it comprises a corresponding step for generating training data for unsupervised or self-supervised training of the fMLM.

[0026] Generally speaking, a trained MLM can mimic the cognitive functions that humans associate with other humans' thinking. In particular, by being trained based on training data, the MLM can adapt to new situations and detect and extrapolate patterns. Another term for a trained MLM is a "trained function."

[0027] In general, the parameters of the MLM can be adjusted or updated through training. In particular, supervised training, semi-supervised training, unsupervised training, reinforcement learning (English: reinforcement learning) and / or active learning can be used. In addition, representation learning (English: representation learning), which is also called feature learning (English: feature learning), can be used. In particular, the parameters of the MLM can be iteratively adjusted through multiple training steps. In particular, a specific loss function, also called a cost function, can be minimized during training. When training an artificial neural network (ANN), the backpropagation algorithm can be used in particular.

[0028] In particular, the MLM may include an ANN, a support vector machine, a decision tree, and / or a Bayesian network, and / or the MLM may be based on k-means clustering, Q-learning, a genetic algorithm, and / or association rules. In particular, the ANN may be or include a deep neural network, a convolutional neural network, a CNN (convolutional neural network), or a convolutional deep neural network. Furthermore, the ANN may be an adversarial network, a deep adversarial network, and / or a generative adversarial network.

[0029] Here and hereinafter, acquiring an MLM or a portion of an MLM may be understood to mean, for example, acquiring it in a computer-readable form, for example, by storing it on a data carrier. In particular, acquiring does not include any further method steps for generating an MLM or a portion of an MLM, for example, steps for training an MLM or a portion of an MLM, unless otherwise indicated.

[0030] Obtaining an fMLM in an untrained state can be understood as only presetting the architecture of the fMLM or providing the fMLM only in an initialized manner. If an fMLM in a partially trained state is obtained, this can be understood as providing the fMLM in a pretrained manner, wherein the pretraining can be performed in a supervised, unsupervised and / or self-supervised manner, but is not part of the method according to the present invention, unless otherwise specified. Pretraining can also be performed based on medical data and / or based on other data. The fMLM can be a known MLM architecture, in particular a known ANN architecture. It is also feasible that, although xMLM and at least one nMLM are known MLM architectures, the combination of xMLM and at least one nMLM is not already known. In both cases, unsupervised and / or self-supervised training can be performed according to known training methods, in particular using known loss functions.

[0031] In particular, ANNs are considered for the xMLM and at least one nMLM, such as the above-mentioned autoencoders, U-Net, ResNet, Segment Anything, Vision-Transformer, Meta-Transformer, other transformer networks and transformer networks described in the publications of Ramesh et al. and Ghesu et al.

[0032] In supervised training (or supervised learning), the training data used for training is manually, automatically, or partially automatically provided with labels, which are so-called true values, also called true labels. The loss function used can then effectively compare the corresponding predictions of the MLM to be trained with the corresponding true values, and the MLM can be adjusted based on the results.

[0033] In unsupervised training (or unsupervised learning), the MLM is trained or adapted without the need for labeled training data, i.e., without prior knowledge of the true values, and without rewards obtained from the environment, as used in reinforcement learning.

[0034] In the relevant literature, the terms "self-supervised training" (English: self-supervised training or self-supervised learning) are not used in a completely consistent manner. The term "self-supervised training" is used here and hereinafter in such a way that it includes training methods that correspond neither to supervised training nor to unsupervised training, but in which, as in unsupervised training, no labeling of the training data is required, i.e. no ground truth values ​​are required to be known in advance. This includes in particular training methods in which the MLM to be trained and / or the auxiliary MLM produce labeling or valid labeling, in particular without human intervention. It also includes training methods in which the training data are intentionally modified, in particular degraded, masked, noisy, etc., and the MLM to be trained is adjusted so that it reconstructs the unmodified training data, for example in ANNs of the autoencoder class.

[0035] By applying an fMLM to input data, the input data is converted, in particular with the aid of an xMLM, into encoded features, also referred to as embeddings or latent features, or features in the latent space of the fMLM. At least one nMLM is applied to each of the encoded features in order to perform a corresponding predefined task. Depending on the architecture of the fMLM, the xMLM may also be referred to as an encoder module or an embedding module. In some fMLM architectures, the at least one nMLM may also be referred to as at least one decoder module.

[0036] The tasks associated with at least one nMLM can be of different nature. In particular, it is not necessary for these tasks to correspond to the subsequent purpose of the xMLM or the purpose of the features extracted thereby, although this is in principle possible. For example, the tasks of the nMLM can be the most accurate reconstruction of the input data, a segmentation task, a classification task, an object recognition task, etc. These tasks are also referred to as downstream tasks in this article because they are performed after the feature extraction by the xMLM.

[0037] Medical data can be various types of data from a medical or clinical environment, particularly generated in a clinic or hospital for one or more specific individuals (particularly patients). Medical data can particularly include image data, such as two-dimensional or three-dimensional image data, which may also include a second dimension, such as time. Image data can include, for example, X-ray projection images, raw data from X-ray detectors, reconstructions from computed tomography (CT), raw data from magnetic resonance tomography (MRI), pre-processed raw data from MRI, MRI reconstructions, image data from ultrasound imaging, image data from positron emission spectroscopy (PET), and / or other medical imaging methods. Image data can also include planning and / or reference data for performing imaging methods and / or surgical procedures or other medical treatments or diagnostic methods. Medical data can also particularly include text data, particularly unstructured text data, such as doctor's reports, medical history text data, and the like, as well as tabular or otherwise structured text data and / or numerical data, such as laboratory reports, test results, patient identification data, and the like. Generally speaking, medical data may contain, in particular, data that is specific to an individual and therefore important, for example, under data protection law.

[0038] A trained xMLM can be understood as the result of the method according to the first aspect of the present invention. The trained and stored xMLM can be used for various purposes, for example, for training an MLM (sMLM) for processing and / or analyzing medical data. The sMLM can be one of the nMLMs of the fMLM or another nMLM of the aforementioned type. To train the sMLM, it can be combined with the trained xMLM so that the trained xMLM is applied to other training data to generate other coding features, and the sMLM can then be applied to the other coding features to perform corresponding predefined tasks associated with the sMLM.

[0039] Because the xMLM has been trained by the method according to the first aspect of the invention, it has in particular learned to extract the most important features in the medical context from the medical data. As a result, the training effort for the sMLM is significantly reduced, in particular with respect to the amount of training data required and / or the number of training iterations or training epochs required. Another advantage is that the fMLM is trained unsupervised or self-supervised, i.e. in particular without manual labeling. The medical data generated in large quantities in a clinical environment, in particular in routine hospital operations, can already be used to train the fMLM and thus the xMLM, without having to leave the hospital or its data processing infrastructure. No specialists are required to prepare and label the medical data, and no additional time is required.

[0040] In particular, from a data protection perspective, no restrictions are expected, since the medical data can be processed where they are generated, in particular in a hospital. If a trained xMLM is provided for the first time, it can be easily and without data protection considerations extracted from the hospital environment, for example in order to train an sMLM. For this purpose, a large number of other coding features can be generated, in particular, by applying the trained xMLM to the second medical data. The other coding features are anonymous data, since the underlying second medical data cannot be easily reconstructed from the anonymous data, in particular not by means of the trained xMLM alone. Subsequently, the anonymous data can be used to train the sMLM outside the clinical infrastructure, without the second medical data themselves. This makes it possible, in particular, to train the sMLM in a simple and risk-free manner from a data protection perspective, in research and / or development institutions, etc., or even in commercial enterprises or manufacturers of medical technology devices, software and solutions.

[0041] According to at least one embodiment, the fMLM is an ANN. Parts of an ANN can also be referred to as ANNs or modules of an ANN, if necessary. In particular, the xMLM and at least one nMLM are ANNs. In particular, the xMLM is a deep ANN, i.e., having one or more, preferably multiple, hidden layers.

[0042] According to at least one embodiment, the fMLM has an architecture that is trainable based on multimodal data by means of unsupervised or self-supervised training, and the first medical data is first multimodal medical data.

[0043] Multimodal data is understood here and below to mean data sets that include two or more data sets in different formats and / or from different sources, in particular imaging modalities. For example, data that includes not only image data but also text data is an example of multimodal data. For example, data that includes not only X-ray projection images but also MRI reconstructions is also an example of multimodal data. In this case, the multimodal data may, for example, at least partially relate to the same person.

[0044] In medical contexts, multimodal data is frequently generated, such as image data from imaging modalities, text data from physician reports, tabular data from laboratory reports, and so on. Therefore, it is particularly advantageous to train an fMLM based on multimodal medical data, thereby utilizing all available types of information. By combining data in different formats and / or from different sources, the individual components of the first medical data are contextualized, allowing the xMLM to be trained particularly efficiently for feature extraction of particularly important features in the medical data, either through unsupervised or self-supervised training. In other words, as described above, the quality of the features extracted by the xMLM is significantly improved for continued use in medical contexts.

[0045] For example, the architectures mentioned above for MLMs, especially ANNs, especially autoencoders, SegmentAnything, Vision-Transformer, Meta-Transformer, and the architectures described in the publications by Ramesh et al. and Ghesu et al. are suitable for processing multimodal data.

[0046] According to at least one embodiment, the first medical data include first image data from an imaging method.

[0047] For example, an imaging method is performed, in particular by means of an imaging modality, in order to generate at least a portion of first medical data, in particular first image data from the imaging method. In this embodiment, the method is not, as mentioned above, purely computer-implemented.

[0048] Imaging methods can be, for example, X-ray imaging methods, CT methods, MRT methods, PET methods, ultrasound imaging methods or the like.

[0049] Image data from imaging methods is particularly valuable and important information as a basis for feature extraction in medical contexts. This is particularly true because subsequent potential applications of a trained xMLM can often be based on image data from the imaging method as input data. Therefore, unsupervised or self-supervised training allows for particularly efficient training of xMLMs for feature extraction of particularly important features in medical data. As described above, the quality of features extracted by xMLM is improved for continued use in medical contexts.

[0050] For example, the imaging modality can be configured to transmit the first image data to the first data processing system after its generation, in particular directly, and the first data processing system can then train the fMLM as described, at least based on the first image data. This can be done, in particular, during regular use of the imaging modality, for example, also after the fMLM has been initially trained and / or for the purpose of refining or further training the fMLM. Thus, in particular, a "lifelong learning" approach can also be implemented.

[0051] According to a second aspect of the present invention, a method, in particular a computer-implemented method, for anonymizing medical data is provided. The method according to the first aspect of the present invention is performed to provide a trained xMLM. Second medical data, in particular second multimodal medical data, containing personal data is acquired. The second medical data is anonymized by encoding the second medical data by applying the trained xMLM to the second medical data.

[0052] In other words, according to the second aspect of the present invention, it is proposed to anonymize medical data using the xMLM provided by the method according to the first aspect of the present invention.

[0053] Unless otherwise stated, all steps of the computer-implemented method for anonymizing medical data can be performed by the first data processing system or a further data processing system. The statements regarding the computer-implemented method according to the first aspect of the invention apply analogously.

[0054] The encoded second medical data as the output of the trained xMLM are anonymous data, because the second medical data as the basis cannot be reconstructed in a simple way from the anonymous data, in particular, it cannot be reconstructed only by the trained xMLM. The anonymous data can therefore be further used in a risk-free manner from a data protection perspective. The continued use can include training the sMLM as mentioned above. However, it is also feasible that the anonymous data is only transmitted or transferred in a secure manner and the second medical data is reconstructed again by a trustworthy department with a corresponding nMLM. The nMLM can therefore be regarded as a key for data reconstruction. For example, in this case, the xMLM can form an automatic encoder or other MLM for data reconstruction together with the nMLM.

[0055] According to at least one embodiment, the second medical data include second image data from an imaging method.

[0056] For example, an imaging method is performed in order to generate at least a portion of second medical data, in particular second image data resulting from the imaging method.

[0057] According to at least one embodiment, the personal data include image data from an imaging method and / or text data relating to a medical assessment or judgment of the patient and / or tabular data relating to a medical assessment or judgment of the patient and / or numerical data relating to a medical assessment or judgment of the patient and / or data relating to the patient's identity.

[0058] According to a third aspect of the present invention, a method, in particular a computer-implemented method, is provided for providing a trained MLM (sMLM) for processing and / or analyzing medical data. The method according to the second aspect of the present invention is performed to generate encoded second medical data. The sMLM is obtained, in particular, in an untrained or partially trained state. The sMLM is trained unsupervised or self-supervised based on the encoded second medical data.

[0059] In other words, according to the third aspect of the present invention, it is proposed to provide a trained sMLM using the xMLM provided by the method according to the first aspect of the present invention.

[0060] Unless otherwise stated, all steps of the computer-implemented method for providing a trained sMLM can be performed by the first data processing system or the second data processing system. The explanations regarding the computer-implemented method according to the first aspect of the invention apply analogously.

[0061] In the method according to the third aspect of the present invention, the sMLM is trained directly based on the anonymous data, i.e., the encoded second medical data, without requiring the second medical data itself. This can be done, for example, by applying the xMLM to the second medical data in a hospital to generate anonymous data, which is then transferred from the hospital to another department, such as a research and / or development institute or a service company, and the sMLM is trained there based on the anonymous data.

[0062] The sMLM is one of the at least one nMLM of the fMLM, which is used when training the fMLM, in particular the xMLM. In this case, the sMLM is provided, for example, already pre-trained. However, the sMLM can also be different from the at least one nMLM, in particular, the sMLM can be associated with different tasks than the at least one nMLM.

[0063] Because the xMLM is already trained on the first medical data according to the first aspect of the present invention before the sMLM is trained on the encoded second medical data, significantly less training data is required to train the sMLM than in conventional methods. This is due to the fact that the xMLM can already extract features that are particularly important for medical applications at the beginning of the sMLM training.

[0064] In some embodiments, it is also possible to use further training data, for example publicly available or conventionally anonymized data, in addition to the encoded second medical data, for training the sMLM.

[0065] To train the sMLM, methods known per se can be used. For example, methods such as those used to train the fMLM can be used, wherein the xMLM is, however, frozen, i.e., no longer changed. However, it is also possible to further train the xMLM together with the sMLM.

[0066] According to at least one embodiment, to train an sMLM based on the encoded second medical data, a processing and / or analysis result is predicted by applying the sMLM to the encoded second medical data. A predetermined loss function is evaluated in relation to the predicted processing and / or analysis result. The sMLM is updated in relation to the result of the loss function evaluation.

[0067] If the sMLM is an ANN, then the updating of the sMLM in particular includes updating the weights of the sMLM, in particular using a corresponding algorithm, such as a back-propagation algorithm.

[0068] According to at least one embodiment, the xMLM remains unchanged while the sMLM is trained based on the encoded second medical data.

[0069] In other words, the xMLM is not updated depending on the result of the loss function evaluation. The xMLM is therefore particularly frozen during the training of the sMLM. It is also possible to train the sMLM without the involvement of the xMLM if a corresponding amount of second medical data is first anonymized by the xMLM and then provided for training the sMLM.

[0070] In this way, the training effort for training the sMLM can be further reduced.

[0071] According to at least one embodiment, the training of the sMLM is performed based on the encoded second medical data by means of a second data processing system, which has no access to the second medical data.

[0072] Accordingly, it is advantageously possible to generate and anonymize the second medical data within a hospital or another secure area and to train the sMLM outside the hospital or secure area without risk from the perspective of data protection of the second medical data.

[0073] The second data processing system not having access to the second medical data can be achieved, for example, by a hard separation of the second data processing system from the memory storing the second medical data or by protecting the second medical data by means of encryption, password protection or the like.

[0074] According to a fourth aspect of the present invention, a method for processing and / or analyzing medical data is provided. The method according to the third aspect of the present invention is performed to provide a trained sMLM. Further coded data is acquired. Alternatively, third medical data, particularly third bimodal medical data, is acquired and further coded data is generated by applying the xMLM to the third medical data. Processing and / or analysis results are generated by applying the trained sMLM to the further coded data.

[0075] In other words, according to the fourth aspect of the present invention, it is proposed to use the sMLM provided by the method according to the third aspect of the present invention to process and / or analyze medical data, in particular to use the xMLM provided by the method according to the first aspect of the present invention and the sMLM provided by the method according to the third aspect of the present invention to process and / or analyze medical data.

[0076] Unless otherwise stated, all steps of the computer-implemented method for processing and / or analyzing medical data can be performed by the first data processing system, the second data processing system, or the third data processing system. The explanations regarding the computer-implemented method according to the first aspect of the invention apply analogously.

[0077] According to at least one embodiment, the third medical data include third image data from an imaging method.

[0078] For example, an imaging method is performed in order to generate at least a portion of third medical data, in particular third image data resulting from the imaging method.

[0079] According to at least one embodiment, the sMLM is an MLM for anatomical image classification. In particular, the third medical data contain further image data from the imaging method.

[0080] The output data of sMLM can then, for example, contain one of a plurality of predefined first-category additional image data and / or an object shown by the additional image data. Algorithms for medical image processing are generally optimized for specific anatomical content. However, in daily clinical work, it is not always guaranteed that the correct protocol is used in Dicom header files, etc., and thus the correct labels are used. Therefore, it is advantageous to classify the additional image data to ensure that the algorithm is correctly used for medical image processing. Different first categories or anatomical contents can, for example, specifically specify body parts such as "hand", "foot", "chest", etc., or organs such as "liver", "heart", "brain", etc.

[0081] According to at least one embodiment, the sMLM is an MLM for detecting diseases or abnormalities, in particular medical or anatomical abnormalities. In particular, the third medical data include further image data from an imaging method.

[0082] The output data of the sMLM can then, for example, include a plurality of predefined second-category additional image data and / or one of the objects depicted by the additional image data and / or the position of the object in the image data. In particular, for such applications, only very small amounts of training data are typically available. Therefore, the present invention is particularly advantageous in this context.

[0083] According to at least one embodiment, the sMLM is an MLM for medical image segmentation. In particular, the third medical data contain further image data from an imaging method.

[0084] The output data of sMLM can, for example, include segmented images in which regions or pixels of other image data belonging to one or more predefined object classes are correspondingly masked, or in which one or more object classes are associated with pixels or regions. Object classes can, for example, specify specific organs, blood vessels, bone structures, medical tools such as catheters, stents, or guidewires, etc. For such applications in particular, very little training data is often available. Therefore, the present invention is particularly advantageous in this regard.

[0085] According to at least one embodiment, the sMLM is an MLM for image processing. In particular, the third medical data contain further image data from the imaging method.

[0086] The output data of the sMLM then contain, in particular, processed versions of the other image data. The processing can include, for example, noise reduction, geometric transformations such as rotation or distortion correction, contrast enhancement, other digital filtering, etc.

[0087] For each of the above-mentioned aspects of the present invention, other embodiments of the corresponding method can be derived from different design solutions of the method according to other aspects of the present invention. In particular, the various features and corresponding explanations and advantages of the different embodiments of the method according to one aspect of the present invention can be similarly transferred to the corresponding embodiments of the method according to other aspects of the present invention.

[0088] According to a fifth aspect of the present invention, an infrastructure system is provided for performing the method according to the fourth aspect of the present invention and the method according to the present invention for providing a trained sMLM. The infrastructure system comprises a first data processing system adapted to perform the method according to the first aspect of the present invention for providing a trained xMLM and the method according to the second aspect of the present invention for generating anonymous data. The infrastructure system comprises a second data processing system adapted to train the sMLM based on the anonymous data.

[0089] In this disclosure, the terms "data processing system" and "at least one data processing device" are used interchangeably. A data processing device is particularly understood to include a data processing device that includes processing circuitry. A data processing device therefore processes data in particular for performing computational operations. This also includes operations performed for indexed access to data structures, such as translation tables or LUTs (English: "look-up tables"), as well as signal processing processes implemented in hardware.

[0090] The data processing device may include, in particular, one or more computers, one or more microcontrollers, and / or one or more integrated circuits, such as one or more application-specific integrated circuits (ASICs), one or more field-programmable gate arrays (FPGAs), and / or one or more systems-on-chips (SoCs). The data processing device may also include one or more processors, such as one or more microprocessors, one or more central processing units (CPUs), one or more graphics processing units (GPUs), and / or one or more signal processors, in particular one or more digital signal processors (DSPs). The data processing device may also include a physical or virtual combination of computers or the other mentioned units.

[0091] In various embodiments, the data processing device comprises one or more hardware and / or software interfaces and / or one or more memory units.

[0092] The memory cell can be designed as a volatile data memory, such as a random access dynamic memory (DRAM) or a random access static memory (SRAM), or a non-volatile data memory, such as a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a flash memory or a flash EEPROM, a random access ferroelectric random access memory (FRAM), a random access magnetoresistive random access memory (MRAM), or a random access phase-change random access memory (PCRAM).

[0093] According to at least one embodiment, the second data processing system has no access to the first medical data and has no access to the second medical data.

[0094] According to at least one embodiment, the first data processing system is suitable for acquiring third medical data and generating further anonymous data by applying the trained xMLM to the third medical data. The first data processing system in this case particularly stores the trained xMLM.

[0095] According to at least one embodiment, the first data processing system is suitable for generating processing and / or analysis results by applying the trained sMLM to further anonymous data. The first data processing system in this case in particular stores the trained sMLM.

[0096] According to at least one embodiment, the infrastructure system has a third data processing system that is suitable for generating processing and / or analysis results by applying the trained sMLM to further anonymous data. The third data processing system in this case particularly stores the trained sMLM.

[0097] According to at least one embodiment, the infrastructure system has a hospital which contains the first data processing system and / or the third data processing system.

[0098] According to at least one embodiment, the infrastructure system has a research and / or development facility which contains the second data processing system and is physically separated from the hospital.

[0099] According to a sixth aspect of the present invention, a computer program is provided.

[0100] According to at least one embodiment, the computer program comprises instructions which, when executed by at least one data processing system, cause the at least one data processing system to perform the method for providing a trained xMLM according to the first aspect of the invention and / or the method for anonymizing medical data according to the second aspect of the invention.

[0101] According to at least one embodiment, the computer program comprises first instructions which, when executed by a first data processing system, cause the first data processing system to perform the method for providing a trained xMLM according to the first aspect of the invention and the method for anonymizing medical data according to the second aspect of the invention. The computer program comprises second instructions which, when executed by a second data processing system, cause the second data processing system to train an sMLM based on the anonymized data.

[0102] According to at least one embodiment, a computer program includes first instructions that, when executed by a first data processing system, cause the first data processing system to perform the method for providing a trained xMLM according to the first aspect of the present invention and the method for anonymizing medical data according to the second aspect of the present invention. The computer program includes second instructions that, when executed by a second data processing system, cause the second data processing system to train an sMLM based on anonymous data. The computer program includes third instructions that, when executed by a third data processing system, cause the third data processing system to generate processing and / or analysis results by applying the trained sMLM to further anonymous data.

[0103] The instruction, the first instruction, the second instruction and / or the third instruction can each exist as a program code, for example. The program code can be provided, for example, as binary code or assembly language and / or source code of a programming language, such as C language and / or as a program script, such as Python.

[0104] According to another aspect of the present invention, a computer-readable storage medium is provided, which stores the computer program according to the present invention.

[0105] The computer program and the computer-readable storage medium are respectively a computer program product having instructions, first instructions, second instructions and / or third instructions.

[0106] In the foregoing and hereinafter, the solution according to the invention is described not only with respect to the claimed system but also with respect to the claimed method. Features, advantages, or alternative embodiments can be associated with other claimed subject matter, and vice versa. In other words, system claims and embodiments can be improved by incorporating corresponding method descriptions or claimed features. In such cases, the functional features of the method are implemented by the physical elements of the system.

[0107] Furthermore, above and below, solutions according to the invention are described with respect to methods and systems for applying trained MLMs and with respect to methods and systems for providing trained MLMs. Features, advantages or alternative embodiments can be associated with other claimed subject matter and vice versa. In other words, the claims and embodiments for providing trained MLMs can be improved with the aid of features described or claimed in conjunction with the application of trained MLMs. In particular, the data sets used in the methods and systems can have the same properties and features as the corresponding data sets used in the methods and systems for providing trained MLMs, and the trained MLMs provided by the corresponding methods and systems can be used in the methods and systems.

[0108] Further features and combinations of features of the invention are apparent from the drawings and their description, as well as from the claims. In particular, further embodiments of the invention do not necessarily have to include all features of one of the claims. Further embodiments of the invention may have features or combinations of features not mentioned in the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0109] The present invention will be described in detail below based on specific embodiments and the associated schematic diagrams. In the figures, identical or functionally identical elements are provided with identical reference numerals. The description of identical or functionally identical elements may not necessarily be repeated for different figures.

[0110] In the accompanying drawings:

[0111] Figure 1 A schematic diagram showing an exemplary embodiment of an infrastructure system according to the present invention;

[0112] Figure 2 A schematic block diagram showing an exemplary embodiment of a method for providing a trained MLM for feature extraction according to the present invention;

[0113] Figure 3 A schematic block diagram showing an exemplary embodiment of a method for anonymizing medical data according to the present invention;

[0114] Figure 4 A schematic block diagram shows an exemplary embodiment of a method according to the present invention for providing a trained MLM for processing and / or analyzing medical data or an exemplary embodiment of a method according to the present invention for processing and / or analyzing medical data;

[0115] Figure 5 A schematic diagram showing an artificial neural network;

[0116] Figure 6 A schematic diagram showing a convolutional neural network; and

[0117] Figure 7 Schematic diagram showing a convolutional neural network based on the U-Net architecture. DETAILED DESCRIPTION

[0118] Figure 1 A schematic diagram shows an exemplary embodiment of an infrastructure system 1 according to the invention.

[0119] The infrastructure system 1 has a first data processing system 3, which is arranged, for example, in a hospital 2. The first data processing system 3 is suitable for carrying out the method according to the first aspect of the invention for providing a trained MLM, xMLM 7, for feature extraction and for carrying out the method according to the second aspect of the invention for anonymizing medical data. The infrastructure system 1 has a second data processing system 6, which is arranged, in particular, separately from the hospital 2 and is suitable for training an MLM for processing and / or analyzing medical data based on encoded second medical data.

[0120] Figure 2 A schematic block diagram of an exemplary embodiment for providing a trained xMLM 7 according to the present invention is shown.

[0121] For this purpose, basic machine learning models fMLM 7, 8 are provided in an untrained or partially trained state. fMLM 7, 8 have an architecture that can be trained using unsupervised or self-supervised training. fMLM includes an untrained or partially trained xMLM and an untrained or partially trained downstream machine learning model nMLM 8 for performing downstream tasks. First medical data 10 is acquired. The first medical data 10 can, for example, include image data 10a of one or more imaging modalities 4 of a hospital 2 and / or text data 10b from a database 5 of the hospital 2 and / or tabular data 10c from the database 5.

[0122] The fMLMs 7 , 8 including the xMLM 7 and the nMLM 8 are trained unsupervisedly or self-supervised based on the first medical data 10 , and the trained xMLM 7 is stored, for example on a storage device of the first data processing system 3 .

[0123] For example, xMLM 7 can encode first medical data 10 and nMLM 8 can reconstruct the first medical data 10 based on the encoded first medical data 9. The loss function can then be evaluated based on the reconstructed first medical data 11 and based thereon xMLM 7 and nMLM 8 can be updated.

[0124] Figure 2 A schematic block diagram of an exemplary embodiment of the method for anonymizing medical data according to the present invention is shown. In this case, a trained xMLM 7 is generated, as described with respect to Figure 2 As explained above, second medical data 10' is acquired, which contains personal data. The personal data can include, for example, image data 10a' from an imaging method and / or text data 10b' related to a medical evaluation or diagnosis of the patient and / or tabular data 10c' and / or numerical data and / or data related to the patient's identity.

[0125] The second medical data 10' are anonymized by encoding them by applying the trained xMLM 7 to the second medical data 10'. The encoded second medical data 9' can then be removed from the hospital 2 in compliance with data protection for further use.

[0126] Figure 4 A schematic block diagram shows an exemplary embodiment of a method according to the present invention for providing a method for processing and / or analyzing medical data sMLM 12 .

[0127] Here, a trained xMLM 7 is generated, as described with respect to Figure 2 As explained. Further medical data 10 ″ is acquired, which for example include personal data. The personal data can for example include image data 10 a ″ from an imaging method and / or text data 10 b ″ relating to a medical assessment or decision about the patient and / or tabular data 10 c ″ and / or numerical data and / or data relating to the patient's identity.

[0128] The further medical data 10 ″ are encoded by applying the trained xMLM 7 to the further medical data 10 ″, in particular by the first data processing system 3 . The sMLM 12 is trained unsupervisedly or self-supervised based on the encoded further medical data 9 ″, in particular by the second data processing system 6 .

[0129] In order to train the sMLM 12 based on the encoded further medical data 9″, a processing and / or analysis result 13 is predicted by applying the sMLM 12 to the encoded further medical data 9″. A predefined loss function is evaluated in relation to the predicted processing and / or analysis result 13, and the sMLM 12 is updated in relation to the result of the evaluation of the loss function.

[0130] The trained sMLM 12 can then be stored, for example, on a storage device of the first data processing system 3. The first data processing system 3 can then execute the method according to the present invention for processing and / or analyzing medical data. To this end, third medical data are acquired, in particular from an imaging modality 4 and / or a database 5, and encoded data are acquired by applying the trained xMLM 7 to the third medical data. Processing and / or analysis results are generated by applying the trained sMLM 12 to the third encoded data.

[0131] Different implementations according to aspects of the present invention propose the use of a multimodal xMLM 7 that is capable of computing universal embeddings, such as image embeddings, text embeddings for clinical reports, patient data, etc. For example, the xMLM 7 can be trained unsupervised or self-supervised, resulting in a powerful model adapted to the medical domain for different potential tasks, which enables the training of the sMLM 12 in a very short learning phase.

[0132] According to various implementations of aspects of the present invention, such a universal multimodal xMLM 7 adapted to the medical field can be used to process medical data. This can utilize the fact that embeddings calculated using deep ANNs are anonymous, low-dimensional, and shareable, while simultaneously meeting requirements for personal data protection.

[0133] According to various implementations of aspects of the present invention, the xMLM is trained with a large amount of data, enabling it to compute universal embeddings that are not fixed for a specific downstream task. The training data is medical data, which can include, for example, image data from various imaging modalities 4, such as X-ray, MRT, CT, ultrasound, etc., but can also include patient data, such as gender, age, etc., as well as diagnostic data, clinical reports, etc.

[0134] In order to train xMLM in particular, different approaches can be followed in order to collect the required training data. For example, "offline learning" can be used, where a large data set is collected and then used to train all data sets.

[0135] On the other hand, "continuous learning" can also be used, either additionally or alternatively. Here, new data is used for further training, as soon as it is available, without having to use previously collected data, for example in the form of large databases. This also includes the method of "lifelong learning," in which learning continues continuously based on new data, similar to human learning.

[0136] According to various implementations of aspects of the present invention, lifelong learning can be used, in particular, to continuously refine fMLMs, and in particular xMLMs. For this purpose, corresponding medical data accumulated in clinical settings can be directly transferred and used for training, and conventional clinical procedures can then be circumvented. This allows for particularly efficient use of accumulated medical data.

[0137] According to different implementations of aspects of the present invention, image data generated by means of an imaging modality can, for example, be transmitted via a corresponding data transmission infrastructure to a data processing system, which trains or continues to train or refine the fMLM based on it, in particular in the sense of a lifelong learning approach.

[0138] Different implementations according to aspects of the present invention enable the training of the sMLM to be performed in several steps, since it only has to be trained to perform the correspondingly set task starting from a universal embedding, while the already trained xMLM is either frozen or only slightly adapted to the downstream task.

[0139] According to different implementations of aspects of the present invention, training xMLM based on medical data can achieve good generalization and reduce the amount of data required to train sMLM.

[0140] According to different implementations of aspects of the present invention, by computing universal embeddings with the help of trained xMLM in the hospital 2, medical data can be anonymized and subsequently forwarded for different purposes. The embeddings can be used for further training in different frameworks, such as through cloud learning, federated learning, or centralized learning.

[0141] In particular, the use of low-dimensional embeddings offers various advantages, for example with regard to reducing the amount of data to be transmitted and the required computing power.

[0142] In particular, the generation of the embedding by xMLM can be realized on different hardware devices, for example also on the medical modality 4 itself, on a PC within the hospital, on a mobile device and / or in the cloud.

[0143] Examples of useful ANN architectures include, in some embodiments, a transformer network that can be trained unsupervised, for example, to fill artificially created gaps in images and text, or a generator network that can be trained unsupervised, for example, to generate image content from noise and text input by inverting a sequential noise addition process.

[0144] Figure 5 One embodiment of an artificial neural network ANN 800 is shown. ANN 800 includes nodes 820, ..., 832 and edges 840, ..., 842, wherein each edge 840, ..., 842 is a directed connection from a first node 820, ..., 832 to a second node 820, ..., 832. Typically, the first node 820, ..., 832 and the second node 820, ..., 832 are different nodes 820, ..., 832. However, it is also possible that the first node 820, ..., 832 and the second node 820, ..., 832 are the same. For example, in Figure 5, edge 840 is a directed connection from node 820 to node 823, and edge 842 is a directed connection from node 830 to node 832. Edges 840, ..., 842 from a first node 820, ..., 832 to a second node 820, ..., 832 are also referred to as in-edges to the second node 820, ..., 832 and out-edges to the first node 820, ..., 832.

[0145] In this example, the nodes 820, ..., 832 of the ANN 800 can be arranged in layers 810, ..., 813, wherein the layers can have an intrinsic order introduced by edges 840, ..., 842 between the nodes 820, ..., 832. In particular, edges 840, ..., 842 exist only between adjacent layers of nodes. In the example shown, there is an input layer 810 consisting only of nodes 820, ..., 822 and no incoming edges, an output layer 813 consisting only of nodes 831, ..., 832 and no outgoing edges, and hidden layers 811 and 812 between the input layer 810 and the output layer 813. In general, the number of hidden layers 811 and 812 can be selected arbitrarily. In a multilayer perceptron (MLP), the number is at least one. , 822 within the input layer 810 generally refers to the number of input values ​​of the artificial neural network 800 , while the number of nodes 831 , 832 within the output layer 813 generally refers to the number of output values ​​of the artificial neural network 800 .

[0146] In particular, a real number can be assigned as a value to each node 820, ..., 832 of the artificial neural network 800. (n) i represents the value of the i-th node 820, ..., 832 of the n-th layer 810, ..., 813. The values ​​of the nodes 820, ..., 822 of the input layer 810 correspond to the input values ​​of the artificial neural network 800. The values ​​of the nodes 831, 832 of the output layer 813 correspond to the output values ​​of the artificial neural network 800. In addition, each edge 840, ..., 842 can have a weight, which is a real number. In particular, the weight is a real number within the interval [-1, 1] or within the interval [0, 1]. Here, w (m,n) i,j represents the weight of the edge between the i-th node 820, ..., 832 of the m-th layer 810, ..., 813 and the j-th node 820, ..., 832 of the n-th layer 810, ..., 813. (n,n+1) i,j Definition of the abbreviation w (n) i,jIn order to calculate the output value of the neural network 800, the input value is propagated through the neural network 800. In particular, the values ​​of the nodes 820, ..., 832 of the (n+1)th layer 810, ..., 813 can be calculated based on the values ​​of the nodes 820, ..., 832 of the nth layer 810, ..., 813 as follows:

[0147]

[0148] The function f is called a transfer function or activation function. Known transfer functions include step functions, sigmoid functions, such as the logistic function, the generalized logistic function, the hyperbolic tangent, the inverse tangent function, the error function, a smoothed step function, or a rectifier function. Transfer functions are primarily used for normalization. In particular, the values ​​are propagated layer by layer through the neural network 800, wherein the values ​​of the input layer 810 are derived from the inputs of the neural network 800, wherein the values ​​of the first hidden layer 811 can be calculated based on the values ​​of the input layer 810 of the neural network 800, wherein the values ​​of the second hidden layer 812 can be calculated based on the values ​​of the first hidden layer 811, and so on.

[0149] To determine the value of edge w (m,n) i,j , the neural network 800 must be trained with the help of training data. The training data includes in particular training input data and training output data (referred to as t i ). In the training step, the neural network 800 is applied to the training input data in order to generate the calculated output data. In particular, the training data and the calculated output data include a plurality of values ​​corresponding to the number of nodes in the output layer. In particular, the comparison between the calculated output data and the training data is used to recursively adjust the weights within the neural network 800 (back propagation algorithm). In particular, the weights vary according to the following formula

[0150]

[0151] Where γ is a predefined learning rate, and if the (n+1)th layer is not the output layer 813, then based on δ (n+1) j , we can use δ (n) j The recursive calculation is:

[0152]

[0153] If the (n+1)th layer is the output layer 813, then

[0154]

[0155] Where f' is the first-order derivative of the activation function, and t (n+1) jis the comparison training value for the j-th node of the output layer 813.

[0156] A convolutional neural network (CNN) is an ANN that uses convolution operations rather than general matrix multiplication in at least one of its layers. These layers are called convolutional layers. In particular, a convolutional layer performs a dot product of one or more convolution kernels with the input data of the convolutional layer, where the entries of one or more convolution kernels are parameters or weights that can be adjusted through training. In particular, Frobenius inner products and ReLU activation functions can be used. Convolutional neural networks can include additional layers, such as pooling layers, fully connected layers, and / or normalization layers.

[0157] By using convolutional neural networks, inputs can be processed very efficiently because convolution operations based on different kernels can extract different image features, so that by adjusting the weights of the convolution kernels, important image features can be determined during training. In addition, due to the shared use of weights in the convolution kernels, fewer parameters need to be trained, which prevents overfitting during the training phase and enables faster training or more layers in the network, thereby improving network performance.

[0158] Figure 6 An exemplary embodiment of a convolutional neural network 700 is shown. In the illustrated embodiment, the convolutional neural network 700 includes an input node layer 710, a convolutional layer 711, a pooling layer 713, a fully connected layer 714, and an output node layer 716, as well as hidden node layers 712 and 714. Alternatively, the convolutional neural network 700 may also include multiple convolutional layers 711, multiple pooling layers 713, and / or multiple fully connected layers 715, as well as other types of layers. The order of the layers can be selected arbitrarily; typically, the fully connected layer 715 is used as the last layer before the output layer 716.

[0159] In particular, in the convolutional neural network 700, the nodes 720, 722, 724 of the node layers 710, 712, 714 can be viewed as a d-dimensional matrix or a d-dimensional image. In particular, in the two-dimensional case, the value of the node 720, 722, 724 indexed by i and j in the n-th node layer 710, 712, 714 can be referred to as x(n)[i,j]. However, the arrangement of the nodes 720, 722, 724 of the node layers 710, 712, 714 itself has no effect on the calculations performed within the convolutional neural network 700, because these calculations are determined only by the structure and weights of the edges.

[0160] The convolution layer 711 is a connection layer between the previous node layer 710 with node values ​​x(n-1) and the subsequent node layer 712 with node values ​​x(n). The convolution layer 711 is characterized in particular by the structure and weights of the input edges, which form a convolution operation based on a specific number of kernels. In particular, the structure and weights of the edges of the convolution layer 711 are selected so that the value x(n) of the node 722 of the subsequent node layer 712 is calculated as the convolution x(n)=K*x(n-1) based on the value x(n-1) of the node 720 of the previous node layer 710, where the convolution is defined as

[0161]

[0162] Here, the kernel K is a d-dimensional matrix, in this example a two-dimensional matrix, which is typically small compared to the number of nodes 720, 722, for example a 3×3 matrix or a 5×5 matrix. This means, in particular, that the weights of the edges in convolutional layer 711 are not independent, but are chosen so that they produce the convolution equation. In particular, for a kernel of 3×3 matrix, there are only nine independent weights, with each entry of the kernel matrix corresponding to an independent weight, regardless of the number of nodes 720, 722 in the preceding node layer 710 and the following node layer 712.

[0163] In general, the convolutional neural network 700 uses node layers 710, 712, 714 with multiple channels, especially due to the use of multiple kernels in the convolution layer 711. In this case, the node layer can be viewed as a (d+1)-dimensional matrix, where the first dimension indicates the channel. The role of the convolution layer 711 is then defined in the two-dimensional example as

[0164]

[0165] in corresponding to the ath channel of the previous node layer 710, corresponds to the bth channel of the next node layer 712, and K a,b Corresponding to one of the kernels. When the convolution layer 711 acts on the previous node layer 710 with A channels and outputs the next node layer 712 with B channels, there are A·B independent d-dimensional kernels K a,b .

[0166] Generally, an activation function may be used in the convolutional neural network 700. In the embodiment described, a ReLU (rectified linear unit) is used, where R(z)=max(0,z), so that in the two-dimensional example, the convolution layer 711 has the following effect:

[0167]

[0168] It is also possible to use other activation functions, such as ELU (Exponential Linear Unit), LeakyReLU, Sigmoid, Tanh or Softmax.

[0169] In the illustrated embodiment, the input layer 710 includes 36 nodes 720 arranged in a two-dimensional 6×6 matrix. The first hidden node layer 712 contains 72 nodes 722, which are arranged as a two-dimensional 6×6 matrix, where each of the two matrices is the result of convolving the values ​​of the input layer with a 3×3 kernel within the convolution layer 711. Equivalently, the nodes 722 of the first hidden node layer 712 can be interpreted as a three-dimensional 2×6×6 matrix, where the first dimension corresponds to the channel dimension.

[0170] The advantage of using convolutional layers 711 is that the spatial local correlation of the input data can be exploited by enforcing local connectivity patterns between nodes in adjacent layers, especially by having each node only connected to a small region of nodes in the previous layer.

[0171] Pooling layer 713 is a connection layer between the previous node layer 712 with node values ​​x(n-1) and the next node layer 714 with node values ​​x(n). Pooling layer 713 can be characterized in particular by the structure and weights of the edges and the activation function, which form a pooling operation based on the nonlinear pooling function f. For example, in the two-dimensional case, the value x(n) of node 724 in the next node layer 714 can be calculated based on the value x(n-1) of node 722 in the previous node layer 712 as follows:

[0172]

[0173] In other words, the number of nodes 722, 724 can be reduced by using a pooling layer 713 by replacing the number d1-d2 of adjacent nodes 722 in the previous node layer 712 with a unique node 722 in the next node layer 714, which is calculated as the value of the aforementioned number of adjacent nodes. The pooling function f can, in particular, be a Max function, a mean, or an L2 norm. In particular, in the pooling layer 713, the weights of the input edges are fixed and do not change due to training.

[0174] The advantage of using the pooling layer 713 is that the number of nodes 722, 724 and the number of parameters are reduced, which leads to a reduction in computational cost in the network and control of overfitting.

[0175] In the embodiment shown, the pooling layer 713 is a max pooling layer, in which four adjacent nodes are replaced by only one node, where the value is the maximum of the values ​​of the four adjacent nodes. Max pooling is applied to each d-dimensional matrix of the previous layer. In the embodiment shown, max pooling is applied to each of the two-dimensional matrices, thereby reducing the number of nodes from 72 to 18.

[0176] Generally, the last layer of the convolutional neural network 700 may be a fully connected layer 715. The fully connected layer 715 is a connection layer between the previous node layer 714 and the next node layer 716. The fully connected layer 713 is characterized in that most edges, in particular, all edges, exist between the nodes 714 of the previous node layer 714 and the nodes 716 of the next node layer, and the weight of each of the edges can be adjusted separately.

[0177] In this embodiment, nodes 724 of the preceding node layer 714 of the fully connected layer 715 are not only shown as a two-dimensional matrix, but also as disconnected nodes displayed as lines between nodes, wherein the number of nodes is reduced for better visibility. This process is also called flattening. In this embodiment, the number of nodes 726 in the subsequent node layer 716 of the fully connected layer 715 is smaller than the number of nodes 724 in the preceding node layer 714. Alternatively, the number of nodes 726 may be equal to or greater.

[0178] In addition, in this embodiment, the Softmax activation function is used in the fully connected layer 715. By applying the Softmax function, the sum of the values ​​of all nodes 726 of the output layer 716 is equal to 1, and all values ​​of all nodes 726 of the output layer 716 are real numbers between 0 and 1. In particular, when the convolutional neural network 700 is used to classify input data, the value of the output layer 716 can be interpreted as the probability that the input data belongs to one of the various categories.

[0179] In particular, the convolutional neural network 700 can be trained based on a back-propagation algorithm. To prevent overfitting, regularization methods can be used, such as deleting nodes 720, ..., 724, random pooling, using artificial data, weight decay based on L1 or L2 norm, or maximum norm restriction.

[0180] exist Figure 7A CNN with a U-Net structure is schematically shown in FIG. In the example shown, the input data for the CNN is a two-dimensional medical image with 512×512 pixels, where each pixel contains an intensity value. The CNN comprises convolutional layers indicated by solid horizontal arrows, pooling layers indicated by solid downward arrows, and upsampling layers indicated by solid upward arrows. The number of corresponding nodes is given in the box. Within the U-Net structure, the input image is first downsampled (English: downsampling), in particular by reducing the image size and increasing the number of channels. It is then upsampled (English: upsampling), in particular by enlarging the image size and reducing the number of channels, in order to produce a transformed image.

[0181] Except for the last convolutional layers L1, L2, L4, L5, L7, L8, L10, L11, L13, L14, L16, L17, L19, and L20, all convolutional layers use a 3×3 kernel with padding of 1, a ReLU activation function, and a number of filters or convolution kernels corresponding to the number of channels of the corresponding node layer, as shown in Figure 6 As shown in Figure 2. The final convolutional layer uses a 1×1 kernel with no padding and a ReLU activation function.

[0182] Pooling layers L3, L6, and L9 are max pooling layers, which replace four adjacent nodes with a single node whose value is the maximum of the values ​​of these four adjacent nodes. Upsampling layers L12, L15, and L18 are transposed convolutional layers with a 3×3 kernel and a stride of 2, effectively quadrupling the number of nodes. The dashed horizontal arrows correspond to concatenation operations, where the outputs of convolutional layers L2, L5, and L8 of the downsampling branch of the U-Net structure are used as additional inputs to convolutional layers L13, L16, and L19 of the upsampling branch of the U-Net structure. The additional input data is treated as additional channels in the input node layer of the convolutional layers L13, L16, and L19 of the upsampling branch.

[0183] To train the CNN, a database of 500 first medical images was used, for which corresponding segmentation masks were created based on annotations by expert radiologists. In particular, for each of the 500 first medical images, the expert determined a segmentation mask for the structure of interest, in which pixels corresponding to the structure of interest were associated with a value of 1, and pixels not corresponding to the structure of interest were associated with a value of 0. The database was divided into training data (320 datasets), validation data (80 datasets), and test data (100 datasets). To train the CNN, a backpropagation algorithm based on a binary cross entropy cost function was used:

[0184]

[0185] in

[0186] BCE(a,b):=-alog(b)(b)-(1-a)log(1-b),

[0187] where x represents the first medical image, y determines the corresponding segmentation mask created by the radiologist, and M(x) represents the result of applying the CNN to the first medical input image x. Alternatively, other cost functions such as weighted binary cross entropy, focal loss, or dice loss can also be used.

[0188] Based on a validation set of 80 datasets and their corresponding annotations, the best-performing model was selected from multiple machine learning models (with different hyperparameters, such as the number of layers, kernel size and number, padding, etc.). Specificity and sensitivity were determined based on a test set containing 100 datasets and their corresponding annotations.

[0189] Regardless of the grammatical gender of a particular term, persons with male or female gender identities are included.

Claims

1. A method for providing a trained machine learning model xMLM (7) for extracting features from medical data (10, 10', 10"), wherein - obtaining a base machine learning model fMLM (7, 8) in an untrained or partially trained state, wherein the fMLM (7, 8) has an architecture that can be trained by means of unsupervised or self-supervised training; - the fMLM (7, 8) includes the xMLM (7) and at least one downstream machine model nMLM (8), wherein the at least one machine model nMLM (8) is used to perform at least one corresponding downstream task; - acquiring first medical data (10, 10', 10"); and - unsupervisedly or self-supervisedly training the fMLM (7, 8) including the xMLM (7) and the at least one nMLM (8) based on the first medical data (10, 10', 10"), and storing the trained xMLM (7).

2. The method according to claim 1, The fMLM (7, 8) has an architecture that can be trained based on multimodal data by means of unsupervised or self-supervised training, and the first medical data (10, 10', 10") is first multimodal medical data (10, 10', 10").

3. The method according to any one of the preceding claims, An imaging method is performed to generate at least a portion of the first medical data (10, 10', 10").

4. A method for anonymizing medical data, The method according to any of the preceding claims is performed in order to provide a trained xMLM (7), to acquire second medical data (10, 10', 10"), the second medical data containing personal data, and to anonymize the second medical data (10, 10', 10") by encoding the second medical data (10, 10', 10") by applying the trained xMLM (7) to the second medical data (10, 10', 10").

5. The method according to claim 4, The personal data include: - image data (10a') resulting from an imaging method; and / or - text data (10b') and / or tabular data (10c') and / or numerical data relating to a medical assessment or judgment of a patient; and / or - Data relating to the patient's identity.

6. A method for providing a trained machine learning model sMLM (12) for processing and / or analyzing medical data, wherein the method according to any one of claims 4 or 5 is performed to generate encoded second medical data (9'), and the sMLM (12) is trained unsupervisedly or self-supervised based on the encoded second medical data (9').

7. The method according to claim 6, wherein in order to train the sMLM (12) based on the encoded second medical data (9'), - predicting a treatment and / or analysis result (13) by applying said sMLM (12) to said encoded second medical data (9'); - evaluating a predefined loss function in relation to the predicted processing and / or analysis result (13); and - Updating the sMLM (12) in relation to the result of the evaluation of the loss function.

8. The method according to any one of claims 6 or 7, When the sMLM (12) is trained based on the encoded second medical data (9'), the xMLM (7) remains unchanged.

9. The method according to any one of claims 6 to 8, The training of the sMLM (12) based on the encoded second medical data (9') is performed by means of a data processing system (3, 6), which does not have access to the second medical data (10, 10', 10").

10. A method for processing and / or analyzing medical data, wherein - performing a method according to any one of claims 6 to 9 in order to provide a trained sMLM (12); - acquiring third medical data (10, 10', 10") and generating further encoded data (9') by applying the trained xMLM (7) to the third medical data (10, 10', 10"); - generating processing and / or analysis results (13) by applying said trained sMLM (12) to said further coded data (9').

11. The method according to claim 10, wherein the sMLM (12) is a machine learning model, - for anatomical image classification; and / or - for the identification of disease or abnormality; and / or - for medical image segmentation; and / or - Used for image processing.

12. An infrastructure system (1) for carrying out the method according to any one of claims 6 or 7, the infrastructure system (1) comprising: - a first data processing system (3) suitable for carrying out the method according to any one of claims 1 or 2 and the method according to any one of claims 4 or 5; and - a second data processing system (6) suitable for training a machine learning model sMLM (12) for processing and / or analyzing medical data based on the encoded second medical data (9').

13. The infrastructure system (1) according to claim 12, has an imaging modality (4), which is configured to perform an imaging method to generate image data and transmit the image data to the first data processing system (3), wherein the first data processing system (3) is suitable for using the image data for training the fMLM (7, 8) and / or using the image data after training the fMLM (7, 8) to further train the fMLM (7, 8).

14. Infrastructure system (1) according to any one of claims 12 or 13, wherein the first data processing system (3) is adapted to acquire third medical data (10, 10', 10") and to generate further coded data (9') by applying the trained xMLM (7) to the third medical data (10, 10', 10"), and wherein - said first data processing system (3) is adapted for producing processing and / or analysis results by applying said trained sMLM (12) to said further coded data (9'); or - the infrastructure system (1) has a third data processing system adapted for generating processing and / or analysis results (13) by applying the trained sMLM (12) to the further coded data (9').

15. The infrastructure system (1) according to any one of claims 12 to 14, comprising: - a hospital (2), said hospital (2) comprising said first data processing system (3); and / or A research and / or development facility, which contains the second data processing system (6) and is spatially separated from the hospital (2).

16. A computer program product comprising: - instructions which, when executed by at least one data processing system (3, 6), cause the at least one data processing system (3, 6) to carry out the method according to any one of claims 1 or 2 or the method according to any one of claims 4 to 11, or - first instructions which, when executed by a first data processing system (3), cause the first data processing system (3) to perform the method according to any one of claims 1 or 2 and the method according to any one of claims 4 or 5; and second instructions which, when executed by a second data processing system (6), cause the second data processing system (6) to train an sMLM (12) based on encoded second medical data (9'); or - a first instruction which, when executed by a first data processing system (3), causes the first data processing system (3) to perform the method according to any one of claims 1 or 2 and the method according to any one of claims 4 or 5; a second instruction which, when executed by a second data processing system (6), causes the second data processing system (6) to train an sMLM (12) based on encoded second medical data (9'); and a third instruction which, when executed by a third data processing system, causes the third data processing system to generate a processing and / or analysis result (13) by applying the trained sMLM (12) to other encoded data (9').