In-situ training of machine learning algorithms for generating synthetic imaging data

The machine learning algorithm for training image synthesis generates synthetic imaging data, solves the problem of scarcity of medical imaging data, improves the accuracy of algorithm training and data privacy protection, and achieves the satisfaction of domain-specific training needs.

CN115081637BActive Publication Date: 2025-05-16SIEMENS HEALTHINEERS AG
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210237004.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-03-12
Filing Date
2022-03-11
Publication Date
2025-05-16
Estimated Expiration
2042-03-11

AI Technical Summary

Technical Problem

In the medical field, it is challenging to obtain enough training data to train machine learning algorithms to achieve the required accuracy, especially for algorithms that process medical imaging data.

Method used

Synthetic imaging data is generated by training image synthesis machine learning algorithms, and other machine learning algorithms are trained using these synthetic data to achieve effective processing of medical imaging data.

Benefits of technology

This method can improve the training accuracy of machine learning algorithms in the case of scarcity of data, especially for domain-specific training needs, reduce dependence on real training data, protect data privacy, and reduce the bandwidth of data transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115081637B_ABST
    Figure CN115081637B_ABST
Patent Text Reader

Abstract

Techniques for training image synthesis ML algorithms are disclosed. The image synthesis ML algorithm can be used to generate synthetic imaging data. The synthetic imaging data can in turn be used to train another ML algorithm. The other ML algorithm can be configured to perform image processing tasks on the corresponding imaging data.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related Applications

[0002] This application claims the benefit of 21162327.7, filed on March 12, 2021, which is incorporated herein by reference in its entirety. Technical Field

[0003] Various examples of the present invention relate to facilitating the training of a machine learning algorithm configured to process medical imaging data. Various examples of the present disclosure relate specifically to the training of an additional machine learning algorithm configured to provide synthetic training imaging data for training the machine learning algorithm. Background Art

[0004] There are various use cases in the medical field for applying machine learning algorithms, such as deep neural networks or support vector machines. Examples include segmentation and detection of anomalies, or classification.

[0005] It has been observed that obtaining sufficient training data to train machine learning (ML) algorithms to perform the corresponding tasks with the required accuracy can be challenging. This is especially true for ML algorithms that process medical imaging data. Typically, suitable medical imaging data is available only to a limited extent. Summary of the invention

[0006] Therefore, there is a need for advanced techniques that facilitate the training of ML algorithms. In particular, there is a need for advanced techniques that facilitate the training of ML algorithms that process medical imaging data.

[0007] The following describes a technique for training an image synthesis ML algorithm. The image synthesis ML algorithm can be used to generate synthetic imaging data. The synthetic imaging data can then be used to train another ML algorithm. The other ML algorithm can be configured to perform image processing tasks on the corresponding imaging data.

[0008] A computer-implemented method of performing a first training of a first ML algorithm is provided. The first ML algorithm is used to generate synthetic imaging data of an anatomical target region of at least one patient. The method includes obtaining multiple instances of the imaging data of the anatomical target region of at least one patient. The multiple instances of the imaging data are acquired at an imaging facility. The method also includes performing a first training of the first ML algorithm for generating the synthetic imaging data on-site at the imaging facility and based on the multiple instances of the imaging data. The method also includes, upon completion of the first training: providing parameter values ​​of at least the first ML algorithm to a shared repository. Thereby, enabling a second training of a second ML algorithm based on additional synthetic imaging data generated by the first ML algorithm using the parameter values.

[0009] A computer program, a computer program product, or a non-transitory computer-readable storage medium includes program code. The program code can be loaded and executed by at least one processor. When loading and executing the program code, at least one processor performs a method for performing a first training of a first ML algorithm. The first ML algorithm is used to generate synthetic imaging data of an anatomical target area of ​​at least one patient. The method includes obtaining multiple instances of imaging data of an anatomical target area of ​​at least one patient. The multiple instances of imaging data are acquired at an imaging facility. The method also includes performing a first training of the first ML algorithm for generating synthetic imaging data on-site at the imaging facility and based on the multiple instances of imaging data. The method also includes, upon completion of the first training: providing parameter values ​​of at least the first ML algorithm to a shared repository. Thereby, based on additional synthetic imaging data generated by the first ML algorithm using the parameter values, a second training of a second ML algorithm is enabled.

[0010] A device comprising a processor and a memory. The processor is configured to load a program code from the memory and execute the program code. When executing the program code, the processor performs a method for performing a first training of a first ML algorithm. The first ML algorithm is used to generate synthetic imaging data of an anatomical target area of ​​at least one patient. The method includes obtaining multiple instances of imaging data of the anatomical target area of ​​at least one patient. The multiple instances of the imaging data are acquired at an imaging facility. The method also includes performing a first training of the first ML algorithm for generating synthetic imaging data on-site at the imaging facility and based on the multiple instances of the imaging data. The method also includes, upon completion of the first training: providing parameter values ​​of at least the first ML algorithm to a shared repository. Thereby, based on additional synthetic imaging data generated by the first ML algorithm using the parameter values, a second training of a second ML algorithm is enabled.

[0011] A computer-implemented method for generating synthetic imaging data of an anatomical target region includes establishing at least one latent space representation associated with the anatomical target region of at least one patient. The method also includes applying a trained first ML algorithm to the at least one latent space representation and generating synthetic imaging data by the trained first ML algorithm to enable second training of a second ML algorithm based on the synthetic imaging data.

[0012] A computer program, a computer program product, or a non-transitory computer-readable storage medium includes program code. The program code can be loaded and executed by at least one processor. When the program code is loaded and executed, the at least one processor performs a method for generating synthetic imaging data of an anatomical target area. The method includes establishing at least one latent space representation associated with the anatomical target area of ​​at least one patient. The method also includes applying the trained first ML algorithm to the at least one latent space representation, and generating synthetic imaging data by the trained first ML algorithm to enable second training of a second ML algorithm based on the synthetic imaging data.

[0013] A device including a processor and a memory. The processor is configured to load a program code from the memory and execute the program code. When executing the program code, the processor performs a method for generating synthetic imaging data of an anatomical target region. The method includes establishing at least one latent space representation associated with the anatomical target region of at least one patient. The method also includes applying a trained first ML algorithm to the at least one latent space representation, and generating synthetic imaging data by the trained first ML algorithm to enable second training of a second ML algorithm based on the synthetic imaging data.

[0014] It is to be understood that the features mentioned above and those yet to be explained below can be used not only in the respective combination indicated but also in other combinations or alone, without departing from the scope of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 schematically illustrates one embodiment of a first ML algorithm and a second ML algorithm, the first ML algorithm providing synthetic imaging data for training the second ML algorithm according to various examples (image synthesis ML algorithm);

[0016] Figure 2 Schematically illustrates details of an image synthesis ML algorithm according to one embodiment;

[0017] Figure 3 are schematic illustrations of systems according to various examples;

[0018] Figure 4 are flow charts of methods according to various examples;

[0019] Figure 5 are flow charts of methods according to various examples;

[0020] Figure 6 Workflows according to various examples are schematically illustrated. DETAILED DESCRIPTION

[0021] Some examples of the present invention generally provide multiple circuits or other electrical devices. All references to circuits and other electrical devices and the functions they each provide are not intended to be limited to the content only contained in the illustrations and descriptions herein. Although specific labels may be assigned to the various circuits or other electrical devices disclosed, such labels are not intended to limit the scope of operation of circuits and other electrical devices. Such circuits and other electrical devices can be combined and / or separated from each other in any way based on the desired specific type of electrical implementation. It should be recognized that any circuit or other electrical device disclosed herein may include any number of microcontrollers, graphics processor units (GPUs), integrated circuits, memory devices (e.g., flash memory, random access memory (RAM), read-only memory (ROM), electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM) or other suitable variants thereof), and software that cooperates with each other to perform (one or more) operations disclosed herein. In addition, any one or more electrical devices may be configured to execute a program code embodied in a non-transitory computer-readable medium programmed to perform any number of functions as disclosed.

[0022] Hereinafter, embodiments of the present invention will be described in detail with respect to the accompanying drawings. It should be understood that the following description of the embodiments should not be understood in a restrictive sense. The scope of the present invention is not intended to be limited by the embodiments or drawings described below, which are considered to be illustrative only.

[0023] The accompanying drawings are considered to be schematic representations, and the elements illustrated in the accompanying drawings are not necessarily shown to scale. On the contrary, the various elements are represented so that their functions and general purposes become clear to those skilled in the art. Any connection or coupling between the functional blocks, devices, components or other physical or functional units shown in the accompanying drawings or described herein may also be achieved by indirect connection or coupling. Coupling between components may also be established by wireless connection. Functional blocks may be implemented with hardware, firmware, software or a combination thereof.

[0024] Various techniques disclosed herein generally relate to facilitating the training of ML algorithms. The ML algorithm may be configured to process imaging data. For example, the ML algorithm may be configured to process medical imaging data, such as delineating an anatomical target region of a patient such as the heart, liver, brain, etc. In other examples, other kinds of imaging data may be processed, such as projection imaging data, such as for security scanners or material inspection.

[0025] According to the present disclosure, various kinds and types of imaging data can be processed. As a general rule, it will be possible for ML algorithms to process 2-D images or raw data acquired in K-space. ML algorithms can process 3-D depth data, such as point clouds or depth maps. ML algorithms can process time-varying data, where one dimension is stored as an image or volume representation at different points in time.

[0026] In the following, various examples will be described in the context of an ML algorithm configured to process medical imaging data. However, similar techniques can be easily applied to other kinds and types of semantic contexts of imaging data. For simplicity, the ML algorithm will be referred to as a medical imaging ML algorithm in the following.

[0027] As a general rule, various kinds and types of medical imaging data may be subject to the techniques described herein. To give a few examples, it would be possible to use magnetic resonance imaging (MRI) imaging data, such as raw data or pre-reconstructed images in K-space. Another example would be about computed tomography (CT) imaging data, such as projection views or pre-reconstructed images. Yet another example would be about positron emission tomography (PET) imaging data. Other examples include ultrasound images.

[0028] As a general rule, various kinds and types of medical imaging ML algorithms can benefit from the techniques described herein. For example, it would be possible to use a deep neural network, such as a convolutional neural network with one or more convolutional layers that perform convolutions between the input data and a kernel. It would also be possible to use a support vector machine, to name a few examples. The U-net architecture can be used, see, for example, Ronneberger, O., Fischer, P., and Brox, T., October 2015, U-net: Convolutional networks for biomedical image segmentation ( International Conference on Medical image computing and computer-assisted intervention (pp. 234-241). Springer, Cham).

[0029] In the examples described herein, the medical imaging ML algorithm may be configured to perform various tasks when processing medical imaging data. For example, the medical imaging ML algorithm may be configured to perform segmentation of medical imaging data. For example, it would be possible to segment predefined anatomical features. In another example, the medical imaging ML algorithm may be configured to perform object detection. For example, a bounding box may be drawn around a predefined object detected in the medical image data. The predefined object may be a predefined anatomical feature, such as a certain organ or blood vessel, a tumor site, etc. It will also be possible to detect abnormalities. Yet another task will be filtering or denoising. For example, it will be possible to remove artifacts such as blur or speckle from an imaging modality. Background noise may be suppressed or removed. Yet another task will be upsampling, thereby increasing the resolution of the medical imaging data. Yet another task of the medical imaging ML algorithm may be image reconstruction. For example, for magnetic resonance imaging data, the raw imaging data is available in K space. Sometimes, K space is undersampled (e.g., relative to a certain field of view, taking into account the Nyquist theorem), i.e., a direct Fourier transform will result in aliasing artifacts. In such a scenario, an appropriate medical imaging ML algorithm may be used to achieve image reconstruction. For example, techniques using unfolded neural networks are known; here, a medical imaging ML algorithm can be used to implement the regularization operation. As should be appreciated from the above, a specific type of medical imaging ML algorithm is not germane to the functionality of the techniques described herein. Rather, medical imaging ML algorithms of all kinds and types can benefit from the techniques described herein, i.e., can be accurately trained.

[0030] Various techniques are based on the discovery that well-adapted training imaging data may be required to accurately train a medical imaging ML algorithm. For example, it has been observed that the amount of training imaging data and the variability of the training imaging data may be correlated with the accuracy of the medical imaging ML algorithm; that is, the more training imaging data with sufficient variability (to capture domain-specific features) is available, the more accurate the training of the medical imaging machine learning algorithm is. Second, it has been observed that domain-specific training imaging data may help to obtain an accurate training state for the medical imaging ML algorithm. For example, depending on a specific imaging facility or the goal of the observation, the imaging data input to the medical imaging ML algorithm may vary. Then, domain-specific training, i.e., training specific to the observation goal and / or a specific imaging facility, may be required to achieve higher accuracy. For example, it has been observed that even after the medical imaging ML algorithm has been trained on a large dataset of generic (i.e., non-domain-specific) training imaging data, further fine-tuning of the training may still be required to adapt the medical imaging ML algorithm to a specific patient population (e.g., adults to children), imaging modality (e.g., 1.5 Tesla MRI scanner versus 3 Tesla MRI scanner, different undersampling acceleration factors, etc., to name just a few examples). That is, retraining or fine-tuning with domain-specific training imaging data may be helpful. According to the reference technology, obtaining corresponding domain-specific training imaging data may be challenging.

[0031] According to various examples, such training, in particular domain-specific training, may be facilitated. Customized training imaging data may be provided; that is, domain-specific training imaging data may be provided.

[0032] Various techniques are based on the discovery that an additional ML algorithm (hereinafter referred to as an image synthesis ML algorithm) can be used to generate / infer synthetic training imaging data of an anatomical target region. The synthetic training imaging data can then be used to enable training of a medical imaging ML algorithm.

[0033] The general concept of generating synthetic images using ML algorithms has been shown to produce realistic results including medical imaging. Synthetic image generation can often help improve the accuracy of machine learning models when data is scarce.

[0034] For example, the generative adversarial network architecture can be used to implement image synthesis ML algorithms. More generally, there are different approaches for synthetic data generation in the literature. For example, the GauGan approach, where a neural network learns to generate an image given a segmentation of an image as input; see Park, Taesung, et al., “Semantic image synthesis with spatially-adaptive normalization” Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition , 2019.

[0035] For example, an image synthesis ML algorithm may include a decoder that decodes at least one latent space representation (ie, a particular sample of the latent space defined by the associated imaging data) to obtain synthesized imaging data. Feature expansion may be performed.

[0036] According to various examples, it is possible to perform training of an image synthesis ML algorithm. The training is based on multiple instances of imaging data of an anatomical target region for which the image synthesis ML algorithm generates synthetic training imaging data. The training imaging data may be acquired at an imaging facility.

[0037] During inference, it is possible to sample the latent space. That is, multiple latent space representations (i.e., different feature vectors that sample the latent space) may be used to determine the synthetic imaging data.

[0038] As a general rule, an image synthesis ML algorithm may take as input a segmentation map representing an anatomical target region present in the imaging data. Synthesized imaging data may then be provided for the anatomical target region that is separately shaped.

[0039] It would also be possible to randomly sample the latent space, for example, to determine a plurality of latent space representations, and provide these as input to an image synthesis algorithm. An example of a latent space representation may include a feature vector containing a set of numerical values ​​representing image features corresponding to anatomical properties associated with an anatomical target region of at least one patient. Other examples are possible, such as image appearance, presence of pathology, anatomical properties, etc.

[0040] The image synthesis ML algorithm can be trained as a variational autoencoder to learn the statistical distribution of the training image data. When performing training of the image synthesis ML algorithm, the image synthesis ML algorithm learns how to reproduce a real image by first encoding it into a latent space representation (i.e., a latent vector z using an encoder network) and then expanding the latent vector z into an image similar to the input image using a decoder network that implements the image synthesis ML algorithm. Based on the distribution of the training imaging data (e.g., by considering multiple instances of the training imaging data), various latent space representations distributed across the latent space can be determined.

[0041] During training of the image synthesis ML algorithm, it is possible to update parameter values ​​of the image synthesis ML algorithm based on a comparison between the imaging data and the corresponding generated imaging data. The comparison can be implemented as a loss function. Iterative numerical optimization can be used to update the parameter values. Back propagation can be used. That is, it can be determined how well the image synthesis ML algorithm can reconstruct the original imaging data based on the latent space representation; and the reconstruction can be optimized.

[0042] The image synthesis ML algorithm thus learns a latent space representation of multiple instances of the imaging data on which it is trained, but it does not learn the individual images of the imaging data pixel by pixel. Therefore, the synthesized images are never identical to the real patient data, and they cannot be traced back to a specific patient.

[0043] Then, to generate synthetic training imaging data for the medical imaging ML algorithm, the decoder network acts as a synthetic image generator. During inference of the image synthesis ML algorithm, corresponding to the training of the medical imaging ML algorithm, the (one or more) latent spaces of the trained image synthesis ML algorithm can be randomly sampled to produce synthetic training imaging data. In other words, multiple possible latent space representations associated with the anatomical target region can be sampled.

[0044] More specifically, it will be possible to establish at least one latent space representation associated with the patient's anatomical target region, for example by random sampling and / or user input and / or based on some reference imaging data and / or based on an instance vector z of the latent space representation and / or based on a segmentation of the anatomical target region. Thereby, the trained image synthesis ML algorithm can generate synthetic medical imaging data (synthetic training medical imaging data) that enables training of the medical imaging ML algorithm.

[0045] When the training of the medical imaging ML algorithm is completed, the medical imaging ML algorithm can be provided to the imaging facility. At the imaging facility, the trained medical imaging ML algorithm can support medical staff in analyzing or processing medical imaging data.

[0046] According to the techniques described herein, it is possible to perform training of an image synthesis ML algorithm on-site at an imaging facility used to acquire multiple instances of imaging data of an anatomical target region. Here, "on-site" can mean that the corresponding processing equipment performing the training of the image synthesis ML algorithm is located in the same local area network or virtual private network as the imaging facility. "On-site" can mean that the corresponding processing equipment is located at the same venue as the imaging facility, for example, in the same hospital or radiography center. "On-site" can mean that the training is performed before the training imaging data is output and sent to a picture archiving and communication system (PACS). "On-site" can even mean that the processor performing the training is an integrated processor of the imaging facility, for example, also providing control of the imaging routine.

[0047] Once the image synthesis ML algorithm is trained, it is possible to provide at least the parameter values ​​of the image synthesis ML algorithm to the shared repository. For example, the weights of the convolutional neural network model implementing the image synthesis ML algorithm set by the training can be provided to the shared repository. It will also be possible to provide the entire image synthesis ML algorithm to the shared repository.

[0048] The shared repository may be located off-site relative to the imaging facility. That is, it will be possible that the shared repository is not located in the same local area network or virtual private network as the imaging facility. For example, the shared repository may be located in a data cloud. The shared repository may be in a server farm. Off-site storage will be possible. Specifically, the data connection between the processing device performing the training of the image synthesis ML algorithm and the shared repository may have limited bandwidth. For example, the data connection may be achieved via the Internet.

[0049] Through such a technique, it is possible to limit the amount of data provided to a shared repository. Specifically, it will be possible not to provide training imaging data to a shared repository. The training imaging data implements payload data, i.e., a large volume of data used to perform training. By providing an image synthesis ML algorithm, or at least its parameter values, to a shared repository, the amount of data can be significantly reduced. The privacy of the training imaging data may not be mitigated. In addition to providing the training imaging data itself, it is also possible to provide inferred information in the form of parameter values ​​of the image synthesis ML algorithm itself to the shared repository. Therefore, there is no need to reveal the identity or characteristics of individual imaging data used to train the image synthesis ML algorithm. For example, imaging data used to train an image synthesis ML algorithm never have to leave a hospital or imaging facility. Only the model weights leave the hospital, and the imaging data used for training can be discarded. It will be possible to retain / store only at least one of its latent space representations.

[0050] By using such techniques, medical imaging ML algorithms can be (re)trained for a specific patient distribution and image appearance of the respective imaging facility before deployment. Domain-specific training is possible. The algorithm is personalized according to the needs of a specific radiology site.

[0051] No additional effort is required to collect disparate data from different locations for algorithm optimization. No additional effort is required to retrospectively collect and organize data for the purpose of fine-tuning ML models.

[0052] Figure 1 Aspects related to training a medical imaging ML algorithm 202 are schematically illustrated. The medical imaging ML algorithm 202 is trained based on synthetic training imaging data 211. For example, based on the synthetic training imaging data 211, it will be possible to determine an output 212 having a ground truth value 213; the ground truth value 213 may be determined based on manual labeling or semi-automatic labeling. A corresponding loss function may be defined, and a loss value may be determined by such comparison. Then, based on the corresponding loss value feedback 215, one or more parameter values ​​of the medical imaging ML algorithm 202 may be adjusted. For example, a gradient descent optimization technique may be used. Back propagation may be employed.

[0053] Various techniques described herein relate to determining synthetic training imaging data 211. Specifically, the synthetic training imaging data 211 may be generated by an image synthesis ML algorithm 201. The image synthesis ML algorithm may be applied to one or more vector instances of a latent space, i.e., a latent space representation 311 of a patient's anatomical target region. Each time the image synthesis ML algorithm 201 is applied to a given latent space representation 311, a corresponding instance of synthetic training medical imaging data is generated.

[0054] For example, a given latent space representation 311 of an anatomical target region may include a feature vector containing a set of values ​​representing a set of image features corresponding to different characteristics of the corresponding anatomical target region, e.g., defining the circumference of the heart or a particular blood vessel, etc. It would alternatively or additionally be possible to specify other parameters of the anatomical target region using one or more latent features (values ​​in the vector instance of the latent space), e.g., size, pathology, etc., of the latent space representation.

[0055] As a general rule, the latent space representation 311 may indicate the configuration of the imaging facility used to acquire the corresponding imaging data. The configuration may pertain to, for example, the protocol used, such as MRI sequences, parameter settings of the imaging modality, and the like.

[0056] It would be possible to randomly sample the latent space, e.g., determine at least one latent space representation using randomization.

[0057] The image synthesis ML algorithm 201 may then be applied to the at least one latent space representation 311 to generate synthetic training medical imaging data 211 to enable training of the medical imaging ML algorithm 202. Figure 1 In the illustrated example of , multiple instances of synthetic training medical imaging data are generated, for example by using different feature vector instances from a latent space and input images to an image synthesis ML algorithm 201 .

[0058] Figure 2 Schematically illustrating various aspects of an image synthesis ML algorithm 201. In the illustrated example, the image synthesis ML algorithm 201 comprises a decoder network for decoding a latent space representation 311 to generate synthetic medical imaging data 211. That is, a feature vector used as input 321 may be expanded to provide, for example, a 2D image as output 322. The dimensionality is increased.

[0059] To learn the latent space representation 311 of the true data distribution, the additional algorithm 309 may be used during training. During inference, the additional algorithm 309 may no longer be needed (during inference, the latent space 311A ​​may be randomly sampled, or the latent space representation 311 used for inference may be determined in other ways). The additional algorithm 309 may be an image encoder network that learns the latent space representation 311 of the true image used as input 301. For example, imaging data 210 acquired at an imaging facility - e.g., MRI images, CT images, etc. - may be provided as input 301 to the algorithm 309, and the output 302 of the algorithm 309 may then be the latent space representation 311 of the anatomical target region included in the imaging data 210.

[0060] Figure 3 A system 70 according to various examples is schematically illustrated. The system 70 comprises an imaging facility 91, such as an MRI scanner or a CT scanner, to give just two examples. The imaging facility 91 may be located in a hospital.

[0061] The system 70 also includes a device (computer) 80. The device 80 is located on-site at an imaging facility 91. For example, in the illustrated scenario, the device 80 and the imaging facility 90 are part or both parts of the same local area network 71. The device 80 includes a processor 81 and a memory 82. The device 80 also includes an interface 83. For example, the processor 81 can receive medical imaging data 210 from the imaging facility 91 via the interface 83. The processor 81 can load program code from the memory 82 and execute the program code. When loading and executing the program code, the processor 81 can perform the techniques described herein, for example, training a first ML algorithm based on imaging data from the imaging facility 91, filtering the imaging data before training the first ML algorithm, and so on.

[0062] The processor 81, upon completing the training of the first machine learning algorithm 201, may then provide the first machine learning algorithm 201, or at least its parameter values ​​201-1 (the parameter values ​​201-1 were set during the training), to the shared repository 89. This facilitates training of the second machine learning algorithm 202 at the further device 60. Thus, the first ML algorithm may be labeled as an image synthesis ML algorithm.

[0063] The further device (computer) 60 is located off-site relative to the imaging facility 91, i.e., outside the local area network 71. The further device 60 comprises a processor 61, a memory 62, and an interface 63. The processor 61 can load parameter values ​​201-1 or the entire training machine ML algorithm 201 from the shared repository 89, and then generate synthetic medical imaging data using the image synthesis ML algorithm 201 appropriately configured according to the parameter values ​​201-1 (see Figure 1The synthetic training medical imaging data 211 in FIG. 1 ). The synthetic imaging data may then be used to train a second machine learning algorithm, for example at a further device 60 . The processor 61 may load and execute corresponding program code from the memory 62 .

[0064] Figure 4 is a flow chart of a method according to various embodiments. Figure 4 The method relates to training of a first ML algorithm for generating synthetic imaging data of an anatomical target region.

[0065] Synthetic imaging data may be generated for at least one patient's anatomical target region. The synthetic imaging data may depict the anatomical target region as acquired using a specific imaging facility with a specific imaging protocol.

[0066] Once the first ML algorithm has been trained, the first ML algorithm can be used to generate further synthetic imaging data - i.e., synthetic training imaging data - and a further second ML algorithm can be trained based on the synthetic training imaging data. That is, the training phase of the second ML algorithm belongs to the inference phase of the first ML algorithm. Therefore, the first ML algorithm can also be referred to as an image synthesis ML algorithm.

[0067] When program code is loaded, it can be executed by at least one processor Figure 4 For example, when loading program code from memory 82, processor 81 of device 80 may execute Figure 4 method.

[0068] Use a dotted line to mark selectable boxes.

[0069] In block 3005, multiple instances of imaging data are obtained. Block 3005 may include obtaining multiple instances of imaging data. Block 3005 may include sending a control instruction to an imaging facility to obtain multiple instances of imaging data. Block 3005 may include loading multiple instances of imaging data from a buffer memory. For example, during use of an imaging facility (e.g., an MRI or CT scan), it will be possible to buffer acquired images in so-called mini-batches. Once a mini-batch is full, multiple instances of imaging data may be obtained from the corresponding buffer memory.

[0070] Multiple instances of imaging data may be acquired for multiple patients. Inter-patient variability may then be taken into account when training the first ML algorithm. Multiple instances of imaging data may be acquired using multiple configurations of an imaging facility. For example, different parameters for image acquisition may be selected, such as exposure time, MRI scan protocol, CT contrast, etc. It would be possible to also obtain, at block 3005, corresponding configuration data indicating a corresponding configuration of the imaging facility for a given instance of imaging data. As a general rule, a latent space representation of the medical imaging data may indicate the configuration.

[0071] At optional block 3010, multiple instances of imaging data are filtered, for example, to enable a quality check. For example, filtering may be based on quality. Various quality metrics are possible, such as sharpness, contrast, motion blur, to name just a few examples. It would be possible to check whether the correct anatomy and / or view is present in the corresponding imaging data. It may be checked whether the imaging data depicts a calibration phantom (in which case it may be helpful to discard the corresponding instance of the imaging data).

[0072] Various examples of determining quality are possible. For example, another ML algorithm can be used to determine a quality score, which can then be compared to a predefined threshold. Imaging data with a quality above the predefined threshold can then be retained; other imaging data can be discarded.

[0073] In box 3015, training of the first ML algorithm is performed. This can include learning a latent space for each of the anatomical regions, imaging protocols, or configurations. That is, for each latent space, multiple latent space representations in the corresponding latent space can be determined, for example, based on multiple instances of the corresponding image data. An encoder network can be used for this purpose. Then, for a given latent space, the latent space representations of multiple instances of image data in the latent space can be processed by a decoder network that implements the first ML algorithm to obtain synthetic imaging data. Then, a comparison can be made between the corresponding input (i.e., a given instance of the multiple instances of imaging data) and the output (i.e., the synthetic imaging data) of the first ML algorithm. A loss function can be defined based on the comparison, and based on the value of the loss function, the parameter values ​​of the first ML algorithm can be updated.

[0074] The described combination of encoder network and autoencoder network corresponds to a variational autoencoder architecture.

[0075] The above has been combined Figure 2 and image synthesis ML algorithm 201 explain various aspects of the latent space representation and operation relative to the first ML algorithm.

[0076] As a general rule, when executing block 3015, it will likely be that the first ML algorithm is in a pre-trained state. For example, the first ML algorithm may be pre-trained based on offline imaging data associated with an anatomical target region of another patient. Thus, the training of block 3015 may add domain-specific training to the general training state.

[0077] Once the training of the first ML algorithm is completed in block 3015, the first ML algorithm may be provided to a shared repository (see Figure 3 : shared repository 89). More generally, at block 3020, parameter values ​​of at least a first ML algorithm may be provided to a shared repository.

[0078] At optional block 3025, it will then be possible to discard multiple instances of imaging data used for training of the image synthesis ML network. It will be possible to not provide multiple instances of imaging data for training to the shared repository.

[0079] Optionally, latent space representations of multiple instances of imaging data may be stored, for example, in a shared repository. The latent space representations may then be used as seed values ​​for determining additional latent space representations for which synthetic training medical imaging data may be determined. For example, an encoder network used to determine the latent space representation of imaging data during training at block 3015 may not be provided to the shared repository.

[0080] Once the parameter values ​​of the first ML algorithm (or even the entire first ML algorithm) have been provided to the shared repository, a second training of the second ML algorithm based on the synthetic imaging data that may be generated by the first ML algorithm is enabled. That is, the first ML algorithm may be used to implement inference on the synthetic training medical imaging data. Next, in conjunction with Figure 5 The corresponding details of such an application of the first ML algorithm are explained.

[0081] Figure 5 is a flow chart of a method according to various embodiments. Figure 5 The method is associated with inferring synthetic training medical imaging data using a first ML algorithm. The synthetic training medical imaging data may depict an anatomical target region of at least one patient. Figure 4 Methods to train the first ML algorithms.

[0082] When program code is loaded, it can be executed by at least one processor Figure 5 For example, when loading program code from memory 62, Figure 5 The method may be executed by the processor 61 of the device 6.

[0083] Figure 5 In block 3105, from a shared repository (see Figure 3 : Shared repository 89) loads the first ML algorithm.

[0084] At block 3110 , at least one latent space representation of an anatomical target region of at least one patient is established.

[0085] There are various options. For example, the latent space - that is, the set of allowable latent space representations - can be learned during training in block 3015. It will then be possible to determine that at least one latent space representation is within the latent space. For example, the latent space can be sampled randomly or based on a predefined scheme.

[0086] The first ML algorithm may then be applied to the at least one latent space representation established at block 3110, i.e., the corresponding feature vector (sometimes also referred to as the instance vector z), thereby generating synthetic training medical imaging data at block 3115. The synthetic imaging data enables training of a second ML algorithm.

[0087] Thus, at block 3120, it is optionally possible to train a second ML algorithm. The second machine learning algorithm may be generally configured to process medical imaging data. Thus, the second machine learning algorithm may be labeled as a medical imaging ML algorithm. The second machine learning algorithm may implement various tasks such as image segmentation, denoising, upsampling, reconstruction, etc.

[0088] At block 3125, it may then be possible to optionally provide a second ML algorithm, or at least its parameter values, to the imaging facility. Figure 3 , it will be possible to provide the parameter values ​​of the second ML algorithm to the device 80. The image processing task can then be performed on the spot. As should be appreciated, Figure 4 and Figure 5 The combined approach facilitates domain-specific training. Nevertheless, the imaging data remains on-site, saving bandwidth and preserving privacy.

[0089] Figure 6 is a schematic illustration of a workflow for training a medical imaging ML algorithm 202. The workflow is generally structured into two stages 501, 502. The workflow steps of stage 501 are performed on-site using the imaging facility; therefore, they can be implemented Figure 4 The workflow steps of stage 502 are performed off-site, such as at a central / shared server; thus, they can achieve Figure 5 method.

[0090] New medical imaging data is acquired at 511. For example, a cardiac CT scan may be performed.

[0091] At 512, image quality control may be applied to the corresponding medical imaging data acquired at 511. This may be accomplished by Figure 4 This may be accomplished by filtering of block 3010 of . For example, a check may be made to see if the correct anatomy and view are present in the acquired medical imaging data. A check may be made to exclude calibration / phantom depictions.

[0092] Depending on whether a certain quality score exceeds a predefined threshold, it would be possible to selectively discard previously acquired medical imaging data instances.

[0093] It is then possible, at 513 , to add the corresponding instance of the imaging data to the mini-batch buffer.

[0094] 511 - 513 may then be performed multiple times until the mini-batch buffer has been filled with multiple instances of imaging data.

[0095] At 514, it is possible to perform training of the image synthesis ML algorithm 201. 514 implements box 3015 accordingly.

[0096] At 515, upon completion of training of the image synthesis ML algorithm, at least parameter values ​​of the image synthesis ML algorithm 201 are provided to the shared repository 89. Then, at 521, the image synthesis ML algorithm 201 may be loaded, and at 522, the image synthesis ML algorithm 201 is used to generate synthetic training imaging data. More specifically, the image synthesis ML algorithm may be applied to at least one latent space representation of the anatomical target region—i.e., a feature vector that samples the latent space—to generate synthetic training medical imaging data that enables training of the medical imaging ML algorithm 202. The medical imaging ML algorithm 202 is trained at 523.

[0097] As should be appreciated, the trained ML algorithm 201 is thus used offline to generate synthetic imaging data 211, e.g., a distribution of patients with image appearance / texture consistent with the clinical state. Segmented anatomy or synthetically generated masks from any patient can be used as input to the image synthesis ML algorithm to obtain synthetic imaging data. This means that the medical imaging ML algorithm can be fine-tuned / retrained on synthetic imaging data before deployment to the respective field.

[0098] In summary, techniques have been described that facilitate the generation of realistic synthetic training medical image data having a desired feature distribution. At the same time, the acquired real medical imaging data can be retained at the hospital's imaging facility. The patient data can remain anonymous. The neural network used to generate synthetic training medical image data can be trained on-site, or even directly at the imaging facility such as an MRI scanner. Medical imaging MR algorithms - for example, for segmentation tasks, classification tasks, etc. - can be retrained / fine-tuned for a specific patient distribution and image appearance of the imaging device before deployment.

[0099] Although the invention has been shown and described with respect to certain preferred embodiments, equivalents and modifications will occur to others skilled in the art upon the reading and understanding of the specification. The present invention includes all such equivalents and modifications and is limited only by the scope of the appended claims.

[0100] For illustration, various scenarios for generating synthetic training medical imaging data are described above, as well as a central server, or more generally, an off-site scenario relative to an imaging facility. It will also be possible to use a trained machine learning algorithm on-site at an imaging facility to generate synthetic training medical imaging data.

Claims

1. A computer-implemented method of performing a first training of a first machine learning algorithm for generating synthetic imaging data of an anatomical target region of at least one patient, the method comprising: obtaining a plurality of instances of imaging data of an anatomical target region of the at least one patient, the plurality of instances of imaging data being acquired at an imaging facility and representing a patient-specific distribution and image appearance of the imaging data acquired at the imaging facility; On-site at the imaging facility and based on the plurality of instances of the imaging data: performing a first training of a first machine learning algorithm for generating synthetic imaging data; and Upon completion of the first training: providing parameter values ​​of at least the first machine learning algorithm instead of imaging data to a shared repository, thereby enabling domain-specific training of a second machine learning algorithm based on additional domain-specific synthetic imaging data generated by the first machine learning algorithm using the domain-specific parameter values, wherein the second machine learning algorithm is tuned for image appearance of imaging data for a specific patient distribution and imaging facility by training with the domain-specific synthetic imaging data before deployment.

2. The computer-implemented method of claim 1 , wherein performing the first training comprises: generating at least one latent space representation for each of a plurality of instances of imaging data using an encoder network; processing the latent space representation of the plurality of instances of imaging data through a first machine learning algorithm to obtain synthetic imaging data; and Parameter values ​​of the first machine learning algorithm are updated based on a comparison between the imaging data and the corresponding synthetic imaging data.

3. The method of claim 1, wherein an input to the first machine learning algorithm comprises a segmentation map of an anatomical target region, the segmentation map representing anatomical characteristics associated with the anatomical target region of the at least one patient.

4. The method according to claim 1, further comprising: Upon completion of the domain-specific training, performing domain-specific training of a second machine learning algorithm: providing the second machine learning algorithm to at least one of the imaging facility or one or more additional imaging facilities.

5. The method of claim 1, wherein domain-specific training comprises sampling a latent space representation associated with an anatomical target region used as input to a first machine learning algorithm.

6. The method according to claim 1, further comprising: The multiple instances of imaging data are filtered based on a quality of the multiple instances of imaging data such that imaging data having a quality above a predefined threshold is used to perform a first training of a first machine learning algorithm.

7. The method according to claim 2, further comprising: providing the at least one latent space representation to a shared repository; and / or Multiple instances of imaging data were discarded.

8. The method according to claim 2, wherein: The first machine learning algorithm includes a decoder for decoding the at least one latent space representation to generate synthetic imaging data.

9. The method according to claim 1, further comprising: The first machine learning algorithm is pre-trained based on offline imaging data associated with an anatomical target region of an additional patient.

10. A computer-implemented method for generating domain-specific synthetic imaging data of an anatomical target region, the method comprising: establishing at least one latent space representation associated with at least one patient's anatomical target region based on imaging data acquired at an imaging facility; applying the trained first machine learning algorithm to the at least one latent space representation; and Domain-specific synthetic imaging data is generated by the trained first machine learning algorithm to enable second training of a second machine learning algorithm based on the domain-specific synthetic imaging data, wherein the second machine learning algorithm is adjusted for the image appearance of the imaging data for the anatomical target region and a specific patient distribution and imaging facility before deployment.

11. The method according to claim 10, wherein: The at least one latent space representation is established based on a segmentation map of the anatomical target region.

12. The method according to claim 10, further comprising: performing a second training of a second machine learning algorithm, Upon completion of the second training: providing a second machine learning algorithm to the equipment at the imaging facility site.

13. The method of claim 10, wherein the second training comprises sampling a latent space of at least one latent space representation associated with an anatomical target region used as input to the trained first machine learning algorithm.

14. A system for performing a first training of a first machine learning algorithm for generating domain-specific synthetic imaging data of an anatomical target region of at least one patient, the system comprising: A memory and at least one processor, the at least one processor being configured to load a program code from the memory and execute the program code, the at least one processor being configured to: obtaining a plurality of instances of domain-specific imaging data of an anatomical target region of the at least one patient, the plurality of instances of the domain-specific imaging data being acquired at an imaging facility, the plurality of instances of the domain-specific imaging data being representative of a patient-specific distribution and image appearance of the imaging data acquired at the imaging facility; On-site at the imaging facility and based on a plurality of instances of the domain-specific imaging data: performing a first training of a first machine learning algorithm for generating domain-specific synthetic imaging data that is specific to a particular patient distribution and image appearance of the imaging data acquired at the imaging facility; and Upon completion of the first training: providing at least parameter values ​​of the first machine learning algorithm instead of imaging data to a shared repository, thereby enabling a second training of a second machine learning algorithm based on additional synthetic imaging data generated by the first machine learning algorithm using the parameter values ​​to fine-tune the second machine learning algorithm for the image appearance of the imaging data for a specific patient distribution and imaging facility.

15. A system for generating synthetic imaging data of an anatomical target region, the system comprising: A memory and at least one processor, the at least one processor being configured to load a program code from the memory and execute the program code, the at least one processor being configured to: establishing at least one latent space representation associated with an anatomical target region of at least one patient from a plurality of imaging data from an imaging facility, wherein the at least one latent space representation is specific to a particular patient distribution and image appearance of the imaging data acquired at the imaging facility; applying the trained first machine learning algorithm to the at least one latent space representation; and Synthetic imaging data is generated by the trained first machine learning algorithm to enable second training of a second machine learning algorithm based on the synthetic imaging data, wherein the second machine learning algorithm is tuned to an anatomical target region of at least one patient prior to deployment.

Citation Information

Patent Citations

  • Modality-agnostic method for medical image representation

    US20190332900A1

  • A system and method for automated labeling and annotating unstructured medical datasets

    US20200286614A1