Training data for 3D image cleaning ML models

By using natural high-quality videos as training data and creating training datasets through downsampling and cropping, the problem of small amount and poor quality of data in the MRI scanning AI model is solved, and the image quality is improved and the versatility of the model is enhanced.

CN120693633APending Publication Date: 2025-09-23KONINKLIJKE PHILIPS NV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480009922.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-02-01
Filing Date
2024-01-24
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Training AI models for magnetic resonance imaging (MRI) scanning faces the problem of limited data volume and poor quality, which makes model improvement difficult. In addition, high-quality MR data acquisition is time-consuming and difficult to obtain.

Method used

Using natural high-quality videos as the training data source, we create a training dataset by downsampling, cropping, and changing video segments for training 3D AI image quality improvement models.

Benefits of technology

The model's image quality enhancement effects, including denoising, anti-ringing, and super-resolution, are improved, extended to other modalities and tasks, and enhanced the model's versatility and data representativeness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120693633A_ABST
    Figure CN120693633A_ABST
Patent Text Reader

Abstract

A method for creating training data for a machine learning model is disclosed. The machine learning model is configured to improve the quality of the acquired 3D medical image. The method comprises: selecting a video having desired values of one or more image parameters; video of a data set is pre-processed. The pre-processing includes down-sampling the video with a predefined reduction factor, clipping the down-sampled video in space and / or time to obtain video tiles, referred to herein as clean video tiles, changing the clean video tiles to reduce quality of the clean video tiles, creating a training data set, and generating a training data set. The training data set includes changed video tiles as inputs and associated clean video tiles as learning targets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to magnetic resonance imaging, and in particular to creating and using data for training machine learning models for cleaning 3D medical image data. Background Art

[0002] Selecting training data for healthcare artificial intelligence (AI) models is a fundamental and important part of the overall workflow. However, training AI models on magnetic resonance (MR) scans can be problematic because the provided training data may not be suitable for obtaining an improved model. Summary of the Invention

[0003] The invention provides a medical system, a computer program, a data structure and a method according to the independent claims. Embodiments are given in the dependent claims.

[0004] Training AI models on MR scans can be problematic because the training data is typically small in size and the quality and representation of the data can be poor. This is because the acquisition of high-quality MR data can be challenging and time consuming, and sometimes may not be possible. An embodiment may use natural high-quality videos as a data source for training three-dimensional (3D) AI image quality improvement models for 3D MR scans. An advantage of this approach may lie in the diversity of the data. The videos may be of high quality, without compression, noise, or interpolation artifacts, which can serve as good reference data. Special pre-processing and data augmentation techniques may be provided to enable such data to be used for 3D image quality improvement. This approach may be advantageous for denoising, anti-ringing, and super-resolution tasks of MR scans in Image Quality (IQ)Boost 3D, and it may be extended to other modalities and tasks.

[0005] In one aspect, the present invention provides a medical system comprising a memory storing machine-executable instructions and a computing system, wherein execution of the machine-executable instructions causes the computing system to: receive videos from one or more sources, select videos having desired values ​​of one or more image parameters from the received videos; create a dataset (initial dataset) comprising the selected videos, and preprocess the videos of the dataset. The preprocessing comprises: downsampling the videos by a predefined reduction factor; cropping the downsampled videos in space and / or time to obtain video patches, referred to herein as clean video patches; and altering the clean video patches to reduce the quality of the clean video patches. A training dataset may be created. The training dataset comprises altered video patches as input and associated clean video patches as learning targets.

[0006] For example, the video can be downsampled along the spatial dimension. Alternatively, the video can be downsampled along the temporal dimension. This can be particularly advantageous in the case of slow-motion video.

[0007] The machine learning model can be configured to enhance the quality of acquired 3D image data (such as 3D MR image data). The quality of the image data can include the noise level in the image data, the resolution of the image data, the number of artifacts in the image data, etc. The enhancement can include denoising, anti-ringing, super-resolution, motion artifact reduction, ghosting artifact removal, or undersampling artifact removal. The machine learning model can be, for example, a fully convolutional neural network with a 3D kernel, a ResNet, a DenseNet, an EfficientNet, or a Vision Transformer with a tuned kernel for 3D processing.

[0008] Video can be received from one or more sources. Video can be viewed as a three-dimensional volume, where the third dimension is time. The source can include an online source, such as a social media site or a public database, e.g., with appropriate licensing policies. The received video can represent a non-medical domain. For example, the received video does not include medical images, such as magnetic resonance (MR) images. For example, the received video can be natural, high-quality video (NV). Training on natural videos can increase the versatility of machine learning models. Machine learning models can learn more valuable features from non-MRI content. Using the received videos, machine learning models are less likely to overfit to any anatomical structures; therefore, this can increase confidence that anatomical structures will remain unaffected during inference. Furthermore, the received videos (such as NV) may be available in large quantities, and there may be many high-resolution videos that can serve as good reference for the machine learning model. In addition, the received videos can be easily used for research purposes because the risk of data privacy issues is lower compared to medical domain data. Because the machine learning model can be trained with many high-quality non-medical examples, it can perform well on data from other domains during inference. Furthermore, using videos from non-medical domains can be advantageous for the following reasons. Acquiring high-quality medical images (e.g., 3DMR images) to train AI models can be very expensive and sometimes impossible. For example, acquiring completely noise-free medical images may be impossible. By acquiring and averaging multiple medical images, noise can be reduced to some extent. However, for in-vivo scans, this may result in subject motion and, therefore, reduce image quality. This video can overcome these issues associated with 3DMR images.

[0009] This subject matter can further improve the training of a machine learning model based on the received video by selecting relevant videos, downsampling the selected videos, and creating video patches based on the downsampled videos. The downsampling step followed by a cropping step can form a preprocessing (or processing) step.

[0010] Downsampling of a video can be performed along the spatial dimension of the video or along the temporal dimension of the video. Downsampling of a video can, for example, include downsampling each frame of the video to a smaller size. For example, a frame can have an initial size of X0xY0 and can be downsampled to a size of X1xY1, where X1 < X0 and Y1 < Y0. This can save the memory space required to store each frame of the video while providing 3D image data having a quality that is better than or at least similar to the quality of the acquired medical image data. Another purpose of downsampling can be to match the scale of the image content to the MRI-like image content. For example, a patch of size 50x50 cropped from a 4k resolution image can contain less structure and content than a 50x50 cropped from a full high-definition resolution image. Thus, a downsampling factor can be selected such that the resulting image content and scale are similar to the MRI image content.

[0011] Each downsampled video can be cropped. This can produce a set of video patches, referred to as a set of clean video patches. Video patches can be randomly cropped from the corresponding downsampled videos. A video patch is a video. A video patch can be fully contained within its source video. Video patches obtained from the same video can overlap or not overlap. In one example, the overlap between video patches obtained from the same video can be no more than 20%. This can enable coverage of the entire 3D space occupied by the video and thus coverage of all features in the video. For example, three different types of video patches can be obtained from each downsampled video in a dataset. A spatial video patch can be obtained by cropping a given video in space such that the spatial video patch has the same duration as the given video but is cropped to, for example, 30% of the spatial dimension. A temporal video patch can be obtained by cropping a given video in time such that the temporal video patch has the same spatial size as the given video but is cropped to, for example, 30% of the duration. A spatio-temporal video patch can be obtained by cropping a given video in both time and space such that the cropping is performed to, for example, 30% along all three dimensions of the given video. Video patches can enable augmentation of training data and capture local quality features of the video. This can enable improvement of medical image cleaning because patch-based training can result in a more generalized AI model.

[0012] Thus, the cropping of the downsampled video can provide the set of clean video patches. Due to their improved quality, these clean video patches can be advantageously used as targets or labels for training a machine learning model. The input for training the machine learning model can be obtained by changing the set of clean video patches. The preprocessing can also include a change step. The change of the set of clean video patches can result in a set of changed video patches associated with the corresponding clean video patches. The present subject matter can change the video patches based on the purpose of the machine learning model. For example, if the machine learning model is used to denoise medical images, the change can be performed by adding noise to the video patches. A training dataset can be created using the set of changed video patches as input and the associated set of clean video patches as targets.

[0013] The present subject matter can be advantageous because it can achieve image quality enhancement, for example, by more accurately denoising images using a machine learning model after training using the present training dataset. That is, when used, the present training dataset can enable improved cleaning of 3D medical images. The present subject matter can provide a reliable training dataset in the desired amount without relying on medical images, whose acquisition can be time-consuming and may not provide the very high-quality medical images that may be required for training. Cleaning of 3D medical images can, for example, include denoising, super-resolution, sharpening, motion artifact reduction, ghosting artifact removal, or undersampling artifact removal.

[0014] This topic prevents the use of 2D AI image quality improvement models and applies them slice by slice to 3D MRI scans. In fact, due to the very small slice thickness of 3D MRI scans, applying 2D models slice by slice can lead to inconsistencies between slices, which reduces quality and degrades the scan. Therefore, IQ improvement models for 3D MRI acquisitions may require 3D models trained on the 3D data provided by this topic.

[0015] According to one embodiment, execution of the machine-executable instructions causes a computing system to alter dimensions in selected video tiles from the set of clean video tiles prior to altering the set of clean video tiles. For example, before altering the set of clean video tiles, a subset of one or more video tiles may be randomly selected from the set of clean video tiles. The dimensions of each selected video tile may be altered. The resulting set of clean video tiles includes altered video tiles and unaltered video tiles. The resulting set of clean video tiles may be altered. This may result in a set of altered video tiles, respectively, associated with the set of clean video tiles. This embodiment may enable blending of dimensions so that a machine learning model is not biased towards the temporal dimension. This may also enable all data to be processed equally across different dimensions.

[0016] According to one embodiment, execution of the machine-executable instructions causes the computing system to flip selected video slices from the set of clean video slices before modifying the set of clean video slices. For example, before modifying the set of clean video slices, a subset of one or more video slices may be randomly selected from the set of clean video slices. Each selected video slice may be flipped. The resulting set of clean video slices includes flipped video slices and non-flipped video slices. The resulting set of clean video slices may be modified.

[0017] By randomly shuffling and / or flipping, the machine learning model can be prevented from overfitting to the properties of the time dimension. Shuffling and / or flipping can make the video more similar to 3D MR data properties, where each image dimension has similar properties.

[0018] In one example, execution of the machine-executable instructions causes the computing system to alter and flip a subset of selected video tiles in the resulting set of clean video tiles before altering the resulting set of clean video tiles. Alternatively, execution of the machine-executable instructions causes the computing system to alter a first subset of selected video tiles in the resulting set of clean video tiles before altering the resulting set of clean video tiles, and flip a second, different subset of selected video tiles in the resulting set of clean video tiles. These examples can further improve data representation of training datasets.

[0019] According to one embodiment, each video in the initial dataset may have at least one of the following characteristics: a resolution higher than a minimum resolution, missing artifacts from compression or interpolation, an amount of noise less than a maximum amount, and a gradient change greater than a minimum amount. For example, the embodiment may define the following image parameters: resolution, amount of artifacts, amount of noise, and gradient in each direction. The target (or expected) value of the resolution parameter may be a resolution higher than the minimum resolution. The target value for the amount of artifacts may be zero. The target value for the amount of noise may be an amount of noise less than a maximum amount. The gradient change in each direction may be greater than a minimum amount. The last point means that the video should contain some motion and not include static images over time. This initial selection of videos can enable a reliable training dataset to be obtained for accurate training of machine learning models. This can enable the use of such trained models to improve medical image cleaning. For example, cleaning can include denoising, super-resolution, sharpening, motion artifact reduction, ghosting artifact removal, or undersampling artifact removal.

[0020] According to one embodiment, execution of the machine executable instructions causes the computing system to modify the video slice by adding random Gaussian noise to the video slice or reducing the resolution of the video slice. The reduction in resolution can be performed, for example, by applying cropping in k-space of the video slice, wherein a specific percentage of the information of the k-space is maintained. The k-space cropping of the video slice can be performed to obtain a correspondence between the properties of the video and the k-space of the acquired MR data. For example, the k-space cropping can achieve a mapping between the data interval in k-space and the cropped video slice. The application of k-space cropping can result in low resolution video and ringing artifacts in the video, which is used as input to the machine learning model. The resolution reduction can also be caused by downsampling and interpolation (e.g., using bspline).

[0021] The changes are performed based on the purpose of the machine learning model. As an example, if the machine learning model is a 3D AI denoiser, the machine learning model can be trained by adding random Gaussian noise to a preprocessed video, which is then used as input to the machine learning model. The learning target is a clean preprocessed video without noise. In this case, training can be performed with a standard gradient descent method. In another example, if the machine learning model is a 3D AI anti-ringing and super-resolution model, the model can be trained by applying k-space cropping to the preprocessed video, thereby producing a low-resolution video with ringing artifacts, which is used as input to the model. The learning target is a clean preprocessed video.

[0022] According to one embodiment, execution of the machine-executable instructions causes the computing system to randomly perform cropping so that each video tile has the same size, which is smaller than a maximum size. This embodiment can systematically handle large volume sizes using small but fixed tile sizes. Random cropping can prevent machine learning models from overfitting to specific features.

[0023] According to one embodiment, execution of the machine-executable instructions causes the computing system to apply an interpolator to downsample the video. This can enable good video quality to be maintained after downsampling. The interpolator can, for example, be a high-quality interpolator such as Lanczos or other high-quality interpolator that can be used to downsample the received video.

[0024] According to one embodiment, execution of the machine-executable instructions causes a computing system to augment an initial dataset by resampling each video in the initial dataset to obtain a different version of the video, where each version of the video has a different size. The resulting videos are then added to the initial dataset before preprocessing the videos in the dataset. This can further increase the size and diversity of the training dataset, thereby enabling accurate medical image cleanup. For example, cleanup can include denoising, super-resolution, sharpening, motion artifact reduction, ghosting artifact removal, or undersampling artifact removal.

[0025] According to one embodiment, a reduction factor is defined based on the resolution of the medical image used by the machine learning model. For example, a video can be downsampled to the size of a 3D medical image used by the machine learning model during inference. The 3D medical image can be, for example, an MRI image or a computed tomography (CT) image. For example, the reduction factor is 2, 3, or 4. For example, these values ​​can be user-defined, such as empirically defined values ​​that provide the best downsampling results. Having a predefined fixed reduction factor can enable systematic downsampling of the video and save resources required to dynamically calculate it.

[0026] According to one embodiment, execution of the machine-executable instructions causes a computing system to train a machine learning model using a training dataset. In one example, the trained machine learning model can be used to infer the quality of an input 3D medical image. The input 3D medical image can be an MR image, an ultrasound image, or a CT image.

[0027] According to one embodiment, the medical system further includes a magnetic resonance imaging system, wherein the memory further includes pulse sequence commands configured to control the magnetic resonance imaging system to acquire k-space data according to a magnetic resonance imaging protocol, wherein execution of the machine-executable instructions further causes the computing system to: acquire k-space data by controlling the magnetic resonance imaging system; reconstruct 3D magnetic resonance images based on the k-space data; and enhance the quality of the 3D magnetic resonance images using the trained model. The use of the trained machine learning model includes inference of the machine learning model using the reconstructed 3D MR images. This can achieve a reliable and accurate MR imaging system.

[0028] Natural images (e.g., received videos) may not have complex parts, while MR scans do. To overcome this during inference, the real and imaginary parts of the MR scan can be viewed and processed as separate images by the model, and the two outputs are combined to obtain the final result.

[0029] According to one embodiment, the received video is from a non-medical field. This can be advantageous compared to medical images for the following reasons. The received video can provide representative, high-quality, and sufficient data for training. However, in the medical field of MRI, existing MRI datasets may lack quality and quantity, and they may often be biased. MRI data collected at high resolution may not be available at all, as no scanners are capable of collecting data at, for example, 4K resolution. MR scans are affected by noise and Gibb's ringing artifacts, resulting in low resolution and low perceived sharpness, which reduces image quality. Training on MRI data can also be unsafe, as the model can learn anatomical structures (such as brain / knee structures) and perform on them during inference, introducing unexpected core changes. These issues can be overcome by using videos such as natural video (NV). The data has a higher resolution, so a single data point contains more information. The data volume can be 10-100 times that of MRI data. NV can contain every possible structure and geometry, not just MRI-specific ones. Therefore, the data is much more representative.

[0030] In another aspect, the present invention relates to a computer program comprising machine-executable instructions for execution by a computing system, wherein execution of the machine-executable instructions causes the computing system to: receive videos from one or more sources; select videos having desired values ​​of one or more image parameters from the received videos; create a dataset comprising the selected videos; pre-process the videos of the dataset, wherein the pre-processing comprises: downsampling the videos by a predefined reduction factor; cropping the downsampled videos in space and / or time to obtain video patches, referred to herein as clean video patches; and altering the clean video patches to reduce the quality of the clean video patches, and creating a training dataset comprising the altered video patches as input and associated clean video patches as learning targets. For example, the video may be downsampled along a spatial dimension of the video or along a temporal dimension of the video.

[0031] In another aspect, the present invention relates to a method for creating training data for a machine learning model configured to improve the quality of acquired 3D medical images. The method comprises:

[0032] receiving video from one or more sources;

[0033] selecting a video having desired values ​​of one or more image parameters from the received videos;

[0034] Creating a dataset including the selected videos;

[0035] Preprocessing the video of the data set, the preprocessing comprising:

[0036] downsampling the video by a predefined reduction factor;

[0037] cropping the downsampled video in space and / or time to obtain video patches, referred to herein as clean video patches;

[0038] changing the clean video slice to reduce the quality of the clean video slice;

[0039] A training dataset is created, comprising the altered video patches as input and the associated clean video patches as learning targets. For example, the video may be downsampled along a spatial dimension of the video or along a temporal dimension of the video.

[0040] In another aspect, the present invention relates to a computer-implemented data structure for training a machine learning model for cleaning 3D medical images, the data structure comprising altered video patches and associated clean video patches, wherein the altered video patches have alterations obtained according to a cleaning purpose of the machine learning model. Cleaning of the 3D medical images may include, for example, denoising, super-resolution, sharpening, motion artifact reduction, ghosting artifact removal, or undersampling artifact removal from the 3D medical images.

[0041] It should be understood that one or more of the above-described embodiments of the present invention may be combined as long as the combined embodiments are not mutually exclusive.

[0042] As will be appreciated by those skilled in the art, aspects of the present invention may be embodied as apparatus, methods, or computer program products. Thus, aspects of the present invention may take the form of entirely hardware embodiments, entirely software embodiments (including firmware, resident software, microcode, etc.), or embodiments combining software and hardware aspects, all of which may generally be referred to herein as "circuits," "modules," or "systems." Furthermore, aspects of the present invention may take the form of computer program products embodied in one or more computer-readable media having computer executable code embodied thereon.

[0043] Any combination of one or more computer-readable media may be utilized. A computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. As used herein, a computer-readable storage medium includes any tangible storage medium that can store instructions that can be executed by a processor or a computing system of a computing device. A computer-readable storage medium may be referred to as a computer-readable non-transitory storage medium. A computer-readable storage medium may also be referred to as a tangible computer-readable medium. In some embodiments, a computer-readable storage medium may also be capable of storing data that can be accessed by a computing system of a computing device. Examples of computer-readable storage media include, but are not limited to, floppy disks, magnetic hard disk drives, solid-state hard disks, flash memory, USB thumb drives, random access memory (RAM), read-only memory (ROM), optical disks, magneto-optical disks, and register files of a computing system. Examples of optical disks include compact disks (CDs) and digital versatile disks (DVDs), such as CD-ROMs, CD-RWs, CD-Rs, DVD-ROMs, DVD-RWs, or DVD-R disks. The term computer-readable storage medium also refers to various types of recording media that can be accessed by a computing device via a network or communication link. For example, data can be retrieved via a modem, via the Internet, or via a local area network. Computer executable code embodied on a computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0044] A computer-readable signal medium may include a propagated data signal having computer-executable code embodied therein, for example, in baseband or as part of a carrier wave. Such a propagated signal may take any of a variety of forms, including but not limited to electromagnetic, optical, or any suitable combination thereof. A computer-readable signal medium may be any computer-readable medium that is not a computer-readable storage medium and that can convey, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.

[0045] Computer memory or storage is an example of a computer-readable storage medium. Computer memory is any memory directly accessible to a computing system. Computer storage or storage is another example of a computer-readable storage medium. Computer storage is any non-volatile computer-readable storage medium. In some embodiments, computer storage can also be computer memory, and vice versa.

[0046] As used herein, a computing system includes an electronic component capable of executing a program or machine-executable instructions or computer-executable code. References to computing systems including examples of "computing systems" should be interpreted as potentially including more than one computing system or processing core. A computing system may, for example, be a multi-core processor. A computing system may also refer to a collection of computing systems within a single computer system or distributed across multiple computer systems. The term computing system should also be interpreted as potentially referring to a collection or network of computing devices, each of which includes a processor or computing system. Machine-executable code or instructions may be executed by multiple computing systems or processors, which may be within the same computing device or even distributed across multiple computing devices.

[0047] Machine executable instructions or computer executable code can include instructions or programs that cause a processor or other computing system to perform an aspect of the present invention. The computer executable code for performing the operations of various aspects of the present invention can be written in any combination of one or more programming languages, including object-oriented programming languages ​​(such as Java, Smalltalk, C++, etc.) and conventional procedural programming languages ​​(such as "C" programming language or similar programming languages), and compiled into machine executable instructions. In some cases, the computer executable code can be in the form of a high-level language or in a precompiled form and is used in conjunction with an interpreter that generates machine executable instructions on the fly. In other cases, the machine executable instructions or computer executable code can be the programming form of a programmable logic gate array.

[0048] The computer-executable code may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the latter case, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0049] Aspects of the present invention are described with reference to the flow chart and / or block diagram of the method, device (system) and computer program product according to an embodiment of the present invention.It should be understood that each frame or the part of frame of flow chart, diagram and / or block diagram can be implemented by the computer program instruction of computer executable code form when applicable.It should also be understood that, when not mutually exclusive, the combination of the frames in different flow charts, diagrams and / or block diagrams can be combined.These computer program instructions can be provided to the computing system of general-purpose computer, special-purpose computer or other programmable data processing device to produce machine, so that the instruction executed via the computing system of computer or other programmable data processing device creates the module for implementing the function / action specified in one or more frames of flow chart and / or block diagram.

[0050] These machine-executable instructions or computer program instructions may also be stored in a computer-readable medium that can direct a computer, other programmable data processing apparatus, or other device to work in a specific manner so that the instructions stored in the computer-readable medium produce an article of manufacture including instructions for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0051] Machine-executable instructions or computer program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable apparatus provide a process for implementing the functions / actions specified in one or more boxes of the flowchart and / or block diagram.

[0052] As used herein, a 'user interface' is an interface that allows a user or operator to interact with a computer or computer system. A user interface may also be referred to as a human-machine interface device. A 'user interface' can provide information or data to an operator and / or receive information or data from an operator. A user interface enables input from an operator to be received by a computer and can provide output from the computer to a user. In other words, a user interface can allow an operator to control or manipulate a computer, and the interface can allow the computer to indicate the effects of the operator's controls or manipulations. The display of data or information on a display or graphical user interface is an example of providing information to an operator. Receiving data via a keyboard, mouse, trackball, touchpad, pointing stick, graphics tablet, joystick, game controller, webcam, headset, pedal, wired gloves, remote control, and accelerometer are all examples of user interface components that enable receiving information or data from an operator.

[0053] As used herein, a 'hardware interface' includes an interface that enables a computing system of a computer system to interact with and / or control an external computing device and / or apparatus. A hardware interface may allow a computing system to send control signals or instructions to an external computing device and / or apparatus. A hardware interface may also enable a computing system to exchange data with an external computing device and / or apparatus. Examples of hardware interfaces include, but are not limited to, a universal serial bus, an IEEE 1394 port, a parallel port, an IEEE 1284 port, a serial port, an RS-232 port, an IEEE-488 port, a Bluetooth connection, a wireless LAN connection, a TCP / IP connection, an Ethernet connection, a control voltage interface, a MIDI interface, an analog input interface, and a digital input interface.

[0054] As used herein, a 'display' or 'display device' includes an output device or user interface suitable for displaying images or data. A display may output visual, audio, and / or tactile data. Examples of displays include, but are not limited to, computer monitors, television screens, touch screens, tactile electronic displays, Braille screens, cathode ray tubes (CRTs), memory tubes, bi-stable displays, electronic paper, vectorscopes, flat panel displays, vacuum fluorescent displays (VFs), light emitting diode (LED) displays, electroluminescent displays (ELDs), plasma display panels (PDPs), liquid crystal displays (LCDs), organic light emitting diode displays (OLEDs), projectors, and head-mounted displays.

[0055] K-space data is defined herein as measurements of radio frequency signals emitted by atomic spins recorded using the antenna of a magnetic resonance apparatus during a magnetic resonance imaging scan.Magnetic resonance data is an example of tomographic medical image data.

[0056] A magnetic resonance image or MR image is defined herein as a reconstructed two- or three-dimensional visualization of anatomical data contained within k-space data. This visualization can be performed using a computer. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Preferred embodiments of the present invention will now be described, by way of example only, with reference to the accompanying drawings, in which:

[0058] Figure 1 An example of a medical system is illustrated;

[0059] Figure 2 Shows the use Figure 1 A flowchart of a method of a medical system;

[0060] Figure 3 An example of a medical system is illustrated;

[0061] Figure 4 Shows the use Figure 3A flowchart of a method of a medical system;

[0062] Figure 5 Shows the use Figure 1 A flowchart of a method of a medical system;

[0063] Figure 6 Illustrated are examples of several different medical images before and after being processed by a trained machine learning model. DETAILED DESCRIPTION

[0064] Like numbered elements in these figures are equivalent elements or perform the same function. Elements that have been discussed previously will not necessarily be discussed in subsequent figures if the function is equivalent.

[0065] Figure 1 An example of a medical system 100 is illustrated. In this example, the medical system 100 includes a computer 102 having a computing system 104. Computer 102 is intended to represent one or more computers or computer systems. Computing system 104 is intended to represent one or more computing systems or computing cores. Computer 102 is also shown as including an optional hardware interface 106 that connects to computing system 104. If other components of the medical system 100 are present or included, such as a magnetic resonance imaging system, hardware interface 106 can be used to exchange data and commands with these other components. Medical system 100 is also shown as including an optional user interface 108, which can provide various modules for relaying data and receiving data and commands from an operator.

[0066] Medical system 100 is also shown as including memory 110 connected to computing system 104. Memory 110 is intended to represent various types of memory accessible by computing system 104. Memory 110 is shown as containing machine-executable instructions 120. Machine-executable instructions 120 are instructions that enable computing system 104 to perform various control, data processing, and image processing tasks. Memory 110 is also shown as containing a machine learning model 122 and video 124 obtained from one or more network sources.

[0067] The machine learning model 122 may be configured to receive a 3D medical image as input and clean the 3D medical image. Cleaning may include, for example, denoising, anti-ringing, or super-resolution.

[0068] Figure 2 is a flow chart of a method for creating training data for a machine learning model according to an example of the present subject matter. The machine learning model is configured to obtain quality-enhanced images from acquired 3D medical images. The method may be, for example, Figure 1The medical system may be configured to perform a training process. In step 201, videos may be received from one or more sources, such as a network source. In step 203, videos having target values ​​for image parameters may be selected from the received videos. In step 205, a dataset including the selected videos may be created. This may be referred to as an initial dataset. In step 207, the videos of the initial dataset may be downsampled by a predefined reduction factor. The downsampling may be performed, for example, along a spatial dimension of the video or along a temporal dimension of the video. In step 209, the downsampled videos of the initial dataset may be cropped in space and / or time to obtain video patches, referred to herein as clean video patches. In step 211, the clean video patches may be altered to reduce the quality of the video patches. In step 213, a training dataset may be created. The training dataset includes the altered video patches as input and the associated clean video patches as learning targets. Step 213 may also, for example, store the training dataset, for example, in the medical system 100 or in an accessible remote system.

[0069] Figure 3 Another example of a medical system 300 is illustrated. Figure 3 The medical system in Figure 1 The medical system in FIG. 1 further comprises a magnetic resonance imaging system 302 controlled by the computing system 104 . Figure 3 The example shown is a magnetic resonance imaging system 302 .

[0070] The magnetic resonance imaging system 302 comprises a magnet 304. The magnet 304 is a superconducting cylindrical magnet having a bore 306 passing through it. The use of different types of magnets is also possible; for example, split cylindrical magnets and so-called open magnets can also be used. A split cylindrical magnet is similar to a standard cylindrical magnet, except that the cryostat has been divided into two parts to allow access to the isoplane of the magnet, which magnet can be used, for example, in conjunction with charged particle beam therapy. An open magnet has two magnet sections, one above the other, with a space in between that is large enough to receive the subject: the arrangement of the two section areas is similar to that of Helmholtz coils. Open magnets are popular because the subject is less restricted. Inside the cryostat of the cylindrical magnet, there is a collection of superconducting coils.

[0071] Within the bore 306 of the cylindrical magnet 304 lies an imaging zone 308, wherein the magnetic field is sufficiently strong and uniform to perform magnetic resonance imaging. A field of view 309 is shown within the imaging zone 308. K-space data is typically acquired for the field of view 309. The region of interest may be the same as the field of view 309, or it may be a subvolume of the field of view 309. An object 318 is shown supported by an object support 320 such that at least a portion of the object 318 is within the imaging zone 308 and the field of view 309.

[0072] Within the bore 306 of the magnet, there is also a set of magnetic field gradient coils 310 for acquiring preliminary k-space data to spatially encode magnetic spins within the imaging zone 308 of the magnet 304. The magnetic field gradient coils 310 are connected to a magnetic field gradient coil power supply 312. The magnetic field gradient coils 310 are intended to be representative. Typically, the magnetic field gradient coils 310 include three separate sets of coils for spatial encoding in three orthogonal spatial directions. The magnetic field gradient power supply supplies current to the magnetic field gradient coils. The current supplied to the magnetic field gradient coils 310 is controlled as a function of time and can be ramped or pulsed.

[0073] Adjacent to the imaging zone 308 is a radio frequency coil 314, which is used to manipulate the orientation of magnetic spins within the imaging zone 308 and to receive radio transmissions from spins also within the imaging zone 308. The radio frequency antenna may include multiple coil elements. An radio frequency antenna may also be referred to as a channel or antenna. The radio frequency coil 314 is connected to a radio frequency transceiver 316. The radio frequency coil 314 and the radio frequency transceiver 316 may be replaced by separate transmit and receive coils, and separate transmitters and receivers. It should be understood that the radio frequency coil 314 and the radio frequency transceiver 316 are representative. The radio frequency coil 314 is also intended to represent a dedicated transmit antenna and a dedicated receive antenna. Similarly, the transceiver 316 may also represent a separate transmitter and receiver. The radio frequency coil 314 may also have multiple transmit / receive elements, and the radio frequency transceiver 316 may have multiple transmit / receive channels.

[0074] The transceiver 316 and gradient controller 312 are shown connected to the hardware interface 106 of the computer system 102 .

[0075] The memory 110 is also shown as containing pulse sequence commands 330. The pulse sequence commands 330 are commands or data that can be converted into commands that enable the computing system 104 to control the magnetic resonance imaging system to acquire k-space data 332. The memory 110 shows k-space data 332 that has been acquired by controlling the magnetic resonance imaging system using the pulse sequence commands 330.

[0076] Figure 4 is a flow chart of a method for acquiring MR image data according to an example of the present subject matter. The method may be, for example, Figure 3 For example, you can use Figure 2 The training dataset obtained by the method is used to train a machine learning model. The machine learning model can be configured to clean 3D MR images.

[0077] K-space data may be acquired by controlling a magnetic resonance imaging system in step 401. A 3D magnetic resonance image may be reconstructed based on the k-space data in step 403. The trained machine learning model may be used in step 405 to enhance the quality of the 3D magnetic resonance image. Figure 6 An example of a reconstructed image 601 before denoising and a corresponding image 605 denoised using a trained machine learning model is depicted. That is, the denoised image is obtained through inference using the trained machine learning model. Natural images (e.g., received videos) may lack complex features, while MR scans do. To overcome this for inference, the real and imaginary parts of the MR scan can be treated as separate images, and the two outputs are summed to obtain the final result.

[0078] Figure 5 is a flowchart of a method for creating training data for a machine learning model according to an example of the present subject matter.

[0079] Natural high-quality video (initial video) can be obtained from a web source and used to train a machine learning model. The machine learning model can be integrated in the IQ Boost 3D system, for example. The machine learning model can be, for example, a convolutional neural network, which can be trained for denoising, anti-ringing and super-resolution tasks. Data selection can be performed in step 501. For example, a video can be selected from the initial video. The video can be selected so that it has a high resolution, it should not contain artifacts from compression or interpolation, and it should contain a limited amount of noise. Also, the gradient in each direction in the selected video should not be zero, which means that the video should contain some motion and not include static images over time. The data can also be selected based on its licensing policy so that it can be used for product development purposes.

[0080] Preprocessing of the selected videos may be performed. In particular, the following steps may be embodiments of preprocessing; however, variations of the preprocessing may occur. The videos are viewed as three-dimensional volumes, where the third dimension is time. As an example, the preprocessing steps may be used for IQ Boost3D, but the preprocessing is not limited to these specific steps. All videos may be downsampled to a smaller matrix size in step 503, which results in a structural appearance that is more similar to an MRI scan. The time dimension remains intact. To maintain good quality, a high-quality interpolator such as Lanczos may be used. As additional data augmentation, more versions of the video may be generated in step 505 by resampling to different matrix sizes.

[0081] In step 507, random patches of small size from each dimension are cropped from the normalized video. This step can allow processing of large volume sizes. In some embodiments, a small but fixed patch size can be used. In step 509, random permutation and flipping of the dimensions are applied to reduce the influence of the temporal dimension. This can be critical because the temporal dimension is very different from the spatial dimension. By random permutation and flipping, the AI ​​model can be prevented from overfitting to the properties of the temporal dimension. For example, these steps make the video more similar to the MR data properties.

[0082] The preparation of input data for the machine learning model may depend on the purpose of the model, as it includes a modification step. In step 511, the video segments may be modified according to the purpose of the machine learning model. A training dataset may be created from the video segments, and the training dataset may be used to train the machine learning model in step 513.

[0083] The machine learning model can be referred to as a 3D AI model. The 3D AI model can be a fully convolutional neural network with a 3D kernel or a ResNet, DenseNet, EfficientNet, or VisionTransformer with a tuned kernel for 3D processing. As an example of an ML model, a 3D AI denoiser can be trained by adding random Gaussian noise to a clean preprocessed video, which is then used as input to the model. The learning target is a clean preprocessed clean video without noise. Training is performed using a standard gradient descent method. Another example is a 3D AI anti-ringing and super-resolution model. The model can be trained by applying k-space cropping to a clean preprocessed video, resulting in a low-resolution video with ringing artifacts, which is used as input to the model. The learning target is a clean preprocessed video.

Claims

1. A medical system for creating training data for a machine learning model configured to improve the quality of acquired 3D medical images, the medical system comprising: memory that stores machine-executable instructions; and a computing system, wherein execution of the machine-executable instructions causes the computing system to: receiving video from one or more sources; selecting a video having desired values ​​of one or more image parameters from the received videos; Creating a dataset including the selected videos; Preprocessing the video of the data set, the preprocessing comprising: downsampling the video by a predefined reduction factor; cropping the downsampled video in space and / or time to obtain video patches, referred to herein as clean video patches; changing the clean video slice to reduce the quality of the clean video slice; A training dataset is created that includes the altered video patches as input and the associated clean video patches as learning targets.

2. The medical system according to claim 1, wherein Execution of the machine-executable instructions causes the computing system to alter dimensions in selected ones of the clean video tiles prior to the changing of the clean video tiles.

3. The medical system according to any one of the preceding claims, wherein Execution of the machine-executable instructions causes the computing system to flip selected ones of the clean video tiles prior to the changing of the clean video tiles.

4. The medical system according to any one of the preceding claims, wherein The desired value of the image parameter includes at least any of: a resolution higher than a minimum resolution, missing artifacts from compression or interpolation, an amount of noise less than a maximum amount, and a gradient change greater than a minimum amount for each direction in the video.

5. The medical system according to any one of the preceding claims, wherein Execution of the machine-executable instructions causes the computing system to alter the video tiles by adding random Gaussian noise to the video tiles and / or reducing the resolution of the video tiles.

6. The medical system according to any one of the preceding claims, wherein Execution of the machine-executable instructions causes the computing system to randomly perform the cropping such that each video tile has a size that is smaller than a maximum size. 7 . The medical system according to claim 1 , wherein the received video is a video in a non-medical field.

8. The medical system according to any one of the preceding claims, wherein Execution of the machine-executable instructions causes the computing system to apply an interpolator to perform the downsampling of the video.

9. The medical system according to any one of the preceding claims, wherein Execution of the machine-executable instructions causes the computing system to: augment the dataset by resampling each video in the dataset to obtain a different version of the video, each version of the video having a different size; and add the obtained videos to the dataset before the preprocessing.

10. The medical system of any one of the preceding claims, wherein the reduction factor is defined such that the resolution obtained after downsampling the video is similar to or equal to the resolution of the 3D medical image of the machine learning model.

11. The medical system according to any one of the preceding claims, wherein Execution of the machine-executable instructions causes the computing system to train the machine learning model using the training dataset.

12. The medical system according to claim 11, wherein The medical system also includes a magnetic resonance imaging system, wherein the memory further includes pulse sequence commands, the pulse sequence commands being configured to control the magnetic resonance imaging system to acquire k-space data according to a magnetic resonance imaging protocol, wherein execution of the machine-executable instructions further causes the computing system to: acquire the k-space data by controlling the magnetic resonance imaging system; reconstruct a 3D magnetic resonance image based on the k-space data; and enhance the quality of the 3D magnetic resonance image using the trained model.

13. A computer program comprising machine-executable instructions, wherein: Execution of the machine-executable instructions causes the computing system to: receiving video from one or more sources; selecting a video having desired values ​​of one or more image parameters from the received videos; Creating a dataset including the selected videos; Preprocessing the video of the data set, the preprocessing comprising: downsampling the video by a predefined reduction factor; cropping the downsampled video in space and / or time to obtain video patches, referred to herein as clean video patches; changing the clean video slice to reduce the quality of the clean video slice; A training dataset is created that includes the altered video patches as input and the associated clean video patches as learning targets.

14. A method for creating training data for a machine learning model configured to improve the quality of acquired 3D medical images, the method comprising: receiving video from one or more sources; selecting a video having desired values ​​of one or more image parameters from the received videos; Creating a dataset including the selected videos; Preprocessing the video of the data set, the preprocessing comprising: downsampling the video by a predefined reduction factor; cropping the downsampled video in space and / or time to obtain video patches, referred to herein as clean video patches; changing the clean video slice to reduce the quality of the clean video slice; A training dataset is created that includes the altered video patches as input and the associated clean video patches as learning targets.

15. A computer-implemented data structure for training a machine learning model for cleaning 3D medical images, the data structure comprising altered video patches and associated clean video patches, wherein: The changed video tiles have changes obtained according to the cleaning purpose of the machine learning model.

16. The data structure of claim 15, the changed video slice comprising a noisy video slice or a low-resolution video slice with ringing artifacts.