Training a deep learning model
By introducing a loss value that penalizes the expected rate of change into the longitudinal data to train the machine learning model, the problem of noisy ground truth data in the longitudinal data is solved, and accurate inference and longitudinal consistency at a single time point are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- COMMONWEALTH SCI & IND RES ORG
- Filing Date
- 2024-12-19
- Publication Date
- 2026-07-31
AI Technical Summary
When training machine learning models on longitudinal data, the ground-based data is often noisy, especially due to measurements taken at different time intervals using different equipment and calibration settings, making it difficult to train accurate models.
By calculating the rate of change of biomedical factors and introducing a penalty value to reflect the expected rate of change, the weights of the machine learning model are updated to minimize the loss value, and the model is trained using the inherent properties of longitudinal data.
Under noisy ground-based real-world data conditions, it is possible to train a model capable of inference at a single time point, while improving longitudinal consistency in training and testing and reducing the impact of noise.
Smart Images

Figure CN122497957A_ABST
Abstract
Description
Cross-references to related applications
[0001] This application claims priority to Australian Provisional Patent Application No. 2023904194, filed on 22 December 2023, and Australian Provisional Patent Application No. 2024903443, filed on 23 October 2024, the contents of which are incorporated herein by reference in their entirety. Technical Field
[0002] This disclosure relates to training machine learning models on longitudinal data. More specifically, but not exclusively, this disclosure relates to training machine learning models on longitudinal medical scan images from patient studies. Background Technology
[0003] Machine learning models, such as deep neural networks, are powerful tools for solving challenging problems and providing predictions. For example, machine learning models can be used to determine the presence of a disease from an image or collection of images, such as medical images. Such machine learning models are trained and can then be used to make predictions. Training involves comparing the model's output with ground truth information (i.e., information known to be true or real and which is the model's target output) and updating the machine learning model based on the comparison.
[0004] Machine learning models, particularly those utilizing deep learning techniques, are trained to minimize a loss value (or loss function) based on an indicator or combination of indicators derived from some ground-based measurements. The loss value is typically calculated based on the difference between the output and the ground truth, and the goal of training is to minimize this loss so that the machine learning model produces an output that is nearly identical to the ground truth. In one example, for a model used for image segmentation, the loss value indicates the overlap between the predicted mask from the model and a manually defined mask of the structure to be segmented in the image. In this example, the loss value would be formulated in a way that maximizes the overlap by minimizing the loss value. Other examples of loss values include the mean squared error (MSE) measure between the predicted volumetric structure and its ground-based true volume.
[0005] However, ground truth data can often contain inherent noise due to various reasons, such as equipment used to measure the data or other errors. This makes training machine learning models on such ground truth data difficult. In particular, when the training data is longitudinal data, measurements are often taken over several years at intervals using different machines and calibration settings for each measurement. This introduces noise because of inconsistencies between consecutive measurements. Furthermore, noise makes it difficult to train models based on ground truth data from multiple sources. This situation often arises when training machine learning models on medical images, where databases of large amounts of training data are amassed because a very large amount of training data is required to accurately train the model.
[0006] Any discussion of documents, actions, materials, devices, articles, etc. included in this specification shall not be construed as an admission that any or all of these matters constitute part of the prior art or common general knowledge in the field relating to this disclosure that existed prior to the priority date of each appended claim.
[0007] Throughout this specification, the word “comprise” or variations thereof (such as “comprises” or “comprising”) shall be understood to imply inclusion of the stated element, integer or step, or group of elements, integers or steps, but does not exclude any other element, integer or step, or group of elements, integers or steps. Summary of the Invention
[0008] A method for training a machine learning model on longitudinal medical scan images from patient studies, the method comprising: To compute two or more predictions for biomedical factors, perform the following operations on each of two or more medical images: Applying machine learning models to medical images to compute outputs; and The prediction is calculated using the output of a machine learning model for medical images; Calculate the rate of change between two or more predictions for two or more medical images; and A trained machine learning model is created by updating the weights of the machine learning model to minimize the loss value, where the loss value includes a penalty value that reflects the difference between the rate of change and the expected rate of change.
[0009] The advantage is that the loss value includes a penalty reflecting the difference between the rate of change and the expected rate of change, since the expected rate of change can be forced into the machine learning model during training. More specifically, the training of the machine learning model utilizes information not available during inference. This is ideal for situations with noisy ground truth data. Furthermore, inference of the trained machine learning model can be performed by applying the model to a single image.
[0010] In some embodiments, the method further includes determining a polynomial function of the rate of change of a biomedical factor and using the polynomial function to calculate the expected rate of change.
[0011] In some embodiments, calculating the expected rate of change using a polynomial function includes applying a polynomial function to the average of a biomedical factor calculated from the averages of two or more predictions.
[0012] In some embodiments, determining the polynomial function includes calculating the rate of change and the mean for each pair of medical images in a patient study.
[0013] In some embodiments, the loss value further includes a change penalty value that reflects the difference between two or more predictions, the change penalty value penalizing the unexpected time trend of biomedical factors associated with the expected rate of change.
[0014] In some embodiments, the loss value further includes a correction penalty value that reflects the difference between two or more predictions and two or more corresponding ground truth values, the correction penalty value penalizing the two or more predictions for deviating from the two or more corresponding ground truth values.
[0015] In some embodiments, the loss value further includes a correction penalty value that reflects the difference between two or more outputs and the overcorrection factor, which penalizes the two or more outputs for deviating from the overcorrection factor.
[0016] In some embodiments, the loss value further includes an additional penalty value calculated by one or more of the following: the difference between each of two or more outputs of the machine learning model and the corresponding ground truth measurement; and the mean square error between each of two or more outputs and the corresponding ground truth measurement.
[0017] In some embodiments, the loss value is calculated as a weighted sum of the penalty value, the change penalty value, the correction penalty value, and the additional penalty value.
[0018] In some embodiments, each of the medical images in a patient study has a timestamp, and at least two of two or more medical images have non-adjacent timestamps.
[0019] In some embodiments, the method further includes computing two or more additional predictions of a biomedical factor by performing the following operations for each of two or more additional medical images of a patient study: applying a trained machine learning model to the medical images to compute an output; and using the output to compute a prediction; computing an additional rate of change between the two or more predictions; and creating an additional trained machine learning model by updating the weights of the trained machine learning model to minimize an additional loss value based on the additional rate of change.
[0020] In some embodiments, the method further includes computing two or more additional predictions of a biomedical factor by performing the following operations for each of two or more medical images from another patient study different from the patient study: applying a trained machine learning model or an additional trained machine learning model to the medical images to compute an output; and using the output to compute a prediction; computing another rate of change between the two or more additional predictions; and creating another trained machine learning model by updating the weights of the trained machine learning model or the additional trained machine learning model to minimize another loss value based on the other rate of change.
[0021] In some embodiments, the method further includes applying a trained machine learning model to a patient's medical images to aid in the diagnosis of the patient based on the output of the trained machine learning model.
[0022] In some embodiments, each of the medical images is one of the following: magnetic resonance imaging; positron emission tomography (PET) images; or computed tomography (CT) images.
[0023] In some embodiments, the biomedical factor is one of the following: medical image quantification factor; biological structure value; or biomarker.
[0024] In some embodiments, the medical image quantization factor is a Centiloid or normalized uptake value to scaling factor, and the biological structure value is the volume of the biological structure visible in the medical image.
[0025] In some embodiments, each of the medical images is a PET image, and the method further includes performing PET quantization by applying a trained machine learning model to the patient's PET images.
[0026] In some embodiments, the medical image quantization factor is a Centiloid or a normalized uptake ratio scaling factor; and the method further includes applying a trained machine learning model to the patient's medical images to aid in the diagnosis of Alzheimer's disease based on the Centiloid or normalized uptake ratio scaling factor as the output of the trained machine learning model.
[0027] In some embodiments, the method further includes generating one or more masks corresponding to corrections determined by a trained machine learning model.
[0028] In some embodiments, the machine learning model is a neural network that includes one or more convolutional layers.
[0029] Software that, when installed on and executed by a computer, causes the computer to perform the methods described above.
[0030] A system for training a machine learning model on longitudinal medical scan images from patient studies, the system comprising: The processor is configured as follows: To compute two or more predictions for biomedical factors, perform the following operations for each of two or more medical images: Applying machine learning models to medical images to compute outputs; and The prediction is calculated using the output of a machine learning model for medical images; Calculate the rate of change between two or more predictions for two or more medical images; and A trained machine learning model is created by updating the weights of the machine learning model to minimize the loss value, where the loss value includes a penalty value that reflects the difference between the rate of change and the expected rate of change.
[0031] The optional features provided in relation to this method also apply to optional features of the software and system. Attached Figure Description
[0032] The following figures will be used as a reference for the example description: Figure 1 A graph showing the calculated Centiloid value over time, calculated using longitudinal PET scan images.
[0033] Figure 2 The graph shows the rate of change of the Centiloid value versus the average Centiloid value.
[0034] Figure 3 A system for training machine learning models on longitudinal medical scan images from patient studies is shown.
[0035] Figure 4 A method for training a machine learning model on longitudinal medical scan images from patient studies is shown.
[0036] Figure 5 It is an example machine learning model with a convolutional neural network (CNN) architecture.
[0037] Figure 6 This illustrates the progression of the rate of change of the calculated scaled Centiloid value as the training of the machine learning model increases in an exemplary embodiment.
[0038] Figure 7 This shows the progress of the standardized uptake ratio (SUVR) scaling factor as the optimization process described in this paper continues.
[0039] Figure 8aThe initial target mask used to generate the optimized target mask is shown.
[0040] Figure 8b An optimized target mask with an SUVR scaling factor similar to that provided by the trained machine learning model disclosed herein is shown.
[0041] Figure 9a The initial reference mask used to generate the optimized reference mask is shown.
[0042] Figure 9b An optimized reference mask with an SUVR scaling factor similar to that provided by the trained machine learning model disclosed herein is shown. Detailed Implementation
[0043] The disclosed systems and methods relate to training machine learning models on longitudinal data to address the problem of noisy ground truth training data. The disclosed systems and methods utilize inherent or anticipated properties of the longitudinal data as a source of additional information to reduce the impact of noisy ground truth data. More specifically, longitudinal data exhibiting anticipated rates of change of factors such as biomedical factors can be used to train machine learning models by formulating training in a manner that reflects that anticipated rate of change.
[0044] Longitudinal data includes continuous measurements of the same subject over a period of time. For example, longitudinal data could be repeated positron emission tomography (PET) or magnetic resonance (MR) images of the brain on the same participant annually for N years. This is typically performed on patients at risk of developing Alzheimer's disease to determine disease progression by studying the amount of amyloid plaques in the brain over time. Because Alzheimer's disease is a progressive condition, an increase in the amount of amyloid plaques in the brain over time is expected, a trend anticipated from longitudinal data of patients with the disease.
[0045] To train machine learning models using inherent or anticipated properties, the disclosed systems and methods use a class of penalty values (called longitudinal constraints) in the loss value. More specifically, penalty values utilize the anticipated rate of change in longitudinal data by penalizing unexpected rates of change. A simple longitudinal constraint might be to penalize negative (positive) changes between the baseline and subsequent values when the anticipated change is positive (negative). These changes could be biomedical factors, such as PET quantization within a mask or the volume of a segmented mask, or other (more complex) factors. For example, with respect to Centiloid values, since the anticipated trend is a positive rate of change (i.e., the Centiloid value increases over time), penalty values can be formulated to penalize training machine learning models on negative rates of change.
[0046] This disclosure also provides a framework for incorporating these constraints onto a single point-in-time inference model. In other words, while the model is trained using longitudinal information inherent in the longitudinal data (e.g., multiple images), the trained model does not require longitudinal information during inference. This means that the trained model can be used to compute predictions at a single point in time (e.g., a single image) during inference. One of the main advantages of the disclosed systems and methods is that they enable the inclusion of longitudinal constraints without relying on multiple points in time during inference.
[0047] The proposed constraint or penalty value can also complement the existing loss by ensuring longitudinal consistency in the training set, rather than relying solely on noisy "ground truth" data. With proper training, this also translates to better longitudinal consistency in the test set, even if each image is inferred (tested) independently of all other time points. In particular, since longitudinal data can be captured over many years using machines with different calibration settings each time, or even entirely different machines together, or, in the case of PET imaging, using different PET tracers, the proposed penalty value enables the normalization of longitudinal data, which in turn allows machine learning models to be trained on such noisy data.
[0048] For example, the overall trajectory of the expected change in the mean measure can be calculated using a large population analysis. This overall trajectory can use the error between the actual change and the expected change as an additional loss. Since any information can be added to the loss, additional information, such as demographic data, can also be included. In another example, where different trajectories are expected for different clinical diagnoses, these different trajectories can be used to define different losses for different clinical groups. Because this information is not used during inference (in other words, it is only used to calculate the loss during training), this limits the risk of the model overfitting the data.
[0049] The disclosed systems and methods are applicable to many different machine learning (specifically, deep learning) applications where longitudinal data is available and where known longitudinal constraints exist that can be applied based on expected trends in the longitudinal data. For example, a machine learning application could be image segmentation, where there is a model regarding the expected change in structural volume over time as it increases / decreases. In another example, this information can be used in a loss function combined with other standard losses when training a model to predict structural volumes that are expected to shrink over time, and in this case, the volume that increases over time is penalized. This is particularly useful for problems where ground truth values may be unreliable. For example, this is particularly useful for PET quantization, as PET images can be very noisy.
[0050] Specifically, PET quantification is used to analyze PET images and extract quantification information from them. PET quantification for Alzheimer's disease research involves normalizing PET images to a Centiloid scale, which is described as a standard scale for PET images acquired using amyloid PET tracers. Generally, PET images are normalized using a normalized uptake ratio, or SUVR. SUVR is defined as the ratio of the area in a PET image containing specific binding (the target area where the PET tracer will bind to the target) to the area in a PET image containing non-specific binding (called the reference area). In the case of amyloid imaging, these areas are defined as areas containing amyloid, typically the neocortex, and areas not containing amyloid, typically the cerebellum. Because different PET tracers exist that bind to amyloid, each with different pharmacokinetic properties, the Centiloid scale was developed to unify the quantification (SUVR) from all amyloid PET tracers to the same scale using a linear transformation based on head-to-head paired data. In other words, while SUVR is commonly used for most PET tracer quantifications, the Centiloid scale is typically used in the specific case of amyloid PET imaging because it allows the same scale to quantify different PET tracers. Due to the progressive nature of Alzheimer's disease, it is expected that the units of the Centiloid scale will increase over time as the amount of amyloid deposits in the brain increases. However, plotting calculated Centiloid values over time can demonstrate that this is not always the trend observed in the data.
[0051] Figure 1 This graph shows the calculated Centiloid value over time, calculated using longitudinal PET scan images. The slope of the line represents the "rate of change" of the Centiloid value. The blue line, for example at 101, represents the expected trend of the Centiloid value increasing over time, while the red line, for example at 102, represents an unexpected trend of the Centiloid value decreasing over time, which is a result of noise in the PET scan images. This unexpected trend in Centiloid values is often overlooked because Centiloid value calculations are typically performed in the background during typical PET quantization. For example, noise causing the unexpected trend in Centiloid values could be due to noise in the PET images, noise from the reference area used to calculate the Centiloid value, and variations in the PET tracer or scanner.
[0052] Figure 2The graph shows the rate of change of the Centiloid values versus the average Centiloid values. From the expanding trend of the Centiloid values, there should be no negative rate of change (i.e., each plotted point should be above the horizontal line 201). Figure 2 Curve 202 is also shown, representing a calculated curve of the mean versus the rate of change. Curve 202 can be obtained by calculating the mean and rate of change for some or each data segment in a set of longitudinal data. More specifically, it can be applied to... Figure 2 A regression algorithm was used to generate curve 202 from the data. Expected... Figure 2 The data points should not deviate too far from curve 202, as this curve represents the expected trend of the data.
[0053] For example, considering the expected trend of the longitudinal data (in this example, the expected increase in the Centiloid value, making the rate of change positive), the training of the machine learning model can be designed to estimate the Centiloid value by penalizing negative rates of change, penalizing overcorrection, and penalizing the distance from curve 202. The training of the machine learning model is penalized in these ways by introducing constraints or penalty values, which are formulated in a way that represents the penalties for these factors.
[0054] In addition to the advantages of the disclosed systems and methods in combating noisy ground-based data during training, other benefits exist. For example, the volumes of structures present in the training images may not be well defined (such as volumes of high-signal white matter with poor inter-evaluator reproducibility), which are compensated for during training using the disclosed methods. Another advantage is that training can be performed using temporally sparse (long-interval acquisition) longitudinal data. This is particularly advantageous for training machine learning models on medical images such as MRI or CT images, since the data is collected only once every 1 or 2 years. It should be noted that, for example, machine learning models can also be trained on less sparse data using the disclosed methods, such as longitudinal data collected daily, weekly, or monthly.
[0055] While this disclosure is generally directed toward the application of medical images and machine learning models in the medical field, it should be noted that the disclosed systems and methods are not limited to applications in the medical field and medical images. More specifically, this disclosure relates more generally to training machine learning models on longitudinal data, where longitudinal data may include medical images. However, in other cases, longitudinal data may include other images. For example, any model used to estimate quantizations / values from images or any other input can be trained using the disclosed systems and methods, provided that some prior knowledge of the expected changes over time exists.
[0056] In other cases, longitudinal data may not include images at all, but rather some other form of data. For example, due to the predicted increase in drought in the region, the rate of change of water flowing through a dam can be expected to decrease over time. However, it may be difficult to accurately measure the precise volume of water flowing through the dam, which introduces a noise source. Unexpected fluctuations in volume due to sudden heavy rainfall can also introduce noise. The disclosed system and method can compensate for noise in the data by utilizing the expected rate of change of the data during training.
[0057] System for training Figure 3 An example system 300 for training machine learning models on longitudinal medical scan images of patients is shown. Figure 3 This is one example of the configuration of system 300. However, system 300 is not strictly limited to this configuration, and this may be one possible embodiment of system 300. Medical scan images, or simply medical images, may be magnetic resonance (MR) images, positron emission tomography (PET) images, or computed tomography (CT) images. Throughout this disclosure, medical scan images or medical images may be simply referred to as images or image data, and it should be noted that system 300 is equally applicable to images other than medical images.
[0058] Patient studies are studies of a single patient, in which medical images of the patient's brain are captured, for example, at regular time intervals (such as annually). It should be noted that machine learning models can be trained on longitudinal medical scans from multiple patient studies, thus training on data from multiple patients. This is because a penalty value indicating the expected rate of change reduces noise due to inconsistencies in the data between patients. Machine learning models typically require large amounts of training data, so the ability to train them on longitudinal data from multiple patient studies is an advantage. In embodiments outside the medical field, patient studies can be studies of subjects over time.
[0059] System 300 includes device 301, which may be a smartphone, computer, or similar device. Device 301 includes processor 302 connected to program memory 303 and data memory 304. In other examples, device 301 is implemented in a distributed computing architecture, such as a cloud computing provider. Processor 302 can receive data through various interfaces, including memory access to volatile memory (such as cache or RAM) or non-volatile memory (such as optical disc drives, hard disk drives, storage servers, or cloud storage). Program memory 303 is a non-transitory computer-readable medium, such as a hard disk, solid-state drive, or CD-ROM.
[0060] The software, specifically the executable program stored in program memory 303, causes processor 302 to execute methods for training a machine learning model on longitudinal medical scan images of a patient study. For example, once executed, the software causes processor 302 to compute two or more predictions of a biomedical factor by performing the following operations for each of two or more medical images: (i) applying the machine learning model to the medical image to compute an output; and (ii) using the output of the machine learning model for the medical image to compute a prediction, compute the rate of change between the two or more predictions for the two or more medical images, and create a trained machine learning model by updating the weights of the machine learning model to minimize a loss value.
[0061] Data storage 304 can store and retrieve image data for later use. The image data can indicate two-dimensional images, which can be stored on data storage 304, for example, in Joint Image Experts Group (JPEG) format, RAW image format, or similar / equivalent image formats. In some examples, the image data can be RGB images, multispectral images, hyperspectral images, infrared images, or two-dimensional cross-sections. Image data can indicate three-dimensional images, such as PET, MR, or CT images. For example, data storage 304 can store image data indicating three-dimensional image data as a single-file DICOM (Digital Imaging and Communications in Medicine) format, a multi-file DICOM format, or a similar / equivalent image format.
[0062] The machine learning models described herein (such as machine learning models before, during, and after training) can be stored on data storage 304. Data storage 304 can also store predictions computed by processor 302 performing any of the disclosed methods, or any other variables or data required to perform such methods. For example, when processor 302 performs the disclosed methods, predictions, machine learning model outputs, loss values, and rates of change are computed and can be stored on data storage 304.
[0063] It should be understood that processor 302 may determine or calculate the data to be received later before any receiving step. For example, processor 302 may then store the image data in data memory 304 (such as RAM or processor registers). Processor 302 may then request data from data memory 304, for example, by providing a read signal and a memory address. Data memory 304 provides data as voltage signals on physical bit lines, and processor 302 receives image data as input images via a memory interface.
[0064] System 300 may also include a scanner 305, a server 306, and a monitor 307, each communicating with device 301 via input / output (I / O) ports 308. In the example, scanner 305 scans the patient by incrementally capturing two-dimensional cross-sectional images of the patient. The set of cross-sectional images represents image data indicating the scan and representing a 3D scan of the patient. Once scanner 305 has created the patient's image data, processor 302 receives the image data via I / O port 308. Processor 302 then uses the image data received from scanner 305 to execute the disclosed methods and train a machine learning model.
[0065] If image data indicates an MRI scan, an MRI scanner can be used to create the image data. An MRI scanner consists of a stage that slides into a cylinder. Inside the cylinder are magnets that generate a strong magnetic field. This strong magnetic field is used to align protons in the water molecules within the patient's body, as protons have magnetic moments due to their inherent spin. Once the protons are aligned with the strong magnetic field, the MRI scanner emits a radio frequency signal, which excites the protons and aligns them with the magnetic field. Once the radio frequency signal is turned off, the protons relax and realign with the strong magnetic field, emitting electromagnetic radiation that is detected by the MRI scanner.
[0066] The patient lies on a motorized table that moves incrementally horizontally through a cylinder, capturing a single MR image at each increment. This image corresponds to a two-dimensional cross-sectional image (virtual "slice") of the patient's body. Each two-dimensional cross-sectional image corresponds to a different region of the patient's body as they move incrementally horizontally through the cylinder. The collection of two-dimensional cross-sectional images also has an associated spatial sequence in which the cross-sectional images are arranged. The horizontal direction of the patient's movement within the cylinder indicates the spatial dimension along the spatial sequence of two-dimensional cross-sectional images. The collection of two-dimensional cross-sectional images provides a three-dimensional representation of the patient's interior, which helps physicians diagnose internal diseases and conditions.
[0067] Two-dimensional cross-sectional images consist of black and white regions and gray shades. These regions correspond to the rate at which excited protons return to their equilibrium state after exposure to radiofrequency (equilibrium state corresponds to the spin state aligned with a strong external magnetic field), and the amount of energy released, which depends on the environment and the chemistry of the molecules. These factors can have different weights in MR images to provide more contrast for certain tissues in the image. For example, in a T1-weighted (T1w) image, high-signal areas representing white regions may correspond to protein-rich fluids, while in a T2-weighted (T2w) image, these regions may correspond to water-rich fluids in the patient's body.
[0068] If image data indicates a CT scan, a CT scanner can be used to create the image data. A CT scanner uses a rotating X-ray tube and a row of detectors placed in the machine gantry to measure the attenuation of X-rays in different tissues within the body. X-ray attenuation refers to the absorption of X-ray photons by tissues as they pass through the patient's body in the wavelength range of 10 picometers to 10 nanometers.
[0069] The patient lies on a motorized table that moves horizontally through a rotating X-ray tube in increments, capturing multiple X-ray measurements from different angles at each increment. These measurements are then processed on a computer using a reconstruction algorithm to produce tomographic two-dimensional cross-sectional images (virtual “slices”) of the body. Each two-dimensional cross-sectional image corresponds to a different region of the patient's body as they move horizontally through the rotating X-ray tube.
[0070] Similar to MR images, the collection of two-dimensional cross-sectional images also possesses an associated spatial sequence in which the cross-sectional images are ordered. The horizontal direction of the patient's movement within the cylinder indicates the spatial dimension along the spatial sequence of the two-dimensional cross-sectional images. Like MR images, the collection of two-dimensional cross-sectional images provides a three-dimensional representation of the patient's interior, which can aid physicians in diagnosing internal diseases and conditions.
[0071] Two-dimensional cross-sectional (X-ray) images consist of black and white areas and gray shades. These areas correspond to body regions that attenuate X-rays. For example, structures such as bone readily absorb X-rays and therefore produce high contrast on an X-ray detector. Consequently, bone structures appear whiter than other tissues in an X-ray image. Conversely, X-rays travel more easily through tissues with lower radiation density (such as fat and muscle) and through air-filled cavities (such as the lungs). These structures are shown as gray shades in X-ray images. Areas that experience little attenuation will appear as black areas in an X-ray image.
[0072] If the image data indicates a PET scan, a PET scanner can be used to create the image data. Before placing the patient in the PET scanner, a radiopharmaceutical (a radioactive isotope attached to a drug) is injected into the body as a radioactive tracer. When the radiopharmaceutical undergoes beta decay, it emits positrons, and when these positrons interact with ordinary electrons present in the body, the two particles annihilate and emit gamma rays. These gamma rays are detected by a gamma camera in a similar way to how X-ray images are captured, thus forming a three-dimensional image.
[0073] Some parts of the body absorb more of a particular radioactive tracer than others. For example, diseased cells (such as cancer cells) in a patient's body absorb more radioactive tracers than healthy cells. In another example, the brain is a frequent user of glucose, where the radioactive tracer could be glucose containing a radioactive isotope. Areas with higher gamma-ray detection (called "hot spots") will appear as warmer colors, such as yellow and red, on a PET image, while areas with lower gamma-ray detection will appear as blue. Therefore, the size and location of these "hot spots" make it possible to diagnose diseases using PET images.
[0074] Processor 302 may also receive additional longitudinal data from other sources. For example, system 300 may include an RGB camera (not shown) configured to capture RGB image data over a period of time. In another example, system 300 may include an inertial measurement unit (IMU) that measures the position and attitude data of the machine over a period of time. Furthermore, system 300 may include a set of sensors that provide measurement results to processor 302 over a period of time. Therefore, the disclosed systems and methods are not limited to applications involving medical imaging.
[0075] Before and after the processor 302 executes the disclosed method, images received from the scanner 305 may be stored on the data storage 304. In another example, the server 306 may store the image data. Additionally, the image data may be transmitted to the server 306, and the server 306 may then transmit the image data to the device 301. The image data may be received from a source outside the system 300, such as another system located remotely from the system 300. The remote system may be located within the same facility as the system 300 or may be completely remote from the system 300. The image data may be stored in the server 306 in single-file DICOM or multi-file DICOM format.
[0076] Server 306 can communicate with scanner 305, or via a communication network. Similarly, processor 302 receives image data via input / output port 308 and executes the disclosed methods on the image data. After processor 302 executes the disclosed methods on the image data, a trained machine learning model can be transmitted to server 306 by transmitting the parameters of the trained machine learning model. Server 306 can then transmit these parameters to an external system or store them.
[0077] System 300 may also include a monitor 307, which can be configured to display image data from scanner 305 and server 306. In this way, processor 302 transmits image data from scanner 305 and server 306 to the monitor via input / output port 308. Monitor 307 may be further configured to apply the results of a trained machine learning model of new medical images during inference. In this way, processor 302 transmits inference results to the monitor via I / O port 308.
[0078] Although Figure 3 Only one device 301 is depicted, but a network of many devices capable of communicating with server 306 can exist. The device network can also communicate directly with each other via the Internet, any other wireless communication method, or a wired connection. In this example, multiple devices can each capture image data of a patient. Each instance of image data from the multiple devices can then be stored on server 306 for later use. In some embodiments, a database of image data captured from the multiple devices can be stored on server 306.
[0079] It should be noted that server 306 may have similar functionality to apparatus 301. Server 306 may execute multiple parts of the disclosed methods described herein. For example, program memory may include software, i.e., an executable program stored in program memory, and enable the processor to train a machine learning model on longitudinal data.
[0080] The software can provide a user interface presented to the user on device 301. The user interface is configured to accept input from the user (via buttons or text fields, etc.) via a touchscreen or a device attached to device 301 (such as a keyboard or computer mouse). These devices may also include touchpads, externally connected touchscreens, joysticks, buttons, and dials. In this example, device 301 can display multiple instances of image data, and the user can select one of the multiple instances of image data (such as a medical image) via processor 302. The user can select one of the multiple instances of image data by interacting with the touchscreen or by inputting selection using a keyboard or computer mouse.
[0081] Processor 302 can receive or send data, such as image data, from data storage 304 and from I / O port 308. In one example, processor 302 may send image data from device 301 to server 306 via I / O port 308 using a Wi-Fi network according to IEEE 802.11. The Wi-Fi network can be a decentralized, self-organizing network, eliminating the need for dedicated management infrastructure such as routers, or a centralized network with routers or access points for management. System 300 can be further implemented within a cloud computing environment, such as a management group of interconnected servers hosting a dynamic number of virtual machines.
[0082] Although I / O port 308 is shown as a single entity, it should be understood that any type of data port can be used to receive data, such as network connections, memory interfaces, pins of the chip package of processor 302, or logic ports, such as IP sockets or parameters of functions stored in program memory 303 and executed by processor 302. Function parameters can be stored in data memory 304 and can be handled by value or by reference (i.e., as pointers in the source code).
[0083] Methods for training Figure 4 A method 400 for training a machine learning model on longitudinal medical scan images of a patient study is shown. Figure 4 It should be understood as a blueprint for a software program that can be implemented step by step, enabling... Figure 4 Each step in the process is represented by a function in a programming language such as Python, C++, or Java. The resulting source code is then compiled and stored as computer-executable instructions in program memory 303, which causes processor 302 to execute method 400. Before the process of training the machine learning model begins, processor 302 can initialize the weights of the machine learning model (which defines the machine learning model). For example, processor 302 can use random weights at the beginning of the training process and then update those random weights during the training process.
[0084] In some examples, the machine learning model can be a neural network. In further examples, the machine learning model can be a neural network that includes one or more convolutional layers. This type of machine learning is called a convolutional neural network (CNN). CNNs are ideal for applications involving images because they can depict the location and shape of objects in image data. However, other types of machine learning models are equally suitable for this purpose. For example, machine learning models can be implemented as K-nearest neighbor models, decision trees, and support vector machines.
[0085] Mathematically speaking, convolution is an integral function, representing a function... gIn another function f The amount of overlap during upward movement. Intuitively, convolution acts as a mixer, blending one function with another to reduce data space while preserving information. CNNs perform convolutions using filters (matrices / vectors), and these filters contain learnable parameters used to extract low-dimensional features from the input data. They have the property of preserving the spatial or positional relationships between input data points. CNNs leverage spatial-local correlations by enforcing local connectivity patterns between neurons in adjacent layers. In this example, a CNN architecture may include an optimal number of convolutional layers, filter size, and stride.
[0086] Intuitively, convolution is the process of applying the concept of a sliding window (a filter with learnable weights) to the input and producing a weighted sum (of the weights and the input) as the output. The weighted sum is the feature space used as input for the next layer. More specifically, each convolutional layer includes a filter that computes a weighted sum of the values of neighboring pixels. For example, a 2x2 filter contains four weights, which are the coefficients of the filter. The filter starts from an initial position in the image data structure, multiplies each pixel value in the image data structure by the corresponding filter coefficient, and sums the results. Finally, the filter stores the resulting number in the output pixel. In this sense, the output pixel value of each of one or more convolutional layers contains a weighted sum of the input pixel values. The weights in the weighted sum correspond to the coefficients of the filter. The filter then moves one pixel in one direction along the data structure and repeats the computation for the next voxel of the output image. For example, if the stride is 2, the filter will move two pixels in one direction. This direction can be either the x-dimensional or y-dimensional dimension.
[0087] A single convolutional layer can use multiple filters, each corresponding to a different feature. The corresponding feature for each filter is determined during CNN training. More specifically, each filter may contain weights, which are adjusted during training via backpropagation. Filters are used to quantitatively determine the contribution of a particular feature to the CNN output. For a convolutional layer appearing at the beginning of a CNN architecture, the filters of that convolutional layer might correspond to simple features. However, convolutional layers appearing later in the CNN architecture might exhibit more complex or abstract features. As an example, complex or abstract features could be a combination of simple features from previous convolutional layers.
[0088] CNN architectures can also include batch normalization layers and rectified linear unit (ReLU) activation functions. Furthermore, the output of the last convolutional layer can undergo global average pooling (GAP). Even further, fully connected layers with sigmoid activation can be used on the output layer. However, other activation functions can be used. These activation functions include, but are not limited to, the binary step function, the tanh function, the ReLU function, or the softmax function.
[0089] Figure 5 This is Example CNN Architecture 500. This Example CNN Architecture 500 is designed to predict parameters from MR images to dynamically transform the intensity contrast of the input MR image, thereby adapting it to a segmentation task. More specifically, the Example CNN predicts parameters for power function transformations and piecewise linear transformations. Figure 5 As can be seen from the diagram, the example CNN architecture 500 includes three convolutional blocks 511, 512, and 513, and three fully connected (FC) blocks 521 and 522 (523 and 524 together form one FC block). This CNN architecture 500 is used to provide the results of the disclosed methods presented later in this disclosure.
[0090] To initiate method 400, processor 302 may select 401 two or more images from longitudinal data of a patient study. Preferably, the two or more images are from the same patient, rather than from different patients. Two or more images may be randomly selected from the patient study. For example, the two or more images do not necessarily have to be consecutively captured images.
[0091] Then, processor 302 computes two or more predictions 420 for biomedical factors for each of two or more medical images in the medical imagery. The biomedical factors have expected trends, are known in advance, and are used to train a machine learning model. For each of the two or more medical images in the medical imagery, processor 302 applies the machine learning model to medical images 421, 426 to compute outputs 422, 427, and then uses the outputs of the machine learning model for the medical images to compute predictions 423, 428. Preferably, processor 302 uses the same machine learning model to compute the two or more predictions 420, and more specifically, processor 302 uses a model with the same architecture and weights to compute the two or more predictions 420.
[0092] In some embodiments, processor 302 calculates two or more predictions 420 for factors that are not biomedical. This factor can be any variable with some expected trend (e.g., expected to change over time). For example, this factor can be a physical factor, such as the volume of water in a dam. It should be noted that while this disclosure describes embodiments related to biomedical factors, these embodiments are equally applicable to factors other than biomedical factors.
[0093] It should also be noted that the computation of predictions for each of two or more medical images does not have to be performed sequentially; that is, it is not necessary to compute the first prediction before computing the second, and so on. Predictions for each of two or more medical images can be computed in parallel because the computations do not necessarily depend on each other. For example, each prediction can be computed in parallel across many CPUs and multiple cores. More specifically, parallel programming with shared memory across multiple platforms, such as OpenMP which utilizes multiple computational cores of a CPU, can efficiently compute each prediction on a single CPU.
[0094] In some embodiments, processor 302 computes predictions 423, 428 by applying a function to the output of a machine learning model. For example, the function may be a polynomial function, an exponential function, or a trigonometric function, or a combination thereof. In other embodiments, processor 302 computes predictions 423, 428 by multiplying the output by a factor or measurement result. For example, the output may be a scaling factor used to correct a particular measurement result, and thus the prediction corresponds to the corrected measurement result. In a further embodiment, the output may simply correspond to the prediction, which can be considered as processor 302 computed predictions 423, 428 by multiplying the output by 1.
[0095] Biomedical factors can be one of the following: medical image quantification factors, biological structure values, or biomarkers. For example, a biomedical factor can be a medical image quantification factor, which could be a Centiloid value or a normalized uptake-to-scaling factor. In other examples, a biomedical factor can be a biological structure value, which could be the volume of a biological structure visible in a medical image, such as gray matter in the brain. For example, a biomarker could be blood pressure or heart rate.
[0096] After calculating two or more predictions 420, processor 302 calculates the rate of change 430 between the two or more predictions for two or more medical images. In some embodiments, processor 302 calculates the rate of change 430 based on the difference between the two or more predictions. In other embodiments, processor 302 calculates the rate of change 430 based on the difference between the two or more predictions and the difference between the variables on which the rate of change is based. For example, if only two predictions exist, processor 302 can calculate the rate of change by dividing the difference between the predictions by the difference between the time points of each corresponding medical image.
[0097] The calculated rate of change indicates the trend of longitudinal data. Thus, processor 302 can explicitly calculate the rate of change by calculating the ratio of the predicted change to the change in the independent variable (such as time). However, processor 302 can also implicitly calculate the rate of change by calculating the difference between two or more predictions, because medical images can inherently contain information about change (such as temporal information). For example, two medical images in a patient study may have been taken a year apart, and therefore, given their temporal correlation, these two medical images possess inherent temporal information. Thus, since medical images inherently contain temporal information, the rate of change can be calculated from the difference between two predictions for the corresponding medical image.
[0098] After calculating the rate of change 430, processor 302 calculates a loss value 440 to train the machine learning model. Specifically, the loss value contains a penalty value that reflects the difference between the rate of change and the expected rate of change. Reflecting the difference typically means that the penalty value indicates the difference. Reflecting the difference can mean that the penalty value is correlated with the difference, such as being proportional to the difference. Reflecting the difference can also mean that processor 302 uses the difference to calculate the penalty value, such as by applying a mathematical function to the difference. By using this penalty value to calculate the loss value, the training of the machine learning model utilizes information not available during inference.
[0099] In some examples, the expected rate of change can be a function of time (such as linear, polynomial, exponential, etc.) or it can be a second-order rate (i.e., a rate of change), and processor 302 calculates the second-order rate of change. In other examples, the expected rate of change may be zero, indicating that the biomedical factor remains constant over time. Therefore, a penalty value can be specified in a way that penalizes the rate of change that deviates from zero.
[0100] Calculating the loss value can be referred to as applying a loss function. In other words, the result of applying the loss function is the loss value. The loss value and loss function can also be called the cost value and cost function, respectively. The loss function can be applied to the output and other factors, such as the timestamps of two or more medical images. The loss value can incorporate many types of additional information that are not available at inference time. For example, difference curves based on clinical diagnosis, age, or any other relevant information can be used to train a machine learning model using the loss value.
[0101] Then, processor 302 updates the weights 450 of the machine learning model to minimize the loss value, thereby creating a trained machine learning model. This process is commonly referred to as "training." Longitudinal data (i.e., training data) contains labeled data, such as image data already labeled with categories, where the label for each item in the training data can be manually provided by the user. Training data can also be obtained, or otherwise obtained, from a database containing labeled data, thus eliminating the need for manual labeling. Training a machine learning model in this way is called "supervised learning," which involves validating the results of the machine learning model and adjusting the weights accordingly. Weights can be updated via a backpropagation process. Weights can also be updated via a gradient descent process.
[0102] It should be noted that "supervised learning" is not the only method that can be used to train machine learning models. For example, "semi-supervised learning" can be used when only a relatively small amount of training data is available. In some examples, the training set can be augmented to generate more training data, which is advantageous when large amounts of training data are difficult to obtain. For example, the training set can be augmented by randomly resizing, cropping, flipping, or rotating the image data in the training set.
[0103] It should be noted that during Method 400, the rate of change (e.g., the rate of change between two or more predictions) can be used only during training to update the machine learning model, rather than during inference, because the predictions provided by the machine learning model can correspond to a single point-in-time measurement. Therefore, the rate of change may not be used at inference to obtain predictions from the trained machine learning model. Instead, a single point-in-time measurement can be used at inference to generate the corresponding prediction using the trained machine learning model.
[0104] Repeat the training method Preferably, method 400 is repeated more than 100 times using two or more different images. Generally, the more times method 400 is repeated, the higher the accuracy of the machine learning model when making predictions during inference. Furthermore, training the machine learning model on a series of different images or training data usually also makes the machine learning model more accurate when making predictions during inference. In this disclosure, references to "trained machine learning model" mean that the initial machine learning model has undergone at least one iteration of method 400. However, "trained machine learning model" may also mean that the initial machine learning model has undergone multiple iterations of method 400. Other terms such as "further trained machine learning model," "another trained machine learning model," and "fully trained machine learning model" may also be used throughout this disclosure to indicate that the machine learning model has undergone multiple training iterations.
[0105] Since processor 302 can execute method 400 multiple times using two or more different images, processor 302 can compute two or more additional predictions of the biomedical factor by performing the following operations for each of two or more additional medical images from the patient study: (i) applying a trained machine learning model to the medical image to compute an output; and (ii) using the output to compute a prediction. Processor 302 can then compute the rate of change between the two or more additional predictions; and create an additional trained machine learning model by updating the weights of the trained machine learning model based on the additional rate of change to minimize an additional loss value. Thus, the machine learning model is trained on different images from the same patient study. This process can be repeated using each permutation of two or more images from the patient study.
[0106] In some embodiments, processor 302 may further use images from another patient study, different from the patient study used to initially train the machine learning model, to train the machine learning model. More specifically, processor 302 may compute two or more additional predictions for a biomedical factor by performing the following operations on each of two or more medical images from another patient study different from the patient study: (i) applying the trained machine learning model or an additional trained machine learning model to the medical images to compute an output; and (ii) using the output to compute a prediction. Processor 302 may then compute another rate of change between the two or more additional predictions; and create another trained machine learning model by updating the weights of the trained machine learning model or the additional trained machine learning model based on the other rate of change to minimize another loss value.
[0107] Machine learning models can be independent of the order of two or more medical images studied by a patient. Therefore, machine learning models cannot learn the order because each image is processed in the same way; the only difference lies in how the loss is calculated. It is beneficial for machine learning models not to learn the order of input images because it is desirable for the model to learn the expected trend (i.e., the expected rate of change) in the data, rather than the sequence of images. This provides better predictions during inference.
[0108] To further reinforce this, the order of the inputs (i.e., two or more medical images) can be randomly switched. Thus, in some embodiments, each of the medical images in a patient study has a timestamp, and at least two of the two or more medical images have non-adjacent timestamps. Non-adjacent timestamps are discontinuous. For example, if a patient study contains longitudinal data of brain CT images captured each year, the non-adjacent timestamps would be for discontinuous years.
[0109] Calculate the expected rate of change Since the penalty reflects the difference between the rate of change and the expected rate of change, processor 302 can calculate the expected rate of change. Processor 302 can calculate the expected rate of change based on longitudinal data. For example, processor 302 can calculate the rate of change for each pair of images in a patient study. Processor 302 can then average the calculated rates of change to calculate the expected rate of change. Processor 302 can also use other techniques such as regression to calculate the expected rate of change based on longitudinal data.
[0110] In some embodiments, processor 302 determines a polynomial function of the rate of change of a biomedical factor and uses this polynomial function to calculate the expected rate of change. This polynomial function represents the expected trend of longitudinal data and can therefore be used to calculate the expected rate of change of a biomedical factor based on two or more medical images used to train a machine learning model. For example, and referring to… Figure 2 Curve 202 represents a polynomial function and is based on longitudinal data used to train the machine learning model. It should be noted that this function is not limited to a polynomial function. While a polynomial function is preferred, it can be a polynomial function, an exponential function, a trigonometric function, or a combination thereof.
[0111] In some embodiments, the processor 302 uses a polynomial function to calculate the expected rate of change by applying the polynomial function to the average of a biomedical factor calculated from the average of two or more predictions. Specifically, the average can be calculated by dividing the difference between two or more predictions by the number of the two or more predictions. The polynomial function can be a function of the rate of change with respect to the average, and therefore a function of the average, which, when applied to the average, produces a rate of change corresponding to that average.
[0112] In some embodiments, processor 302 determines the polynomial function by calculating the rate of change and mean of each pair of medical images for a patient study. Furthermore, processor 302 can determine the polynomial function by calculating the rate of change and mean of each pair of medical images that form the training data for each patient study. Processor 302 can determine the polynomial function by determining the best representation of the training data. For example, processor 302 can use regression (e.g., best-fit curve) or other techniques, such as applying a machine learning model.
[0113] Additional penalty value In some embodiments, the loss value further includes a change penalty value reflecting the difference between two or more predictions, wherein the change penalty value penalizes unintended time trends of biomedical factors associated with the expected rate of change. For example, the expected time trend of longitudinal data could be a positive rate of change over time. Therefore, the change penalty value would penalize cases where a negative rate of change occurs. In some examples, the change penalty value may include a clamping function to retain only positive or negative changes.
[0114] In some embodiments, the loss value further includes a correction penalty value reflecting the difference between two or more predictions and two or more corresponding ground truth values. The correction penalty value can penalize deviations of two or more predictions from two or more corresponding ground truth values. For example, the correction penalty value can include two penalty values reflecting the slope and intercept between the two or more predictions and the ground truth measurements. Thus, the correction penalty value can be based on the difference between the rate of change of the two or more predictions and the rate of change of the two or more corresponding ground truth values. In other examples, the slope and intercept can be based on a regression line between the ground truth measurements and the two or more predictions.
[0115] In some examples, it can be expected that the slope between the predicted output and the actual ground measurement should not deviate too far from 1. In other examples, it can be expected that the intercept between the predicted output and the actual ground measurement should not deviate too far from 0. Therefore, during the training of the machine learning model, two correction penalty values reflecting this expectation can be used to penalize overcorrection. This ensures that no bias is introduced during the training process. In some examples, the correction penalty values can be the sum of the differences between each output and the overcorrection factor. In further examples, the correction penalty values can be the sum of the absolute values of these differences.
[0116] For example, a regression line can be computed by determining the slope and intercept based on each prediction and its corresponding ground truth. More specifically, a graph can be plotted using data points corresponding to the predictions and their corresponding ground truths, and then a linear regression of that graph can be determined. Essentially, a linear function can be computed where the input to the linear function corresponds to the prediction and the output (i.e., the value of the function evaluated at the input) corresponds to the ground truth. Therefore, the slope can be given as:
[0117] Since the goal is to predict values similar to the ground truth, it's logical that the slope of the linear function should be around 1, and the intercept around 0. Therefore, the correction penalty can be based on the difference between the slope and 1, and the intercept and 0, minimizing this difference during training. For example, the correction penalty value can be given by:
[0118] However, it should be noted that other changes to the correction penalty value are also possible.
[0119] In some embodiments, the loss value further includes a correction penalty value reflecting the differences between two or more outputs and the overcorrection factor, wherein the correction penalty value penalizes two or more outputs for deviating from the overcorrection factor. In some examples, it can be known that the outputs of the machine learning model should not deviate too far from the overcorrection factor. Therefore, a correction penalty value can be used during the training of the machine learning model to penalize overcorrection. In some examples, the correction penalty value can be the sum of the differences between each output and the overcorrection factor. In further examples, the correction penalty value can be the sum of the absolute values of these differences.
[0120] In some embodiments, the loss value further includes an additional penalty value calculated by one or more of the following: (i) the difference between each of two or more outputs of the machine learning model and the corresponding ground truth measurement, and (ii) the mean squared error between each of two or more outputs and the corresponding ground truth measurement. Therefore, the loss value may include an additional (standard) loss term. Essentially, the loss value may include any number of additional loss terms.
[0121] In some embodiments, the processor 302 calculates the loss value by a weighted sum. For example, the loss value may include a penalty value, a change penalty value, a correction penalty value, and an additional penalty value. Thus, the processor 302 calculates the loss value by a weighted sum of the penalty value, the change penalty value, the correction penalty value, and the additional penalty value. The weights of the weighted sum can be considered as a “learning rate,” which is predetermined before training occurs, such as through user input. In some examples, one or more of the weights may be 1.
[0122] Help with diagnosis It should be noted that biomedical factors can then be used to help determine a patient's diagnosis. For example, clinicians can apply a trained machine learning model to images of a new patient. Therefore, method 400 can further include applying a trained machine learning model to medical images of a patient to aid in the diagnosis based on the output of the trained machine learning model.
[0123] In one example, a trained machine learning model can calculate a metric, such as the amount of amyloid plaques on the brain. The method can then apply one or more thresholds to that calculated metric to determine whether a patient has Alzheimer's pathology or a low / moderate / high risk of developing Alzheimer's disease in the future. This use of trained machine learning models could also be applied to other medical indications, including progressive diseases or general phenotypes.
[0124] The diagnosis and research of Alzheimer's disease often utilizes PET quantization of patient brain PET images. PET quantization can provide quantitative information, such as, for example, the amount of amyloid plaques in the brain. PET quantization can also involve processor 302 creating a three-dimensional rendering of the patient's brain based on the PET images. Thus, in some embodiments, each medical image is a PET image and method 400 includes performing PET quantization by applying a trained machine learning model to the patient's PET images.
[0125] Therefore, to aid in the diagnosis of Alzheimer's disease, the biomedical factor will be a medical image quantization factor, and more specifically, the medical image quantization factor will be a Centiloid or a normalized uptake ratio scaling factor. Furthermore, to aid in the diagnosis of Alzheimer's disease, method 400 further includes applying the trained machine learning model to the patient's medical images to aid in the diagnosis of Alzheimer's disease based on the Centiloid or normalized uptake ratio scaling factor as the output of the trained machine learning model.
[0126] Exemplary embodiments Exemplary embodiments of the disclosed methods will now be presented for illustrative purposes, but are not intended to limit the disclosed systems and methods to these embodiments. These exemplary embodiments are further illustrated by experimental results obtained from training machine learning models using the disclosed methods.
[0127] In this exemplary embodiment, the biomedical factor is a medical image quantification factor, and more specifically, the biomedical factor is a Centiloid value. The output of the machine learning model being trained is a scaling factor for the SUVR value, which is then converted to a Centiloid value, and the prediction is a scaled Centiloid value. Therefore, the medical images used during training and inference will be PET images, since the SUVR and Centiloid scaling are associated with PET images.
[0128] For PET quantization, amyloid PET images are standardized using centiloid values by normalizing the PET SUVR of different amyloid PET tracers to a standard and universal unit system, where each PET tracer has its own linear transformation to normalize the SUVR to centiloid. However, centiloid values are often noisy due to noise in PET images and the use of different PET tracers / scanners that may have varying degrees of noise and bias. Therefore, PET quantization of PET images relying on the centiloid scale may be inaccurate due to noise in PET images and variability introduced by using different tracers and scanners. Therefore, machine learning techniques can be used to determine corrections to the calculated SUVR values before converting them to centiloid values, in an effort to reduce the impact of noise in PET images and increase the accuracy of PET quantization.
[0129] In this exemplary embodiment, the machine learning model is trained on two medical images at a time (i.e., two images from the same patient study for each training iteration). Thus, the processor 302 computes two predictions of the Centiloid value (i.e., the scaled Centiloid value) by applying the machine learning model individually to each of the two medical images, computing the corresponding output, and using the corresponding output to compute the corresponding prediction.
[0130] In this embodiment, processor 302 calculates the prediction by first calculating the SUVR value for each of the two medical images. Processor 302 can calculate the SUVR value by calculating the ratio between a reference region of the corresponding PET image and a “hotspot” region (such as amyloid) of the corresponding PET scan, and then convert them to Centiloid. Processor 302 can also use tracer and / or scanner information as part of the prediction, allowing for different corrections for different tracers, which allows for better control over their varying noise levels. Then, before converting the scaled SUVR to Centiloid (i.e., scaled Centiloid values), processor 302 calculates the prediction by multiplying the calculated SUVR value by the output of a machine learning model (i.e., a scaling factor).
[0131] Then, processor 302 calculates the rate of change between the two predictions by dividing the difference between the two predictions (i.e., the scaled Centiloid value) by the difference between the timestamps of the two medical images, given by the following formula:
[0132] The processor 302 then calculates the loss value. It should be noted that in some embodiments, the processor 302 calculates the loss value by applying a loss function. For example, in an exemplary embodiment, the processor 302 applies the loss function to the calculated Centiloid value, the output of the machine learning model, and the timestamps of the two medical images. In other words, the calculated Centiloid value, the output of the machine learning model, and the timestamps of the two medical images are all input into the loss function. Thus, the processor 302 can calculate the rate of change by applying the loss function to the input.
[0133] In this exemplary embodiment, the loss value includes three penalty values: (i) a penalty value indicating the rate of change, (ii) a penalty value penalizing changes in the Centiloid value due to unexpected time trends, and (iii) a correction penalty value penalizing two or more outputs deviating from the overcorrection factor. In this embodiment, the processor 302 uses a weighted sum of the three penalty values to calculate the loss value. For experimental results, the loss value is calculated using the following formula:
[0134] Where α and β are predetermined parameters or learning rates.
[0135] Penalty value for indicating rate of change To calculate the penalty value, processor 302 calculates the rate of change and mean for each pair of medical images for each patient study. Then, processor 302 determines a polynomial function by applying a regression algorithm to the calculated rate of change. The calculated rate of change and the determined polynomial function in this exemplary embodiment are as follows: Figure 2 As shown. Essentially, the expected rate of change does not deviate from the polynomial. Therefore, the penalty value indicating the rate of change penalizes large distances from the polynomial.
[0136] Processor 302 uses a polynomial function to calculate the expected rate of change by applying the polynomial function to the average of the Centiloid values calculated from the averages of two or more predictions. More specifically, the average is given by:
[0137] Then, processor 302 calculates the penalty value by calculating the difference between the calculated rate of change and the expected rate of change. More specifically, processor 302 calculates the absolute value of the difference between the calculated rate of change and the expected rate of change. Therefore, the penalty value is given by the following formula:
[0138] in It is a polynomial function.
[0139] Change penalty value The change penalty value penalizes any unexpected time trend in the Centiloid value related to the expected rate of change. In an exemplary embodiment, the Centiloid value is expected to increase over time, and therefore, the rate of change of the Centiloid value is expected to be positive (greater than zero). Therefore, the change penalty value is formulated to penalize decreasing Centiloid values and negative rates of change. In an exemplary embodiment, the change penalty value is based on a "baseline / subsequent" concept, where a "baseline" medical image scan will be performed on the patient, which will be compared to future "subsequent" medical image scans.
[0140] The change penalty value is given by the following formula:
[0141] Here, "clamp" is the clamping function, which restricts a given value between an upper and lower bound. More specifically, the clamping function is defined by the following formula:
[0142] The clamping function restricts the predicted difference to be positive only, thereby forcing the expansion trend of the rate of change to always be positive during training.
[0143] Correction penalty value The correction penalty value penalizes two or more outputs of the machine learning model for deviations from the overcorrection factor. In an exemplary embodiment, the overcorrection factor is 1 because the output of the machine learning model is a scaling factor, and therefore, ideally, the output should be 1 in the absence of noise in the PET image.
[0144] The correction penalty value is given by the following formula:
[0145] result The results of training the machine learning model are as follows Figure 6 As shown. To produce these results, the machine learning model was trained using 5,000 PET images from 1,500 patients. The machine learning model used to generate these results is... Figure 5 The CNN architecture shown should be noted that the input to the machine learning model is PET images, not MR images. Figure 6 This shows the progression of the calculated rate of change of the scaled Centiloid value as the machine learning model's training increases. Specifically, it can be shown that the calculated rate of change shifts towards a polynomial function as training continues. Furthermore, the calculated rate of change becomes more positive than the expected rate of change of the Centiloid value.
[0146] Generate masks for machine learning interpretability Machine learning models (especially deep machine learning models) are often known as black boxes, meaning we don't know how they compute predictions based on the input. For example, while a model might be able to detect the presence of a disease from image data, it may not be clear how the model makes the prediction, or which features the model considers important when making predictions. The black-box nature of machine learning models means that interpretability can be challenging. Understanding the decisions made by black-box algorithms is a challenge, and assessing their fairness and impartiality is an important step in deploying them across many industries.
[0147] As previously discussed, the trained machine learning model disclosed herein can output predictions of biomedical factors, which can be a standardized uptake ratio (SUVR) scaling factor (which may be simply referred to as SUVR). Recall that SUVR can be defined as the ratio of regions in a PET image containing specific binding (the target region where the PET tracer will bind to the target) to regions in a PET image containing non-specific binding (referred to as reference regions). Thus, a target mask highlighting the target region in the PET image can be generated. Similarly, a reference mask highlighting non-specific binding regions in the PET image can be generated. In the case of centiloids in the PET image, the reference mask can be referred to as a centiloid reference mask, etc. Therefore, generating a mask corresponding to the correction provided by the trained machine learning model can provide insight into which regions of the PET image the trained machine learning uses in determining its corrected SUVR.
[0148] The corrections provided by the trained machine learning model disclosed herein are thought to reflect variations in a reference mask, a target mask, or both. This idea can be used to generate masks with SUVR corresponding to the corrections provided by the trained machine learning model. In other words, a new mask can be optimized for each tracer such that its correlation with the SUVR value provided by the trained machine learning model is optimized (e.g., maximized). This can provide insight into features that the trained machine learning model deems important, for example, features that may not be apparent in the original PET image.
[0149] Thus, in some embodiments, processor 302 can generate one or more masks corresponding to corrections determined by a trained machine learning model. It should be noted that in this context, "mask" can refer to an image that highlights relevant regions used for decision-making by the machine learning model, such as a heatmap, saliency map, etc. Processor 302 can generate one or more masks by adjusting one or more medical images (which can be input into a trained machine learning model) based on the corresponding outputs of the trained machine learning model. For example, processor 302 can apply algorithms, mathematical operations, or another machine learning model to one or more medical images based on the corresponding outputs to generate one or more masks.
[0150] Obtaining these new masks can be formulated as an optimization problem, where, for each tracer, a new reference and target mask can be optimized across the entire dataset such that the resulting SUVR maximizes the Pearson correlation coefficient using a corrected SUVR provided by a trained machine learning model. The Pearson correlation coefficient is a correlation coefficient that measures the linear correlation between two sets of data.
[0151] This optimization problem can be solved using gradient descent or another type of optimization technique. The loss function may involve the following two types of losses: 1. Penalize low Pearson correlation coefficients. New SUVRs can be computed across the entire dataset using an updated mask. The correlation between these SUVRs and the corrected SUVRs obtained from the trained machine learning model is used to calculate the Pearson coefficients. The resulting loss is defined as:
[0152] 2. Facilitate binary masks. The two masks being optimized need to have continuous values to ensure they can be derived throughout the optimization process. To force the masks to tend towards binary, for each mask... M Values that are not 1 or 0 will be penalized as follows:
[0153] Experiments and Results An experiment was conducted to generate reference and target masks for the interpretability of machine learning models. Experimental details are now provided. However, it is worth noting that substitutions and modifications are possible during the generation of these masks.
[0154] The reference and target masks are initialized using the original Centiloid reference and target masks. The masks are then smoothed using a Gaussian kernel (specifically a 4mm FWHM Gaussian kernel). At each iteration, the mask is updated with gradient information and smoothed again using the same Gaussian kernel to reduce sparsity. The mask is also mirrored to generate a symmetric mask. In this experiment, optimizations utilize PyTorch's autograd engine to automatically compute gradients.
[0155] Optimization stops once the Pearson loss no longer improves within 10 epochs or reaches the maximum number of epochs (2000). After obtaining the final masks, the new SUVR derived from these masks is converted to a corrected SUVR using a linear transformation, and then converted to Centiloid using a standard transformation.
[0156] Figure 7 The results of this optimization process are shown. Specifically, Figure 7 This shows the progress of the SUVR derived from the optimization mask as the optimization process continues (i.e., after multiple iterations). Figure 7As can be seen, as the optimization process continues, the ratio of the corrected SUVR provided by the trained machine learning model (labeled as DeepSUVR SUVR on the x-axis) to the SUVR determined based on the optimized mask approaches 1 (as shown by the data points near the diagonal). This indicates that through the above optimization process, the reference mask and the target mask are indeed optimized such that the SUVR determined from these masks is correlated with the corresponding SUVR provided by the trained machine learning model.
[0157] Figure 8a The initialized target mask is displayed, while Figure 8b The optimized target mask is displayed. Similarly, Figure 9a The initialized reference mask is shown, while Figure 9b The optimized reference mask is shown. It is noteworthy that, for both the reference and target masks, the optimized mask still includes salient features that can be seen in the initial mask. However, the optimization process provides more salient features and a better definition of these features. These results provide a better indication of the features that the machine learning model trained by the paper considers important when making predictions. Therefore, these results contribute to the interpretability of the trained machine learning model. These optimized masks may also have the potential to replace the machine learning model in directly generating the corrected SUVR.
[0158] Those skilled in the art will understand that many variations and / or modifications can be made to the above embodiments without departing from the broad general scope of this disclosure. Therefore, these embodiments should be considered illustrative rather than restrictive in all respects.
Claims
1. A method for training a machine learning model on longitudinal medical scan images from a patient study, the method comprising: Two or more predictions of biomedical factors are calculated by performing the following operations on each of two or more medical images: The machine learning model is applied to the medical image to compute the output; and The prediction is calculated using the output of the machine learning model applied to the medical image; Calculate the rate of change between the two or more predictions for the two or more medical images; as well as A trained machine learning model is created by updating the weights of the machine learning model to minimize the loss value, wherein the loss value includes a penalty value that reflects the difference between the rate of change and the expected rate of change.
2. The method of claim 1, wherein the method further comprises determining a polynomial function of the rate of change of the biomedical factor and using the polynomial function to calculate the expected rate of change.
3. The method of claim 2, wherein calculating the expected rate of change using the polynomial function comprises applying the polynomial function to the average of the biomedical factors calculated from the averages of the two or more predictions.
4. The method of claim 2 or 3, wherein determining the polynomial function comprises calculating the rate of change and average value for each pair of medical images in the patient study.
5. The method according to any one of the preceding claims, wherein the loss value further comprises a change penalty value reflecting the difference between the two or more predictions, the change penalty value penalizing an unexpected time trend of the biomedical factor associated with the expected rate of change.
6. The method according to any one of the preceding claims, wherein the loss value further comprises a correction penalty value reflecting the difference between the two or more predictions and two or more corresponding ground truth values, the correction penalty value penalizing the two or more predictions for deviating from the two or more corresponding ground truth values.
7. The method according to any one of the preceding claims, wherein the loss value further comprises an additional penalty value calculated by one or more of the following: The difference between each of the two or more outputs of the machine learning model and the corresponding ground-based measurements; and The mean square error between each of the two or more outputs and the corresponding ground-based actual measurement result.
8. The method of claim 7, wherein the loss value is calculated by a weighted sum of the penalty value, the change penalty value, the correction penalty value, and the additional penalty value.
9. The method according to any one of the preceding claims, wherein each of the medical images studied by the patient has a timestamp, and at least two of the two or more medical images have non-adjacent timestamps.
10. The method according to any one of the preceding claims, wherein the method further comprises: Two or more additional predictions of the biomedical factor are calculated by performing the following operations on each of two or more additional medical images from the study of the patient: The trained machine learning model is applied to the medical image to compute an output; and Use the output to calculate the prediction; Calculate the additional rate of change between the other two or more predictions; as well as An additional trained machine learning model is created by updating the weights of the trained machine learning model to minimize an additional loss value based on the additional rate of change.
11. The method according to any one of the preceding claims, wherein the method further comprises: Two or more predictions of the biomedical factor are calculated by performing the following operations on each of two or more medical images from another patient study different from the stated patient study: The trained machine learning model or the additional trained machine learning model is applied to the medical image to compute an output; and Use the output to calculate the prediction; Calculate another rate of change between the other two or more predictions; as well as Another trained machine learning model is created by updating the weights of the trained machine learning model or the other trained machine learning model to minimize another loss value based on the other rate of change.
12. The method according to any one of the preceding claims, wherein the method further comprises applying the trained machine learning model to a patient's medical images to aid in the diagnosis of the patient based on the output of the trained machine learning model.
13. The method according to any one of the preceding claims, wherein each of the medical images is one of the following: Magnetic resonance imaging; Positron emission tomography (PET) images; or Computed tomography (CT) images.
14. The method according to any one of the preceding claims, wherein the biomedical factor is one of the following: Medical image quantification factor; Biological structure value; or Biomarkers.
15. The method of claim 14, wherein the medical image quantization factor is a Centiloid or a normalized uptake ratio scaling factor, and the biological structure value is the volume of a biological structure visible in the medical image.
16. The method of any one of claims 11 to 15, wherein each of the medical images is a PET image, and the method further comprises performing PET quantization by applying the trained machine learning model to the PET images of the patient.
17. The method according to claim 15 or 16, wherein The medical image quantization factor is a Centiloid or a normalized uptake ratio scaling factor; and The method further includes applying the trained machine learning model to a patient's medical images to aid in the diagnosis of Alzheimer's disease based on the Centiloid or standardized uptake ratio scaling factor, which is the output of the trained machine learning model.
18. The method according to any one of the preceding claims, further comprising generating one or more masks corresponding to the corrections determined by the trained machine learning model.
19. Software that, when installed on and executed by a computer, causes the computer to perform the method according to any one of the preceding claims.
20. A system for training a machine learning model on longitudinal medical scan images from a patient study, the system comprising: Processor, the processor being configured to: Two or more predictions of biomedical factors are calculated by performing the following operations for each of two or more medical images in the aforementioned medical images: The machine learning model is applied to the medical image to compute the output; and The prediction is calculated using the output of the machine learning model applied to the medical image; Calculate the rate of change between the two or more predictions for the two or more medical images; as well as A trained machine learning model is created by updating the weights of the machine learning model to minimize the loss value, wherein the loss value includes a penalty value that reflects the difference between the rate of change and the expected rate of change.