3D Object Segmentation in Medical Images via Object Detection and Localization

By adopting deep learning networks in medical images combined with derivative contrast mechanisms, using multiple medical images with different characteristics for object detection and segmentation, the challenge of object segmentation in low-contrast and low-resolution images is solved, and more accurate and fine-grained object segmentation is achieved.

CN114503159BActive Publication Date: 2025-05-13F HOFFMANN LA ROCHE & CO AG
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202080057028.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-08-14
Filing Date
2020-08-13
Publication Date
2025-05-13
Estimated Expiration
2040-08-13

AI Technical Summary

Technical Problem

In medical images, especially in low-contrast and low-resolution images, there are challenges in object segmentation, including blurred object definitions, excessive noise, loss of semantic information after repeated pooling operations in deep learning models, difficulty in learning boundary knowledge, and background effects.

Method used

Deep learning network combined with a derivative contrast mechanism is used to detect and segment objects by using multiple medical images of different features. The specific method includes: positioning an object in a first medical image with the first feature, projecting onto the second medical image with the second feature using a bounding box or segmentation mask of the first image, and then inputting a portion of the second image to a three-dimensional neural network model for volume segmentation.

Benefits of technology

By reducing background effects and reducing input data complexity, deep learning models focus on object edges, improving the accuracy and fine-grainedness of object segmentation, and improving segmentation performance in low-contrast and low-resolution images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114503159B_ABST
    Figure CN114503159B_ABST
Patent Text Reader

Abstract

This disclosure relates to techniques for segmenting objects within medical images using deep learning networks, which locate objects through object detection based on a derived contrast mechanism. Specifically, aspects of the invention relate to locating a target object within a first medical image having a first feature, projecting a bounding box or segmentation mask of the target object onto a second medical image having a second feature to define a portion of the second medical image, and inputting said portion of the second medical image into a deep learning model configured to use a detector capable of segmenting said portion of the second medical image and generating a weighted loss function around the segmentation boundary of the target object. The segmentation boundary can be used to calculate the volume of the target object for determining the subject's diagnosis and / or prognosis.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to and the benefit of U.S. Provisional Application No. 62 / 886,844, filed on August 14, 2019, the contents of which are incorporated herein by reference in their entirety for all purposes. Technical Field

[0003] The present disclosure relates to automated object segmentation of medical images, and in particular to techniques for segmenting objects within medical images using a deep learning network that localizes via object detection based on a derived contrast mechanism. Background Art

[0004] Computer vision involves using digital images and videos to infer some understanding of what is in those images and videos. Object recognition is associated with computer vision and refers to a set of related computer vision tasks that involve identifying objects present in an image frame. These tasks include image classification, object localization, object detection, and object segmentation. Image classification involves predicting the category of one or more objects in an image frame. Object localization refers to identifying the location of one or more objects in an image frame and drawing a bounding box around its extent. Object detection combines these two tasks, locating and classifying one or more objects in an image frame. Object segmentation involves highlighting specific pixels (generating a mask) that highlight the located or detected object instead of a coarse bounding box. Object recognition techniques are generally divided into machine learning-based methods or deep learning-based methods. For machine learning-based object localization and detection methods, features within the image are initially defined using feature descriptors such as Haar-like features, scale-invariant feature transforms, or histograms of oriented gradients (HOG), and then the object of interest is detected based on this feature descriptor using techniques such as support vector machines (SVM). On the other hand, deep learning techniques are able to perform end-to-end object detection and segmentation in the absence of well-defined features and are typically based on convolutional neural networks (CNNs) such as region-based networks (R-CNN, Fast R-CNN, Faster R-CNN, and Cascade R-CNN). Summary of the invention

[0005] In some embodiments, a computer-implemented method for segmenting an object within a medical image is provided. The method includes: obtaining a medical image of a subject, the medical image including a first image having a first feature and a second image having a second feature, wherein the medical image is generated using one or more medical imaging modalities; using a localization model to locate and classify the object within the first image into a plurality of object categories, wherein the classification assigns a set of pixels or voxels of the first image to one or more of the plurality of object categories; using the localization model, determining a bounding box or segmentation mask for a target object within the first image based on a set of pixels or voxels assigned to an object category in the plurality of object categories; transferring the bounding box or segmentation mask to a second image to define a portion of the second image including the target object; inputting a portion of the second image into a three-dimensional neural network model constructed for volume segmentation using a weighted loss function; using the three-dimensional neural network model to generate an estimated segmentation boundary around the target object; and using the three-dimensional neural network to output a portion of the second image having an estimated segmentation boundary around the target object.

[0006] In some embodiments, the one or more medical imaging modalities include a first medical imaging modality and a second medical imaging modality different from the first medical imaging modality, and wherein the first image is generated by the first medical imaging modality and the second image is generated by the second medical imaging modality.

[0007] In some embodiments, the one or more medical imaging modalities include a first medical imaging modality and a second medical imaging modality that is identical to the first medical imaging modality, and wherein the first image is generated by the first medical imaging modality and the second image is generated by the second medical imaging modality.

[0008] In some embodiments, the first image is an image of a first type and the second image is an image of a second type, and wherein the image of the first type is different from the image of the second type.

[0009] In some embodiments, the first image is an image of a first type and the second image is an image of a second type, and wherein the image of the first type is the same as the image of the second type.

[0010] In some embodiments, the first characteristic is different from the second characteristic.

[0011] In some embodiments, the first feature is the same as the second feature.

[0012] In some embodiments, the first medical imaging modality is magnetic resonance imaging, diffusion tensor imaging, computed tomography images, positron emission tomography images, photoacoustic tomography images, X-ray images, ultrasound scans, or a combination thereof, and the second medical imaging modality is magnetic resonance imaging, diffusion tensor imaging, computed tomography images, positron emission tomography images, photoacoustic tomography images, X-ray images, ultrasound scans, or a combination thereof.

[0013] In some embodiments, the first type of image is a magnetic resonance image, a diffusion tensor image or map, a computed tomography image, a positron emission tomography image, a photoacoustic tomography image, an X-ray image, an ultrasound scan image, or a combination thereof, and wherein the second type of image is a magnetic resonance image, a diffusion tensor image or map, a computed tomography image, a positron emission tomography image, a photoacoustic tomography image, an X-ray image, an ultrasound scan image, or a combination thereof.

[0014] In some embodiments, the first feature is partial anisotropy index contrast, mean diffusivity contrast, axial diffusivity contrast, radial diffusivity contrast, proton density contrast, T1 relaxation time contrast, T2 relaxation time contrast, diffusion coefficient contrast, low resolution, high resolution, drug contrast, radiotracer contrast, light absorption contrast, echo distance contrast, or a combination thereof, and wherein the second feature is partial anisotropy index contrast, mean diffusivity contrast, axial diffusivity contrast, radial diffusivity contrast, proton density contrast, T1 relaxation time contrast, T2 relaxation time contrast, diffusion coefficient contrast, low resolution, high resolution, drug contrast, radiotracer contrast, light absorption contrast, echo distance contrast, or a combination thereof.

[0015] In some embodiments, the one or more medical imaging modalities are diffusion tensor imaging, the first image is a partial anisotropy index (FA) map, the second image is a mean diffusion rate (MD) map, the first feature is a partial anisotropy index contrast, the second feature is a mean diffusion rate contrast, and the target object is the kidney of the subject.

[0016] In some embodiments, locating and classifying the object within the first image includes applying one or more clustering algorithms to a plurality of pixels or voxels of the first image.

[0017] In some embodiments, the one or more clustering algorithms include a k-means algorithm that assigns observations to clusters associated with a plurality of object classes.

[0018] In some embodiments, the one or more clustering algorithms further include an expectation-maximization algorithm that calculates the probability of cluster membership based on one or more probability distributions, and wherein the k-means algorithm initializes the expectation-maximization algorithm by estimating initial parameters for each of the multiple object categories.

[0019] In some embodiments, a segmentation mask is determined, and determining the segmentation mask includes: identifying a seed position of a target object using a set of pixels or voxels assigned to an object class; growing the seed position by projecting the seed position to a z-axis representing a depth of the segmentation mask; and determining the segmentation mask based on the projected seed position.

[0020] In some embodiments, determining the segmentation mask further comprises performing morphological closing and filling on the segmentation mask.

[0021] In some embodiments, the method further includes cropping the second image based on the object mask plus a margin before inputting the portion of the second image into the three-dimensional neural network model to generate the portion of the second image.

[0022] In some embodiments, the method further includes inputting the portion of the second image into a deep super-resolution neural network before inputting the portion of the second image into the three-dimensional neural network model to increase the resolution of the portion of the second image.

[0023] In some embodiments, the three-dimensional neural network model includes a plurality of model parameters identified using a training dataset, the training dataset comprising: a plurality of medical images having annotations associated with segmentation boundaries surrounding a target object; and a plurality of additional medical images having annotations associated with segmentation boundaries surrounding the target object, wherein the plurality of additional medical images are artificially generated by matching image histograms from the plurality of medical images with image histograms from a plurality of reference images; and wherein the plurality of model parameters are identified using the training dataset based on minimizing a weighted loss function.

[0024] In some embodiments, the weighted loss function is a weighted Dice loss function.

[0025] In some embodiments, the three-dimensional neural network model is a modified 3D U-Net model.

[0026] In some embodiments, the modified 3D U-Net model includes a total number of between 5,000,000 and 12,000,000 learnable parameters.

[0027] In some embodiments, the modified 3D U-Net model includes a total number of between 800 and 1,700 kernels.

[0028] In some embodiments, the method further includes: determining the size, surface area and / or volume of the target object based on an estimated boundary surrounding the target object; and providing: (i) a portion of the second image having the estimated segmentation boundary surrounding the target object and / or (ii) the size, surface area and / or volume of the target object.

[0029] In some embodiments, the method further includes determining, by a user, a diagnosis of the subject based on (i) a portion of the second image having an estimated segmentation boundary surrounding the object of interest and / or (ii) a size, surface area, and / or volume of the object of interest.

[0030] In some embodiments, the method further includes: acquiring a medical image of the subject by a user using an imaging system, wherein the imaging system uses one or more medical imaging modalities to generate the medical image; determining the size, surface area and / or volume of the target object based on an estimated segmentation boundary surrounding the target object; providing: (i) a portion of the second image having an estimated segmentation boundary surrounding the target object and / or (ii) the size, surface area and / or volume of the target object; receiving by the user (i) a portion of the second image having an estimated segmentation boundary surrounding the target object and / or (ii) the size, surface area and / or volume of the target object; and determining by the user a diagnosis of the subject based on (i) a portion of the second image having an estimated segmentation boundary surrounding the target object and / or (ii) the size, surface area and / or volume of the target object.

[0031] In some embodiments, the method further includes administering, by the user, treatment with the compound based on (i) the portion of the second image having an estimated segmented boundary surrounding the object of interest, (ii) the size, surface area and / or volume of the object of interest, and / or (iii) a diagnosis of the subject.

[0032] In some embodiments, a system is provided that includes one or more data processors and a non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform part or all of one or more methods disclosed herein.

[0033] In some embodiments, a computer program product is provided, which is tangibly embodied in a non-transitory machine-readable storage medium and includes instructions configured to cause one or more data processors to perform a part or all of one or more methods disclosed herein.

[0034] Some embodiments of the present disclosure include a system including one or more data processors. In some embodiments, the system includes a non-transitory computer-readable storage medium containing instructions, which, when executed on the one or more data processors, cause the one or more data processors to perform part or all of one or more methods disclosed herein and / or part or all of one or more processes disclosed herein. Some embodiments of the present disclosure include a computer program product tangibly embodied in a non-transitory machine-readable storage medium, which includes instructions configured to cause one or more data processors to perform part or all of one or more methods disclosed herein and / or part or all of one or more processes disclosed herein.

[0035] The terms and expressions that have been adopted are used as terms of description rather than limitation, and when using these terms and expressions, there is no intention to exclude any equivalents of the features shown and described or parts thereof, but it should be recognized that various modifications are possible within the scope of the invention claimed. Therefore, it should be understood that although the claimed invention has been specifically disclosed through embodiments and optional features, modifications and variations of the concepts disclosed herein may be adopted by those skilled in the art, and such modifications and variations are considered to be within the scope of the invention defined by the appended claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] The present disclosure is described in conjunction with the accompanying drawings:

[0037] Figure 1 An example computing environment for segmenting an instance of an object of interest according to various embodiments is shown;

[0038] Figure 2 shows histogram matching for simulating additional contrasts and increasing the variance of a training data set according to various embodiments;

[0039] Figure 3 An exemplary U-Net according to various embodiments is shown;

[0040] Figure 4 shows a process for segmenting an instance of a destination object according to various embodiments;

[0041] Figure 5A diffusion tensor elements according to various embodiments;

[0042] Figure 5B shows a partial anisotropy index image for expectation maximization (EM) segmentation (12 classes) and object detection steps according to various embodiments;

[0043] Figure 5C shows a super-resolution image in a slice direction according to various embodiments; and

[0044] Figures 6A to 6E The segmentation results obtained using various strategies are shown. Fig. 6A : 3D U-Net. Figure 6B : Use connected component preprocessing to detect foreground. Figure 6C : EM segmentation. Fig.6D : Kidney detection by EM segmentation. Fig. 6E : Kidney detection by EM segmentation on super-resolved images. The first row shows: the ground truth manual labels superimposed on the magnetic resonance imaging (MRI). The second row shows: transparent surface rendering of the ground truth and segmentation mask. Coronal and axial positions are shown in pairs. The third row shows: Dice similarity coefficient (DSC) displayed as violin plots. Example datasets selected based on the average DSC for each segmentation strategy. All segmentation results are based on 3D U-Net, except C, which only uses EM segmentation. The dotted box indicates the object detection area. Scale bar = 4mm.

[0045] In the drawings, similar parts and / or features may have the same reference numerals. In addition, various parts of the same type may be distinguished by following the reference numeral with a dash and a second reference numeral that distinguishes the similar parts. If only the first reference numeral is used in the specification, the description applies to any similar part having the same first reference numeral, regardless of the second reference numeral. DETAILED DESCRIPTION

[0046] I. Overview

[0047] The present disclosure describes automated object segmentation techniques for medical images. More specifically, embodiments of the present disclosure provide techniques for segmenting objects within medical images using a deep learning network that localizes through object detection based on a derived contrast mechanism.

[0048] Medical image segmentation, the identification of pixels of an object (e.g., an organ, lesion, or tumor) from background medical images such as computed tomography (CT) or MRI images, is a fundamental task in medical image analysis and is used to provide information about the shape, size, and volume of an object. Changes in organ size or volume may be a primary feature of the disease process or a pathological manifestation of other parts of the subject. In addition, tumor or lesion size or volume may be an important independent indicator of a subject suffering from cancer (e.g., repeated measurements of size during initial systemic treatment will produce detailed information about the response, which can be used to select the most effective treatment regimen and estimate the prognosis of the subject). In the past, attempts have been made to estimate the size and volume of tumors or organs using various radiological and clinical techniques, but most techniques have limited practicality due to unacceptable accuracy, poor reproducibility, or difficulty in obtaining suitable images for size and volume measurement. Recently, several quantitative schemes have been developed and have demonstrated promising results in measuring the size and volume of various organs, lesions, or tumors by various imaging modalities such as CT and MRI. For this reason, time-consuming manual segmentation methods are often used to obtain size and volume data. However, the great potential of deep learning techniques has made these techniques the dominant choice for image segmentation, especially medical image segmentation, which has greatly improved the practicality of obtaining size and volume data through quantitative schemes.

[0049] Although object segmentation has been facilitated using deep learning techniques and has achieved significant improvements over traditional manual and machine learning based methods, it remains challenging in low-contrast and low-resolution images, which are particularly prevalent in the field of medical images such as MRI and CT images. The main reasons for this challenge stem from the following: (i) the definition of objects (such as organs, lesions, or tumors) is strongly affected by blurred visual recognition and excessive noise in low-contrast and low-resolution images, which can mislead deep learning models in predicting the true contours of objects; (ii) due to the repeated pooling operations in deep learning architectures, object semantics and image structure information are inevitably lost, which is severely missing in low-contrast and low-resolution images, so the results of deep learning models often suffer from inaccurate object shapes and poor localization; (iii) since pixels around object boundaries are concentrated in similar receptive fields and deep learning models only distinguish binary labels of image pixels, it is difficult for deep learning algorithms to learn boundary knowledge; (iv) many medical images have large fields of view or backgrounds to avoid aliasing effects, but when the background occupies a large part of the image, deep learning may not be optimally trained to segment foreground objects (“background effect”); and (v) in the background, there may be objects with similar appearance (e.g., liver, heart, kidneys may sometimes look like tumors), and deep learning and simpler machine learning algorithms may not be optimally trained to distinguish these structures.

[0050] To address these limitations and problems, the automated object segmentation technology of the present embodiment uses various imaging modalities and / or medical image types with different features (features that enable objects (or the way they are presented in images or displays) to be distinguished, such as contrast or resolution) as derived contrast mechanisms to locate the target object, separate the target object, and then segment the target object using a deep learning model. For example, a first image of an object obtained by a first imaging modality may have a first feature (e.g., good contrast) that provides a good general outline of the object, so that the first image can be used for object detection (providing a coarse-grained boundary around the object and classification). However, the first image may be very blurry, making it impossible for the deep learning network to determine the exact object edge location for accurate object segmentation. In contrast, a second image of the object obtained using a second imaging modality or a second image with the same pattern of image features / contrast mechanisms may have a second feature (e.g., high resolution) that provides a good clear boundary of the object, so that the second image can be used for edge detection and fine-grained object segmentation. Once the object is detected in the first image, the coarse-grained boundary of the object is projected onto the second image to locate the object within the second image. The second image is then cropped using the coarse-grained boundaries of the object on the second image before object segmentation. The localization and cropping of the second image mitigates background effects and enables the deep learning model to focus on the edges of the object to learn boundary knowledge for fine-grained object segmentation.

[0051] An exemplary embodiment of the present disclosure relates to a method comprising: initially locating (e.g., using an algorithm such as expectation maximization) a target object, such as an organ, a tumor, or a lesion, within a first medical image having a first feature; projecting a bounding box or segmentation mask of the target object onto a second medical image having a second feature to define a portion of the second medical image that includes the target object; and then inputting the portion of the second medical image into a deep learning model such as a convolutional neural network model, the deep learning model being constructed to use a detector that can segment the portion of the second medical image and generate a weighted loss function around the segmentation boundary of the target object. The segmentation boundary can be used to calculate the volume of the target object for determining diagnosis and / or prognosis. In some cases, the calculated volume can be further associated with a time point. The volume of an object from a certain time point can be compared with the volume of an object from a previous time point to determine the efficacy of a treatment. Time point analysis provides context for changes in an organ or tumor over time. In addition, the specific content within the target object defined by the segmentation boundary may have changes, for example, more necrotic content or an invasive tumor type. In some cases, the segmentation boundary and the corresponding segmented region or volume can be used to quantitatively analyze image metrics, such as image intensity. For example, there is a standardized uptake value (SUV) in PET, or diffusion rate, T2, T1, etc. in MRI, which are related to certain image metrics such as image intensity, so quantitative analysis of image metrics can be used to determine values / metrics such as SUV for specific purpose objects. In other words, the volume within the segmentation boundary itself is a useful metric, and the value or metric within the segmentation boundary and the corresponding segmented region is also a useful metric.

[0052] Advantageously, these methods utilize multiple medical images with different features and object detection techniques to detect the approximate area of ​​the target object before attempting to segment the target object using a deep learning model. This reduces background effects, reduces the complexity of the input data, and focuses the deep learning model on the edges of the object to learn boundary knowledge for object segmentation. In addition, as the complexity of the input data decreases, the complexity of the deep learning model can be reduced (e.g., by reducing the number of kernels in each convolutional layer). In some cases, the deep learning model is constructed using a weighted loss function that minimizes segmentation errors, improves training performance optimization, and further reduces background effects that may still exist in some approximate areas determined by object detection techniques.

[0053] II. Definitions

[0054] As used herein, when an action is "based on" something, it means that the action is at least partially based on at least a portion of the something.

[0055] As used herein, the terms "substantially," "approximately," and "about" are defined as being largely, but not necessarily completely, as specified (and including completely as specified), as understood by one of ordinary skill in the art. In any disclosed embodiment, the terms "substantially," "approximately," or "about" may be replaced with "within [a certain percentage]" as specified, where percentages include 0.1%, 1%, 5%, and 10%.

[0056] As used herein, "mask" refers to an image representing the surface area of ​​the detected object. A mask may include pixels with non-zero intensity to indicate one or more target areas (e.g., one or more detected objects) and pixels with zero intensity to indicate background.

[0057] As used herein, a "binary mask" refers to a mask in which each pixel value is one of two values ​​(e.g., 0 or 1). A zero intensity value may indicate that the corresponding pixel is part of the background, and a non-zero intensity value (e.g., a 1 value) may indicate that the corresponding pixel is part of the target region.

[0058] As used herein, "classification" refers to a process that takes an input (e.g., an image or a portion of an image) and outputs a class (e.g., "organ" or "tumor") or the probability that the input is a particular class. This process may include binary classification (is it a member of a class), multi-class classification (assignment to one or more classes), providing probabilities of membership in each class (e.g., there is a 90% probability that the input is an organ), and similar classification schemes.

[0059] As used herein, "object localization" or "object detection" refers to the process of detecting instances of objects of a particular class in an image.

[0060] As used herein, a "bounding box" refers to a rectangular box that represents the approximate location of an object of a particular class in an image. The bounding box may be defined by the x and y coordinates of the upper left and / or upper right corners of the rectangle and the x and y coordinates of the lower right and / or lower left corners of the rectangle.

[0061] As used herein, "segmentation boundary" refers to the estimated perimeter of an object within an image. The segmentation boundary may be generated during a segmentation process in which features of an image are analyzed to determine the locations of object edges. The segmentation boundary may further be represented by a mask, such as a binary mask.

[0062] As used herein, "segmentation" refers to determining the location and shape of an object within an image. Segmentation may involve determining a set of pixels that delineate an area or perimeter of an object within an image. Segmentation may involve generating a mask of the object, such as a binary mask. Segmentation may further involve processing multiple masks corresponding to an object in order to generate a 3D mask of the object.

[0063] III. Derived contrast mechanisms

[0064] The goal of imaging procedures of a subject's anatomical structures (e.g., organs or other human or mammalian tissues) and physiological processes is to generate image contrast with good spatial resolution. The initial evolution of medical imaging focused on tissue (proton) density functions and tissue relaxation properties to generate signal contrast, which is the main principle behind conventional MRI. MRI detects signals from protons of water molecules, but it can only provide grayscale images in which each pixel contains an integer value. Unless two anatomical regions A and B contain water molecules with different physical or chemical properties, the two regions cannot be distinguished from each other using MRI. Otherwise, no matter how high the image resolution is, region A is indistinguishable from region B. To generate MR contrast based on the physical properties of water molecules, proton density (PD), T1 and T2 relaxation times, and diffusion coefficient (D) are widely used. PD represents water concentration. T1 and T2 are signal relaxation (decay) times after excitation, which are related to environmental factors such as viscosity and the presence of nearby macromolecules. The diffusion term D represents the thermal (or Brownian) motion of water molecules.

[0065] After initially focusing on tissue (proton) density functions and tissue relaxation properties, researchers explored other methods of exploiting other properties of water molecules to create contrast. Diffusion imaging (DI) is the result of those research efforts. In DI, complementary MR gradients are applied during image acquisition. During the application of these gradients, the motion of the protons affects the signal in the image, providing information about molecular diffusion. DI can be performed using a variety of techniques, including diffusion-spectral imaging (DSI) and diffusion-weighted imaging (DWI).

[0066] DWI is a non-invasive imaging method that is sensitive to water diffusion within tissue structures, using existing MRI technology combined with specialized software and without the need for additional hardware equipment, contrast agents, or chemical tracers. To measure diffusion using MRI, complementary MR gradients are used to create an image that is sensitive to diffusion in a specific direction. In DWI, the intensity of each image element (voxel) reflects the best estimate of the water diffusion rate in a specific direction. However, biological tissues are highly anisotropic, meaning that their diffusion rate is different in each direction. For conventional DWI, the anisotropic properties of tissues are often ignored and diffusion is reduced to a single average value, the apparent diffusion coefficient (ADC), which is an oversimplification for many use cases. An alternative approach is to model diffusion in complex materials using a diffusion tensor, which is a [3×3] array of numbers corresponding to the diffusion rate in each combination of directions. The three diagonal elements (Dxx, Dyy, Dzz) represent the diffusion coefficients measured along each of the major (x-, y-, and z-) laboratory axes. The six off-diagonal terms (Dxy, Dyz, etc.) reflect the correlation of the random motion between each pair of principal directions.

[0067] The introduction of diffusion tensor imaging has enabled indirect measurements of the degree of anisotropy and structural orientation that characterizes diffusion tensor imaging (DTI). The basic concept behind DTI is that water molecules diffuse differently along tissues, depending on the type, integrity, structure, and presence of barriers, providing information about their orientation and quantitative anisotropy. With DTI analysis, properties such as the rate of molecular diffusion (mean diffusivity (MD) or apparent diffusion coefficient (ADC)), the directional preference of diffusion (partial anisotropy index (FA)), axial diffusivity (AD) (diffusion rate along the diffusion direction), and radial diffusivity (RD) (diffusion rate in the transverse direction) can be inferred in each voxel. DTI is usually displayed by compressing the information contained in the tensor into one number (scalar) or into 4 numbers (to provide R, G, B color and brightness values, which are called color partial anisotropy indices). Diffusion tensors can also be viewed using glyphs, which are small three-dimensional (3D) representations of the principal eigenvectors or the entire tensor.

[0068] Similar to MRI and DTI, other medical imaging modalities such as CT, X-ray, positron emission tomography (PET), photoacoustic tomography (PAT), ultrasound scanning, combinations thereof (such as PET-CT, PET-MR, etc.) rely on various measurements, algorithms, and agents to generate image contrast and spatial resolution. For example, CT and X-ray use X-ray absorption to distinguish between air, soft tissue, and dense structures such as bones. Dense structures in the body block X-rays and are therefore easy to image and visualize, while soft tissues vary in their ability to block X-rays, so they may be weak and impossible or difficult to image and visualize. One technique for increasing image contrast in X-ray or CT scans is to use contrast agents that contain substances that are better at blocking X-rays, making them more visible on X-ray or CT images, and therefore can be used to better visualize soft tissues such as blood vessels. PET uses small amounts of radioactive substances called radiotracers that can be detected and measured in the scan. Contrast is generated using the measured difference between areas that accumulate or carry radiotracer labels and non-accumulate or unlabeled areas, thereby enabling visualization of structures and functions in the subject's body. PAT is an imaging method based on the photoacoustic (PA) effect. Short pulsed light sources are usually used to illuminate tissues to obtain broadband PA waves. After absorbing light, the initial temperature increase causes an increase in pressure, which propagates in the form of photoacoustic waves and is detected by an ultrasound transducer to achieve light absorption contrast imaging. Ultrasound is a non-invasive diagnostic technique for in vivo imaging. The transducer emits a beam of sound waves into the human body. The sound waves are reflected back to the transducer by the boundaries between tissues in the path of the sound beam (for example, the boundaries between body fluids and soft tissues or between tissues and bones). When these echoes hit the transducer, the echoes generate electrical signals that are sent to an ultrasound scanner. The scanner uses the speed of sound and the time it takes for each echo to return to calculate the distance from the transducer to the tissue. These distances are then used to generate contrast to visualize tissues and organs.

[0069] All of these imaging modalities produce image contrast with sufficiently high spatial resolution, which is sufficient to visualize the internal presentation of the body and the visual presentation of some organ or tissue functions for clinical analysis, medical intervention and / or medical diagnosis. However, as described herein, the image contrast and spatial resolution provided by each of these imaging modalities alone are not sufficient for accurate object segmentation, especially for object segmentation to obtain size and volume data, which is performed by a deep learning network. To overcome this limitation and other limitations, the technology described herein uses a combination of imaging modalities, image types and / or changing features to locate the target object, separate the target object, and then use a deep learning model to segment the target object. Specifically, it is found that some imaging modalities, image types and / or features perform better in object detection compared to object segmentation; while other imaging modalities, image types and / or features perform better in object segmentation. By identifying computer vision tasks (e.g., object detection or object segmentation) that are more suitable for imaging modalities, image types and / or features, these differences can be used as a derived contrast mechanism.

[0070] Image characteristics that can be exploited by derived contrast mechanisms include brightness, contrast, and spatial resolution. Brightness (or luminous brightness) is a measure of the relative intensity values ​​on an array of pixels after an image is acquired using a digital camera or digitized using an analog-to-digital converter. The higher the relative intensity value, the brighter the pixel, and the image generally appears whiter; while the lower the relative intensity value, the darker the pixel, and the image generally appears blacker. Contrast refers to the difference that exists between various image features in analog and digital images. The differences within an image can be in the form of different grayscales, light intensities, or colors. Images with higher contrast levels generally show greater grayscale, color, or intensity variations than images with lower contrast. Spatial resolution refers to the number of pixels used to construct a digital image. Images with higher spatial resolutions are composed of a greater number of pixels than images with lower spatial resolutions.

[0071] The derived contrast mechanism includes: (i) a first imaging method that can obtain an image having features for detecting a target object (e.g., DTI-FA); and (ii) a second imaging method that can obtain an image having features for segmenting a target object (e.g., DTI-MD). Various imaging methods, image types, and / or features can be combined to improve on various computer vision tasks (e.g., object detection or object segmentation). In various embodiments, the imaging methods of the derived contrast mechanism are the same, such as MRI or DTI. In some embodiments, a first image having a first feature and a second image having a second feature are obtained using an imaging method, wherein the first image is different from the second image. For example, MRI can be used to obtain a diffusion tensor parameter map of a subject. The diffusion tensor parameter map may include a first measurement map (such as an FA map) and a second measurement map (such as an MD map). In some embodiments, a first image having a first feature and a second image having a second feature are obtained using an imaging method, wherein the first feature is different from the second feature. For example, CT can be used to obtain multiple CT scans of a subject. The CT scan may include a first CT scan (such as a low-resolution CT scan) and a second CT scan (such as a high-resolution CT (HRCT) scan). Alternatively, MRI may be used to obtain a diffusion tensor parameter map of the subject. The diffusion tensor parameter map may include a first MD map (such as a low-resolution MD map) and a second MD map (such as a high-resolution MD map). In other embodiments, the imaging methods of the derived contrast mechanism are different, such as PAT and ultrasound. PAT can be used to obtain a first type of image having a first feature, and ultrasound can be used to obtain a second type of image having a second feature, wherein the first type of image and the first feature are different from the second type of image and the second feature.

[0072] Specific examples of derived contrast mechanisms using different imaging modality types, image types, and features include the following:

[0073] (A)MRI

[0074] Kidney segmentation: (i) FA measurements for object detection (contrast generated by the fractional anisotropy index); and (ii) MD measurements (contrast generated by the mean diffusion ratio) or T2-weighted anatomical images (contrast generated by the signal relaxation time after excitation) for object segmentation.

[0075] Multiple sclerosis brain lesion segmentation: (i) single-echo T2 images for object detection (contrast created by single-shot echo planar imaging of the signal relaxation time after excitation); and (ii) echo-enhanced or T2-weighted anatomical images for object segmentation (contrast created by low flip angles, long echo times, and long repetition times for highlighting the signal relaxation time after excitation).

[0076] Liver segmentation: (i) MD measurements for object detection (contrast due to mean diffusion rate); and (ii) high-resolution MD measurements (high resolution and contrast due to mean diffusion rate), T2-weighted anatomical images (contrast due to post-excitation signal relaxation time), or PD (contrast due to water concentration) for object segmentation.

[0077] (B)CT

[0078] Lung and liver tumor segmentation: (i) CT scan (low resolution) for object detection of the lung or liver; and (ii) CT scan (HRCT) for object segmentation of tumors.

[0079] Trabecular bone: (i) CT scan (low resolution) for object detection of intertrabecular spaces (non-cortical bone), and (ii) CT scan (HRCT) for object segmentation of trabeculae.

[0080] (C)PET

[0081] Tumor or organ detection: (i) PET high contrast / low resolution (contrast produced by radiotracer measurements) for object detection; and (ii) PET-CT or PET-MR high contrast / high resolution (contrast produced by radiotracer measurements) for object segmentation.

[0082] (D) Photoacoustic tomography (optical ultrasound technology)

[0083] Tumor or organ detection: (i) PAT for object detection (contrast created by light absorption); and (ii) ultrasound for object segmentation (contrast created by echo return distance between transducer and tissue boundary).

[0084] It should be understood that the examples and embodiments described herein with respect to MRI, CT, PAT, PET, etc. are for illustrative purposes only and will suggest to those skilled in the art alternative imaging modalities (e.g., fluoroscopy, magnetic resonance angiography (MRA), and mammography) for implementing the various derived contrast mechanisms described in accordance with aspects of the present disclosure. In addition, the parameters of any of these imaging modalities (e.g., different tracers, angle configurations, wavelengths, etc.) may be modified to capture different structures or regions of the body, and one or more of these types of modified imaging techniques may be combined with one or more other imaging techniques to implement the various derived contrast mechanisms described in accordance with aspects of the present disclosure.

[0085] IV. Techniques for Segmentation of Medical Images

[0086] The segmentation of MRI images is divided into two parts. The first part of the segmentation involves a first visual model, which is constructed to perform localization (object detection) of categories within a first image (e.g., a diffusion tensor parameter map or a CT image). These categories are "semantically interpretable" and correspond to real-world categories (such as liver, kidney, heart, etc.). Localization is performed using EM, You Only Look Once (YOLO) or (YOLOv2) or (YOLOv3) or similar object detection algorithms, which are heuristically initialized with standard clustering techniques (e.g., k-means clustering techniques, Otsu algorithm, density-based clustering with noise (DBSCAN) techniques, small batch K-means techniques, etc.). Initialization is used to provide an initial estimate of the likelihood model parameters for each category. Expectation maximization is an iterative process for finding the (local) maximum likelihood or maximum a posteriori probability (MAP) estimate of parameters in one or more analytical statistical models. The EM iteration alternates between performing an expectation (E) step, in which a function of the expected log-likelihood evaluated using the estimated values ​​of the current parameters is created, and a maximization (M) step, in which the parameters are calculated to maximize the expected log-likelihood found in the E step. These parameter estimates are then used to determine the distribution of the latent variables in the next E step. The result of the localization is a bounding box or segmentation mask around each object and its probability of belonging to each category. The approximate location of the target object (e.g., kidney) is separated using one or more of the categories associated with the target object.

[0087] To mitigate the background effect of medical images, a bounding box or segmentation mask of a target object is projected onto a second image (e.g., a diffusion tensor parameter map or a CT image) along the slice direction (axial, coronal, and sagittal). In the case where positioning is used to determine the segmentation mask, the bounding box (the rectangular box drawn completely surrounds the pixel-level mask associated with the target object) of the approximate position of the target object (e.g., kidney) in the second image is defined by the boundary of the projected segmentation mask. In some cases, the bounding box (defined by positioning or based on the boundary of the projected segmentation mask) is enlarged by a predetermined number of pixels on all sides to ensure that the target object is covered. The area within the bounding box is then cropped from the second image to obtain a portion of the second image with the target object, and the portion of the second image is used as an input to a second visual model to segment the target object. The second part of the segmentation involves a second visual model (deep learning neural network) constructed using a weighted loss function (e.g., Dice loss) to overcome the unbalanced nature between the target object and the background, and thereby focus on training and evaluating the segmentation target object. In addition, the second visual model can be trained using an enhanced data set so that the deep learning neural network can be trained with a limited set of medical images. The trained second visual model uses the cropped portion of the second image as input, and outputs the portion of the second image with the estimated segmentation boundary around the target object. The estimated segmentation boundary can be used to calculate the volume, surface area, axial dimension, maximum axial dimension or other size-related metrics of the target object. Any one or more of these metrics can then be used alone or in combination with other factors to determine the diagnosis and / or prognosis of the experimenter.

[0088] IV.A. Example Computing Environment

[0089] Figure 1 An example computing environment 100 (ie, a data processing system) is shown for use in segmenting an object of interest within an image using a multi-stage segmentation network in accordance with various embodiments. Figure 1 As shown, in this example, the segmentation performed by the computing environment 100 includes several stages: an image acquisition stage 105 , a model training stage 110 , an object detection stage 115 , a segmentation stage 120 , and an analysis stage 125 .

[0090] The image acquisition stage 110 includes one or more imaging systems 130 (e.g., MRI imaging systems) for obtaining images 135 (e.g., MR images) of various parts of the subject. The imaging system 130 is configured to obtain the images 135 using one or more radiographic imaging techniques such as X-ray photography, fluoroscopy, MRI, ultrasound, nuclear medicine functional imaging (e.g., PET), thermal imaging, CT, mammography, etc. The imaging system 130 is capable of determining the differences between various structures and functions within the subject based on the characteristics associated with each of the imaging systems 130 (e.g., brightness, contrast, and spatial resolution), and generates a series of two-dimensional images. Once the series of two-dimensional images are collected by the scanner's computer, these two-dimensional images can be digitally "stacked" together through computer analysis to reconstruct a three-dimensional image of the subject or a portion of the subject. The two-dimensional images and / or reconstructed three-dimensional images 135 allow for easier identification and location of basic structures (e.g., organs) and possible tumors or abnormalities. Each two-dimensional image and / or reconstructed three-dimensional image 135 can correspond to a session time and subject and depict an internal area of ​​the subject. Each of the two-dimensional images and / or reconstructed three-dimensional images 135 may further have a standardized size, resolution, and / or magnification.

[0091] In some embodiments, one or more imaging systems 130 include a DI system (e.g., an MRI system equipped with dedicated software) configured to apply supplemental MR gradients during image acquisition. During the application of these gradients, the motion of protons affects the signal in the image, providing information about molecular diffusion. A DTI matrix is ​​obtained from a series of diffusion-weighted images in various gradient directions. Three diffusion rate parameters or eigenvalues ​​(λ1, λ2, and λ3) are generated using matrix diagonalization. Diffusivity is a scalar metric that describes the diffusion of water in a specific voxel (the smallest volume element in the image) associated with the geometry of the tissue. Diffusion tensor parameter maps and additional image contrast can be calculated from the diffusivity using various diffusion imaging techniques. DTI characteristics or metrics represented by these maps may include (but are not limited to) molecular diffusivity (MD map or ADC map), directional preference of diffusion (FA map), AD map (diffusion rate along the main axis of diffusion), and RD map (diffusion rate in the transverse direction). The diffusivities (λ1, λ2, and λ3) obtained by diagonalizing the DTI matrix can be divided into a component parallel to the tissue (λ1) and a component perpendicular to the tissue (λ2 and λ3). The sum of the diffusivities (λ1, λ2, and λ3) is called the trajectory, and their average (= trajectory / 3) is called the MD or ADC. The partial anisotropy index (FA) is an indicator of the amount of diffusion asymmetry within a voxel and is defined based on its diffusivities (λ1, λ2, and λ3). The FA value varies between 0 and 1. For perfect isotropic diffusion, λ1=λ2=λ3, the diffusion ellipsoid is a sphere, and FA=0. As the diffusion anisotropy increases, the eigenvalues ​​differ more, the ellipsoid becomes more elongated, and FA→1. The axial diffusivity (AD), λ║≡λ1>λ2,λ3, describes the average diffusion coefficient of the diffusion of water molecules parallel to the tract within the target voxel. Similarly, the radial diffusivity (RD), λ┴≡(λ2+λ3) / 2, is defined as the magnitude of the water dispersion perpendicular to the principal eigenvectors.

[0092] Image 135 shows one or more target objects.Target object can be any target "thing" in subject, such as region (e.g., abdominal region), organ (e.g., kidney), lesion / tumor (e.g., malignant liver tumor or brain lesion), metabolic function (e.g., synthesis of plasma protein in liver), etc. In some cases, multiple images 135 show target object, so that each of the multiple images 135 can correspond to the virtual "slice" of target object. Each of multiple images 135 can have the same viewing angle, so that the plane shown in each image 135 is parallel to other planes corresponding to subject and target object shown in other images 135. Each of multiple images 135 can further correspond to different distances along the axis perpendicular to the plane. In some cases, multiple images 135 showing target object are through pre-processing steps to align each image and generate the three-dimensional image structure of target object.

[0093] In some embodiments, the image 135 includes diffusion tensor parameter maps that show one or more target objects of the subject. In some cases, at least two diffusion tensor parameter maps (e.g., a first diffusion tensor parameter map and a second diffusion tensor parameter map) are generated for the target object. The diffusion tensor parameter map may be generated by the DTI system and describes the diffusion rate and / or direction of water molecules to provide more context for the target object. More than one diffusion tensor parameter map may be generated so that each diffusion tensor parameter map corresponds to a different direction. For example, the diffusion tensor parameter map may include an image showing FA, an image showing MD, an image showing AD, and / or an image showing RD. Each of the diffusion tensor parameter maps may additionally have a viewing angle of the same plane and the same distance along the vertical axis of the plane as the corresponding MR image, so that each MR image showing a virtual "slice" of the target object has a corresponding diffusion tensor image that shows the same virtual "slice" of the target object.

[0094] The model training phase 110 constructs and trains one or more models 140a to 140n ("n" represents any natural number) to be used by other phases (which may be individually referred to as models 140 or collectively referred to as models 140 herein). Model 140 may be a machine learning ("ML") model, such as a convolutional neural network ("CNN"), for example, an inception neural network, a residual neural network ("Resnet"), a U-Net, a V-Net, a single shot multibox detector ("SSD") network, or a recurrent neural network ("RNN"), such as a long short-term memory ("LSTM") model or a gated recurrent unit ("GRUs") model, or any combination thereof. Model 140 may also be any other suitable ML model trained in object detection and / or segmentation based on an image, such as a three-dimensional CNN ("3DCNN"), a dynamic time warping ("DTW") technique, a hidden Markov model ("HMM"), etc., or a combination of one or more such techniques, such as a CNN-HMM or an MCNN (multi-scale convolutional neural network). The computing environment 100 can use the same type of model or a different type of model to segment instances of the target object. In some cases, the model 140 is constructed using a weighted loss function that compensates for the unbalanced nature between the large field of view or background and the small foreground target object within each image, as described in further detail herein.

[0095] To train the model 140 in this example, samples 145 are generated by acquiring digital images, segmenting the images into image subsets 145a (e.g., 90%) for training and image subsets 145b (e.g., 10%) for validation, preprocessing the image subsets 145a and 145b, augmenting the image subsets 145a, and in some cases annotating the image subsets 145a with labels 150. The image subsets 145a are acquired by one or more imaging modalities (e.g., MRI and CT). In some cases, the image subsets 145a are acquired by a data storage structure associated with the one or more imaging modalities, such as a database, an imaging system (e.g., one or more imaging systems 130), etc. Each image shows one or more objects of interest, such as a head region, a chest region, an abdominal region, a pelvic region, a spleen, a liver, a kidney, a brain, a tumor, a lesion, etc.

[0096] The segmentation can be performed randomly (e.g., 90% / 10% or 70% / 30%), or the segmentation can be performed according to more complex validation techniques (such as K-fold cross validation, leave-one-out cross validation, leave-one-out cross validation, nested cross validation, etc.) to minimize sampling bias and overfitting. Preprocessing may include cropping images so that each image contains only a single object of interest. In some cases, preprocessing may further include standardization or normalization to place all features on the same scale (e.g., the same size scale, or the same color scale or color saturation scale). In some cases, the image is resized using a minimum size (width or height) of predetermined pixels (e.g., 2500 pixels) or a maximum size (width or height) of predetermined pixels (e.g., 3000 pixels), and the original aspect ratio is maintained.

[0097] Augmentation can be used to artificially expand the size of the image subset 145a by creating modified versions of the images in the dataset. Image data augmentation can be performed by creating transformed versions of the images in the dataset that belong to the same category as the original images. Transformations include a range of operations from the field of image manipulation, such as shifting, flipping, scaling, etc. In some cases, these operations include random erasing, shifting, highlighting, rotating, Gaussian blurring, and / or elastic transformations to ensure that the model 140 can perform in environments other than those available from the image subset 145a.

[0098] Enhancement can be used additionally or alternatively to artificially expand multiple images in the data set, which the model 140 can use as input during the training process. In some cases, at least a portion of the training data set (i.e., image subset 145a) can include a first image set corresponding to an area of ​​one or more subjects and a second image set corresponding to different areas of the same or different subjects. For example, if at least a first subset of images corresponding to the abdominal area is used to detect one or more target objects in the abdominal area, a second subset of images corresponding to the head area can also be included in the training data set. In such cases, the first image set corresponding to the area of ​​one or more subjects is histogram matched with the second image set corresponding to different areas of the same or different subjects in the training data set. For the previous example, the image histogram corresponding to the head area can be processed as a reference histogram, and then the image histogram corresponding to the abdominal area is matched with the reference histogram. Histogram matching can be based on pixel intensity, pixel color and / or pixel brightness, so that the processing of the histogram requires matching the pixel intensity, pixel color and / or pixel brightness of the histogram with the pixel intensity, pixel color and / or pixel brightness of the reference histogram.

[0099] Annotation can be performed manually by one or more people (annotators, such as radiologists or pathologists), confirming that there are one or more target objects in each image of image subset 145a, and providing labels 150 to the one or more target objects, for example, using annotation software to draw a bounding box (true value) or segmentation boundary around the area confirmed by the person to include one or more target objects. In some cases, the bounding box or segmentation boundary can be drawn only for examples with a probability of being a target object greater than 50%. For images annotated by multiple annotators, bounding boxes or segmentation boundaries from all annotators can be used. In some cases, the annotation data can further indicate the type of the target object. For example, if the target object is a tumor or lesion, the annotation data can indicate the type of the tumor or lesion, such as a tumor or lesion in the liver, lung, pancreas and / or kidney.

[0100] In some cases, the image subset 145 may be transmitted to the annotator device 155 to be included in the training data set (i.e., the image subset 145a). The annotator device 155 may be provided (e.g., by a radiologist) using, for example, a mouse, a touchpad, a stylus, and / or a keyboard to provide input indicating whether the image shows a target object (e.g., a lesion, an organ, etc.); the number of target objects shown in the image; and the perimeter (bounding box or segmentation boundary) of each target object shown in the image. The annotator device 155 may be configured to generate a tag 150 for each image using the provided input. For example, the tag 150 may include the number of target objects shown in the image; the type classification of each target object shown; the number of each target object shown; and the perimeter and / or mask of one or more identified target objects in the image. In some cases, the tag 150 may further include the perimeter and / or mask of one or more identified target objects superimposed on the image of the first type and the image of the second type.

[0101] The training process of model 140 includes selecting hyperparameters for model 140 and performing an iterative operation of inputting images from image subset 145a into model 140 to find a set of model parameters (e.g., weights and / or biases) that minimize the loss or error function of model 140. Hyperparameters are settings that can be adjusted or optimized to control the behavior of model 140. Most models explicitly define hyperparameters that control different aspects of the model (such as memory or execution cost). However, additional hyperparameters can be defined to adapt the model to a specific scenario. For example, hyperparameters may include the number of hidden units of the model, the learning rate of the model, the width of the convolution kernel of the model, or the number of convolution kernels. In some cases, the number of model parameters for each convolution and deconvolution layer and / or the number of convolution kernels for each convolution and deconvolution layer is reduced by half compared to a typical CNN, as described in detail herein. Each iteration of training can involve finding a model parameter set (configured with a defined hyperparameter set) for model 140 so that the value of the loss or error function using the model parameter set is less than the value of the loss or error function using a different model parameter set in a previous iteration. The loss or error function may be configured to measure the difference between the output inferred using model 140 (in some cases, a segmentation boundary surrounding one or more instances of a target object as measured using the Dice similarity coefficient) and the true value segmentation boundary annotated to the image using labels 150.

[0102] Once the model parameter set is identified, the model 140 is trained and the image subset 145b (test or validation data set) can be used to validate the model. The validation process includes an iterative operation of inputting images from the image subset 145b into the model 140 to adjust the hyperparameters and ultimately find the optimal set of hyperparameters using validation techniques (such as K-fold cross validation, leave-one-out cross validation, leave-one-out cross validation, nested cross validation, etc.). Once the optimal set of hyperparameters is obtained, a reserved test set of images from the image subset 145b is input into the model 145 to obtain an output (in this example, a segmentation boundary around one or more target objects), and the output is evaluated relative to the true value segmentation boundary using correlation techniques (such as the Bland-Altman method and the Spearman rank correlation coefficient) and computing performance metrics (such as error, accuracy, precision, recall, receiver operating characteristic curve (ROC), etc.).

[0103] It should be understood that other training / validation mechanisms are also contemplated and can be implemented within the computing environment 100. For example, the model can be trained and hyperparameters can be adjusted on images from the image subset 145a, and images from the image subset 145b can be used only for testing and evaluating the performance of the model. In addition, although the training mechanisms described herein focus on training new models 140. These training mechanisms can also be used to fine-tune existing models 140 trained based on other data sets. For example, in some cases, the model 140 may have been pre-trained using images of other objects or biological structures or images of slices from other subjects or studies (e.g., human trials or rodent experiments). In those cases, the model 140 can be used for transfer learning and retrained / validated using the images 135.

[0104] The model training phase 110 outputs a training model including one or more trained object detection models 160 and one or more trained segmentation models 165. A first image 135 is obtained by a positioning controller 170 in the object detection phase 115. The first image 135 shows the target object. In some cases, the first image is a diffusion tensor parameter map having a first feature such as FA or MD contrast. In other cases, the first image is an MR image having a first feature such as single echo T2 contrast or T2-weighted anatomical contrast. In other cases, the first image 135 is a CT image having a first feature such as low resolution or high resolution. In other cases, the first image 135 is a CT image having a first feature such as drug contrast. In other cases, the first image 135 is a PET image having a first feature such as radiotracer contrast or low resolution. In other cases, the first image 135 is a PET-CT image having a first feature such as radiotracer contrast or high resolution. In other cases, the first image 135 is a PET-MR image having a first feature such as radiotracer contrast or high resolution. In other cases, the first image 135 is a PAT image having a first characteristic such as light absorption contrast. In other cases, the first image 135 is an ultrasound image having a first characteristic such as echoes or transducer-tissue boundary distance.

[0105] The positioning controller 170 includes positioning the target object within the image 135 using one or more object detection models 160. Positioning includes: (i) positioning and classifying an object within the first image having a first feature into a plurality of object categories using the object detection model 160, wherein the classification assigns a set of pixels or voxels of the first image to one or more of the plurality of object categories; and (ii) determining a bounding box or segmentation mask for the target object within the first image based on the set of pixels or voxels assigned to the object category in the plurality of object categories using the object detection model 160. The object detection model 160 utilizes one or more object detection algorithms to extract statistical features for locating and labeling the object within the first image, and predicting a bounding box or segmentation mask for the target object.

[0106] In some cases, EM, YOLO, YOLOv2, YOLOv3 or similar object detection algorithms are used for positioning, and these algorithms are heuristically initialized with standard clustering techniques (e.g., K-means or Otsu's algorithm). Initialization is used to provide initial estimates of likelihood model parameters for each category. For example, in the case of using EM and K-means clustering techniques, given a fixed number of k clusters, observations are assigned to k clusters so that the means (for all variables) between clusters are as different as possible from each other. Then, the EM clustering technique calculates the posterior probability and cluster boundaries of cluster members based on one or more prior probability distributions, and the one or more prior probability distributions are parameterized using the initial estimates of the parameters of each cluster (category). Then, the goal of the EM clustering technique is to maximize the overall probability or likelihood of data under a given (final) cluster. The results of the EM clustering technique are different from those calculated using the K-means clustering technique. The K-means clustering technique assigns observations (pixels or voxels, such as pixel or voxel intensity) to clusters to maximize the distance between clusters. The EM clustering technique does not calculate the actual assignment of observations to clusters, but calculates the classification probabilities. In other words, each observation belongs to each cluster with a certain probability. Then, the observations are assigned to clusters based on the (maximum) classification probabilities using the positioning controller 170. The result of the positioning is a bounding box or segmentation mask and its probability of belonging to each category. Based on the set of pixels or voxels assigned to the object class associated with the target object, the approximate location of the target object (e.g., kidney) is separated.

[0107] The bounding box or segmentation mask of the target object may be used in the image processing controller 175 of the object detection stage 115. The second image 135 is obtained by the image processing controller 175 in the object detection stage 115. The second image 135 shows the same target object shown in the first image 135. In some cases, the second image is a diffusion tensor parameter map with a second feature such as FA or MD contrast. In other cases, the second image is an MR image with a second feature such as single echo T2 contrast or T2-weighted anatomical contrast. In other cases, the second image 135 is a CT image with a second feature such as low resolution or high resolution. In other cases, the second image 135 is a CT image with a second feature such as drug contrast. In other cases, the second image 135 is a PET image with a second feature such as radiotracer contrast or low resolution. In other cases, the second image 135 is a PET-CT image with a second feature such as radiotracer contrast or high resolution. In other cases, the second image 135 is a PET-MR image with a second feature such as radiotracer contrast or high resolution. In other cases, the second image 135 is a PAT image having a second characteristic, such as light absorption contrast. In other cases, the second image 135 is an ultrasound image having a second characteristic, such as echoes or transducer-tissue boundary distance.

[0108] The figure processing controller 170 includes a process for superimposing a bounding box or segmentation mask corresponding to the target object detected from the first image on the same target object as shown in the second image. Where the segmentation mask is determined, the segmentation mask is projected onto the second image (e.g., the two-dimensional slice of the second image) so that the boundary of the segmentation mask can be used to define a rectangular bounding box surrounding the target area corresponding to the target object in the second image. In some cases, the bounding box includes additional filling (e.g., filling 5 pixels, 10 pixels, 15 pixels, etc.) to each edge of the perimeter of the segmentation mask to ensure that the entire target area is surrounded. The figure processing controller 175 further includes a process configured to crop the second image so that only the cropping portion 180 corresponding to the bounding box is shown. Where multiple bounding boxes are defined (e.g., for the case where multiple target objects are detected in the image), the cropping portion of the second image will be generated for each bounding box. In some cases, the size of each cropping portion is further adjusted (e.g., using additional filling) to maintain a uniform size.

[0109] The cropped portion 180 of the second image is transmitted to the segmentation controller 185 of the segmentation stage 120. The segmentation controller 185 includes a process of segmenting the target object in the cropped portion 180 of the second image using one or more segmentation models 165. Segmentation includes using one or more segmentation models 165 to generate an estimated segmentation boundary around the target object; and using one or more segmentation models 165 to output the cropped portion of the second image with an estimated segmentation boundary 190 around the target object. Segmentation may include evaluating the change in pixel or voxel intensity of each cropped portion to identify a group of edges and / or contours corresponding to the target object. When identifying the group of edges and / or contours, one or more segmentation models 165 generate an estimated segmentation boundary 190 of the target object. In some embodiments, the estimated segmentation boundary 190 corresponds to a three-dimensional representation of the target object. In some cases, segmentation further includes determining a probability score that the target object is present in the estimated segmentation boundary 190, and outputting the probability score and the estimated segmentation boundary 190.

[0110] The cropped portion of the second image with the estimated segmentation boundary 190 around the target object can be transmitted to the analysis controller 195 of the analysis stage 125. The analysis controller 195 includes a process of obtaining or receiving the cropped portion of the second image with the estimated segmentation boundary 190 around the target object (and optional probability score) and determining the analysis result 197 based on the estimated segmentation boundary 190 around the target object (and optional probability score). The analysis controller 195 may further include a process of determining the size, axial size, surface area and / or volume of the target object based on the estimated segmentation boundary 190 around the target object. In some cases, the estimated segmentation boundary 190 of the target object or its derived parameters (e.g., the size, axial size, volume, etc. of the target object) are further used to determine the diagnosis and / or prognosis of the subject. In other cases, the estimated segmentation boundary 190 of the target object is compared with the estimated segmentation boundary 190 of the same target object imaged at a previous time point to determine the therapeutic effect on the subject. For example, if the object of interest is a lesion, the estimated segmentation boundaries 190 of the subject's lesions can provide information about the cancer type (e.g., the location of the lesions), metastatic progression (e.g., if the number of lesions increases and / or if the number of lesion locations in the subject increases), and drug efficacy (e.g., if the number, size, and / or volume of lesions increases or decreases).

[0111] Although not explicitly shown, it should be understood that the computing environment 100 can also include a developer device associated with a developer. Communications from the developer device to components of the computing environment 100 can indicate the type of input images to be used for the models, the number and type of models to be used, the hyperparameters of each model (e.g., learning rate and number of hidden layers), the manner in which data requests are formatted, the training data to be used (e.g., and the manner in which access to the training data is obtained), and the verification techniques to be used and / or the manner in which the controller processing is to be configured.

[0112] IV.B. Exemplary Data Augmentation for Model Training

[0113] The second part of segmentation involves a second visual model (e.g., a deep learning neural network) constructed using a weighted loss function (e.g., Dice loss). The deep learning neural network is trained on images of one or more target objects from a subject. These images are generated by one or more medical imaging methods. However, the data set of images generated by certain medical imaging methods can be sparse. To address the sparsity of these images, the images of the training data are expanded to artificially increase the number and variety of images in the data set. More specifically, expansion can be performed by performing histogram matching to simulate other contrasts in other areas of the same or different subjects (e.g., areas of the subject where the target object may not be found) and increasing the variance of the training data set.

[0114] For example, each image of the training set or a subset of images from the training set can be histogram matched with one or more reference images to generate a new set of images, thereby artificially increasing the size of the training dataset in a process called data augmentation that reduces overfitting. Figure 2As shown, each image of the training set (left image) can be histogram matched with the reference image (center image) to generate a new image set (right image). Therefore, through histogram matching, the original training set or image subset is basically increased by 2 times in quantity and variety. Histogram matching is to transform the original image so that the histogram of the original image matches the reference histogram. Histogram matching is performed in the following manner: first use histogram equalization (for example, stretch the histogram to fill the dynamic range while trying to keep the histogram uniform) to equalize the original histogram and the reference histogram, and then map the original histogram to the reference histogram based on the equalized image and the conversion function. For example, assuming that the pixel intensity value 20 in the original image is mapped to 35 in the equalized image and assuming that the pixel intensity value 55 in the reference image is mapped to 35 in the equalized image, it can be determined that the pixel intensity value 20 in the original image should be mapped to the pixel intensity value 55 in the reference image. The original image can then be converted to a new image using a mapping from the original image to the equalized image to the reference image. In some cases, one or both datasets (i.e., the original image set and the new image set) can be further augmented using standard techniques such as rotation and flipping (e.g., rotating each image 90°, flipping left to right, flipping upside down, etc.) to further increase the number and variety of images available for training.

[0115] The benefits of using data augmentation based on histogram matching are: (i) the technique utilizes another dataset from the same species / instrument / image intensity, etc.; (ii) the masks (labels) corresponding to the histogram matched images are exactly the same as the original images in the training set; (iii) the number of images in the training set is increased by a multiple of the number of images used as reference; and (iv) the variance of the training set is increased, so the image structure of the training set is preserved; while the intensity of the pixels is changed, making the segmentation framework independent of the pixel intensity and dependent on the structure of the image and the object of interest.

[0116] IV.C. Exemplary 3D Deep Neural Network

[0117] exist Figure 3In the exemplary embodiment shown, the modified 3D U-Net 300 extracts features from an input image (e.g., a cropped portion of a second image) alone, detects a target object in the input image, generates a three-dimensional segmentation mask around the shape of the target object, and outputs the input image with the three-dimensional segmentation mask around the shape of the target object. The 3D U-Net 300 includes a contraction path 305 and an expansion path 310, which gives it a U-shaped architecture. The contraction path 305 is a CNN network that includes repeated applications of convolutions (e.g., 3×3×3 convolutions (unfilled convolutions)), each followed by a rectified linear unit (ReLU) and a maximum pooling operation (e.g., 2×2×2 maximum pooling with a stride of 2 in each direction) for downsampling. The input to the convolution operation is a three-dimensional volume (i.e., an input image of size n×n×channels, where n is the number of input features) and a set of "k" filters (also called kernels or feature extractors), each of which is of size (f×f×f channels, where f is an arbitrary number, such as 3 or 5). The output of the convolution operation is also a three-dimensional volume (also called output image or feature map) of size (m×m×k, where M is the number of output features and k is the convolution kernel size).

[0118] Each block 315 of the contraction path 315 includes one or more convolutional layers (indicated by gray horizontal arrows), and the number of feature channels varies, for example, from 1 to 64 (e.g., depending on the starting number of channels in the first process) as the convolution process increases the depth of the input image. The downward gray arrows between each block 315 are the max pooling processes, which halve the size of the input image. In each downsampling step or pooling operation, the number of feature channels may double. During the contraction process, the spatial information of the image data decreases while the feature information increases. Therefore, before pooling, the information that existed in, for example, a 572x572 image, after pooling, the (almost) same information now exists in, for example, a 284x284 image. Now, when the convolution operation is applied again in a subsequent process or layer, the filters in the subsequent process or layer will be able to see a larger context, i.e., as the input image goes deeper into the network, the size of the input image decreases, while the receptive field increases (the receptive field (context) is the area of ​​the input image covered by the kernel or filter at any given point in time). Once block 315 is executed, two more convolutions are performed, but without max pooling, in block 320. The image after block 320 has been resized to, for example, 28x28x1024 (this size is illustrative only, and the size at the end of process 320 may vary depending on the starting size of the input image - the size is nxnx channels).

[0119] The expansion path 310 is a CNN network that combines features from the contraction path 305 and spatial information (up-sampling of the feature map from the contraction path 305). As described herein, the output of the three-dimensional segmentation is not only a class label or bounding box parameters. Instead, the output (three-dimensional segmentation mask) is a complete image (e.g., a high-resolution image) in which all voxels are classified. If a conventional convolutional network with pooling layers and dense layers is used, the CNN network will lose the "where" information and only retain the "what" information that is unacceptable for image segmentation. In the case of image segmentation, both the "what" and "where" information are used. Therefore, the image is up-sampled to convert the low-resolution image to a high-resolution image, thereby recovering the "where" information. The transposed convolution represented by the upward white arrow is an exemplary up-sampling technique that can be used in the expansion path 310 to up-sample the feature map and expand the size of the image.

[0120] After the transposed convolution at block 325, the image is enlarged from 28×28×1024 to 56×56×512 by a 2×2×2 upconvolution with a stride of 2 in each dimension (upsampling operator), and then the image is concatenated with the corresponding image from the contracting path (see the horizontal gray bar 330 from the contracting path 305) to form an image of size 56×56×1024, for example. The reason for the concatenation is to combine the information from the previous layers (i.e., combining the high-resolution features from the contracting path 305 with the upsampled output from the expanding path 310) to obtain a more accurate prediction. The process continues as a series of upconvolutions that halve the number of channels, concatenation with the corresponding cropped feature maps from the contracting path 305, repeated application of each convolution followed by a convolution with a rectified linear unit (ReLU) (e.g., two 3×3×3 convolutions), and a final convolution in block 335 (e.g., one 1×1×1 convolution) to generate a multi-channel segmentation as a 3D segmentation mask. For localization, U-Net 300 uses the effective part of each convolution without any fully connected layers, i.e., the segmentation map only contains voxels for which the full context in the input image is available and uses skip connections that link the context features learned in the contraction block with the localization features learned in the expansion block.

[0121] In traditional neural networks with 3D U-Net architectures, a softmax function with a cross entropy loss is often used to compare the output and the true value label. While these networks show improved segmentation performance over traditional CNNs, they do not immediately translate to the small foreground objects, small sample sizes, and isotropic resolution found in medical imaging datasets. To address these and other issues, 3D U-Net 300 is constructed to include a reduced number of parameters and / or kernels relative to traditional 3D U-Nets (a total of 2,784 kernels and 19,069,955 learnable parameters, see e.g. A. Abdulkadir, S. Lienkamp, ​​T. Brox, and O. Ronneberger, "3D U-Net: learning dense volumetric segmentation from sparse annotation", in International conference on medical image computing and computer-assisted intervention, 2016: Springer, pp. 424-432). Specifically, the number of weights, the number of layers, and / or the overall width of the network are reduced to reduce network complexity and avoid the over-parameterization problem. In some cases, the number of kernels in each convolution and deconvolution layer is halved compared to the conventional 3D U-Net. The number of learnable parameters for each convolution and deconvolution layer of 3D U-Net 300 is then reduced, so that 3D U-Net 300 has 9,534,978 learnable parameters compared to the conventional 3D U-Net with 19,069,955 learnable parameters. In terms of kernels, the total number of kernels is reduced from 2,784 in the conventional 3D U-Net to 1,392 kernels in the 3D U-Net 300. In some cases, the total number of learnable parameters of the 3D U-Net 300 is reduced to between 5,000,000 and 12,000,000 learnable parameters. In some cases, the total number of kernels of the 3D U-Net 300 is reduced to between 800 and 1,700 kernels. The reduction in parameters and / or kernels is advantageous because it enables the model to more efficiently handle smaller sample sizes (i.e., the cropped portion of the second DTI parameter map).

[0122] In addition, 3D U-Net 300 is constructed to perform volume segmentation using a weighted loss function. Specifically, the metric used to evaluate segmentation performance is the Dice similarity coefficient (DSC, Formula 1). Therefore, in order to train 3D U-Net 300 with the goal of maximizing DSC, the DSC of all images is minimized (Formula 2). In addition, since the background and the target object are unevenly distributed in the foreground of the volume image, a weighted loss function is used, which is referred to as Dice loss in this article (Formula 3), in which the weight of the common background is reduced and the weight of the target object in the foreground is increased to balance the impact of the foreground and background voxels on the loss.

[0123]

[0124]

[0125]

[0126] Where N is the number of images, p i represents the predicted mask, q i Represents the true value mask corresponding to the target object.

[0127] It should be understood by those skilled in the art that 3D U-Net 300 need not be incorporated into Figure 1 The overall computing environment 100 is used to implement the object segmentation according to aspects of the present disclosure. Instead, various types of models can be used for object segmentation (e.g., CNN, Resnet, typical U-Net, V-Net, SSD network or recursive neural network RNN, etc.), as long as the type of model can be learned for object segmentation of medical images.

[0128] V. Volume Segmentation Techniques

[0129] Figure 4 A flowchart showing an exemplary process 400 for segmenting an instance of a target object using the multi-stage segmentation network is shown. Figures 1 to 3 One or more computing systems, models, and networks described herein may be used to perform process 400.

[0130] Process 400 begins at block 405, where a medical image of a subject is acquired. The medical image may show a head region, a chest region, an abdomen region, a pelvic region, and / or a region corresponding to a limb of the subject. The medical image is generated using one or more medical imaging modalities. For example, a user may operate one or more imaging systems using one or more medical imaging modalities to generate a medical image, as described in Section IV with respect to Figure 1 described.

[0131] At box 410, a medical image of the subject is obtained. For example, the medical image acquired in step 405 can be retrieved from a data storage device or one or more medical imaging systems. The medical image includes a first image having a first feature and a second image having a second feature. In some embodiments, the image is a DTI parameter map including a first measurement map (the first image having the first feature) and a second measurement map (the second image having the second feature). The first measurement map is different from the second measurement map. The DTI parameter map is generated by applying a supplemental MR gradient during the acquisition of the MR image. For example, a user may input parameters of one or more diffusion gradients into the imaging system, and the DTI parameter map is generated by applying a supplemental MR gradient based on the parameters of one or more diffusion gradients during the acquisition of the MR image (during the application of the gradient, the movement of protons affects the signal in the image). In some cases, the diffusion gradient is applied in more than one direction during the acquisition of the MR image. In some cases, the first measurement map is a partial anisotropy index map, and the second measurement map is a mean diffusivity map.

[0132] At box 415, the object in the first image is located and classified using the positioning model. Classification assigns the set of pixels or voxels of the first image to one or more of a plurality of object classes. Object classes may include classes corresponding to the target object (e.g., depending on the type of the target object), one or more classes corresponding to different biological structures, one or more classes corresponding to different organs and / or one or more classes corresponding to different tissues. For example, if the target object is a lesion, the object class may be defined as being used to identify lesions, blood vessels and / or organs. Positioning and classification may be performed by the positioning model using one or more clustering algorithms, which assign pixels or voxel sets to one or more object classes in a plurality of object classes. In some cases, the one or more clustering algorithms include a k-means algorithm, which assigns observed values ​​to clusters associated with a plurality of object classes. In some cases, the one or more clustering algorithms further include an expectation maximization algorithm. The expectation maximization algorithm calculates the probability of cluster members based on one or more probability distributions. The k-means algorithm may be used to initialize the expectation maximization algorithm by estimating the initial parameters of each object class in a plurality of object classes.

[0133] At frame 420, a localization model is used to determine a bounding box or a segmentation mask of a destination object in the first image. The bounding box or the segmentation mask of the destination object is determined based on a set of pixels or voxels that are assigned to the object class in the multiple object classes. For determining the segmentation mask, a pixel set that is assigned to the object class corresponding to the destination object is used to identify the seed position of the destination object. The identified seed position is projected to the z axis to increase the seed position and determine the segmentation mask. The z axis represents depth, and increases the seed position to fill the entire space of the object mask on the third or final dimension. In some cases, morphological closure and filling are performed in addition on the segmentation mask.

[0134] At box 425, a bounding box or segmentation mask is transferred to the second image to define the portion of the second image that includes the target object. Transferring the object mask includes projecting the bounding box or segmentation mask to the corresponding area of ​​the image (the portion of the second image that includes the target object) along the slice direction and / or superimposing the bounding box or segmentation mask to the corresponding area of ​​the second image (the portion of the second image that includes the target object). In some cases, the segmentation mask is projected onto the two-dimensional slice of the second image so that the boundary of the segmentation mask within the two-dimensional slice can be used to define a rectangular bounding box that surrounds the target area corresponding to the detected target object. In some cases, the bounding box includes additional padding (e.g., filling 5 pixels, 10 pixels, 15 pixels, etc.) to each edge of the perimeter of the segmentation mask to ensure that the entire target area is surrounded. In some cases, the second image is cropped based on the bounding box or segmentation mask plus an optional margin to generate a portion of the second image. The size of each cropped portion can be further adjusted (e.g., using additional padding) to maintain a uniform size.

[0135] In some embodiments, a portion of the second image is transmitted to a deep super-resolution neural network for pre-processing. The deep super-resolution neural network can be, for example, a convolutional neural network, a residual neural network, an attention-based neural network, and / or a recursive convolutional neural network. The deep super-resolution neural network processes the transmitted portion of the second image to improve the image spatial resolution of the portion of the second image (e.g., amplification and / or refinement of image details).

[0136] At box 430, a portion of the second image is input into a three-dimensional neural network model that is constructed for volume segmentation using a weighted loss function (e.g., a modified 3D U-Net model). In some cases, the weighted loss function is a weighted Dice loss function. The three-dimensional neural network model includes multiple parameters trained using a training data set. The number of the multiple model parameters may be reduced relative to a standard three-dimensional U-Net architecture. The training data set may include: multiple images with annotations associated with a segmentation boundary around a target object; and multiple additional images with annotations associated with a segmentation boundary around a target object. In some cases, the multiple additional images are artificially generated by matching image histograms from multiple medical images with image histograms from multiple reference images (e.g., images obtained from other regions of a subject). The multiple model parameters are identified using the training data set based on minimizing the weighted loss function. The three-dimensional neural network model further includes multiple kernels, and the number of kernels may be reduced relative to a standard three-dimensional U-Net architecture.

[0137] At box 435, the three-dimensional neural network model segments the portion of the second image. The segmentation includes generating an estimated segmentation boundary of the target object using the identified features. For example, the segmentation may include evaluating features (such as intensity changes of each cropped portion) to identify a set of edges and / or contours corresponding to the target object, and generating an estimated segmentation boundary of the target object using the identified set of edges and / or contours. The estimated segmentation boundary may represent the three-dimensional perimeter of the target object. In some cases, the three-dimensional neural network model may also determine the classification of the target object. For example, an object corresponding to a lesion may be classified based on its type or location in the subject, such as, for example, a lung lesion, a liver lesion, and / or a pancreatic lesion. As another example, an object corresponding to an organ and / or tissue may be classified as healthy, inflamed, fibrotic, necrotic, and / or cast filled.

[0138] At block 440, a portion of the second image having an estimated segmentation boundary around the target object is output. In some cases, the portion of the second image is provided. For example, the portion of the second image may be stored in a storage device and / or displayed on a user device.

[0139] At optional frame 445 place, take action based on the segmentation boundary of the estimation around the target object.In some cases, this action comprises determining the size, surface area and / or volume of the target object based on the segmentation boundary of the estimation around the target object.In some cases, provide (i) second image with the part of the segmentation boundary of the estimation around the target object and / or the size, surface area and / or volume of (ii) target object.For example, (i) second image with the part of the segmentation boundary of the estimation around the target object and / or the size, surface area and / or volume of (ii) target object can be stored in the storage device and / or displayed on the user device.The user can receive or obtain (i) second image with the part of the segmentation boundary of the estimation around the target object and / or the size, surface area and / or volume of (ii) target object.In other cases, (i) second image with the part of the segmentation boundary of the estimation around the target object and / or the size, surface area and / or volume of (ii) target object is used for quantitative analysis image measurement, such as image intensity. For example, in PET there is the standardized uptake value (SUV), or in MRI there are diffusion rate, T2, T1, etc., which are related to certain image metrics such as image intensity, so quantitative analysis of image metrics can be used to determine values / metrics such as SUV for specific purposes.

[0140] In some cases, the action includes using the following information to determine the diagnosis of the subject: (i) a portion of the second image with an estimated segmentation boundary around the target object and / or (ii) the size, surface area and / or volume of the target object. In some cases, the action includes administering treatment (e.g., administering to the subject) with a compound by a user based on (i) a portion of the second image with an estimated segmentation boundary around the target object, (ii) the size, surface area and / or volume of the target object and / or (iii) the diagnosis of the subject. In other cases, the action includes determining a treatment plan based on (i) a portion of the second image with an estimated segmentation boundary around the target object, (ii) the size, surface area and / or volume of the target object and / or (iii) the diagnosis of the subject, so that the dose of the drug can be calculated based on the size, surface area and / or volume of the target object. In some cases, the action includes determining whether the treatment is effective or whether the dose of the drug needs to be adjusted based on comparing the size, surface area and / or volume of the target object corresponding to the first time point with the size, surface area and / or volume of the target object corresponding to the second time point.

[0141] VI. Examples

[0142]

[0046] Systems and methods implemented in various embodiments may be better understood by reference to the following examples.

[0143] VI.A. Example 1.—Kidney Segmentation

[0144] Kidney segmentation using 3D U-Net (localization with expectation-maximization).

[0145] VI. Ai Background Technology

[0146] In various diseases such as polycystic kidney disease, lupus nephritis, renal parenchymal disease, and renal transplant rejection, renal function and activity are highly dependent on kidney volume. Automated assessment of the kidneys through imaging can be used to determine the diagnosis, prognosis, and / or treatment plan of a subject. In vivo imaging modalities offer unique advantages and limitations. In particular, MRI does not have ionizing radiation, is operator-independent, and has good tissue contrast, providing information related to kidney segmentation and volume. Traditional methods have been used to assess the kidney more locally, such as manual tracking, stereology, or general image processing. These methods can be laborious or inconsistent. To address these issues, an integrated deep learning model is used to segment the kidney.

[0147] Deep learning segmentation networks have been used for semantic segmentation of large biological image datasets. Although these networks provide advanced performance, they have high computational cost and memory consumption, which limits their field of view and depth. Therefore, these networks are particularly problematic for the segmentation of small objects in limited images commonly found in MRI studies. MRI often includes a large field of view or background to avoid aliasing effects. When the background accounts for a large part, the network may not be optimally trained to segment the foreground object. Therefore, an alternative strategy is needed to reduce the parameters of large 3D segmentation networks, avoid overfitting, and improve network performance.

[0148] First, to address the problem of background effects, a derived MRI contrast mechanism (using DTI) was incorporated into the localization step before the learned segmentation. Second, the 3D U-Net was modified to reduce the number of parameters and the Dice loss function was incorporated for segmentation. Third, augmentation and MRI histogram matching were incorporated to increase the number of training data sets. In addition, in some cases, these techniques were applied to super-resolution images of the data set to determine whether the enhanced images could improve segmentation performance. These techniques were implemented on preclinical MRI using an animal model of lupus nephritis.

[0149] VI.A.ii. Animal Models and Data Acquisition

[0150] Fifteen female mice infected with Friend B virus were used in this study, of which eight were used in the lupus nephritis (LN) disease group and seven were used in the control group. Animals were imaged every 2 weeks starting at 13 weeks of age, for a total of 4 time points. At each time point, multiple MRI datasets were acquired for each animal. A total of 196 3D MR images were acquired for this study. All images were manually segmented by a single user. The kidneys were outlined slice by slice across the entire image volume using Amira (Thermo Fisher Scientific, Hillsboro, OR). During MR imaging, animals were anesthetized with isoflurane, breathing spontaneously, and maintained at 37°C. MRI was performed on a Bruker 7T (Billerica, MA) equipped with a volume transmission and cryogenic surface receiving coil. Custom in vivo scaffolds were constructed using 3D printing (Stratasys Dimension) to provide secure positioning of the brain and spine. MRI diffusion tensor imaging (single-shot EPI) was performed on a single local patch using the following parameters: TR = 4 s, TE = 42 ms, BW = 250 kHz, diffusion directions = 12. FOV = 22 × 22 mm 2 , encoding matrix = 110 × 110, slices = 15, image resolution = 200 × 200 μm 2 , slice thickness = 1 mm, acquisition time = 13 min. Diffusion tensor parameter maps were calculated, including: FA, MD, AD and RD. FA and MD images were used in an integrated semantic segmentation algorithm.

[0151] VI.A.iii. Phase 1: Localization using EM

[0152] The FA images were used for the localization step. The FA images were segmented using EM, and the model was heuristically initialized using K-means (12 classes). The approximate kidney vicinity was isolated using one of the tissue classes and used as the object detected. These parameters were used for the algorithm: number of iterations to converge = 7, Markov random field smoothing factor = 0.05.

[0153] VI.A.iv. Data Augmentation

[0154] The MD images were histogram matched with the mouse brain dataset to generate a new dataset ( Figure 2 ). Both datasets were rotated 90°, flipped left to right, and flipped upside down. Data augmentation was performed only on the training set to ensure that the network was validated on completely unseen data. The total number of datasets acquired was n=196. The training dataset was increased from n=180 to n=1800 by augmentation, while the test dataset was kept at n=16. Separate training and testing was performed per animal, where one animal was not used for testing at a time and the remaining animals were used for training.

[0155] VI.Av Stage 2: Deep Semantic Segmentation

[0156] The metric used to evaluate the segmentation performance is the Dice similarity coefficient (DSC, Formula 1). Therefore, in order to train the 3D U-Net with the goal of maximizing DSC, the DSC of all images is minimized (Formula 2). In addition, due to the imbalanced distribution of background and kidneys in volumetric images, a weighted loss function (Dice loss, Formula 3) is used. To mitigate the background effect, the EM segmentation mask is projected along the slice direction. The boundaries of each projected EM segmentation mask are used to define a rectangular box for object detection. The defined box is enlarged by 5 pixels on all sides to ensure coverage of the kidneys. The 3D U-Net is trained and tested on the MD images within the detected region. The same detected region is used for super-resolution images. Since the cropped objects based on the 2D projection mask have arbitrary sizes in the first two dimensions, the original resolution images of all cropped images are resized to 64×64×16, and the super-resolution images are resized to 64×64×64.

[0157] VI.A.vi. Super-resolution

[0158] The MD images were super-resolved in the through-plane direction to improve spatial resolution. The original matrix of 110 × 110 × 15 was resolved 5 × to obtain a resulting matrix of 110 × 110 × 75. The images were enhanced using a deep super-resolution neural network.

[0159] VI.A.vii. Results

[0160] Figure 5A The six elements of the diffusion tensor are shown. Diffusion contrast changes are most pronounced in the intramedullary and extramedullary regions. Contrast changes are diagonal (D xx ,D yy ,D zz ) and off-diagonal elements (D xy ,D xz ,D yz ). Therefore, there is no change in contrast in the cortex, resulting in very low FA ( Figure 5B ). This low FA allows the kidney to be segmented from the background. MR images are super-resolved in the through-plane direction to improve spatial resolution, e.g. Figure 5C The improvement is most evident in the sagittal and coronal directions. In-plane resolution is minimally affected, as shown in the axial slice ( Figure 5C ). Fig. 6A The results of training a 3D U-Net on MD images without any preprocessing are shown. The DSC plot shows a uniform distribution with a mean of 0.49. Figure 6B In the figure, the abdomen region is detected as foreground by connected component analysis and cropped using the MD image. The DSC plot shows a normal distribution with a mean of 0.52. Figure 6C Results obtained using only EM segmentation are shown. The average DSC obtained was 0.65. Fig.6D Results of a representative ensemble strategy: First, kidneys were detected using EM segmentation on FA images, and then 3D U-Net was trained on the kidney regions detected in MD images. The average DSC of this method was 0.88. DSC plot of semantic segmentation using super-resolution MD images ( Fig. 6E ) is very similar to the semantic segmentation map at the original resolution ( Figure 3 D). Here, the mean DSC was 0.86. Table 1 summarizes the results with other comparative measures such as volume difference (VD) and positive predictive value (PPV).

[0161] Table 1. Means and standard deviations of segmentation results obtained using DSC, VD, and PPV. The best values ​​for each method are shown in bold.

[0162]

[0163] VI.A.viii. Discussion and Conclusion

[0164] This example demonstrates the integration of EM-based localization and 3D U-Net for kidney segmentation. The localization step leads to significantly improved results for deep learning methods. It is also demonstrated that while EM segmentation leads to improved performance for deep learning, the performance is poor when using EM segmentation alone. The EM segmentation method separates the kidney in the center slice, but does not retain a joint representation of the kidney volume. Therefore, the center slice is used for rectangular objects detected in all slices of the entire volume. Weighted Dice loss can be significant for error minimization and balance between objects and background. However, in the absence of the localization step, it was found that the performance was not significantly improved when the weighted Dice loss was included. Therefore, the background contains objects and organs that look similar to the kidney and cannot be distinguished when using 3D U-Net alone.

[0165] The method introduced in this example reduces the background effect and reduces the complexity of the data. Therefore, the complexity of the network can be reduced by reducing the number of kernels in each convolutional layer by at least half. In this study, using a limited MRI dataset (n=196), a DSC of 0.88 was obtained.

[0166] VII. Other considerations

[0167] Some embodiments of the present disclosure include a system including one or more data processors. In some embodiments, the system includes a non-transitory computer-readable storage medium containing instructions, which, when executed on the one or more data processors, cause the one or more data processors to perform part or all of one or more methods disclosed herein and / or part or all of one or more processes disclosed herein. Some embodiments of the present disclosure include a computer program product tangibly embodied in a non-transitory machine-readable storage medium, which includes instructions configured to cause one or more data processors to perform part or all of one or more methods disclosed herein and / or part or all of one or more processes disclosed herein.

[0168] The terms and expressions that have been adopted are used as terms of description rather than limitation, and when using these terms and expressions, there is no intention to exclude any equivalents of the features shown and described or parts thereof, but it should be recognized that various modifications are possible within the scope of the invention claimed. Therefore, it should be understood that although the claimed invention has been specifically disclosed through embodiments and optional features, modifications and variations of the concepts disclosed herein may be adopted by those skilled in the art, and such modifications and variations are considered to be within the scope of the invention defined by the appended claims.

[0169] The following description provides only preferred exemplary embodiments and is not intended to limit the scope, applicability or configuration of the present disclosure. Instead, the following description of the preferred exemplary embodiments will provide a feasible description for implementing various embodiments for those skilled in the art. It should be understood that various changes may be made to the function and arrangement of the elements without departing from the spirit and scope set forth in the appended claims.

[0170] Specific details are given in the following description to provide a thorough understanding of the embodiments. However, it should be understood that the embodiments can be practiced without these specific details. For example, circuits, systems, networks, processes, and other components can be shown as parts in block diagram form to avoid confusing the embodiments in unnecessary details. In other cases, well-known circuits, processes, algorithms, structures, and techniques can be shown without unnecessary details to avoid confusing the embodiments.

Claims

1. A method for segmenting an object in a medical image, comprising: obtaining a medical image of a subject, the medical image comprising a first image having a first characteristic and a second image having a second characteristic, wherein the medical image is generated using one or more medical imaging modalities; localizing and classifying objects within the first image into a plurality of object classes using a localization model, wherein the classification assigns a set of pixels or voxels of the first image to one or more of the plurality of object classes, wherein locating and classifying objects within the first image comprises applying one or more clustering algorithms to a plurality of pixels or voxels of the first image, wherein the one or more clustering algorithms include a k-means algorithm that assigns observations to clusters associated with the plurality of object categories, wherein the one or more clustering algorithms further include an expectation-maximization algorithm that calculates probabilities of cluster membership based on one or more probability distributions, and wherein the k-means algorithm initializes the expectation-maximization algorithm by estimating initial parameters for each of the plurality of object categories; determining, using the localization model, a bounding box or segmentation mask for an object of interest within the first image based on a set of pixels or voxels assigned an object class from the plurality of object classes; transferring the bounding box or the segmentation mask to the second image to define a portion of the second image that includes the target object; inputting the portion of the second image into a three-dimensional neural network model constructed for volume segmentation using a weighted loss function; generating an estimated segmentation boundary around the object of interest using the three-dimensional neural network model, The three-dimensional neural network model includes a plurality of model parameters identified using a training data set, the training data set including: a plurality of medical images having annotations associated with segmented boundaries surrounding an object of interest; as well as a plurality of additional medical images having annotations associated with segmentation boundaries around the object of interest, wherein the plurality of additional medical images are artificially generated by matching image histograms from the plurality of medical images with image histograms from a plurality of reference images; and wherein the plurality of model parameters are identified using the training data set based on minimizing the weighted loss function, The weighted loss function is a weighted Dice loss function; and The portion of the second image having the estimated segmentation boundary surrounding the object of interest is output using the three-dimensional neural network.

2. The method of claim 1 , wherein the one or more medical imaging modalities include a first medical imaging modality and a second medical imaging modality different from the first medical imaging modality, and wherein the first image is generated by the first medical imaging modality and the second image is generated by the second medical imaging modality.

3. The method of claim 1 , wherein the one or more medical imaging modalities include a first medical imaging modality and a second medical imaging modality that is the same as the first medical imaging modality, and wherein the first image is generated by the first medical imaging modality and the second image is generated by the second medical imaging modality. 4 . The method of claim 1 , wherein the first image is an image of a first type and the second image is an image of a second type, and wherein the image of the first type is different from the image of the second type. 5 . The method of claim 1 , wherein the first image is an image of a first type and the second image is an image of a second type, and wherein the image of the first type is the same as the image of the second type. The method of claim 1 , wherein the first characteristic is different from the second characteristic. The method of claim 1 , wherein the first characteristic is the same as the second characteristic.

8. The method of claim 2, wherein the first medical imaging modality is magnetic resonance imaging, diffusion tensor imaging, computed tomography, positron emission tomography, photoacoustic tomography, X-ray, ultrasound scanning, or a combination thereof, and wherein the second medical imaging modality is magnetic resonance imaging, diffusion tensor imaging, computed tomography, positron emission tomography, photoacoustic tomography, X-ray, ultrasound scanning, or a combination thereof.

9. The method of claim 4, wherein the first type of image is a magnetic resonance image, a diffusion tensor image or map, a computed tomography image, a positron emission tomography image, a photoacoustic tomography image, an X-ray image, an ultrasound scan image, or a combination thereof, and wherein the second type of image is a magnetic resonance image, a diffusion tensor image or map, a computed tomography image, a positron emission tomography image, a photoacoustic tomography image, an X-ray image, an ultrasound scan image, or a combination thereof.

10. The method of claim 6, wherein the first feature is fractional anisotropy contrast, mean diffusivity contrast, axial diffusivity contrast, radial diffusivity contrast, proton density contrast, T1 relaxation time contrast, T2 relaxation time contrast, diffusion coefficient contrast, low resolution, high resolution, drug contrast, radiotracer contrast, light absorption contrast, echo distance contrast, or a combination thereof, and wherein the second feature is fractional anisotropy contrast, mean diffusivity contrast, axial diffusivity contrast, radial diffusivity contrast, proton density contrast, T1 relaxation time contrast, T2 relaxation time contrast, diffusion coefficient contrast, low resolution, high resolution, drug contrast, radiotracer contrast, light absorption contrast, echo distance contrast, or a combination thereof.

11. The method of claim 1, wherein the one or more medical imaging modalities are diffusion tensor imaging, the first image is a fractional anisotropy (FA) map, the second image is a mean diffusion rate (MD) map, the first feature is fractional anisotropy contrast, the second feature is mean diffusion rate contrast, and the object of interest is a kidney of the subject.

12. The method of claim 1 , wherein determining the segmentation mask comprises: identifying a seed location of the target object using the set of pixels or voxels assigned the object class; growing the seed position by projecting it onto a z-axis representing the depth of the segmentation mask; as well as The segmentation mask is determined based on the projected seed positions. The method of claim 12 , wherein determining the segmentation mask further comprises performing morphological closing and filling on the segmentation mask.

14. The method of claim 12, further comprising cropping the second image based on the object mask plus a margin before inputting the portion of the second image into the three-dimensional neural network model to generate the portion of the second image.

15. The method of claim 1, further comprising inputting the second image into a deep super-resolution neural network before inputting the portion of the second image into the three-dimensional neural network model to increase the resolution of the portion of the second image.

16. The method according to claim 1, wherein the three-dimensional neural network model is a modified 3D U-Net model.

17. The method of claim 16, wherein the modified 3D U-Net model comprises a total number of between 5,000,000 and 12,000,000 learnable parameters.

18. The method of claim 16, wherein the modified 3D U-Net model includes a total number of between 800 and 1,700 kernels.

19. The method of claim 1, further comprising: determining a size, surface area, and / or volume of the target object based on the estimated boundary surrounding the target object; as well as Providing: (i) the portion of the second image having the estimated segmentation boundary around the target object and / or (ii) a size, surface area and / or volume of the target object.

20. The method of claim 19, further comprising: A diagnosis of the subject is determined by a user based on (i) the portion of the second image having the estimated segmentation boundary surrounding the object of interest and / or (ii) a size, surface area, and / or volume of the object of interest.

21. The method of claim 1, further comprising: acquiring, by a user, the medical image of the subject using an imaging system, wherein the imaging system uses the one or more medical imaging modalities to generate the medical image; determining a size, surface area, and / or volume of the target object based on the estimated segmentation boundary surrounding the target object; providing: (i) the portion of the second image having the estimated segmentation boundary around the target object and / or (ii) the size, surface area and / or volume of the target object; receiving, by the user, (i) the portion of the second image having the estimated segmentation boundary surrounding the target object and / or (ii) the size, surface area and / or volume of the target object; as well as A diagnosis of the subject is determined by the user based on (i) the portion of the second image having the estimated segmentation boundary surrounding the target object and / or (ii) a size, surface area and / or volume of the target object.

22. The method of claim 21, further comprising administering treatment with a compound by the user based on (i) the portion of the second image having the estimated segmentation boundary surrounding the object of interest, (ii) the size, surface area and / or volume of the object of interest, and / or (iii) the diagnosis of the subject.

23. A system for segmenting an object in a medical image, comprising: one or more data processors; as well as A non-transitory computer-readable storage medium comprising instructions that, when executed on the one or more data processors, cause the one or more data processors to perform actions including: obtaining a medical image of a subject, the medical image comprising a first image having a first characteristic and a second image having a second characteristic, wherein the medical image is generated using one or more medical imaging modalities; localizing and classifying objects within the first image into a plurality of object classes using a localization model, wherein the classification assigns a set of pixels or voxels of the first image to one or more of the plurality of object classes, wherein locating and classifying objects within the first image comprises applying one or more clustering algorithms to a plurality of pixels or voxels of the first image, wherein the one or more clustering algorithms include a k-means algorithm that assigns observations to clusters associated with the plurality of object categories, wherein the one or more clustering algorithms further include an expectation-maximization algorithm that calculates probabilities of cluster membership based on one or more probability distributions, and wherein the k-means algorithm initializes the expectation-maximization algorithm by estimating initial parameters for each of the plurality of object categories; determining, using the localization model, a bounding box or segmentation mask for an object of interest within the first image based on a set of pixels or voxels assigned an object class from the plurality of object classes; transferring the bounding box or the segmentation mask to the second image to define a portion of the second image that includes the target object; inputting the portion of the second image into a three-dimensional neural network model constructed for volume segmentation using a weighted loss function; generating an estimated segmentation boundary around the object of interest using the three-dimensional neural network model, The three-dimensional neural network model includes a plurality of model parameters identified using a training data set, the training data set including: a plurality of medical images having annotations associated with segmented boundaries surrounding an object of interest; as well as a plurality of additional medical images having annotations associated with segmentation boundaries around the object of interest, wherein the plurality of additional medical images are artificially generated by matching image histograms from the plurality of medical images with image histograms from a plurality of reference images; and wherein the plurality of model parameters are identified using the training data set based on minimizing the weighted loss function, The weighted loss function is a weighted Dice loss function; and The portion of the second image having the estimated segmentation boundary surrounding the object of interest is output using the three-dimensional neural network.

24. The system of claim 23, wherein the one or more medical imaging modalities include a first medical imaging modality and a second medical imaging modality different from the first medical imaging modality, and wherein the first image is generated by the first medical imaging modality and the second image is generated by the second medical imaging modality.

25. The system of claim 23, wherein the one or more medical imaging modalities include a first medical imaging modality and a second medical imaging modality that is the same as the first medical imaging modality, and wherein the first image is generated by the first medical imaging modality and the second image is generated by the second medical imaging modality.

26. The system of claim 23, wherein the first image is an image of a first type and the second image is an image of a second type, and wherein the image of the first type is different from the image of the second type.

27. The system of claim 23, wherein the first image is an image of a first type and the second image is an image of a second type, and wherein the image of the first type is the same as the image of the second type.

28. The system of claim 23, wherein the first characteristic is different from the second characteristic.

29. The system of claim 23, wherein the first characteristic is the same as the second characteristic.

30. The system of claim 24, wherein the first medical imaging modality is magnetic resonance imaging, diffusion tensor imaging, computed tomography, positron emission tomography, photoacoustic tomography, X-ray, ultrasound scanning, or a combination thereof, and wherein the second medical imaging modality is magnetic resonance imaging, diffusion tensor imaging, computed tomography, positron emission tomography, photoacoustic tomography, X-ray, ultrasound scanning, or a combination thereof.

31. The system of claim 26, wherein the first type of image is a magnetic resonance image, a diffusion tensor image or map, a computed tomography image, a positron emission tomography image, a photoacoustic tomography image, an X-ray image, an ultrasound scan image, or a combination thereof, and wherein the second type of image is a magnetic resonance image, a diffusion tensor image or map, a computed tomography image, a positron emission tomography image, a photoacoustic tomography image, an X-ray image, an ultrasound scan image, or a combination thereof.

32. The system of claim 28, wherein the first feature is fractional anisotropy contrast, mean diffusivity contrast, axial diffusivity contrast, radial diffusivity contrast, proton density contrast, T1 relaxation time contrast, T2 relaxation time contrast, diffusion coefficient contrast, low resolution, high resolution, drug contrast, radiotracer contrast, light absorption contrast, echo distance contrast, or a combination thereof, and wherein the second feature is fractional anisotropy contrast, mean diffusivity contrast, axial diffusivity contrast, radial diffusivity contrast, proton density contrast, T1 relaxation time contrast, T2 relaxation time contrast, diffusion coefficient contrast, low resolution, high resolution, drug contrast, radiotracer contrast, light absorption contrast, echo distance contrast, or a combination thereof.

33. The system of claim 23, wherein the one or more medical imaging modalities are diffusion tensor imaging, the first image is a fractional anisotropy (FA) map, the second image is a mean diffusion rate (MD) map, the first feature is fractional anisotropy contrast, the second feature is mean diffusion rate contrast, and the object of interest is a kidney of the subject.

34. The system of claim 23, wherein determining the segmentation mask, and determining the segmentation mask comprises: identifying a seed location of the target object using the set of pixels or voxels assigned the object class; growing the seed position by projecting it onto a z-axis representing the depth of the segmentation mask; as well as The segmentation mask is determined based on the projected seed positions.

35. The system of claim 34, wherein determining the segmentation mask further comprises performing morphological closing and filling on the segmentation mask.

36. The system of claim 34, wherein the action further comprises cropping the second image based on the object mask plus a margin before inputting the portion of the second image into the three-dimensional neural network model to generate the portion of the second image.

37. The system of claim 23, wherein the action further comprises inputting the second image into a deep super-resolution neural network before inputting the portion of the second image into the three-dimensional neural network model to increase the resolution of the portion of the second image.

38. The system of claim 23, wherein the three-dimensional neural network model is a modified 3D U-Net model.

39. The system of claim 38, wherein the modified 3D U-Net model comprises a total of between 5,000,000 and 12,000,000 learnable parameters.

40. The system of claim 38, wherein the modified 3D U-Net model includes a total number of between 800 and 1,700 kernels.

41. The system of claim 23, wherein the actions further comprise: determining a size, surface area, and / or volume of the target object based on the estimated boundary surrounding the target object; as well as Providing: (i) the portion of the second image having the estimated segmentation boundary around the target object and / or (ii) a size, surface area and / or volume of the target object.

42. The system of claim 41, wherein the actions further comprise: A diagnosis of the subject is determined by a user based on (i) the portion of the second image having the estimated segmentation boundary surrounding the object of interest and / or (ii) a size, surface area, and / or volume of the object of interest.

43. The system of claim 23, wherein the actions further comprise: acquiring, by a user, the medical image of the subject using an imaging system, wherein the imaging system uses the one or more medical imaging modalities to generate the medical image; determining a size, surface area, and / or volume of the target object based on the estimated segmentation boundary surrounding the target object; providing: (i) the portion of the second image having the estimated segmentation boundary around the target object and / or (ii) the size, surface area and / or volume of the target object; receiving, by the user, (i) the portion of the second image having the estimated segmentation boundary surrounding the target object and / or (ii) the size, surface area and / or volume of the target object; as well as A diagnosis of the subject is determined by the user based on (i) the portion of the second image having the estimated segmentation boundary surrounding the target object and / or (ii) a size, surface area and / or volume of the target object.

44. The system of claim 42, wherein the actions further comprise administering, by the user, treatment with a compound based on (i) the portion of the second image having the estimated segmentation boundary surrounding the object of interest, (ii) the size, surface area and / or volume of the object of interest, and / or (iii) the diagnosis of the subject.

45. A computer program product tangibly embodied in a non-transitory machine-readable storage medium, comprising instructions configured to cause one or more data processors to perform actions comprising: obtaining a medical image of a subject, the medical image comprising a first image having a first characteristic and a second image having a second characteristic, wherein the medical image is generated using one or more medical imaging modalities; localizing and classifying objects within the first image into a plurality of object classes using a localization model, wherein the classification assigns a set of pixels or voxels of the first image to one or more of the plurality of object classes, wherein locating and classifying objects within the first image comprises applying one or more clustering algorithms to a plurality of pixels or voxels of the first image, wherein the one or more clustering algorithms include a k-means algorithm that assigns observations to clusters associated with the plurality of object categories, wherein the one or more clustering algorithms further include an expectation-maximization algorithm that calculates probabilities of cluster membership based on one or more probability distributions, and wherein the k-means algorithm initializes the expectation-maximization algorithm by estimating initial parameters for each of the plurality of object categories; determining, using the localization model, a bounding box or segmentation mask for an object of interest within the first image based on a set of pixels or voxels assigned an object class from the plurality of object classes; transferring the bounding box or the segmentation mask to the second image to define a portion of the second image that includes the target object; inputting the portion of the second image into a three-dimensional neural network model constructed for volume segmentation using a weighted loss function; generating an estimated segmentation boundary around the object of interest using the three-dimensional neural network model, The three-dimensional neural network model includes a plurality of model parameters identified using a training data set, the training data set including: a plurality of medical images having annotations associated with segmented boundaries surrounding an object of interest; as well as a plurality of additional medical images having annotations associated with segmentation boundaries around the object of interest, wherein the plurality of additional medical images are artificially generated by matching image histograms from the plurality of medical images with image histograms from a plurality of reference images; and wherein the plurality of model parameters are identified using the training data set based on minimizing the weighted loss function, The weighted loss function is a weighted Dice loss function; and The portion of the second image having the estimated segmentation boundary surrounding the object of interest is output using the three-dimensional neural network.

46. ​​A computer program product according to claim 45, wherein the one or more medical imaging modalities include a first medical imaging modality and a second medical imaging modality different from the first medical imaging modality, and wherein the first image is generated by the first medical imaging modality and the second image is generated by the second medical imaging modality.

47. A computer program product according to claim 45, wherein the one or more medical imaging modalities include a first medical imaging modality and a second medical imaging modality that is the same as the first medical imaging modality, and wherein the first image is generated by the first medical imaging modality and the second image is generated by the second medical imaging modality.

48. The computer program product of claim 45, wherein the first image is an image of a first type and the second image is an image of a second type, and wherein the image of the first type is different from the image of the second type.

49. The computer program product of claim 45, wherein the first image is an image of a first type and the second image is an image of a second type, and wherein the image of the first type is the same as the image of the second type.

50. The computer program product of claim 45, wherein the first characteristic is different from the second characteristic.

51. The computer program product of claim 45, wherein the first characteristic is the same as the second characteristic.

52. The computer program product of claim 46, wherein the first medical imaging modality is magnetic resonance imaging, diffusion tensor imaging, computed tomography, positron emission tomography, photoacoustic tomography, X-ray, ultrasound scanning, or a combination thereof, and wherein the second medical imaging modality is magnetic resonance imaging, diffusion tensor imaging, computed tomography, positron emission tomography, photoacoustic tomography, X-ray, ultrasound scanning, or a combination thereof.

53. The computer program product of claim 48, wherein the first type of image is a magnetic resonance image, a diffusion tensor image or map, a computed tomography image, a positron emission tomography image, a photoacoustic tomography image, an X-ray image, an ultrasound scan image, or a combination thereof, and wherein the second type of image is a magnetic resonance image, a diffusion tensor image or map, a computed tomography image, a positron emission tomography image, a photoacoustic tomography image, an X-ray image, an ultrasound scan image, or a combination thereof.

54. The computer program product of claim 50, wherein the first feature is fractional anisotropy contrast, mean diffusivity contrast, axial diffusivity contrast, radial diffusivity contrast, proton density contrast, T1 relaxation time contrast, T2 relaxation time contrast, diffusion coefficient contrast, low resolution, high resolution, drug contrast, radiotracer contrast, light absorption contrast, echo distance contrast, or a combination thereof, and wherein the second feature is fractional anisotropy contrast, mean diffusivity contrast, axial diffusivity contrast, radial diffusivity contrast, proton density contrast, T1 relaxation time contrast, T2 relaxation time contrast, diffusion coefficient contrast, low resolution, high resolution, drug contrast, radiotracer contrast, light absorption contrast, echo distance contrast, or a combination thereof.

55. The computer program product of claim 45, wherein the one or more medical imaging modalities are diffusion tensor imaging, the first image is a fractional anisotropy (FA) map, the second image is a mean diffusion rate (MD) map, the first feature is fractional anisotropy contrast, the second feature is mean diffusion rate contrast, and the object of interest is a kidney of the subject.

56. The computer program product of claim 45, wherein determining the segmentation mask, and wherein determining the segmentation mask comprises: identifying a seed location of the target object using the set of pixels or voxels assigned the object class; growing the seed position by projecting it onto a z-axis representing the depth of the segmentation mask; as well as The segmentation mask is determined based on the projected seed positions.

57. The computer program product of claim 56, wherein determining the segmentation mask further comprises performing morphological closing and filling on the segmentation mask.

58. The computer program product of claim 56, wherein the actions further comprise cropping the second image based on the object mask plus a margin before inputting the portion of the second image into the three-dimensional neural network model to generate the portion of the second image.

59. The computer program product of claim 45, wherein the action further comprises inputting the second image into a deep super-resolution neural network before inputting the portion of the second image into the three-dimensional neural network model to increase the resolution of the portion of the second image.

60. The computer program product of claim 45, wherein the three-dimensional neural network model is a modified 3DU-Net model.

61. The computer program product of claim 60, wherein the modified 3D U-Net model comprises a total of between 5,000,000 and 12,000,000 learnable parameters.

62. The computer program product of claim 60, wherein the modified 3D U-Net model comprises a total number of between 800 and 1,700 kernels.

63. The computer program product of claim 45, wherein the actions further comprise: determining a size, surface area, and / or volume of the target object based on the estimated boundary surrounding the target object; as well as Providing: (i) the portion of the second image having the estimated segmentation boundary around the target object and / or (ii) a size, surface area and / or volume of the target object.

64. The computer program product of claim 63, wherein the actions further comprise: A diagnosis of the subject is determined by a user based on (i) the portion of the second image having the estimated segmentation boundary surrounding the object of interest and / or (ii) a size, surface area, and / or volume of the object of interest.

65. The computer program product of claim 45, wherein the actions further comprise: acquiring, by a user, the medical image of the subject using an imaging system, wherein the imaging system uses the one or more medical imaging modalities to generate the medical image; determining a size, surface area, and / or volume of the target object based on the estimated segmentation boundary surrounding the target object; providing: (i) the portion of the second image having the estimated segmentation boundary around the target object and / or (ii) the size, surface area and / or volume of the target object; receiving, by the user, (i) the portion of the second image having the estimated segmentation boundary surrounding the target object and / or (ii) the size, surface area and / or volume of the target object; as well as A diagnosis of the subject is determined by the user based on (i) the portion of the second image having the estimated segmentation boundary surrounding the target object and / or (ii) a size, surface area and / or volume of the target object.

66. The computer program product of claim 64, wherein the actions further comprise administering, by the user, treatment with a compound based on (i) the portion of the second image having the estimated segmentation boundary surrounding the object of interest, (ii) the size, surface area and / or volume of the object of interest, and / or (iii) the diagnosis of the subject.

Citation Information

Patent Citations

  • Symmetric brain tumor segmentation method based on neural network

    CN108447052A

  • Method for Extracting Airways and Pulmonary Lobes and Apparatus Therefor

    US20160189373A1

  • Deep Image-to-Image Network Learning for Medical Image Analysis

    US20180330207A1