Inference for cartilage segmentation in computed tomography
By combining training sets of CT and MR images with 3D transformation technology, the accuracy of soft tissue segmentation in CT images was solved, the accuracy of cartilage marking was improved, and the effectiveness of diagnosis and treatment planning was enhanced.
Patent Information
- Application Number
- CN202480077053.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-20
- Filing Date
- 2024-12-19
- Publication Date
- 2026-07-14
AI Technical Summary
Existing technologies struggle to accurately segment soft tissue structures, particularly cartilage, in computed tomography (CT) images, limiting the accuracy of diagnosis and treatment planning.
By training a neural network, a training set of soft tissue markers is generated by combining computed tomography (CT) images and magnetic resonance (MR) images. The accuracy of soft tissue markers is improved by applying 3D transformation and iterative closest point (ICP) registration.
This enables more accurate labeling of soft tissues, especially cartilage, in CT images, improving the accuracy of diagnosis and treatment planning and reducing reliance on ionizing radiation.
Smart Images

Figure CN122397036A_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application claims the benefit of U.S. Provisional Patent Application Serial No. 63 / 612,470, filed December 20, 2023, entitled “INFERENCE OF CARTILAGESEGMENTATION IN COMPUTED TOMOGRAPHY SCANS”. The contents of that provisional patent application are incorporated herein by reference as if copied entirely herein. Background Technology
[0003] Medical imaging of anatomical structures using computed tomography (CT) or magnetic resonance imaging (MRI) techniques can provide detailed information about the anatomical structures and can be used for diagnosis, modeling, and research.
[0004] Medical image segmentation refers to the process of distinguishing different anatomical structures (such as different soft tissue structures and different hard tissue structures) captured in a given medical image from each other and from the background. Segmentation can be performed manually by an experienced radiologist or other personnel while reviewing medical images, or it can be performed at least partially automatically by a computer system configured to segment medical images. Medical image segmentation is generally designed to result in each pixel of the medical image being labeled as either a part of one anatomical structure or as part of the background.
[0005] For example, in a medical image of a knee joint encoded as pixels, after successful segmentation, some pixels can be labeled as part of the tibia, some other pixels as part of the tibial cartilage, some other pixels as part of the femur, some other pixels as part of the femoral cartilage, and some other pixels can simply be labeled as the background. By distinguishing pixels based on the anatomical structures they belong to, the anatomical structures themselves can be distinguished from each other and clearly depicted in the medical image. Accurate depiction of anatomical structures in individual image slices of a set of medical images is useful for injury or pathological diagnosis, and also useful in generating accurate 3D models of anatomical structures and their interrelationships to aid in injury diagnosis, for repair planning, and for surgical guidance during repair.
[0006] Some medical imaging techniques, or "modalities," are better suited than others to capture the detail of certain types of anatomical structures. For example, MRI can be excellent at capturing the detail of soft tissue anatomy such as cartilage, and hard tissue anatomy such as bone. However, while CT is superb at capturing the detail of hard tissue anatomy, it is often much less effective at capturing the detail of soft tissue anatomy. Because CT is less efficient at capturing the detail of soft tissue anatomy, it can be difficult or impractical to fully segment CT images (whether manually or automatically) to accurately label soft tissue anatomy. Summary of the Invention
[0007] One example is a method for training a processor implementation of a neural network for soft tissue labeling, the method comprising: creating a training set of computed tomography (CT) images of soft tissue labeling, including: collecting multiple pairs of image sets for each of a plurality of individuals from at least one database, each pair of image sets comprising both a CT image set and a magnetic resonance (MR) image set of the individual's anatomical structure; and for each pair: using the MR image set to create a first 3D model of both bone and soft tissue of the individual's anatomical structure; using the CT image set to create a second 3D model of the bone of the individual's anatomical structure; generating a 3D transformation between the bone in the first 3D model and the bone in the second 3D model; labeling soft tissue in at least one CT image in the CT image set based on applying the 3D transformation to the labeled soft tissue in the first 3D model, thereby creating at least one CT image of soft tissue labeling; and augmenting the training set using the at least one CT image of soft tissue labeling; and using the training set to train a neural network to label soft tissue in the CT image.
[0008] In the method implemented by the example processor, labeling soft tissue in at least one CT image in a CT image set based on the labeled soft tissue in a first 3D model by applying a 3D transformation to the soft tissue in the first 3D model may include: applying a 3D transformation to the soft tissue in the first 3D model to label the contents of a second 3D model as soft tissue, thereby creating an enhanced second 3D model; and using the enhanced second 3D model to create a CT image with at least one soft tissue label.
[0009] In the method implemented by the example processor, generating a 3D transformation between bones in a first 3D model and bones in a second 3D model may include: applying a global registration process to determine an initial alignment transformation between bones in the first 3D model and bones in the second 3D model; and applying an iterative nearest point (ICP) process to the initial alignment transformation to generate the 3D transformation.
[0010] In the example processor implementation, applying a global registration process to determine the initial alignment transformation between bones in the first 3D model and bones in the second 3D model may include applying a 2PNS (2-Point-Normal Set) process to determine the initial alignment transformation between bones in the first 3D model and bones in the second 3D model.
[0011] In the example processor implementation, the neural network can be a U-net neural network.
[0012] In the method implemented by the example processor, creating a first 3D model of both bone and soft tissue of an individual anatomical structure using a set of MR images may include: receiving at least bone and soft tissue markers corresponding to MR images in the set of MR images; and applying a moving cube procedure to the bone and soft tissue markers in the MR images.
[0013] In the method implemented by the example processor, creating a second 3D model of the bones of an individual anatomical structure using a set of CT images may include: receiving bone markers corresponding to CT images in the set of CT images; and applying a moving cube process to the bone markers in the CT images.
[0014] In the method implemented by the example processor, the individual anatomical structure may be the knee joint, and the soft tissue in at least one CT image in the set of labeled CT images may include the femoral cartilage and tibial cartilage labeled the knee joint.
[0015] Another example is a non-transitory processor-readable medium containing processor-readable program code that can be executed by at least one processor to perform one or more of the processor-implemented methods described above.
[0016] Another example is a neural network trained according to one or more of the processor implementation methods described above.
[0017] In the example, the neural network could be a U-net neural network.
[0018] Another example is a processor implementation method for soft tissue labeling of computed tomography (CT) images. This method may include: providing an input CT image as input to the neural network described above; and receiving soft tissue labels from the input CT image from the neural network.
[0019] Another example is a system for labeling soft tissue in computed tomography (CT) images. The system may include: at least one processor; and a memory coupled to the at least one processor, the memory storing instructions that, when executed by the at least one processor, cause the at least one processor to: provide an input CT image as input to the neural network described above; and receive soft tissue labels from the input CT image from the neural network.
[0020] In one or more processor implementations described above, soft tissue can be cartilage.
[0021] Another example is a processor implementation method for training a neural network for soft tissue labeling. This processor implementation method may include: creating a training set of computed tomography (CT) images for soft tissue labeling, comprising: collecting multiple pairs of images for each of a plurality of individuals from at least one database, each pair including both a CT image and a magnetic resonance (MR) image of the individual's anatomy; and for each of the multiple pairs of images: generating a 2D transformation between bone in the MR image and bone in the CT image; labeling the soft tissue in the CT image based on the 2D transformation applied to the labeled soft tissue in the MR image, thereby creating a soft tissue-labeled CT image; and augmenting the training set with the soft tissue-labeled CT image; and using the training set to train a neural network to label soft tissue in the CT image.
[0022] In the processor-implemented method, labeling soft tissue in a CT image based on applying a 2D transformation to the soft tissue labeled in the MR image may include: applying a 2D transformation to the soft tissue in the MR image to label the contents of the CT image as soft tissue, thereby creating a soft tissue-labeled CT image.
[0023] In the processor-implemented method, generating a 2D transformation between bones in an MR image and bones in a CT image may include: applying a global registration process to determine an initial alignment transformation between bones in the MR image and bones in the CT image; and applying an iterative nearest point (ICP) process to the initial alignment transformation to generate the 2D transformation.
[0024] In the processor implementation method, the neural network can be a U-net neural network.
[0025] In the processor-implemented method, the individual anatomical structure can be the knee joint, and the soft tissues in the labeled CT images can include the femoral cartilage and tibial cartilage labeled with the knee joint.
[0026] Another example is a non-transitory processor-readable medium containing processor-readable program code that can be executed by at least one processor to perform one or more of the processor-implemented methods described above.
[0027] Another example is a neural network trained according to one or more of the processor implementation methods described above.
[0028] In one example, the neural network could be a U-net neural network.
[0029] Another example is a processor implementation method for soft tissue labeling of computed tomography (CT) images. The method may include: providing an input CT image as input to a neural network; and receiving soft tissue labels from the input CT image from the neural network.
[0030] Another example is a system for labeling soft tissue in computed tomography (CT) images. The system may include: at least one processor; and a memory coupled to the at least one processor, the memory storing instructions that, when executed by the at least one processor, cause the at least one processor to: provide an input CT image as input to a neural network; and receive soft tissue labels from the input CT image from the neural network.
[0031] In one or more processor implementations described above, soft tissue can be cartilage.
[0032] Another example is a processor implementation method for soft tissue labeling of input CT images of individual anatomical structures, the method comprising: generating corresponding output synthetic magnetic resonance (MR) images of the individual anatomical structure from the input CT images by a neural network, wherein the neural network is trained using a generative adversarial method and at least using CT training images and MR training images to receive CT images of the individual anatomical structure and generate corresponding output synthetic MR images of the individual anatomical structure; receiving at least soft tissue labels corresponding to the corresponding output synthetic MR images; and applying at least the soft tissue labels from the corresponding output synthetic MR images to the input CT images, thereby performing soft tissue labeling on the input CT images.
[0033] In one example, receiving at least the soft tissue markers corresponding to the output synthetic MR image may include: providing the output synthetic MR image to an automatic image segmentation system; and receiving at least the soft tissue markers corresponding to the output synthetic MR image from the automatic image segmentation system.
[0034] In one example, applying soft tissue labeling from the output synthetic MR image to the input CT image, such that soft tissue labeling of the input CT image may include resampling the segmented volume of the output synthetic MR image using the input CT image as a reference.
[0035] The method implemented by the processor may further include: configuring a neural network as a first generative adversarial network (GAN), the first GAN being interconnected with a second GAN, wherein the first GAN is trained to generate a synthetic output MR image from the first GAN input, and the second GAN is trained to generate a synthetic CT image from the second GAN input; and for each of a plurality of loops, training the neural network by: providing a synthetic CT image generated by the second GAN as input to the first GAN, thereby generating a corresponding synthetic MR image as output to the first GAN; providing a synthetic MR image generated by the first GAN as input to the second GAN, thereby generating a corresponding synthetic CT image as output to the second GAN; determining a first difference metric between the second GAN output and the first GAN input; determining a second difference metric between the first GAN output and the second GAN input; and modifying the first GAN and the second GAN at least based on the first difference metric and the second difference metric, wherein the modification is made to reduce the magnitude of the first difference metric and the second difference metric in successive loops.
[0036] In one example, the method implemented by the processor may further include: during the training of the neural network: receiving bone markers of synthetic CT images generated by a second GAN; and using the bone markers as regularization during modification.
[0037] In one example, the neural network is the U-net neural network.
[0038] In one example, the individual anatomical structure could be the knee joint, and the soft tissue markers could include markers for the femoral cartilage and the tibial cartilage.
[0039] Another example is a non-transitory processor-readable medium containing processor-readable program code that can be executed by at least one processor to perform one or more of the processor-implemented methods described above.
[0040] Another example is a neural network trained according to one or more of the processor implementation methods described above.
[0041] In one example, the neural network is the U-net neural network.
[0042] Another example is a processor implementation method for soft tissue labeling of computed tomography (CT) images. The method may include: providing an input CT image as input to a neural network; and receiving soft tissue labels from the input CT image from the neural network.
[0043] Another example is a system for labeling soft tissue in computed tomography (CT) images, the system comprising: at least one processor; and a memory coupled to the at least one processor, the memory storing instructions that, when executed by the at least one processor, cause the at least one processor to: provide an input CT image as input to a neural network; and receive soft tissue labels from the input CT image from the neural network.
[0044] In one example, soft tissue could be cartilage. Attached Figure Description
[0045] For a detailed description of the example implementation, reference will now be made to the accompanying drawings, in which:
[0046] Figure 1 shows an example MR image of the knee joint of an individual human patient;
[0047] Figure 1B It shows Figure 1A Example CT images of the same knee joint;
[0048] Figure 2A An example unsegmented MR image of the knee joint of an individual human patient is shown;
[0049] Figure 2B The segmented result is shown. Figure 2A MR images;
[0050] Figure 3A Example CT images of the knee joint of an individual human patient without contrast processing are shown;
[0051] Figure 3B This shows the contrast processing after using local windowing technology. Figure 3A Example CT images;
[0052] Figure 4A , Figure 4B and Figure 4C An overall example pipeline is shown that uses a neural network capable of generating synthetic MR images from input CT images;
[0053] Figure 5 Various CT and CT magnified / zoomed images are shown, illustrating a comparison of the input CT image with pseudo-benchmark truth segmentation and each of the three different automatic segments determined according to the examples disclosed herein;
[0054] Figure 6 A computer implementation of a method for training a neural network for soft tissue labeling, based on at least some examples, is shown;
[0055] Figure 7 An example is shown. Figure 6 The steps of the method performed during the creation of a training set of CT images with soft tissue markers;
[0056] Figure 8 The alternative example shown can be used to illustrate this. Figure 6 The steps of the method performed during the creation of a training set of CT images with soft tissue markers;
[0057] Figure 9 This paper illustrates a computer-implemented method for soft tissue labeling based on an example of an input CT image of an individual's anatomical structure; and
[0058] Figure 10 An example computer system is shown that can be configured to perform the processor implementation methods disclosed herein, and can be configured to implement the system disclosed herein.
[0059] definition
[0060] The various terms are used to refer to specific system components. Different companies may use different names to refer to components, and this document does not intend to distinguish components that differ in name rather than function. In the following discussion and in the claims, the terms "comprising" and "including" are used in an open-ended manner and should therefore be interpreted as meaning "including but not limited to...". Furthermore, the terms "couple" or "couples" are intended to indicate indirect or direct connections. Thus, if a first device is coupled to a second device, the connection can be either a direct connection or an indirect connection via other devices and connections.
[0061] This document provides examples of medical images describing and depicting the human knee joint, each of which comprises the tibia, femur, patella, and tibial and femoral cartilages. It should be understood that the images of the human knee joint are provided only as useful examples, and the principles described and depicted in this application are to be understood more broadly to relate to images of other anatomical structures, such as musculoskeletal structures, for which segmentation of soft and hard tissues in the medical images may be desired.
[0062] In this specification, the term "soft tissue" is understood to mean what is commonly used by one of ordinary skill in the art to encompass a variety of anatomical structures typically distinguished from anatomical structures such as bones and teeth, since soft tissue does not harden due to ossification or calcification. Examples of different types of soft tissue may include cartilage, muscle, tendon, ligament, fat, fibrous tissue, lymphatic and vascular systems, fascia, and synovium. It should be understood that due to the different interactions between different types of soft tissue and the imaging modality equipment and their configuration, the imaging modality may be able to represent different types of soft tissue in different ways. For example, an imaging modality configured in a particular way may generate an image representing / depicting the cartilage of a joint in a different way than the muscles of the same joint. Therefore, in this specification, the soft tissue labeling mentioned in the medical image segmentation of anatomical structures is intended to indicate that a particular type of soft tissue is labeled as different from hard tissue, a particular type of hard tissue, and / or another type of soft tissue. For example, labeling soft tissue in medical images could refer to labeling individual pixels as cartilage, or even more specifically, femoral cartilage or tibial cartilage, where other pixels in the medical image are labeled as bone, or even more specifically, femur or tibia, where other pixels in the medical image are labeled as background. In such an example, the other pixels labeled as background could actually correspond to, for example, muscle, but are considered background pixels for image processing and labeling purposes specific to the application, allowing only cartilage and bone to be further processed and / or studied. Alternatives are possible.
[0063] This specification provides examples where the soft tissue to be labeled is cartilage and the hard tissue to be labeled is bone. It should be understood that applications of the methods and systems described and depicted herein are possible where the soft tissue to be labeled is not cartilage, for example, where the soft tissue is muscle and the hard tissue is bone, or some other hard tissue, such as teeth. It should be understood that applications of the methods and systems described and depicted herein may occur where the soft tissue to be labeled is both cartilage and muscle, and the labeling of the cartilage differs from that of the muscle, in order to enable differentiated treatment and / or research on differently labeled muscles and cartilage. Variations are possible. Detailed Implementation
[0064] The following discussion involves various examples. While one or more of these examples may be preferred, the disclosed examples should not be construed as limiting the scope of this disclosure (including the claims). Furthermore, those skilled in the art will understand that the following description has broad application, and the discussion of any example is intended only as an example of that example and is not intended to imply that the scope of this disclosure, including the claims, is limited to that example.
[0065] Computed tomography (CT) is a medical imaging technique that uses X-rays to create detailed cross-sectional images of the body. CT imaging is commonly used in various musculoskeletal (MSK) specialties for diagnostic and surgical planning purposes. CT scans generate CT images, which, due to their excellent ability to depict bone structures with high precision and clarity, contribute to the accurate visualization and diagnosis of bone-related conditions. However, CT scans have several limitations in the visualization and diagnosis of soft tissues. Soft tissues such as cartilage cannot be depicted as clearly as bone in CT images. This can pose challenges to the accurate perception and assessment of such soft tissues. In particular, cartilage is difficult to visualize and diagnose when using CT alone.
[0066] In cases requiring cartilage assessment, magnetic resonance imaging (MRI), a cross-sectional imaging technique that uses magnetic fields, field gradients, and radiofrequency technology to perform MRI scans, is often the standard of clinical practice. MRI scans generate MR images that help to accurately visualize and diagnose soft tissue conditions, such as cartilage conditions, because these MR images have an excellent ability to depict both bone and soft tissue structures with high precision and clarity. In the standard treatment plan (SoC), assessment of knee cartilage conditions is important for diagnosing osteoarthritis (OA), which occurs when the cartilage between bones wears down or breaks due to injury or disease.
[0067] Figure 1A This is an example MR image 10 of the knee joint of an individual human patient, including the femur 20, patella 22, tibia 24, femoral cartilage 26, tibial cartilage 28, and portions of other features. White arrows are overlaid on the MR image and point to some of the femoral cartilage 26 within the dark area located distal to the femur 20. The femoral cartilage 26, indicated by the white arrows, is clearly visible and well-defined in MR image 10, allowing for a clear visual contrast with the femur 20 and other anatomical features of the knee joint. Figure 1B This is an example CT image 15 of the same knee joint from the same patient. Although the femur 20, patella 22, and tibia 24 are clearly visible and well-defined in CT image 15, it is difficult to visually identify or distinguish soft tissue parts (such as femoral cartilage), making it difficult to clearly identify or locate them.
[0068] Automatic segmentation in CT and MRI involves depicting and extracting anatomical structures and tissues captured in CT or MR images, thereby allowing for a better understanding of the musculoskeletal system. Figure 2A The image 40 is an example unsegmented and therefore unlabeled MR image of the patient's knee joint, including portions of the femur 50, patella 52, tibia 54, femoral cartilage 56, tibial cartilage 58, and other features. Figure 2BIt is the segmented MR image 40, and therefore associated with its pixels being labeled as part of an anatomical structure (or background). Specifically, in Figure 2B In this model, segmentation labels corresponding to each pixel are used to represent pixels of different anatomical structures in different colors. For example, a pixel labeled femur 50 is presented in green, a pixel labeled femoral cartilage 56 in yellow, a pixel labeled tibia 54 in red, and a pixel labeled tibial cartilage 58 in blue. Assigning labels corresponding to anatomical structures to at least these pixels aids in diagnosis, treatment planning, supports image-guided interventions, and advances research. This can improve clinical decision-making, enhance patient outcomes, and contribute to advancements in the MSK profession.
[0069] At the time of writing, the state-of-the-art process for automated medical image segmentation utilizes deep learning (DL) techniques. This means that image segmentation is not only performed manually by radiologists or other personnel, but also involves image processing of medical images using at least one computer processor. With DL techniques, a large amount of labeled data is typically used as training data to train a neural network to enable it to perform automated segmentation. In medical image segmentation, such training data typically consists of medical images that have already been manually labeled (i.e., “marked”) by healthcare professionals or individuals with appropriate medical training.
[0070] However, manual annotation of medical images, even when used to provide training data for training deep learning systems to perform automatic segmentation, is typically a time-consuming and error-prone process. Furthermore, manual annotators, even experienced ones, often struggle to segment soft tissues, such as cartilage, in CT images with high accuracy, as previously discussed... Figure 1A As discussed, perceiving the boundaries of such cartilage from CT images can be very difficult or even practically impossible. Therefore, those familiar with annotation generally understand that manual annotation of soft tissue in CT images can usually only be done approximatingly, making it impossible to depict the boundaries of such soft tissue with high confidence. Consequently, training a deep learning system using only such approximate labeled CT training images may result in the system failing to develop sufficient capability for accurate and useful automatic segmentation.
[0071] Various methods have been tried to improve the efficiency and accuracy of training the DL system.
[0072] Such methods include cross-modal synthesis, where CT is synthesized from MRI or MRI from CT. As described in this paper, CT and MRI are two widely used imaging modalities in medical diagnostics. Both modalities provide valuable information about the internal structures of the body. However, they differ in several ways, including in their imaging principles and associated risks.
[0073] For example, CT imaging involves the use of X-rays, a form of ionizing radiation. X-rays are guided through anatomical structures, and detectors measure the amount of radiation passing through different tissues. This information is processed by a computer to generate cross-sectional images of the anatomical structures. The ionizing radiation used during CT scans carries potential risks, especially in cases of repeated or excessive exposure, as radiation can damage DNA and increase the risk of cancer.
[0074] On the other hand, MRI utilizes different imaging mechanisms that do not involve ionizing radiation, relying instead on magnetic fields and radio frequency waves to create detailed images of body structures. MRI is considered a safe imaging modality. However, despite its superior safety profile compared to CT, CT imaging remains necessary in certain situations due to its unique ability to visualize specific anatomical details and pathologies. Specifically, bone boundaries, injuries, and deformities are very clear in CT images. However, due to its ionizing radiation, researchers have been exploring techniques to compute CT-like images from MRI data. Such techniques aim to bridge the gap between the two modalities and provide an alternative to CT scans without the risk of ionizing radiation. The use of Generative Adversarial Networks (GANs) to convert MRI into CT images has been proposed.
[0075] Generating MRI-like images from CT images is advantageous for patients who are unsuitable for MRI due to claustrophobia, pacemakers, and / or artificial joints. Furthermore, CT imaging is generally less expensive than MRI imaging. Proposals to achieve CT-to-MRI translation by combining GANs with dual-cycle consistency loss and voxel-level loss have shown very promising results in brain imaging. Adversarial domain-adaptive deep learning methods have been used for automatic tumor segmentation from T2-weighted MRI. Due to the scarcity of labeled MRI data for tumor segmentation, unsupervised cross-domain adaptation for tumor perception from CT to MRI was used, followed by semi-supervised tumor segmentation using a U-Net model trained with synthetic and a limited number of original MRI images.
[0076] An emerging research direction aims to train segmentation networks for target imaging modalities without requiring manual labeling of specific modalities. As part of this research, the SynSeg-Net model has been proposed and has achieved encouraging performance in two experiments. In the first experiment, a model trained with unpaired CT and MRI images was used for CT segmentation of the spleen, where only manual annotation of the MRI images was available. In the second experiment, a model trained with unpaired MRI and CT images was used for intracranial volume segmentation of MRI images, where only manual annotation of the CT images was available. A similar problem has been addressed by improving the annotation problem by learning a segmentation model for a less labeled target modality using the richly labeled source modality. A network architecture called Diversified Data Augmentation Generative Adversarial Network (DDA-GAN) has been proposed, and its effectiveness has been demonstrated in two experiments. In the first experiment, MRI was used as the target domain and CT was used as the source domain, incorporating manual segmentation, for segmentation of craniofacial bone structures. In the second of these experiments, cardiac substructure segmentation was performed, with CT as the target domain and MRI as the source domain containing manual annotations.
[0077] Other research and development efforts have focused primarily on the automatic segmentation of cartilage in MRI images.
[0078] Contrast agents are known to be used when capturing CT images. These agents help improve the contrast of cartilage in CT images, thus aiding in cartilage segmentation. Specifically, CT with contrast agents refers to the application of a contrast agent or dye during a CT scan. A contrast agent is a substance that enhances the visibility and differentiation of certain tissues and blood vessels, allowing for a more detailed assessment of specific anatomical structures or pathological conditions. However, the use of contrast agents has several disadvantages. These include the risk of contrast agent-related reactions, increased imaging appointment time, potential kidney damage, potential allergic reactions, and so on. The decision to perform contrast-enhanced CT imaging is made based on the specific clinical situation, requiring a trade-off between the benefits and potential risks.
[0079] Methods using MR and CT images of the same knee can include the registration of MR and CT images, where information from both modalities is aligned and merged to create a comprehensive and integrated view of the imaged anatomical structures. Known work primarily focuses on registering MR and CT images of an individual knee to automatically assess osteoarthritis (OA) after manual segmentation of bone and cartilage.
[0080] The study also proposed assessing OA on CT instead of MRI, based on the theory that since OA also affects trabecular bone, joint space narrowing, and causes subchondral cysts and bone sclerosis, these aspects can be well imaged using CT modalities, thus OA can be accurately assessed using CT alone.
[0081] Furthermore, although cartilage itself is not easily visible in knee CT scans, literature indicates that certain bone features in CT images can provide clues for inferring cartilage structures. Additionally, window function techniques can potentially alter the local contrast of CT images, thereby better highlighting certain cartilage structures. Figure 3A Example CT image 60A of the knee joint of an individual human patient without contrast treatment is shown, with the arrow pointing to the femoral cartilage. Figure 3B It shows the relationship with Figure 3A An example of a contrast-processed CT image 60B corresponding to CT image 60A, where the contrast processing was performed using a local window function technique, with the arrow pointing towards the femoral cartilage. The local window function technique involves locally manipulating the contrast of CT image intensity, potentially improving the visualization of cartilage in certain areas.
[0082] However, even with local window function contrast, manually annotating the cartilage on a knee CT scan can only target certain anatomical locations. Therefore, in most cases, manual annotation alone cannot segment the complete cartilage. Furthermore, using local window functions is both time-consuming (because different contrasts are required for each region) and error-prone (because the cartilage boundaries are often still not clearly defined).
[0083] The methods and systems presented in this paper address the shortcomings of existing techniques by improving the training of neural networks to at least accurately label soft tissues (such as cartilage) captured by anatomical structures in CT images.
[0084] A training set of soft tissue-labeled CT images is generated using MR and CT images of the same anatomical structure, and the training is performed using the generated training set.
[0085] In the first method, multiple pairs of image sets are processed for each of multiple individuals, each pair including a set of CT images and a set of MR images of the individual's anatomy, to generate a corresponding three-dimensional (3D) model of the individual's anatomy, each model including the bones of the individual's anatomy. For each individual, the contents labeled as bones in the two resulting 3D models are aligned / registered, thereby generating a 3D transformation between the bones in the 3D models. Then, based on the 3D transformation applied to the labeled soft tissue in the 3D model generated using the MR image set, soft tissue in at least one CT image is labeled to create a corresponding soft tissue-labeled CT image. The soft tissue-labeled CT image is then used to augment the training set of labeled CT images. For example, the soft tissue-labeled CT images can simply be added to the training set. Using the training set created in this way, a neural network is trained to label soft tissue in CT images, utilizing each of the actual CT and MR image sets from the actual CT and MR image sets of multiple individuals. Therefore, according to this method, hard tissue labeled in two different imaging modalities of the same anatomical structure in the same individual provides a link between the two modalities. This link can be used to accurately “transfer” (i.e., convert) the labeling of soft tissue from a first modality that can accurately label soft tissue to a second modality that usually cannot accurately label such soft tissue, thereby generating a training set of accurately labeled soft tissue images of the second modality, which is used to train a neural network to label soft tissue in images of the second modality.
[0086] When training a deep learning (DL) system to segment images, a supervised strategy is typically used. A supervised strategy refers to a training dataset (or "training set") containing images with corresponding segmentation masks. That is, each image in the training set has an associated mask that assigns a specific label of interest to each pixel. These labeled masks are usually obtained manually by experts.
[0087] According to this method, in a CT image set (which contains multiple CT images, each taken along a corresponding cross-section or slice of the anatomical structure), only bone markers are manually labeled, while in an MR image set (which contains multiple MR images, each taken along a corresponding cross-section or slice of the anatomical structure), both bone and cartilage regions are manually labeled. Since the same anatomical structure (e.g., the same knee of the same individual) is imaged using both modalities, the 3D models of the bones from each modality should ideally be identical. In one example, the 3D model of the bone itself can be generated separately by applying a moving cubes procedure to the contents of the CT image set and the MR image set. Once the 3D bone models (a first 3D model of both bone and soft tissue created using the MR image set, and a second 3D model of the bone created using the CT image set) are generated, the bones of the 3D models can be registered / aligned to generate a 3D transformation T between them. This 3D transformation can then be applied to the soft tissue in the first 3D model to label the contents of the second 3D model as the same soft tissue, thereby creating an enhanced second 3D model. The enhanced second 3D model can then be used to create CT images with at least one soft tissue marker.
[0088] The alignment of bones in the first and second 3D models can be referred to as 3D registration. In one example, 3D registration is performed in two steps. In the first step, called global registration, a global registration algorithm is applied to obtain an initial alignment between the bones of the first and second 3D models. In one example, a method called 2PNS (Two-Point Normal Set) determines this initial alignment. In the second step, called alignment refinement, the initial alignment is refined. In one example, a method called Iterative Closest Point (ICP) is used for alignment refinement, iteratively refining the alignment obtained during the initial alignment. The 3D registration process produces a 3D transformation T, which can then be applied to the 3D cartilage in the first 3D model to determine the location of the cartilage in the second 3D model. In one example, the cartilage markers are obtained by resampling the 3D model (generated from a set of CT images) using the Slicer3D process with the second 3D model as a reference volume.
[0089] By applying this method to multiple CT-MR image sets, a training set of soft tissue-labeled CT images is generated and used to train an automatic deep learning segmentation model (i.e., a trained neural network) capable of soft tissue labeling of CT images of anatomical structures. More specifically, the neural network trained according to this method can be provided and thus receive input CT images, and the neural network can then provide soft tissue labels for the input CT images. The format of the labels can vary depending on the implementation. For example, the trained neural network can provide only the labels for the pixels to be applied to the input CT image by downstream processing, or it can provide the entire output CT image, which corresponds precisely to the input CT image, but also has at least some of the generated soft tissue labels integrated with the CT image file.
[0090] In one example, the neural network is the U-net neural network.
[0091] Experimental results of an automated cartilage segmentation system trained according to the description herein using a training set created from multiple pairs of images of corresponding individual anatomy are compared with baseline truth using two metrics, as illustrated herein. In this example, baseline truth is defined as a model of bone and cartilage obtained from CT, where cartilage markers are determined by aligning the CT with the MRI of the same patient using the bone surface as a reference. It should be understood that baseline truth itself may be affected by small errors caused by registration errors in the MR images themselves and errors in manual labeling.
[0092] The first of the two metrics, called Global Average Symmetric Surface Distance (GASSD), measures the average symmetric surface distance (in millimeters) between the reference truth surface (i.e., a model of bone and cartilage, where cartilage markers are obtained through registration with the MR images explained in this paper) and the reconstructed surface from the automatic segmentation. For each point in the reference surface, the closest triangle on the other surface is determined. This is performed using both the reference truth surface and the reconstructed surface as references. The average of these distances is then calculated, with lower values corresponding to better segmentation.
[0093] The second of the two measures, called the mean symmetry surface distance of the cartilage (CASSD), measures the average symmetry surface distance (in millimeters) between the reference truth surface and the reconstructed surface only at points where the cartilage typically lies (e.g., at the condyle of the femur and the plateau of the tibia). A lower value corresponds to better segmentation.
[0094] Table 1 below provides an overview of two different pairs of CT and MRI image sets used to generate training datasets for training two different neural networks and test datasets for testing the corresponding neural networks.
[0095]
[0096] Table 1.
[0097] Specifically, two different DL models were developed. The first of these two different DL models, referred to in this paper as model_Bioskills, was trained using the training set of dataset 1. The second of these two different DL models, referred to in this paper as model_SoC, was trained using the training dataset of dataset 2. PD Sag no FS refers to no fat saturation in the proton density sagittal plane. More specifically, proton density (PD) MRI sequences are specific MRI imaging sequences primarily used for SoC, Sag is sagittal and relates to the acquisition direction, and no FS means no fat saturation, which means that a fat-free suppression technique is used during MRI.
[0098] Table 2 below presents the experimental results of the two DL models in the two datasets detailed in Table 1.
[0099]
[0100] Table 2.
[0101] Table 2 shows the median GASSD and CASSD between the baseline truth and the inference model.
[0102] To quantify the difference between the baseline true bone model and the baseline true model containing both bone and cartilage, the difference in CASSD was found to be + / - 0.5 mm in the median term.
[0103] In the examples described herein, bone registration in the first and second 3D models created using MR and CT image sets can be used to generate a 3D transformation that transfers soft tissue markers from MR images to CT images for use in generating a training set of soft tissue-marked CT images. In another example, a similar approach registers / aligns CT and MR image pairs of corresponding individual anatomy to generate a 2D transformation between the bones in the MR images and the CT images of the pair. This 2D transformation can then be applied to the marked soft tissue in the MR images to mark the soft tissue in the CT images, thereby creating soft tissue-marked CT images. This can be done individually for multiple CT-MR image pairs acquired from corresponding CT and MR image sets of an individual anatomy, resulting in a 2D transformation and thus producing a soft tissue-marked CT image for each CT-MR image pair. Using this method, registration / alignment is performed between bones in 2D CT and MR images, rather than between bones in 3D CT and MR bone models.
[0104] Alignment of bones in CT and MR images can be referred to as 2D registration. In one example, 2D registration is performed in two steps. In the first step, called global registration, a global registration algorithm is applied to obtain an initial alignment between the bones in the CT and MR images. In one example, a method called 2PNS (Two-Point Normal Set) determines this initial alignment. In the second step, called alignment refinement, the initial alignment is refined. In one example, a method called Iterative Closest Point (ICP) is used for alignment refinement, iteratively refining the alignment obtained during the initial alignment. The 2D registration process produces a 2D transform T2, which can then be applied to the 2D cartilage in the MR image to determine the location of the cartilage in its paired CT image.
[0105] By applying this method to multiple CT-MR image pairs, a training set of soft tissue-labeled CT images is generated and used to train an automatic deep learning segmentation model (i.e., a trained neural network) capable of soft tissue labeling of CT images of anatomical structures. More specifically, the neural network trained according to this method can be provided and thus receive the input CT image, and the neural network can then provide soft tissue labels for the input CT image. The format of the labels can vary depending on the implementation. For example, the trained neural network can provide only the labels for the pixels to be applied to the input CT image by downstream processing, or it can provide the entire output CT image, which corresponds precisely to the input CT image, but also has at least some of the generated soft tissue labels integrated with the CT image file.
[0106] In one example, the neural network is the U-net neural network.
[0107] A training set of soft tissue-labeled CT images is generated by transferring soft tissue labels from synthesized MR images to input CT images.
[0108] Since capturing and / or obtaining multiple pairs of CT and MR images of the same body anatomical structure can be costly and time-consuming, alternative methods for soft tissue labeling CT images have attracted much attention.
[0109] In the following example, a generative adversarial method is used to train a neural network to receive CT images of individual anatomical structures and generate corresponding output synthetic MR images of those structures. These output synthetic MR images have useful accuracy, making them usable for MRI segmentation instead of actual MR images captured using MRI techniques.
[0110] Specifically, the neural network receives an input CT image and generates a corresponding output synthetic MR image. This output synthetic MR image is then processed for soft tissue labeling. This processing for soft tissue labeling can be performed by an automatic image segmentation system configured to label soft and hard tissues in the MR image provided for processing. The soft tissue labels can then be applied to at least the input CT image used to generate the synthetic output MR image, thereby creating a CT image with soft tissue labeling.
[0111] To generate an output synthetic MR image corresponding to an input CT image, in this example, the neural network is configured as a first generative adversarial network (GAN), which is interconnected with a second GAN in a CycleGAN configuration for training. In this example, the first GAN is trained to generate a synthetic output MR image from its input, and the second GAN is trained to generate a synthetic CT image from its input. For each of the multiple loops, the entire neural network is trained by performing multiple steps. A first step of these steps includes using the synthetic CT image generated by the second GAN as input to the first GAN, thereby generating a corresponding synthetic MR image as the first GAN output. A second step of these steps includes providing the synthetic MR image generated by the first GAN as input to the second GAN, thereby generating a corresponding synthetic CT image as the second GAN output. A third step of these steps includes determining a first difference metric between the second GAN output and the first GAN input. A fourth step of these steps includes determining a second difference metric between the first GAN output and the second GAN input. A fifth step of these steps includes modifying the first GAN and the second GAN based at least on the first and second difference metrics. Modifications were made to reduce the values of the first and second difference measures within consecutive cycles.
[0112] This CycleGAN approach, performed over multiple loops, enables the neural network to converge to cycle consistency, allowing it to generate useful and accurate synthetic MR images based on input real CT images, including soft tissue present in the MR images when the anatomical structures are actually captured using MRI technology. It should be understood that, thanks to the CycleGAN architecture and method, the neural network can be trained without using corresponding pairs of real CT and MR images (which must be image pairs of the same real individual anatomical structure) for training the first and second GANs themselves.
[0113] Figure 4A , Figure 4B and Figure 4C A general example pipeline using a neural network capable of generating synthetic MR images from input CT images is shown. Figure 4AIn this process, a neural network, identified by reference numeral 100, receives real, unlabeled CT images 70 and generates synthetic MR images 75. Figure 4B In this example, a synthetic MR image 75 generated by a neural network 100 is provided to an automated MRI segmentation system 200. This system infers soft and hard tissues, such as bone, to generate soft and hard tissue markers, thereby creating a labeled synthetic MR image 80. Specifically, the automated MRI segmentation system 200 infers the femur 82, femoral cartilage 84, tibia 86, and tibial cartilage 88, and thus generates markers for these tissues. Finally, using Slicer3D's resampling technique, the inferred markers are transferred to a CT image via a resampling process 300, with the CT image serving as a reference volume. It should be understood that segmentation of hard tissues, such as bone, is not necessary directly on the input CT image, as resampling can be used to transfer the markers generated by the resampling process 300 from the synthetic MR image for both hard and soft tissues to the CT image. Therefore, the CT image 70 is set with markers for the femur 82, femoral cartilage 84, tibia 86, and tibial cartilage 88.
[0114] according to Figure 4A , Figure 4B and Figure 4C The experimental results of the pipeline are shown in Table 3 below. The experimental results of the proposed solution combining CycleGAN and the automatic PRIME MRI segmentation algorithm are presented. The median GASSD and CASSD between the baseline truth and the inference model are shown.
[0115]
[0116] Table 3.
[0117] In another example, during the training of the neural network, bone segmentation / labeling of synthetic CT images generated by a second GAN is received and used as regularization during modification. More specifically, the neural network is configured as a synthetic segmentation network (SynSeg-Net) to be trained using bone CT segmentation as regularization to generate synthetic MR images. During testing with the SynSeg-Net configuration, it was found that the additional segmentation loss improved the quality of the generated synthetic images. Experimental results (including median GASSD and CASSD between real data and the inference model) are shown in Table 4 below.
[0118]
[0119] Table 4.
[0120] Figure 5Various CT and CT magnified / zoomed images are illustrated, illustrating (a) a comparison of input CT images with (b) pseudo-benchmark truth segmentation; (c) automatic segmentation using the image pairing method described herein; (d) automatic segmentation using the CycleGAN method described herein; and (e) automatic segmentation using the SynSeg-Net method described herein. It should be understood that the benchmark truth for cartilage segmentation obtained from MRI includes some errors, such as due to difficulties in manual cartilage depiction, inaccurate registration, or other reasons.
[0121] To account for the error, Tables 5 and 6 present preliminary results demonstrating the performance of automatic segmentation of cartilage in CT images for intraoperative registration using the PRIME (now called TESSA) application provided by Smith & Nephew, Andover, MA.
[0122]
[0123] Table 5.
[0124] Table 5 shows the simulation results of PRIME intraoperative registration using 3D models from different sources (one ACL femoral case). Normal intraoperative data acquisition was performed, with most data corresponding to bone and some to cartilage. [CT bone] refers to automatic bone segmentation in CT, [CT bone + cartilage (GAN)] refers to automatic segmentation of bone and cartilage in CT using the CycleGAN method described herein, and [CT bone + cartilage (SynSeg)] refers to automatic segmentation of bone and cartilage in CT using the SynSeg-Net modified method described herein. [CT bone + cartilage (CT-MRI Pairs SoC)] refers to automatic segmentation of bone and cartilage in CT using model_SoC, [CT bone + cartilage (CT-MRI Pairs BioSkills)] refers to automatic segmentation of bone and cartilage in CT using model_bioskills, [PRIME MRI bone] refers to automatic segmentation of bone in PRIME MRI, and [PRIME MRI bone + cartilage] refers to automatic segmentation of bone and cartilage in PRIME MRI. Down hash coloring, grid hash coloring, and up hash coloring correspond to the best, second best, and third best measures in a given column, respectively.
[0125]
[0126] Table 6.
[0127] Table 6 shows the simulation results of PRIME intraoperative registration using 3D models from different sources. Intraoperative data acquisition was performed, and bone and cartilage points were collected (1 ACL femoral case). [CT bone] refers to automatic bone segmentation in CT, [CT bone + cartilage (GAN)] refers to automatic segmentation of bone and cartilage in CT using the CycleGAN method described in this paper, [CT bone + cartilage (SynSeg)] refers to automatic segmentation of bone and cartilage in CT using SynSeg-Net modification, [CT bone + cartilage (CT-MRI PairsSoC)] refers to automatic segmentation of bone and cartilage in CT using model_SoC, [CT bone + cartilage (CT-MRI PairsBioSkills)] refers to automatic segmentation of bone and cartilage in CT using model_bioskills, [PRIME MRI bone] refers to automatic segmentation of bone in PRIME MRI, and [PRIME MRI bone + cartilage] refers to automatic segmentation of bone and cartilage in PRIME MRI. Down-hash coloring, grid-hash coloring, and up-hash coloring correspond to the best, second-best, and third-best metrics in a given column, respectively. Green, blue, and yellow correspond to the best, second-best, and third-best metrics in a column, respectively.
[0128] like Figure 5 As shown, the baseline truth segmentation shown in the CT image at (b) obtained from MRI has some errors. That is, the cartilage is not correctly aligned with the bone as might be expected. In contrast, the automatic segmentation method shown in the CT images at (c), (d), and (e) provides more reasonable results. These results were discussed with the radiologist, who examined the results of the automatic segmentation shown at (c), (d), and (e) and concluded that these results were reasonable but needed improvement for some reason and were difficult to quantify.
[0129] The various methods for segmenting CT images described herein offer numerous advantages. For example, by enabling CT images to be labeled with soft tissue, 3D anatomical modeling of a given joint or other anatomical structure can be enhanced to include not only hard tissues such as bone in a joint, but also soft tissues such as cartilage in a joint. Thus, intraoperative digitization of a joint can be accomplished using structures in the joint, whether hard tissues such as bone or soft tissues such as cartilage, allowing the joint to be registered with a 3D model of the joint for surgical planning and navigation. Since preliminary studies have shown that (1) digitizing only bone is very difficult, and surgeons sometimes digitize some cartilage structures (as with bone), and (2) digitizing cartilage intraoperatively and using a 3D model that represents not only bone but also cartilage improves registration results, providing a 3D model representing the anatomical structures of both bone and cartilage holds promise for improving the overall performance of surgical navigation. At the time of this application, the situation was that 3D models of bone and cartilage were generally only available from MRI imaging because, as discussed herein, cartilage had not previously been satisfactorily visualized in CT. Because the techniques described in this paper allow for the inference and corresponding labeling of cartilage using only CT images, surgical navigation solutions (such as the PRIME ACL solution provided by Smith & Nephew, Andover, MA, USA) can use standard treatment CT to obtain 3D models of bone and cartilage and / or other types of soft tissue.
[0130] The current standard approach for analyzing and diagnosing cartilage diseases, including osteoarthritis (OA), is the use of MRI techniques. By considering how OA affects trabecular bone, joint space narrowing, and / or causes subchondral cysts and bone sclerosis, the various methods presented in this article can be used to aid in the study of OA progression using CT images.
[0131] Software and hardware
[0132] Figure 6 A computer implementation of a method for training a neural network for soft tissue labeling, based on at least some examples, is shown. Specifically, the method begins (box 1000) and includes: creating a training set of computed tomography (CT) images for soft tissue labeling (box 1100), and using the training set to train a neural network to label soft tissue in the CT images (box 1200). The method then ends (box 1300).
[0133] Figure 7 Example 1100A shows that it can be used in Figure 6The steps of the method performed during the creation of a training set of soft tissue-labeled CT images, as shown in box 1100. Specifically, these steps include: collecting multiple pairs of image sets for each of a plurality of individuals from at least one database, each pair of image sets including both a CT image set and a magnetic resonance (MR) image set of the individual's anatomical structure (box 1102), and for each pair of image sets: using the MR image set to create a first 3D model of both bone and soft tissue of the individual's anatomical structure (box 1104), using the CT image set to create a second 3D model of the bone of the individual's anatomical structure (box 1106), generating a 3D transformation between the bone in the first 3D model and the bone in the second 3D model (box 1108), labeling soft tissue in at least one CT image in the CT image set based on applying the 3D transformation to the labeled soft tissue in the first 3D model, thereby creating at least one soft tissue-labeled CT image (box 1110); and enhancing the training set with the at least one soft tissue-labeled CT image (box 1112).
[0134] Figure 8 It shows that, according to another example, it is possible to Figure 6 An alternative step of the method performed during the creation of a training set of soft tissue-labeled CT images, as shown in box 1100. In this example, a 2D transformation is performed between pairs of CT and MR images, rather than a 3D transformation between bones in a first 3D model and a second 3D model. Specifically, these steps include: collecting image pairs for each of a plurality of individuals from at least one database, each pair including both a CT image and a magnetic resonance (MR) image of the individual's anatomy (box 1152), and for each pair of images: generating a 2D transformation between bones in the MR image and bones in the CT image (box 1154), labeling soft tissue in the CT image based on applying the 2D transformation to labeled soft tissue in the MR image to create a soft tissue-labeled CT image (box 1156), and using the soft tissue-labeled CT image to augment the training set (box 1158).
[0135] Figure 9A computer-implemented method for soft tissue labeling of an input CT image of an individual anatomical structure, based on an example, is shown. Specifically, the method begins (box 1500) and includes: generating a corresponding output synthetic magnetic resonance (MR) image of the individual anatomical structure from an input CT image using a neural network, wherein the neural network is trained using a generative adversarial method and at least using CT training images and MR training images to receive the CT image of the individual anatomical structure and generate a corresponding output synthetic MR image of the individual anatomical structure (box 1502); receiving at least soft tissue labels corresponding to the corresponding output synthetic MR image (box 1504); and applying at least the soft tissue labels from the corresponding output synthetic MR image to the input CT image, thereby performing soft tissue labeling on the input CT image (box 1506).
[0136] Figure 10 An example computer system 2000 is illustrated, which can be configured to perform the processor implementations of the methods disclosed herein and can be configured to implement the systems disclosed herein. In one example, computer system 2000 may correspond to a surgical controller, a separate computing device, or any other system implementing any or all of the various methods discussed in this specification. Computer system 2000 may be connected (e.g., networked) to other computer systems in a local area network (LAN), intranet, and / or extranet (e.g., the device cart 402 network), or connected to the Internet at some point (e.g., when not in use during surgery). Computer system 2000 may be a server, a personal computer (PC), a tablet computer, or any device capable of executing a set of instructions (sequential instructions or other instructions) specifying the actions to be taken by it. Furthermore, although only a single computer system is shown, the term "computer" should also be considered to include any collection of computers that individually or collectively execute a set (or more sets) of instructions to perform any or one of the methods discussed herein.
[0137] The computer system 2000 includes a processing device 2002 that communicates with each other via a bus 2010, a main memory 2004 (e.g., read-only memory (ROM), flash memory, dynamic random access memory (DRAM), such as synchronous DRAM (SDRAM), a static memory 2006 (e.g., flash memory, static random access memory (SRAM)), and a data storage device 2008.
[0138] Processing device 2002 refers to one or more general-purpose processing devices, such as microprocessors, central processing units, etc. More specifically, processing device 2002 may be a Complex Instruction Set Computing (CISC) microprocessor, a Reduced Instruction Set Computing (RISC) microprocessor, a Very Long Instruction Word (VLIW) microprocessor, or a processor implementing other instruction sets or combinations thereof. Processing device 2002 may also be one or more special-purpose processing devices, such as application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), network processors, etc. Processing device 2002 is configured to execute instructions for performing any of the operations and steps discussed herein. Once programmed with specific instructions, processing device 2002, and thus the entire computer system 2000, becomes a special-purpose device, such as surgical controller 418.
[0139] Computer system 2000 may also include a network interface device 2012 for communicating with any suitable network (e.g., device cart 402 network). Computer system 2000 may also include a video display 2014 (e.g., display device 414), one or more input devices 2016 (e.g., microphone, keyboard, and / or mouse), and one or more speakers 2018. In one exemplary example, the video display 2014 and the input devices 2016 may be combined into a single component or device (e.g., LCD-liquid crystal display-touchscreen).
[0140] Data storage device 2008 may include a computer-readable storage medium 2020 serving as memory, on which instructions 2022 embodying any one or more methods or functions described herein (e.g., implementing any method and any function performed by any means and / or component described herein). Instructions 2022 may also reside wholly or at least partially within main memory 2004 and / or processing device 2002 during execution by computer system 2000. Therefore, main memory 2004 and processing device 2002 also constitute computer-readable media. In some cases, instructions 2022 may also be transmitted or received via a network through network interface device 2012.
[0141] Although computer-readable storage medium 2020 is shown as a single medium in the illustrative examples, the terms "computer-readable storage medium" or "processor-readable storage medium" should be considered to include a single medium or multiple media (e.g., a centralized or distributed database and / or associated caches and servers) that store one or more sets of instructions. The terms "computer-readable storage medium" or "processor-readable storage medium" should also be considered to include any medium capable of storing, encoding, or carrying a set of instructions for machine execution and causing the machine to perform any one or more of the methods of this disclosure. Therefore, the terms "computer-readable storage medium" or "processor-readable storage medium" should be understood to include, but are not limited to, solid-state memory, optical media, and magnetic media.
[0142] The foregoing discussion is intended to illustrate the principles and various embodiments of the invention. Many variations and modifications will become apparent to those skilled in the art once the foregoing disclosure is fully understood. The following claims are intended to be interpreted as covering all such variations and modifications.
[0143] For example, while the examples in this paper involve cartilage segmentation in knee CT images, the methods and systems described and depicted herein can be applied to infer segmentation of other musculoskeletal (MSK) soft tissue anatomy that can be well perceived in MRI rather than CT. This can include cartilage as exemplified in this paper, but can also include other soft tissues such as the anterior cruciate ligament (ACL), posterior cruciate ligament (PCL), meniscus, muscles, tendons, medial collateral ligament (MCL), lateral collateral ligament (LCL), and so on.
[0144] Terms and Conditions
[0145] Clause 1. A processor-implemented method for training a neural network for soft tissue labeling, the processor-implemented method comprising:
[0146] A training set of computed tomography (CT) images with soft tissue markers was created, including:
[0147] Multiple image sets are collected for each of a plurality of individuals from at least one database, each image set comprising both a CT image set and a magnetic resonance (MR) image set of the individual's anatomy; and
[0148] For each pair of image sets in the plurality of image sets:
[0149] The MR image set was used to create a first 3D model of both bone and soft tissue of the individual's anatomical structure;
[0150] The CT image set was used to create a second 3D model of the bones of the individual's anatomical structure;
[0151] Generate a 3D transformation between the bones in the first 3D model and the bones in the second 3D model;
[0152] Based on applying the 3D transformation to the labeled soft tissue in the first 3D model, soft tissue in at least one CT image in the CT image set is labeled, thereby creating at least one CT image with labeled soft tissue; and
[0153] The training set is augmented using CT images with at least one soft tissue marker;
[0154] as well as
[0155] The training set is used to train the neural network to label soft tissue in CT images.
[0156] Clause 2. The processor-implemented method as described in Clause 1, wherein labeling soft tissue in at least one CT image in the CT image set based on applying the 3D transformation to labeled soft tissue in the first 3D model comprises:
[0157] The 3D transformation is applied to the soft tissue in the first 3D model to label the content of the second 3D model as soft tissue, thereby creating an enhanced second 3D model; and
[0158] The enhanced second 3D model is used to create CT images of the at least one soft tissue marker.
[0159] Clause 3. The processor-implemented method as described in Clause 1, wherein generating a 3D transformation between bones in the first 3D model and bones in the second 3D model comprises:
[0160] A global registration process is applied to determine the initial alignment transformation between the bones in the first 3D model and the bones in the second 3D model; and
[0161] The Iterative Closest Point (ICP) process is applied to the initial alignment transformation to generate the 3D transformation.
[0162] Clause 4. The processor-implemented method as described in Clause 3, wherein applying a global registration process to determine the initial alignment transformation between the bones in the first 3D model and the bones in the second 3D model includes applying a 2PNS (two-point normal set) process to determine the initial alignment transformation between the bones in the first 3D model and the bones in the second 3D model.
[0163] Clause 5. The processor-implemented method as described in Clause 1, wherein the neural network is a U-net neural network.
[0164] Clause 6. The processor-implemented method as described in Clause 1, wherein using the MR image set to create a first 3D model of both bone and soft tissue of the individual anatomical structure comprises:
[0165] At least bone and soft tissue markers corresponding to MR images in the MR image set are received; and
[0166] The moving cube process is applied to the bone and soft tissue markers in the MR images.
[0167] Clause 7. The processor-implemented method as described in Clause 1, wherein using the CT image set to create a second 3D model of the bone of the individual anatomical structure comprises:
[0168] At least receive bone markers corresponding to CT images in the CT image set; and
[0169] The moving cube procedure is applied to the bone markers in the CT image.
[0170] Clause 8. A processor-implemented method as described in Clause 1, wherein the individual anatomical structure is a knee joint, and wherein soft tissues in at least one CT image in the CT image set that are labeled include femoral cartilage and tibial cartilage that are labeled the knee joint.
[0171] Clause 9. A non-transitory processor-readable medium containing processor-readable program code, said processor-readable program code being executable by at least one processor to perform a processor-implemented method as described in Clause 1.
[0172] Clause 10. A neural network trained according to a processor implementation as described in Clause 1.
[0173] Clause 11. The neural network as described in Clause 10, wherein the neural network is a U-net neural network.
[0174] Clause 12. A processor-implemented method for soft tissue labeling of computed tomography (CT) images, the processor-implemented method comprising:
[0175] Provide an input CT image as input to the neural network as described in Clause 10; and
[0176] The neural network receives soft tissue markers from the input CT image.
[0177] Clause 13. A system for marking soft tissue in computed tomography (CT) images, the system comprising:
[0178] At least one processor; and
[0179] A memory coupled to the at least one processor, the memory storing instructions that, when executed by the at least one processor, cause the at least one processor to:
[0180] Provide an input CT image as input to the neural network as described in Clause 10; and
[0181] The neural network receives soft tissue markers from the input CT image.
[0182] Clause 14. The method implemented by the processor as described in Clause 1, wherein the soft tissue is cartilage.
[0183] Clause 15. A processor-implemented method for training a neural network for soft tissue labeling, the processor-implemented method comprising:
[0184] A training set of computed tomography (CT) images with soft tissue markers was created, including:
[0185] Multiple pairs of images are collected for each individual from a plurality of individuals from at least one database, each pair of images comprising both a CT image and a magnetic resonance (MR) image of the individual's anatomy; and
[0186] For each of the plurality of image pairs:
[0187] Generate a 2D transformation between the bones in the MR image and the bones in the CT image;
[0188] Soft tissue in the CT image is labeled based on the application of the 2D transformation to the labeled soft tissue in the MR image, thereby creating a soft tissue-labeled CT image; and
[0189] The training set is augmented using the CT images labeled with the soft tissue.
[0190] as well as
[0191] The training set is used to train the neural network to label soft tissue in CT images.
[0192] Clause 16. A processor-implemented method as described in Clause 15, wherein labeling soft tissue in the CT image based on applying the 2D transformation to labeled soft tissue in the MR image comprises:
[0193] The 2D transformation is applied to the soft tissue in the MR image to label the contents of the CT image as soft tissue, thereby creating a CT image labeled with the soft tissue.
[0194] Clause 17. A processor-implemented method as described in Clause 15, wherein generating a 2D transformation between the bone in the MR image and the bone in the CT image comprises:
[0195] A global registration process is applied to determine the initial alignment transformation between the bone in the MR image and the bone in the CT image; and
[0196] The Iterative Closest Point (ICP) process is applied to the initial alignment transformation to generate the 2D transformation.
[0197] Clause 18. The processor-implemented method as described in Clause 15, wherein the neural network is a U-net neural network.
[0198] Clause 19. The processor-implemented method as described in Clause 15, wherein the individual anatomical structure is the knee joint, and wherein marking soft tissues in the CT image includes marking the femoral cartilage and tibial cartilage of the knee joint.
[0199] Clause 20. A non-transitory processor-readable medium containing processor-readable program code, said processor-readable program code being executable by at least one processor to perform a processor-implemented method as described in Clause 15.
[0200] Clause 21. A neural network trained according to a processor implementation as described in Clause 15.
[0201] Clause 22. The neural network as described in Clause 21, wherein the neural network is a U-net neural network.
[0202] Clause 23. A processor-implemented method for soft tissue labeling of computed tomography (CT) images, the processor-implemented method comprising:
[0203] Provide an input CT image as input to the neural network as described in Clause 21; and
[0204] The neural network receives soft tissue markers from the input CT image.
[0205] Clause 24. A system for marking soft tissue in computed tomography (CT) images, the system comprising:
[0206] At least one processor; and
[0207] A memory coupled to the at least one processor, the memory storing instructions that, when executed by the at least one processor, cause the at least one processor to:
[0208] Provide an input CT image as input to the neural network as described in Clause 21; and
[0209] The neural network receives soft tissue markers from the input CT image.
[0210] Clause 25. The method implemented by the processor as described in Clause 15, wherein the soft tissue is cartilage.
[0211] Clause 26. A processor-implemented method for soft tissue labeling of input CT images of individual anatomical structures, the processor-implemented method comprising:
[0212] A neural network generates a corresponding output synthetic magnetic resonance (MR) image of the individual anatomical structure based on the input CT image, wherein the neural network is trained using a generative adversarial method and at least using CT training images and MR training images to receive CT images of the individual anatomical structure and generate a corresponding output synthetic MR image of the individual anatomical structure;
[0213] At least receive soft tissue markers corresponding to the corresponding output synthetic MR image; and
[0214] The soft tissue markers from the corresponding output synthetic MR image are applied to the input CT image to perform soft tissue markers on the input CT image.
[0215] Clause 27. The processor-implemented method as described in Clause 26, wherein receiving at least the soft tissue markers corresponding to the output synthesized MR image comprises:
[0216] The output synthesized MR image is provided to an automatic image segmentation system; and
[0217] The automatic image segmentation system receives at least the soft tissue markers corresponding to the output synthetic MR image.
[0218] Clause 28. A processor-implemented method as described in Clause 27, wherein applying at least the soft tissue markers from the output synthesized MR image to the input CT image to perform soft tissue markers on the input CT image comprises:
[0219] The segmented volume of the output synthetic MR image is resampled using the input CT image as a reference.
[0220] Clause 29. The processor-implemented method as described in Clause 26 further includes:
[0221] The neural network is configured as a first generative adversarial network (GAN), the first GAN being interconnected with a second GAN, wherein the first GAN is trained to generate synthetic output MR images from its inputs, and the second GAN is trained to generate synthetic CT images from its inputs; and
[0222] For each of the multiple loops, the neural network is trained in the following manner:
[0223] The synthetic CT image generated by the second GAN is provided as the input of the first GAN, thereby generating a corresponding synthetic MR image as the output of the first GAN;
[0224] The synthetic MR image generated by the first GAN is provided as the input of the second GAN, thereby generating a corresponding synthetic CT image as the output of the second GAN;
[0225] Determine a first difference metric between the second GAN output and the first GAN input;
[0226] Determine a second difference metric between the first GAN output and the second GAN input; and
[0227] The first GAN and the second GAN are modified based at least on the first difference metric and the second difference metric, wherein the modification is made to reduce the magnitude of the first difference metric and the second difference metric in consecutive cycles.
[0228] Clause 30. The processor-implemented method as described in Clause 29 further includes:
[0229] During the training of the neural network:
[0230] Receive bone markers from the synthetic CT image generated by the second GAN; and
[0231] The bone markers are used as regularization during the modification.
[0232] Clause 31. The processor-implemented method as described in Clause 26, wherein the neural network is a U-net neural network.
[0233] Clause 32. The processor-implemented method as described in Clause 26, wherein the individual anatomical structure is the knee joint, and the soft tissue markings include markings for the femoral cartilage and markings for the tibial cartilage.
[0234] Clause 33. A non-transitory processor-readable medium containing processor-readable program code, said processor-readable program code being executable by at least one processor to perform a processor-implemented method as described in Clause 26.
[0235] Clause 34. A neural network trained according to a processor implementation as described in Clause 26.
[0236] Clause 35. The neural network as described in Clause 34, wherein the neural network is a U-net neural network.
[0237] Clause 36. A processor-implemented method for soft tissue labeling of computed tomography (CT) images, the processor-implemented method comprising:
[0238] Provide an input CT image as input to the neural network as described in Clause 34; and
[0239] The neural network receives soft tissue markers from the input CT image.
[0240] Clause 37. A system for marking soft tissue in computed tomography (CT) images, the system comprising:
[0241] At least one processor; and
[0242] A memory coupled to the at least one processor, the memory storing instructions that, when executed by the at least one processor, cause the at least one processor to:
[0243] Provide an input CT image as input to the neural network as described in Clause 34; and
[0244] The neural network receives soft tissue markers from the input CT image.
[0245] Clause 38. The processor implementation of the method described in Clause 26, wherein the soft tissue is cartilage.
Claims
1. A processor-implemented method for training a neural network for soft tissue labeling, the processor-implemented method comprising: A training set of computed tomography (CT) images with soft tissue markers was created, including: Multiple image sets are collected for each of a plurality of individuals from at least one database, each image set comprising both a CT image set and a magnetic resonance (MR) image set of the individual's anatomy; and For each pair of image sets in the plurality of image sets: The MR image set was used to create a first 3D model of both bone and soft tissue of the individual's anatomical structure; The CT image set was used to create a second 3D model of the bones of the individual's anatomical structure; Generate a 3D transformation between the bones in the first 3D model and the bones in the second 3D model; Based on applying the 3D transformation to the labeled soft tissue in the first 3D model, soft tissue in at least one CT image in the CT image set is labeled, thereby creating at least one CT image with labeled soft tissue; and The training set is augmented using CT images with at least one soft tissue marker; as well as The training set is used to train the neural network to label soft tissue in CT images.
2. The processor-implemented method of claim 1, wherein labeling soft tissue in at least one CT image in the CT image set based on applying the 3D transformation to labeled soft tissue in the first 3D model comprises: The 3D transformation is applied to the soft tissue in the first 3D model to label the content of the second 3D model as soft tissue, thereby creating an enhanced second 3D model; and The enhanced second 3D model is used to create CT images of the at least one soft tissue marker.
3. The processor-implemented method of claim 1, wherein generating a 3D transformation between the bones in the first 3D model and the bones in the second 3D model comprises: A global registration process is applied to determine the initial alignment transformation between the bones in the first 3D model and the bones in the second 3D model; as well as The Iterative Closest Point (ICP) process is applied to the initial alignment transformation to generate the 3D transformation.
4. The processor-implemented method of claim 3, wherein applying a global registration process to determine the initial alignment transformation between the bones in the first 3D model and the bones in the second 3D model includes applying a 2PNS (two-point normal set) process to determine the initial alignment transformation between the bones in the first 3D model and the bones in the second 3D model.
5. The processor-implemented method of claim 1, wherein the neural network is a U-net neural network.
6. The processor-implemented method of claim 1, wherein using the MR image set to create a first 3D model of both bone and soft tissue of the individual anatomical structure comprises: At least bone and soft tissue markers corresponding to MR images in the MR image set are received; as well as The moving cube process is applied to the bone and soft tissue markers in the MR images.
7. The processor-implemented method of claim 1, wherein using the CT image set to create a second 3D model of the bone of the individual anatomical structure comprises: At least bone markers corresponding to CT images in the CT image set shall be received; as well as The moving cube procedure is applied to the bone markers in the CT image.
8. The processor-implemented method of claim 1, wherein the individual anatomical structure is a knee joint, and the soft tissues in at least one CT image in the CT image set that are labeled include the femoral cartilage and tibial cartilage that are labeled the knee joint.
9. A non-transitory processor-readable medium comprising processor-readable program code, the processor-readable program code being executable by at least one processor to perform the processor-implemented method of claim 1.
10. A neural network trained according to the processor implementation method of claim 1.
11. The neural network of claim 10, wherein the neural network is a U-net neural network.
12. A processor-implemented method for soft tissue labeling of computed tomography (CT) images, the processor-implemented method comprising: Provide an input CT image as input to the neural network as described in claim 10; as well as The neural network receives soft tissue markers from the input CT image.
13. A system for marking soft tissue in computed tomography (CT) images, the system comprising: At least one processor; as well as A memory coupled to the at least one processor, the memory storing instructions that, when executed by the at least one processor, cause the at least one processor to: Provide an input CT image as input to the neural network as described in claim 10; and The neural network receives soft tissue markers from the input CT image.
14. The processor-implemented method of claim 1, wherein the soft tissue is cartilage.
15. A processor-implemented method for training a neural network for soft tissue labeling, the processor-implemented method comprising: A training set of computed tomography (CT) images with soft tissue markers was created, including: Multiple pairs of images are collected for each individual from a plurality of individuals from at least one database, each pair of images comprising both a CT image and a magnetic resonance (MR) image of the individual's anatomy; and For each of the plurality of image pairs: Generate a 2D transformation between the bones in the MR image and the bones in the CT image; Soft tissue in the CT image is labeled based on the application of the 2D transformation to the labeled soft tissue in the MR image, thereby creating a soft tissue-labeled CT image; and The training set is augmented using the CT images labeled with the soft tissue. as well as The training set is used to train the neural network to label soft tissue in CT images.
16. The processor-implemented method of claim 15, wherein labeling soft tissue in the CT image based on applying the 2D transformation to labeled soft tissue in the MR image comprises: The 2D transformation is applied to the soft tissue in the MR image to label the contents of the CT image as soft tissue, thereby creating a CT image labeled with the soft tissue.
17. The processor-implemented method of claim 15, wherein generating a 2D transformation between the bone in the MR image and the bone in the CT image comprises: A global registration process is applied to determine the initial alignment transformation between the bone in the MR image and the bone in the CT image; as well as The Iterative Closest Point (ICP) process is applied to the initial alignment transformation to generate the 2D transformation.
18. The processor-implemented method of claim 15, wherein the neural network is a U-net neural network.
19. The processor-implemented method of claim 15, wherein the individual anatomical structure is a knee joint, and wherein marking the soft tissue in the CT image includes marking the femoral cartilage and tibial cartilage of the knee joint.
20. A non-transitory processor-readable medium comprising processor-readable program code, the processor-readable program code being executable by at least one processor to perform the processor-implemented method of claim 15.
21. A neural network trained according to the processor implementation method of claim 15.
22. The neural network of claim 21, wherein the neural network is a U-net neural network.
23. A processor-implemented method for soft tissue labeling of computed tomography (CT) images, the processor-implemented method comprising: Provide an input CT image as input to the neural network as described in claim 21; as well as The neural network receives soft tissue markers from the input CT image.
24. A system for marking soft tissue in computed tomography (CT) images, the system comprising: At least one processor; as well as A memory coupled to the at least one processor, the memory storing instructions that, when executed by the at least one processor, cause the at least one processor to: Provide an input CT image as input to the neural network as described in claim 21; and The neural network receives soft tissue markers from the input CT image.
25. The processor-implemented method of claim 15, wherein the soft tissue is cartilage.
26. A processor-implemented method for soft tissue labeling of input CT images of individual anatomical structures, the processor-implemented method comprising: A neural network generates a corresponding output synthetic magnetic resonance (MR) image of the individual anatomical structure based on the input CT image, wherein the neural network is trained using a generative adversarial method and at least using CT training images and MR training images to receive CT images of the individual anatomical structure and generate a corresponding output synthetic MR image of the individual anatomical structure; At least receive soft tissue markers corresponding to the corresponding output synthetic MR image; as well as The soft tissue markers from the corresponding output synthetic MR image are applied to the input CT image to perform soft tissue markers on the input CT image.
27. The processor-implemented method of claim 26, wherein receiving at least the soft tissue markers corresponding to the output synthesized MR image comprises: The output synthesized MR image is provided to an automatic image segmentation system; as well as The automatic image segmentation system receives at least the soft tissue markers corresponding to the output synthetic MR image.
28. The processor-implemented method of claim 27, wherein applying the soft tissue markers from the output synthesized MR image to the input CT image to perform soft tissue markers on the input CT image comprises: The segmented volume of the output synthetic MR image is resampled using the input CT image as a reference.
29. The processor-implemented method of claim 26, further comprising: The neural network is configured as a first generative adversarial network (GAN), the first GAN being interconnected with a second GAN, wherein the first GAN is trained to generate a synthetic output MR image from the input of the first GAN, and the second GAN is trained to generate a synthetic CT image from the input of the second GAN. as well as For each of the multiple loops, the neural network is trained in the following manner: The synthetic CT image generated by the second GAN is provided as the input of the first GAN, thereby generating a corresponding synthetic MR image as the output of the first GAN; The synthetic MR image generated by the first GAN is provided as the input of the second GAN, thereby generating a corresponding synthetic CT image as the output of the second GAN; Determine a first difference metric between the second GAN output and the first GAN input; Determine a second difference metric between the first GAN output and the second GAN input; as well as The first GAN and the second GAN are modified based at least on the first difference metric and the second difference metric, wherein the modification is made to reduce the magnitude of the first difference metric and the second difference metric in consecutive cycles.
30. The processor-implemented method of claim 29, further comprising: During the training of the neural network: Receive bone markers from the synthetic CT image generated by the second GAN; as well as The bone markers are used as regularization during the modification.
31. The processor-implemented method of claim 26, wherein the neural network is a U-net neural network.
32. The processor-implemented method of claim 26, wherein the individual anatomical structure is the knee joint, and the soft tissue markings include markings for the femoral cartilage and markings for the tibial cartilage.
33. A non-transitory processor-readable medium comprising processor-readable program code, the processor-readable program code being executable by at least one processor to perform the processor-implemented method of claim 26.
34. A neural network trained according to the processor implementation method of claim 26.
35. The neural network of claim 34, wherein the neural network is a U-net neural network.
36. A processor-implemented method for soft tissue labeling of computed tomography (CT) images, the processor-implemented method comprising: Provide an input CT image as input to the neural network as described in claim 34; as well as The neural network receives soft tissue markers from the input CT image.
37. A system for marking soft tissue in computed tomography (CT) images, the system comprising: At least one processor; as well as A memory coupled to the at least one processor, the memory storing instructions that, when executed by the at least one processor, cause the at least one processor to: Provide an input CT image as input to the neural network as described in claim 34; and The neural network receives soft tissue markers from the input CT image.
38. The processor-implemented method of claim 26, wherein the soft tissue is cartilage.