Deep learning based organ segmentation quality assurance for medical images
Through a deep learning-based multi-channel reconstruction model and similarity measurement, the quality of medical image segmentation is automatically evaluated, which solves the problem of insufficient accuracy of segmentation models in clinical applications in existing technologies and realizes automated quality control and error visualization of segmentation results.
Patent Information
- Application Number
- CN202510297977.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-03-18
- Filing Date
- 2025-03-13
- Publication Date
- 2025-09-19
AI Technical Summary
Existing deep learning-based medical image segmentation models are difficult to guarantee accuracy and quality in clinical applications, especially when it comes to inter-patient variability and segmentation of small anatomical structures, which can easily lead to inaccurate or insufficient results. Existing technologies lack effective automatic evaluation methods.
A deep learning-based framework is provided to reconstruct and measure the similarity of segmentation masks through a multi-channel reconstruction model, automatically evaluate the quality of segmentation masks, train a neural network model using an unsupervised anomaly detection method, generate corresponding reconstructed versions and evaluate their similarity with the original masks, generate quality assessment results and present them to medical professionals.
It realizes the automated quality assessment of the automatic segmentation results of medical images, reduces the risk of using inaccurate segmentation masks, improves the accuracy and safety of clinical applications, and provides visualization and automatic adjustment functions for errors.
Smart Images

Figure CN120672770A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to medical image processing, and more particularly to a deep learning framework for evaluating the quality of automatic segmentation of medical images. Background Art
[0002] In recent years, deep learning-based methods, particularly convolutional neural network-based methods, have shown great promise in medical image segmentation. Applications include object or lesion classification, organ or lesion detection, organ and lesion segmentation, registration, and other tasks. However, in order to be successfully applied to clinical applications such as intensity modulated radiation therapy (IMRT), the automated segmentation of organs at risk (OARs) must be of sufficient accuracy and quality, which may be difficult to achieve due to inter-patient variability and the large number of anatomical structures to be segmented in a relatively small area. Many deep learning-based segmentation models may also produce inaccurate or insufficient results when the input scans are not acquired using the correct imaging protocol, do not fully cover the entire anatomy, or are affected by severe artifacts. Therefore, there is a need for techniques to automatically evaluate the output accuracy of such segmentation models before using the segmentation results for clinical applications. Summary of the Invention
[0003] The following presents a summary of the invention to provide a basic understanding of one or more embodiments of the invention. This summary is not intended to identify key or important elements, nor is it intended to delineate any scope of the different embodiments or any scope of the claims. Its sole purpose is to present concepts in a simplified form as a prelude to the more detailed description that is presented later. In one or more embodiments described herein, a system, computer-implemented method, apparatus, and / or computer program product that provides a deep learning framework for evaluating the quality of automatic segmentation of medical images is described.
[0004] According to one embodiment, a system is provided, comprising: a memory storing a computer-executable component; and a processor executing the computer-executable component stored in the memory. The computer-executable component may include a receiving component, a model execution component, a quality assessment component, and a presentation component. The receiving component receives a segmentation mask generated from medical image data depicting an anatomical region of a subject via one or more segmentation models, wherein each of the segmentation masks depicts a different anatomical structure from a set of different anatomical structures included in the anatomical region. The model execution component generates a reconstructed version of the segmentation mask based on applying a multi-channel reconstruction model to the segmentation mask, wherein the reconstructed version corresponds to an optimized version of the segmentation mask. The quality assessment component determines a quality assessment of the segmentation mask based on comparing the segmentation mask with the reconstructed version and generates output data regarding the quality assessment, and the presentation component presents the output data via an electronic output device.
[0005] In various implementations, the quality assessment component determines, for each segmentation mask, a similarity metric between the segmentation mask and a reconstructed version of the segmentation mask, wherein the similarity metric represents a quality metric of the segmentation mask as generated via the one or more segmentation models; determines whether the segmentation mask is associated with an error based on whether the similarity metric satisfies a threshold similarity metric; and generates warning data indicating that the segmentation mask is associated with the error based on determining that the segmentation mask is associated with the error, wherein the output data includes the warning data. In one or more implementations, the threshold similarity metric varies for the different anatomical structures.
[0006] In some embodiments, based on determining that the segmentation mask is associated with the error, the quality assessment component determines error information regarding differences between a size and / or geometry of the segmentation mask and the reconstructed version based on comparing the segmentation mask with the reconstructed version, and wherein the presentation component presents the warning data and the error information via an electronic display in association with presenting the segmentation mask and the reconstructed version of the segmentation mask.
[0007] Additionally or alternatively, the quality assessment component determines whether the segmentation masks collectively meet an acceptable quality standard based on the total similarity metric determined for the segmentation masks, and wherein the output data indicates whether the segmentation masks collectively meet the acceptable quality standard. In some specific implementations, the computer-executable component further includes an adjustment component that adjusts use of the segmentation mask by a clinical application based on whether the segmentation masks collectively meet the acceptable quality standard.
[0008] In various embodiments, the multi-channel reconstruction model comprises a neural network model, and the computer-executable component further comprises a training component that trains the multi-channel neural reconstruction model using an unsupervised machine learning method. The unsupervised machine learning method comprises: training the multi-channel reconstruction model to generate a reconstruction mask of the corresponding noise-enhanced segmentation mask as included in the corresponding training dataset using a reference truth segmentation mask for the corresponding noise-enhanced segmentation mask, wherein the reference truth segmentation mask respectively depicts the different anatomical structures as extracted from the training medical image data, and wherein the corresponding noise-enhanced segmentation mask comprises a noise-enhanced version of the reference truth segmentation mask. Using these embodiments, the computer-executable component may further comprise a noise enhancement component that generates the noise-enhanced segmentation mask from the reference truth segmentation mask, wherein for each reference truth segmentation mask, the noise enhancement component integrates a certain amount of noise data into the reference truth segmentation mask that is customized based on the size and geometry of the anatomical structure depicted in the reference truth segmentation mask.
[0009] In some embodiments, the elements described in the disclosed systems and methods may be embodied in a different form, such as a computer-implemented method, a computer program product, or another form. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 An example system that facilitates using a deep learning framework to assess the quality of automated segmentation of medical images in accordance with one or more embodiments of the disclosed subject matter is presented.
[0011] Figure 2 A high-level flow chart is presented of an example computer-implemented method for generating a reconstruction model configured to generate a reconstructed version of an input segmentation mask, according to one or more embodiments of the disclosed subject matter.
[0012] Figure 3 Illustrate one or more embodiments of the disclosed subject matter Figure 2 Example implementations of the methods presented in .
[0013] Figure 4 A high-level flow chart of an example computer-implemented method for automatically evaluating the output quality of a multi-structure segmentation model in accordance with one or more embodiments of the disclosed subject matter is presented.
[0014] Figure 5 Illustrate one or more embodiments of the disclosed subject matter Figure 4 Example implementations of the methods presented in .
[0015] Figure 6 Example graphical output data providing example quality assessment results according to one or more embodiments of the disclosed subject matter is illustrated.
[0016] Figure 7 A flowchart of an example computer-implemented method for evaluating output quality of a multi-structure auto-segmentation model is presented in accordance with one or more embodiments of the disclosed subject matter.
[0017] Figure 8 A high-level flow diagram of an example computer-implemented method for generating a multi-channel segmentation mask reconstruction model according to one or more embodiments of the disclosed subject matter is presented.
[0018] Figure 9 A flowchart of another example computer-implemented method for evaluating output quality of a multi-structure automatic segmentation model in accordance with one or more embodiments of the disclosed subject matter is presented.
[0019] Figure 10A flowchart is presented of an example method for evaluating the output quality of a single structure automatic segmentation model in accordance with one or more embodiments of the disclosed subject matter.
[0020] Figure 11 A block diagram illustrating an example non-limiting operating environment in which one or more embodiments described herein may be facilitated. DETAILED DESCRIPTION
[0021] The following detailed description is merely illustrative and is not intended to limit the application or use of the embodiments and / or the embodiments. In addition, it is not intended to be bound by any express or implied information set forth in the aforementioned "Background Technology" section, "Summary of the Invention" section, or "Detailed Description of the Invention" section.
[0022] The disclosed subject matter relates to systems, computer-implemented methods, apparatus and / or computer program products that facilitate automatically evaluating the output quality of one or more medical image segmentation models. The disclosed technology is advanced by using deep learning-based automatic segmentation of OARs in magnetic resonance (MR) imaging data and computed tomography (CT) data to guide the performance of IMRT. There are many reasons why deep learning multi-organ segmentation in MR and CT may result in inaccurate organ contours. For example, a multi-organ segmentation model may have been trained for a specific MR sequence and may fail when the input is not acquired using the correct imaging protocol. Because it is very challenging to identify all variants of MR sequences, correct segmentation cannot be guaranteed for all variants. The disclosed technology provides for automatic detection and identification of abnormal segmentation results, and accordingly notifies appropriate entities (e.g., oncologists, dosimetrists, etc.) before the results are used in clinical applications such as IMRT.
[0023] To facilitate this purpose, the disclosed technology provides an automated, deep learning-based framework for assessing the quality of medical image segmentation. In various embodiments, medical image segmentation includes segmentation masks (e.g., referred to herein as automatic segmentation masks) automatically generated via one or more segmentation models (such as deep learning-based segmentation models or other types of automatic medical image segmentation models). However, the disclosed technology can also be applied to manual segmentation. In this regard, the solution is intended to detect abnormal anatomical structure automatic segmentation masks and minimize the risk of exposing inaccurate automatic segmentation masks to medical professionals. The anatomical structures (e.g., organs, vessels, tissues, regions of interest (ROIs), etc.) segmented via one or more segmentation models can vary. The disclosed framework utilizes an unsupervised anomaly detection approach. In various embodiments, the disclosed framework trains a deep learning neural network model (referred to herein as a reconstruction model) to reconstruct an input segmentation mask into a target version of the corresponding input segmentation mask, which corresponds to the ground truth (GT) samples used during training. The input segmentation masks correspond to segmentation masks for different anatomical structures, and the selected target automatic segmentation model is configured to automatically generate the segmentation masks for a specific type of medical image data input (e.g., with respect to modality, depicted anatomical region, adopted acquisition parameters, etc.).
[0024] During training, initially only the GT segmentation mask is required. In various embodiments, the GT segmentation mask may correspond to a previously organized GT segmentation mask that was manually annotated (e.g., by one or more medical professionals). In association with training the reconstruction model, noise data tailored to the specific anatomical structure being depicted is applied to each of the GT segmentation masks, resulting in a noise-enhanced version of the GT segmentation mask. The reconstruction model is then trained to accurately reconstruct the GT segmentation mask from the corresponding noise-enhanced version.
[0025] During inference mode (e.g., after training), the segmentation masks for the anatomical structures, as generated by the target segmentation model, are input to a trained version of the reconstruction model, which aims to roughly estimate the ground-truth segmentation mask. To this end, during inference mode, the reconstruction model generates a reconstructed version of the corresponding segmentation mask. As a result of training the reconstruction model, the reconstructed version corresponds to an optimal version of the segmentation mask, such as a correct version without any segmentation errors. The disclosed technology also evaluates the similarity between each individual segmentation mask and its reconstructed version to determine whether each segmentation mask is of sufficient quality (e.g., relative to one or more defined quality criteria). For example, in some embodiments, the similarity assessment may involve calculating a similarity metric (e.g., a Dice coefficient or similarity measure) between each individual segmentation mask and its reconstructed version based on the amount of overlap between the pixels or voxels included in the corresponding segmentation masks. The similarity metric may also reflect differences in size and geometry (or shape) between the corresponding segmentation masks. To this end, the similarity metric is independent of pixel intensity and / or voxel intensity.
[0026] In this regard, the similarity measure between the automatic segmentation mask and its corresponding reconstructed version represents a quality measure or accuracy measure of the automatic segmentation mask as generated via the target segmentation model. If the similarity measure of the automatic segmentation mask falls below a defined similarity threshold, the disclosed system may further identify and classify the automatic segmentation mask as of insufficient quality (and therefore inaccurate, associated with errors, outliers, etc.). In various embodiments, the defined similarity threshold may be customized for the specific anatomical structure depicted in the automatic segmentation mask (e.g., the similarity measure threshold may vary for different types of anatomical structures). In this manner, when applied to evaluating the quality of the automatic segmentation results of a multi-structure segmentation model, the disclosed system may determine a corresponding similarity measure for the segmentation mask of each structure and identify any mask of insufficient quality. The quality assessment may also involve generating an overall quality score for the total segmentation mask based on the total similarity measure determined for each structure.
[0027] The results of the quality assessment can be used to regulate the use of the automated segmentation masks by medical professionals and / or other clinical applications. For example, in some embodiments, the results of the quality assessment can be presented to appropriate medical professionals via a suitable output device and can be used to notify the medical professionals about any automated segmentation results that are considered to be of low or insufficient quality, thereby prompting them to review the results before using them for clinical purposes, such as prescribing radiation doses to organs at risk, such as for IMRT applications. For example, the disclosed system can generate warnings and / or notifications regarding a single segmentation mask that is determined to be associated with an error, and / or warnings and / or notifications regarding a group or cluster of segmentation masks that are determined to be of insufficient overall quality.
[0028] In some embodiments, in association with presenting the quality assessment results, the disclosed system may present the automatically segmented mask determined to be of low quality or insufficient quality via a graphical display following its corresponding reconstructed version, so that the differences between the corresponding masks can be visually viewed and the basis for the automatically segmented mask being classified by the system as having low quality or insufficient quality can be observed. In this way, the system can provide a viewer (e.g., a radiologist or another medical professional) with visual information about the automatically segmented areas or locations associated with errors. In some specific implementations, in association with determining the quality assessment, based on determining that the automatically segmented mask is associated with an error, the system may determine error information about the differences between the size and / or geometry of the automatically segmented mask and its reconstructed version (e.g., based on a comparison of the corresponding masks), and present the error information in association with presenting a comparison view of the automatically segmented mask and its reconstructed version. For example, the error information may indicate areas or locations of the automatically segmented mask associated with errors, such as areas / locations that are over-segmented or under-segmented, etc.
[0029] In some embodiments, the disclosed system may also automatically control and / or regulate the use of the automated segmentation mask by another automated clinical workflow based on the results of the quality assessment (e.g., involving one or more automated methods using the automated segmentation mask, such as automatically calculating doses to organs at risk for IMRT, etc.). For example, in some implementations, the disclosed system may prevent or block an automated clinical workflow from receiving and / or processing the automated segmentation mask based on the results of the quality assessment indicating that the automated segmentation mask fails to meet acceptable quality standards.
[0030] In various embodiments, the reconstruction model includes a multi-channel deep neural network model that is configured to simultaneously process multiple different input segmentation masks via different channels, one channel for each different anatomical structure. However, in other embodiments, the reconstruction model may alternatively include a single-channel model that is configured to process a single type of input segmentation mask corresponding to a single type of anatomical structure. Using these embodiments, separate reconstruction models can be trained for different anatomical structure segmentations. However, multi-channel models provide technical advantages over single-channel variants. For example, compared to using multiple corresponding separate models (each for a different anatomical structure), multi-channel models can have a significantly smaller memory footprint. In addition, as compared to using multiple separate models for each different anatomical structure, multi-channel models can be executed at faster inference speeds, thereby reducing the amount of time it takes to generate corresponding outputs. Multi-channel models are also easier to develop and maintain. In addition, multi-channel models can learn and exploit the relative positioning between different anatomical structures during training, resulting in more accurate reconstructed segmentation masks compared to single model variants.
[0031] The disclosed solution is also modality-independent, so while it was developed for MR and CT segmentation, it can be applied to any other type of medical image. In this regard, the types of medical images processed / analyzed using the techniques described herein can include images captured using various types of image capture modalities. For example, medical images can include, but are not limited to: radiation therapy (RT) images, X-ray (XR) images, digital radiography (DX) X-ray images, X-ray angiography (XA) images, panoramic X-ray (PX) images, computed tomography (CT) images, mammography (MG) images (including tomosynthesis equipment), magnetic resonance imaging (MRI or simply MR) images (including T1-weighted images and T2-weighted images), ultrasound (US) images, color flow Doppler (CD) images, positron emission tomography (PET) images, single photon emission computed tomography (SPECT) images, nuclear medicine (NM) images, optical images, and DWI, among others. Medical images may also include synthesized versions of native medical images, such as synthesized X-ray (SXR) images, modified or enhanced versions of native medical images, enhanced versions of native medical images, and the like generated using one or more image processing techniques. The types of medical image data processed / analyzed herein may include two-dimensional (2D) image data, three-dimensional image data (3D) (e.g., volumetric representations of anatomical regions of the body), and combinations thereof.
[0032] One or more embodiments will now be described with reference to the accompanying drawings, wherein like reference numerals are used throughout to represent like elements. In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a more thorough understanding of one or more embodiments. However, it will be apparent that in various circumstances, one or more embodiments may be practiced without these specific details.
[0033] Now turning to the accompanying drawings, Figure 1 A block diagram of an example, non-limiting computing system 100 is illustrated that facilitates the use of a deep learning framework to assess the quality of automated medical image segmentation, according to one or more embodiments of the disclosed subject matter. Embodiments of the system described herein may include one or more machine-executable or computer-executable components embodied within one or more machines (e.g., embodied in one or more computer-readable storage media associated with the one or more machines). When executed by one or more machines (e.g., processors, computers, computing devices, virtual machines, etc.), such components may cause the one or more machines to perform the operations described.
[0034] For example, the computing system 100 includes several machine / computer executable components, including a receiving component 110, a pre-processing component 112, a noise enhancement component 114, a training component 116, a reconstruction model 118, a model execution component 120, a quality assessment component 122, a presentation component 124, an adjustment component 126, and one or more other clinical applications 128. These computer / machine executable components may be stored in a memory 132 of the computing system 100, which may be coupled to a processing unit 134 for execution thereof. The computing system 100 may also include one or more input / output devices 136 that facilitate receiving user input and / or presenting output data to a user in association with the features and functionality of the machine / computer executable components. The computing system 100 may also include a system bus 130 that communicatively and operably couples the memory 132, the processing unit 134, and the input / output devices 136 to each other. Examples of the described and memory, processing units, input / output devices 136, and other suitable computer or computing-based elements may be referenced. Figure 9 Find and combine to achieve combination Figure 1 or the systems or components shown and described in other figures disclosed herein.
[0035] According to various embodiments, the computing system 100 is configured to process and evaluate the quality of automatically segmented data generated for medical images. To facilitate this purpose, the computing system 100 uses the GT segmentation data 106 to train a deep neural network model (e.g., a reconstruction model 118) (e.g., via a training component 116) to reconstruct a noise-enhanced version of the corresponding GT segmentation mask included in the GT segmentation data 106 (e.g., as generated via the noise enhancement component 114). For example, in various embodiments, the reconstruction model 118 may include or correspond to a convolutional denoising autoencoder. However, the reconstruction model 118 is not limited to this architecture and may include or correspond to other types of neural network models designed to process and reconstruct image data and / or spatial data.
[0036] Once trained, the computing system 100 applies (e.g., via the model execution component 120) the trained version of the reconstruction model 118 to the runtime segmentation data 108 to generate a reconstructed version of the corresponding automatic segmentation mask included in the runtime segmentation data 108. As a result of training the reconstruction model 118, the reconstructed version corresponds to an optimal or correct version of the automatic segmentation mask without any segmentation errors. The quality assessment component 122 also determines an assessment of the quality of the corresponding automatic segmentation masks included in the runtime segmentation data 108 based on the determined similarity measure between the corresponding automatic segmentation masks and their corresponding reconstructed versions. Figure 1The results of the quality assessment represented as quality assessment result data 138 correspond to output data, which can be generated by the quality assessment component 122 and presented via any suitable electronic output device (e.g., via the presentation component 124) (e.g., in an input / output / device 136 such as an electronic display), stored, provided to another system / device and / or used by the regulation component 126 to regulate the use of the runtime segmentation data by one or more other clinical applications 128 (e.g., downstream automated methods, functions, applications, inference models, etc.).
[0037] In this regard, if Figure 1 As illustrated, the computing system 100 can be communicatively and / or operatively coupled (e.g., via one or more wired or wireless communication networks) to a medical image database 102 and an automatic segmentation application 104, from which a receiving component 110 can receive, collect, or otherwise obtain GT segmentation data 106 and runtime segmentation data 108. The automatic segmentation application 104 can include or correspond to a medical image automatic segmentation application configured to execute one or more segmentation models on input medical image data and generate corresponding segmentation mask data (i.e., runtime segmentation data 108) for the medical image data. The segmentation mask data may include or correspond to image data that defines or outlines one or more defined anatomical structures as included in one or more input medical images (including one or more two-dimensional (2D) medical images (e.g., X-ray images, MR slice images, CT slice images, etc.) and / or one or more three-dimensional (3D) or volumetric medical images (e.g., MR volumetric images, CT volumetric images, etc.). Additionally or alternatively, the segmentation mask data for a particular anatomical structure may include or correspond to an extracted portion of the input medical image, the input medical image only being an extracted portion of the anatomical structure. Including particular anatomical structures as separated or segmented from the input medical image data. In embodiments where the input medical image data comprises a 3D or volumetric medical image, the segmentation mask extracted for the anatomical structures depicted in the input medical image may comprise a 3D or volumetric segmentation mask. Likewise, in embodiments where the input medical image data comprises a 2D medical image, the segmentation mask extracted for the anatomical structures depicted in the input medical image may comprise a 2D segmentation mask. As used herein, the term "automatic segmentation mask" is used to refer to a segmentation mask that is automatically generated via a segmentation model (e.g., as executed by an automatic segmentation application 104, etc.).
[0038] In this regard, the particular type of medical image data (e.g., modality, anatomical region depicted, imaging acquisition parameters, etc.) processed by the one or more segmentation models executed by the automatic segmentation application 104 may vary. Likewise, the number and type of anatomical structures automatically segmented by the corresponding segmentation models applied to the input medical image data via the automatic segmentation application 104 may also vary.
[0039] According to various embodiments, the disclosed technology is applied to evaluate multiple (e.g., two or more) anatomical structure segmentation masks corresponding to different anatomical structures that are respectively segmented from the same input medical image data (e.g., including one or more 2D and / or 3D medical images). Utilizing these embodiments, the runtime segmentation data 108 may include two or more automatic segmentation masks segmented from the same input medical image, wherein the two or more automatic segmentation masks respectively depict different anatomical structures (e.g., organs, tissues, vessels, lesions, ROIs, etc.). In one example, the different anatomical structures may respectively correspond to different organs depicted in the input medical image data of a particular anatomical region of the body. For example, according to Figure 3 and Figure 5 In the example illustrated in , the different organs include a defined group of organs of the pelvic anatomy as included in the captured 3D MR image data of the pelvis (e.g., bladder, left and right femoral heads, bulbus of penis, prostate, rectum, and urethra). Using these embodiments, the automatic segmentation model applied by the automatic segmentation application 104 may correspond to a multi-organ segmentation model that is configured to process a specific type of input medical image data (e.g., with respect to modality, depicted anatomical region, acquisition parameters, etc.) and generate corresponding segmentation masks for the different organs. However, the different anatomical structures are not limited to organs and may include other structures such as tissues, blood vessels, lesions, and defined ROIs. Alternatively, the automatic segmentation model applied by the automatic segmentation application 104 may correspond to multiple different single structure segmentation models applied to the same input medical image data, wherein each of the single structure segmentation models generates a segmentation mask for a different anatomical structure depicted in the input medical image.
[0040] However, in other embodiments, the segmentation model applied by the automatic segmentation application 104 may include or correspond to a single structure segmentation model that is configured to generate a single segmentation mask for a single target anatomical structure or ROI depicted in the input medical image data, and the runtime segmentation data 108 may include the single segmentation mask.
[0041] It should be understood that the automatic segmentation application 104 can be deployed and / or executed by any suitable computing device or system that can be communicatively coupled to the computing system (e.g., via one or more wired communication networks or wireless communication networks). In other embodiments, the automatic segmentation application 104 can be stored in the memory 132 and executed by the computing system. Various architectural configurations are contemplated. In some embodiments, the input medical image data that is processed by the automatic segmentation application 104 to generate the runtime segmentation data 108 can be received from the medical image database 102. Additionally or alternatively, the input medical image data processed by the automatic segmentation application can be received directly from a medical image acquisition system (e.g., an X-ray system, a CT system, an MR system, etc.) and / or another suitable source of medical image data.
[0042] The ground truth (GT) segmentation data 106 corresponds to a training dataset of GT segmentation masks whose output accuracy is being estimated by the particular automatic segmentation model that the computing system 100 is configured to generate. In this regard, the reconstruction model 118 can be trained (e.g., via the training component 116) to reconstruct one or more segmentation masks corresponding to those generated by the particular automatic segmentation model executed by the automatic segmentation application 104. Thus, the GT segmentation data 106 will correspond to GT examples of the runtime segmentation data 108 and will vary depending on the particular automatic segmentation model selected for evaluating its output segmentation accuracy via the computing system 100. For example, when applied to a multi-organ segmentation model that is configured to generate 3D segmentation masks for a defined set of different organs included in a pelvic region depicted in 3D MR image data, the GT segmentation data 106 will include GT segmentation masks for different organs as generated for multiple different 3D MR images of the pelvic region of different subjects / patients. Thus, the disclosed embodiments of the computing system 100 assume that the particular automatic segmentation model whose output accuracy is being evaluated by the computing system 100 has been selected and / or otherwise indicated (eg, via user input or via another mechanism).
[0043] However, it should be understood that the computing system 100 can be employed to generate any number of different reconstruction models 118 customized for different automatic segmentation models. In various embodiments, the GT segmentation data 106 includes or corresponds to a manually generated / defined segmentation mask that defines a specific outline of the anatomical structure of interest as applied to the corresponding medical image. Figure 1As illustrated, the GT segmentation data 106 can be included in the medical image database 102 and received from (e.g., collected, received, provided to, etc.) the medical image database by the receiving component 110. In other embodiments, the GT segmentation data 102 can be stored locally 132 and / or provided to the computing system 100 via another source (e.g., a medical imaging annotation application that provides for adding manually annotated segmentation data to medical image data, etc.).
[0044] refer to Figures 2 to 10 Additional features and functionality of the computer-executable components of computing system 100 are described.
[0045] Figure 2 A high-level flow diagram is presented of an example computer-implemented method 200 for generating a reconstruction model (eg, the reconstruction model 118 ) configured to generate a reconstructed version of an input segmentation mask, according to one or more embodiments of the disclosed subject matter.
[0046] refer to Figure 1 and Figure 2 , method 200 begins after a particular segmentation model (hereinafter referred to as a target segmentation model) has been selected for output quality assessment. According to method 200, the target segmentation model may include a multi-structure segmentation model configured to generate segmentation masks for multiple different anatomical structures depicted in an input medical image (or multiple input medical images). Alternatively, the target segmentation model may include multiple separate segmentation models that are each configured to segment different anatomical structures from the same input image (e.g., one structure per segmentation model).
[0047] According to method 200, the reconstruction model 118 corresponds to a multi-channel reconstruction model comprising multiple channels configured to simultaneously (or in parallel) process multiple different input segmentation masks, one for each channel. The different input segmentation masks correspond to different anatomical structures, and the target segmentation model is configured to automatically segment the different anatomical structures from an input medical image (or multiple input medical images). However, in other embodiments, the reconstruction model 118 may alternatively comprise a single-channel model configured to process a single type of input segmentation mask corresponding to a single type of anatomical structure. Using these embodiments, separate reconstruction models 118 can be trained for different anatomical structure segmentation using the same or similar techniques used to train the multi-channel reconstruction model, but applied only to a single channel. However, multi-channel reconstruction models offer technical advantages over single-channel variants. For example, multi-channel models can have a significantly smaller memory footprint and provide significantly faster inference speeds than using multiple corresponding separate models, each for a different anatomical structure. Multi-channel models are also easier to develop and maintain. Furthermore, the multi-channel model can learn and exploit the relative positioning between different anatomical structures during training, leading to more accurate reconstructed segmentation masks relative to single model variants.
[0048] At 202, the receiving component 110 can obtain a GT segmentation mask (e.g., the GT segmentation data 106) for each anatomical structure, and the target segmentation model is configured to automatically segment the GT segmentation mask. This can include or correspond to multiple training sets of GT segmentation masks for different anatomical structures, each accurately extracted from different medical images depicting the same anatomical region of different subjects / patients. For example, in some embodiments, the GT segmentation mask can be manually generated. In other embodiments, the GT segmentation mask can include an automatically generated segmentation mask that has been determined to be accurate (e.g., based on manual review thereof and / or via another mechanism). To this end, the GT segmentation mask corresponds to an anatomically accurate or correct segmentation mask that perfectly defines the boundaries or contours of the corresponding anatomical structure (e.g., without errors, as determined based on manual generation thereof and / or manual review of the automatic segmentation mask, etc.).
[0049] In this regard, in some embodiments, each of the plurality of training sets may correspond to a different medical image (e.g., captured from a different subject / patient) and include a group of segmentation masks, one mask for each of the different anatomical structures as extracted from the corresponding medical image. For example, when applied to a target segmentation model that is a 3D MR organ segmentation model for different defined organs of the pelvic region, one training set of GR segmentation masks may include a group of segmentation masks, the group including different organ segmentation masks, as extracted from a single MR volume image of the pelvis for a specific subject. Another training set of GR segmentation masks may include the group of segmentation masks, the group including different organ segmentation masks, as extracted from another single MR volume image of the pelvis for another specific subject, and so on.
[0050] The term "manual extraction" is used herein to refer to the use of manual interaction to define the outlines of the corresponding anatomical structures as included in the corresponding medical image. This is typically achieved using medical imaging annotation software that provides for viewing the corresponding medical image via a graphical user interface and provides annotation tools that enable expert annotators (e.g., radiologists or other trained medical professionals) to mark and define the outlines of the corresponding anatomical structures on or within the displayed medical image. In some cases, the software may generate segmentation masks for the corresponding anatomical structures based on user input, and these segmentation masks may correspond to GT segmentation masks. In other cases, the user may manually create (e.g., draw and define) the segmentation masks. Regardless of the mechanism by which their generation via manual input is facilitated, it should be understood that the GT segmentation masks may include or correspond to 3D (e.g., extracted from a 3D medical image) or 2D (e.g., extracted from a 2D medical image) image data that accurately defines the outline, size, and geometry of the corresponding target anatomical structure as depicted in the corresponding medical image. In various embodiments, at 202, the receiving component 110 can obtain a GT segmentation mask as having been (manually) applied to (or otherwise extracted from) a corresponding medical image as included in the medical image database 102. In other embodiments, the GT segmentation mask can be manually organized at 202.
[0051] At 204, the pre-processing component 112 may pre-process the GT segmentation masks to normalize the respective GT segmentation masks before they are used to train the reconstruction model 118. This may involve performing one or more image processing functions on the respective GT segmentation masks to adjust and normalize their spacing, size, orientation, and / or visual appearance. The one or more pre-processing functions applied to the GT segmentation masks by the pre-processing component 112 may vary depending on the type of medical image data from which they are extracted (e.g., modality, depicted anatomical region, acquisition protocol used, etc.). For example, in some embodiments in which the GT segmentation masks correspond to 3D MR data covering pelvic anatomy, as extracted from 3D MR images, the pre-processing at 204 may include resampling the respective GT segmentation masks to isotropic spacing (e.g., spacing of 1.5 millimeters (mm) × 1.5 mm × 1.5 mm) and padding the respective GT segmentation masks to a common, predefined spatial resolution (e.g., 240 voxels × 336 voxels × 336 voxels). In other example implementations where the GT segmentation masks correspond to 3D CT data as extracted from a 3D CT image, the preprocessing at 204 may include rotating each GT segmentation mask to a common orientation relative to the subject, thereby resampling the GT segmentation masks to a defined interval (e.g., 1.0 mm × 1.0 mm × 3.0 mm) and padding the GT segmentation masks to a predefined spatial size (e.g., 700 voxels × 700 voxels × 620 voxels). The preprocessed GT segmentation mask generated at 204 corresponds to a target mask used by the reconstruction model 118 during training in association with learning how to reconstruct its corresponding noise-enhanced version to make it similar.
[0052] At 206, the noise enhancement component 114 may generate a noise-enhanced version of the pre-processed GT segmentation mask. This may involve, for each pre-processed GT segmentation mask, inserting or applying noise data to the pre-processed GT segmentation mask that is customized for the depicted anatomical structure, resulting in a noise-enhanced version of the pre-processed GT segmentation mask (i.e., a target mask). Generally, this includes adding an amount of noise data to the corresponding pre-processed GT segmentation mask and / or removing an amount of noise data from the corresponding pre-processed GT segmentation mask in proportion to the size, shape, and / or geometry of the corresponding anatomical structure being depicted. The noise-enhanced segmentation mask corresponds to the input mask used during training of the (multi-channel) reconstruction model (e.g., provided as input to the reconstruction model 118), and the pre-processed GT segmentation mask corresponds to the target mask. In various embodiments, the goal of adding noise to the segmentation mask is to create an input segmentation mask that has errors for the depicted anatomical structure. Such errors may correspond to segmentation errors (e.g., those typically generated by the target segmentation model), but may also correspond to any type of deformation of the depicted anatomical structure (e.g., holes, missing regions, or blocks of added and missing regions, etc.). In this regard, by using such noise-enhanced segmentation masks as input, as well as their error-free corresponding GT versions to train the reconstruction model 118, the reconstruction model 118 is forced to learn the correct physical properties (e.g., size and geometry) of the anatomical structure. Thus, the trained version of the reconstruction model 118 will be able to correct any segmentation errors because it has learned the correct physical properties (e.g., size and geometry or shape) of the anatomical structure.
[0053] The method used by the noise enhancement component 114 to generate the noise-enhanced version of the target mask may vary based on parameters of the image data corresponding to the target mask (including whether the image data is 2D or 3D image data), the modality of the image data (e.g., MR, CT, X-ray, PET, etc.), the acquisition protocol used, and other parameters. In some embodiments, when applied to a 3D segmentation mask (such as, for example, a 3D MR organ segmentation mask), the noise enhancement method may include adding or removing random, binary blocks of pixels or voxels to or from the target mask (e.g., with probabilities 0, 5 and 0, 5) according to defined hyperparameters and customized values of the hyperparameters defined for the corresponding anatomical structure being depicted. Pixel blocks or voxel blocks correspond to groups of two or more pixels or voxels having a 2D or 3D geometric array, depending on whether the image data is 2D or 3D. In some implementations, hyperparameters may include, but are not limited to, a maximum number of blocks, a minimum block size, a maximum block size, and a sampling method (referring to the logic used to sample the centers of the blocks). To this end, the specific values for these hyperparameters can be customized based on the size and / or geometry of the corresponding anatomical structure depicted in the target mask, so that the amount of noise applied to each target mask is proportional to the size / geometry of the corresponding anatomical structure. In one example implementation, for each anatomical structure (e.g., each organ, etc.), the hyperparameters can be optimized according to the following equation 1 so that the signed Dice coefficient (sDC) of the noisy input mask roughly covers the interval [-1, 1], and where and are the input binary segmentation mask and the target binary segmentation mask, respectively, where sgn(·) represents the sign function.
[0054]
[0055] In another example embodiment, when applied to a 3D segmentation mask generated from CT image data (i.e., a 3D CT organ segmentation mask), a noise enhancement method may include adding or removing random, binary blocks of pixels or voxels to or from a target mask (e.g., with probabilities 0, 5 and 0, 5) based on defined hyperparameters and customized values of the hyperparameters defined for the corresponding anatomical structures depicted. In some implementations, the hyperparameters may include, but are not limited to, a maximum number of blocks, a minimum block size, a maximum block size, and a sampling method (referring to the logic used to sample the centers of the blocks). To this end, the specific values for these hyperparameters may be customized based on the size and / or geometry of the corresponding anatomical structures depicted in the target mask, such that the amount of noise applied to each target mask is proportional to the size / geometry of the corresponding anatomical structure. In one example implementation, for each anatomical structure (e.g., each organ, etc.), the hyperparameters may also be optimized according to Equation 1 below so that the signed Dice coefficient (sDC) of the noisy input mask roughly covers the interval [-1, 1].
[0056] At this point in the method 200, the training component 116 now has the training dataset required to train the reconstruction model 118. In this regard, for each anatomical structure, the training dataset includes a plurality of noise-enhanced segmentation masks (i.e., input masks) that are each paired with a (pre-processed) GT segmentation mask (i.e., target mask). At 208, the training component 116 can then proceed to train the (multi-channel) reconstruction model 118 to correctly reconstruct the GT segmentation mask (i.e., target mask) from the corresponding noise-enhanced version (i.e., input mask), wherein each channel of the multi-channel reconstruction model 118 processes a different anatomical structure. In this regard, the (multi-channel) reconstruction model 118 includes a plurality of different channels that are each configured to process a different input mask corresponding to a different anatomical structure. To this end, the number of different channels corresponds to the number of different anatomical features that the target segmentation model is configured to segment. In association with the training reconstruction model 118, the training component 116 can stack the paired input masks and target masks for each of the plurality of training sets on a channel-by-channel basis, such that each different mask pair corresponding to a different anatomical structure is assigned to a designated channel for that structure. In this regard, the first channel will process multiple pairs of input masks and target masks corresponding to the first anatomical structure, the second channel will process multiple pairs of input masks and target masks corresponding to the second anatomical structure, and so on.
[0057] The processing performed by the respective channels of the reconstruction model 118 is performed in parallel or simultaneously. In this regard, each channel of the (multi-channel) reconstruction model 118 independently processes the input mask through its own set of neural network layers, which, depending on the architecture of the reconstruction model 118, may include convolutional layers, recurrent layers, or fully connected layers. The output of each channel includes a reconstructed version of the input mask and a noise-enhanced mask. The training method involves training the reconstruction model 118 to generate reconstructed versions of the respective input masks such that the reconstructed versions accurately resemble their corresponding ground-truth versions. The training component 116 can control the training method such that the respective input masks processed in parallel by the respective channels at each pass through the segmentation model 118 correspond to the same training set; that is, they correspond to segmentation masks extracted from the same medical image and, therefore, reflect different anatomical structures of a particular subject / patient. In this regard, each pass through the segmentation model refers to processing different input masks of the same training set (of multiple training sets) in parallel or simultaneously. In this manner, during training, the reconstruction model 118 may learn and exploit the spatial relationships between different anatomical structures as they are anatomically arranged, shaped, and sized relative to each other within a given medical image and subject / patient.
[0058] During training, backpropagation and one or more optimization algorithms (e.g., gradient descent, stochastic gradient descent (SGD), momentum optimization, root mean square propagation, adaptive motion estimation (Adam optimization), or another optimization function) are used to learn the parameters (e.g., filters, weights, biases, latent space representation, activation functions, etc.) of the neural network (or neural networks) employed by the reconstruction model 118, thereby optimizing the reconstruction model 118 to make accurate predictions based on the average loss (using a suitable loss function) calculated by the training component 116 for each output reconstruction mask across all channels for each pass. In various embodiments, the training component 116 may calculate a loss for each reconstruction mask generated via each channel for a given pass based on a similarity metric between the reconstruction mask and the corresponding target mask (i.e., the corresponding pre-processed GT mask). In this regard, the similarity metric may be based on the amount of overlap between the pixels or voxels included in the corresponding segmentation masks. The similarity metric may also reflect differences in size and geometry between the corresponding segmentation masks. For example, the similarity metric can include or correspond to a Dice score and / or another metric representing a similarity metric between the size and / or geometry of the reconstructed mask and the corresponding target mask. For example, in some implementations, in association with calculating the similarity metric and / or loss metric, the training component 116 can apply one or more image processing functions to the input mask and the target mask that compare the respective masks, determine differences between the size and / or geometry of the respective masks, and quantify the differences using a similarity metric (such as a Dice coefficient / score, etc.). To this end, the similarity comparison does not involve evaluating differences in pixel intensities and / or voxel intensities.
[0059] The training component 116 may also calculate an average loss (e.g., average Dice loss, etc.) across all channels for a given pass based on averaging similarity metrics (e.g., Dice scores, etc.) determined for all reconstructed masks across all channels. For this purpose, the Dice loss is defined as 1 minus the Dice score. During training, the goal is to minimize the average Dice loss or maximize the average Dice score. During the training method, the optimization algorithm employed by the training component 116 iteratively adjusts the parameters of the neural network to minimize the reconstruction error between the input mask and the reconstructed mask. The training method may iteratively proceed using conventional training and validation phases (e.g., splitting the training data accordingly) until one or more defined conditions are met (such as reaching a specified number of epochs, achieving a loss criterion, and / or reaching convergence). To this end, the results or output of the method 200 may include a trained version of the (multi-channel) reconstruction model 118, which may be saved (e.g., stored) in the memory 132 and then applied to the runtime segmentation data 108 by the model execution component 120 in association with using the reconstruction model 118 to evaluate the accuracy or quality of the runtime segmentation data 108.
[0060] In various embodiments, the reconstruction model 118 includes a convolutional denoising autoencoder. Using these embodiments, the multi-channel reconstruction model includes an encoder network and a decoder network. The encoder consists of several convolutional layers followed by maximum pooling layers. These layers gradually reduce the spatial dimension of the input mask while increasing the number of channels or features extracted. The convolutional layers apply filters to extract meaningful features from the input mask. After the convolutional layers, the network may include one or more fully connected layers to further compress the feature representation into a lower-dimensional latent space representation. This latent space representation captures the essence of the input mask features in a more compact form. The decoder portion of the network is responsible for reconstructing the input mask from the latent space representation to generate a reconstruction mask. It mirrors the architecture of the encoder, but in the opposite direction. It includes one or more fully connected layers and subsequent upsampling layers (or transposed convolutional layers) to gradually increase the spatial dimension of the data while reducing the number of channels until the final output matches the dimensions of the input mask.
[0061] Figure 3 An example implementation of method 200 according to one or more embodiments of the disclosed subject matter is illustrated. In this regard, Figure 3An example method 300 is presented for generating a reconstruction model (e.g., reconstruction model 118) configured to generate a reconstructed version of an input segmentation mask, as illustrated using MR segmentation masks for different organs included in different MR images of a pelvic region. For example, the different organs may include, for example, the bladder, left and right femoral heads, bulbus of the penis, prostate, rectum, and urethra (e.g., seven different organs). According to method 300, the target segmentation model corresponds to a multi-organ segmentation model that is configured to generate segmentation masks for the different organs included in the MR volume (or 3D) image of the pelvic region. Thus, in this example embodiment, the segmentation masks (e.g., the GT segmentation mask, the preprocessed GT segmentation mask, the noise enhancement segmentation mask, and the reconstruction segmentation mask) each correspond to a volume or 3D segmentation mask for the corresponding organ. For example, the 3D segmentation mask for the corresponding organ may correspond to an extracted portion of the MR volume image corresponding to the corresponding organ that was separated and removed from the MR volume image.
[0062] refer to Figures 1 to 3 , method 300 illustrates how different GT segmentation masks corresponding to different anatomical structures (e.g., different organs of the pelvic region in this example) can be processed separately via different channels of the (multi-channel) reconstruction model 118. The different channels are depicted as channels 1 to N, where the number of channels N corresponds to the number of different types of segmentation masks corresponding to the different anatomical structures. In this regard, each different organ segmentation mask is processed by a different designated channel of the reconstruction model 118. For ease of illustration, the different organs are generally referred to as organs 1 to N to indicate their corresponding channels, and only two example channels are explicitly illustrated (one for organ 1 and the other for the last organ N). It should be understood that the number of channels N and the number of different types of organ segmentation masks processed thereby may include any number greater than 1 (when applied to a multi-channel reconstruction model). For example, in one example of this specific implementation, the different organ segmentation masks may include segmentation masks for the bladder, left and right femoral heads, bulb of the penis, prostate, rectum, and urethra (e.g., 7 different organs, and therefore N=7). In various embodiments, as a multi-channel reconstruction model, the reconstruction model 118 includes a single convolutional denoising autoencoder. The input to the reconstruction model is a 5D tensor of shape B, N, H, W, D; where B is the batch size, N is the number of channels, H is the height, W is the width, and D is the depth. For each example in the batch, each channel corresponds to an anatomical structure (e.g., a 3D binary volume). The output of the convolutional denoising autoencoder is a 5D tensor again with the same shape as the input. In this regard, the multi-channel reconstruction model processes all organs together at once.
[0063] Method 300 demonstrates the performance of GT mask preprocessing and the generation of noise enhancement masks for corresponding organs as arranged in channels based on performing a single pass through the reconstruction model 118 on corresponding GT segmentation masks 1 to N corresponding to one training set (depicted as set 1) of a plurality of training sets. In some embodiments, these steps can be performed in this manner (e.g., in parallel as arranged in channels). In other embodiments, all GT segmentation masks for all training sets can be preprocessed into target masks and noise enhancement masks in a batch manner and then organized and stacked into their corresponding channels for processing via the reconstruction model 118. It should be understood that although method 300 is demonstrated with respect to one training set (e.g., set 1), the training component 116 iteratively performs method 300 for all training sets.
[0064] To this end, along channel 1, at 302-1, the pre-processing component 112 may pre-process the GT segmentation mask for organ 1 (set 1) to generate a target mask for organ 1 (set 1). The pre-processing performed at 302-1 may correspond to the pre-processing described with respect to step 204 of method 200. At 304-1, the noise enhancement component 114 may insert noise into the target mask specific to organ 1, thereby transforming the target mask into a noise-enhanced mask for organ 1 (set 1), which is the input mask for organ 1. The pre-processing performed at 304-1 may correspond to the noise enhancement method described with respect to step 206 of method 200 and the 3D MR segmentation mask body. Comparing the target mask for organ 1 with the input mask shows that the noise-enhanced input mask includes removed pixels or voxels, thereby providing an example of a segmentation mask with errors similar to the input segmentation mask with segmentation errors. The training component 116 also inputs the input mask for organ 1 into channel 1 of the reconstruction model 118, and the channel 1 neural network generates as output a reconstruction mask for organ 1. In this regard, channel 1 is specifically trained and configured to generate a reconstruction segmentation mask for only one specific type of anatomical structure (and image data input), which in this example corresponds to organ 1 (and its 3D MR segmentation mask).
[0065] According to method 300, the same method described above with respect to organ 1 and channel 1 is performed in parallel for all different organs and channels for the same training set (e.g., set 1). For example, step 302-N corresponds to step 302-1, but is customized for the GT mask of organ k. Similarly, step 304-N corresponds to step 304-1, but is customized for organ N as depicted in the target mask for organ N. The training component 116 also inputs the input mask for organ N into channel N of the reconstruction model 118, and the channel N neural network generates a reconstructed mask for organ N as output. In this regard, channel N is specifically trained and configured to generate a reconstructed segmentation mask for only one specific type of anatomical structure (and image data input), which in this example corresponds to organ N (and its 3D MR segmentation mask).
[0066] After all input masks for training set 1 are passed through the corresponding channels of the reconstruction model 118 once, the output includes reconstructed masks for all input masks. The training component 116 also calculates a similarity metric (e.g., a Dice score, etc.) between the reconstructed masks and their corresponding target masks. As described above, the similarity metric can reflect the amount of overlap between pixels or voxels included in the corresponding masks, the size difference between the corresponding masks, and / or the geometric structure difference between the corresponding masks, respectively. The training component 116 can also calculate the average loss (e.g., average Dice loss) across all channels for the set (e.g., set 1). The training component 116 then optimizes (e.g., tunes) the appropriate parameters of the corresponding channels according to a defined optimization function based on the loss and / or average loss (e.g., corresponding to the average loss across all channels), and repeats the training method for additional training sets.
[0067] In this example illustration, the reconstructed masks for organ 1 and organ N substantially correspond to their respective target masks. It should be understood that this result is to be expected after several training passes or rounds through the reconstruction model 118, and that initially (e.g., only the first few channels), the output reconstructed masks are likely to be inaccurate because the reconstruction model 118 has not yet optimized its parameters based on multiple passes and loss metrics between the reconstructed masks and the corresponding target masks. In this regard, during training, the training component 116 uses the difference between the reconstructed masks and their target masks (e.g., as measured via a Dice score or another similarity metric) to tune the parameters of the model until the reconstructed masks accurately reflect their corresponding target masks (according to defined acceptable loss criteria and / or other training method completion criteria). To this end, once this has been achieved, the training component 116 can end the method 300. The final output of the method 300 may include a trained version of the (multi-channel) reconstruction model 118, which may be saved (e.g., stored) in the memory 132 and then applied to the runtime segmentation data 108 by the model execution component 120 in association with using the reconstruction model 118 to evaluate the accuracy or quality of the runtime segmentation data 108, as described in reference to FIG. Figure 4 and Figure 5 Further described.
[0068] Figure 4 A high-level flow chart of an example computer-implemented method 400 for automatically evaluating the output quality of a multi-structure segmentation model according to one or more embodiments of the disclosed subject matter is presented. The method 400 is described relative to a case where the target segmentation model is the same target segmentation model involved in the method 200 and / or the method 300, and wherein the trained version of the reconstruction model 118 corresponds to the multi-channel reconstruction model described and trained according to the method 200 and / or the method 300. Figures 1 to 4 According to the method 400, at 402, the receiving component can receive the automatic segmentation results (eg, runtime segmentation data 108) from the target segmentation model. In this regard, continuing at Figure 2 and Figure 3 In the embodiment described in , the automatic segmentation result may include a set of automatic segmentation masks for different anatomical structures as extracted from medical image data (e.g., new medical image data other than that used for model training) via the target segmentation model. Figure 3 In the 3D MR organ segmentation implementation illustrated in FIG, the automatic segmentation results received at 402 may include 3D MR segmentation masks for the defined group of different organs of the pelvic region, which are automatically segmented from the 3D MR image of the pelvic region via the corresponding multi-organ segmentation model.
[0069] In some cases, due to errors in the target segmentation model or the input medical image, the result of the automatic segmentation may be missing one or more of the tissue / structure segmentations. Based on the input image data that only covers a portion of the region of interest that can be processed by the target segmentation model, the result may also be missing one or more of the tissue / structure segmentations. Therefore, in some specific implementations, at 402, the set of segmentation masks received at 202 may only include some (e.g., one or more) of the different anatomical structures that the reconstruction model 118 has been configured to process. With these specific implementations, one or more corresponding channels of the reconstruction model 118 for the missing one or more segmentation masks may be deactivated during the inference mode, or error information indicating that no segmentation mask is available for the corresponding structure may be generated.
[0070] At 404, the pre-processing component 112 pre-processes the automatic segmentation mask using the same (or similar) pre-processing operations described with respect to step 204 of the method 200 for the corresponding GT segmentation mask. In this regard, the pre-processing performed at 404 may vary depending on the type of automatic segmentation mask (e.g., modality, 3D or 2D image data, anatomical region scanned, acquisition parameters used, etc.), as described with reference to the method 200.
[0071] At 406, the model execution component 120 applies the trained version of the (multi-channel) reconstruction model 118 to the (pre-processed) automatic segmentation mask to generate a reconstructed version of the automatic segmentation mask. In this regard, the model execution component 120 can stack the corresponding automatic segmentation masks channel by channel, such that the corresponding automatic segmentation masks corresponding to different anatomical structures are input to the designated channel for the corresponding structure. The output of step 406 includes a reconstructed segmentation mask for each of the automatic segmentation masks.
[0072] For example, Figure 5 Illustrated when applied to Figure 3 3D MR segmentation mask body is implemented in the example embodiment of method 500 of steps 404 and 406 of method 400. According to method 500, the target segmentation model corresponds to the same target segmentation model associated with method 300, namely a multi-organ segmentation model configured to segment a group of organs from a 3D MR image of the pelvic region. To this end, Figure 5The reconstruction model 118 illustrated in FIG3 corresponds to a trained version of the reconstruction model 118 as trained according to method 300. According to method 500, the model execution component 120 can stack the automatic segmentation masks for each organ channel by channel, so that each organ segmentation mask is processed in parallel via its corresponding designated channel of the (multi-channel) reconstruction model 118. At 502-1 to 502-N, the pre-processing component 112 pre-processes the corresponding automatic segmentation masks (e.g., automatic segmentation mask 1 to automatic segmentation mask k) in the same manner as the GT segmentation masks at 402-2 and 402-N in method 300. For example, when applied to 3D MR segmentation masks, this may include resampling the corresponding automatic segmentation masks to an isotropic spacing (1.5 mm×1.5 mm×1.5 mm) and then padding them to a common spatial size (240 voxels×336 voxels×336 voxels). The model execution component 120 then inputs the respective pre-processed organ segmentation masks into their respective channels of the reconstruction model 118 , and the reconstruction model generates corresponding reconstruction masks.
[0073] Given that Figures 1 to 3 and Figure 5 Return Reference Figure 4 At 408, the quality assessment component 122 may assess the quality of the automatic segmentation masks (i.e., the automatic segmentation masks received at 402) based on a similarity comparison between the automatic segmentation masks and their corresponding reconstructed versions. In other embodiments, the similarity comparison may be performed between pre-processed versions of the automatic segmentation masks and their corresponding reconstructed versions. In this regard, because the reconstruction model 118 is trained to transform noisy input segmentation masks into their corresponding GT versions (which are manually defined as correct versions), the reconstructed segmentation masks generated from the automatic segmentation masks by the trained version of the reconstruction model 118 will also be similar to the assumed GT segmentation masks for the corresponding automatic segmentation masks. In other words, the reconstructed versions of the automatic segmentation masks correspond to the best versions of the automatic segmentation masks that correctly define the size and geometry of the depicted corresponding anatomical structures (e.g., without any physical or structural errors). Therefore, if the automatic segmentation mask is substantially similar to its reconstructed version generated by the reconstruction model 118, the quality assessment component 122 may deem the automatic segmentation mask to be of sufficient accuracy and / or quality. On the other hand, if the automatic segmentation mask differs substantially from its reconstructed version generated by the reconstruction model 118 , the quality assessment component 122 can deem the automatic segmentation mask to be of insufficient accuracy and / or quality and thus an outlier or otherwise associated with error.
[0074] In various embodiments, in conjunction with performing the quality assessment at 408, the quality assessment component 122 may determine, for each automatic segmentation mask included in the set received at 402, a similarity metric between the automatic segmentation mask and its corresponding reconstructed version. To this end, the similarity metric represents a quality metric and / or accuracy metric of the automatic segmentation mask as generated by the target segmentation model. For example, the similarity metric may correspond to a Dice score or similarity metric that represents a similarity metric between corresponding segmentation masks based on the amount of overlap between pixels or voxels included in the corresponding segmentation masks. The similarity metric may also reflect differences in size and / or geometry between the corresponding segmentation masks. In some embodiments, the similarity comparison performed by the quality assessment component 122 may be the same as or similar to the similarity comparison used to calculate the loss between the corresponding input mask and the target mask during training. Additionally or alternatively, the similarity comparison may involve using one or more geometry-based and / or size-based comparison functions to determine the difference between the geometry and / or size of each automatic segmentation mask and its corresponding reconstructed version. In some embodiments, one or more comparison functions may also generate visual marker data that defines or indicates the location and / or area of the automatic segmentation mask associated with the error (e.g., missing area, auxiliary area, etc.), which visual marker data may be applied to or overlaid on the automatic segmentation mask in association with presenting the automatic segmentation mask, as described below. To this end, the similarity comparison does not involve evaluating differences in pixel intensity and / or voxel intensity. In some specific implementations, the similarity comparison function of the quality assessment component 122 may be integrated within the reconstruction model (e.g., as a final processing module or layer of the reconstruction model 118, etc.).
[0075] In some embodiments, in association with evaluating the quality of each individual automatic segmentation mask, the quality assessment component 122 can determine whether each individual automatic segmentation mask is of sufficient quality based on a defined similarity criterion for the corresponding similarity metric. For example, in some implementations, the defined similarity criterion can include a defined general threshold similarity metric (e.g., a threshold Dice score, etc.) that applies to all individual automatic segmentation masks. Using these implementations, the quality assessment component 122 can determine whether a particular automatic segmentation mask for a particular anatomical structure is of sufficient quality based on whether the similarity metric for the particular automatic segmentation mask for the particular anatomical structure falls above (and therefore is acceptable) or below (and therefore is unacceptable) the general threshold similarity metric.
[0076] In other embodiments, the similarity criteria may define different threshold similarity metrics that are customized for different anatomical structures. In other words, the threshold similarity metric may vary for different anatomical structures. In this regard, the anatomical structure-specific similarity metric threshold may depend on the ability of the reconstruction model 118 to accurately reconstruct the noise-enhanced input mask for the specific structure. For example, in some embodiments, the training component 116 may determine the structure-specific threshold for each anatomical structure by maximizing the F1 score of the outlier detection performance at different thresholds that binarize the outlier scores using a validation dataset in association with the training method. Using these embodiments, the quality assessment component 122 may determine whether a specific automatic segmentation mask for a specific anatomical structure is of sufficient quality based on whether the similarity metric of the specific automatic segmentation mask for the specific anatomical structure falls above (and therefore acceptable) or below (and therefore unacceptable) a specific similarity metric threshold defined for the specific anatomical structure.
[0077] In some implementations, the quality assessment component 122 may classify any automated segmentation mask that fails to meet a general similarity metric threshold or its anatomy-specific similarity metric threshold as an outlier or associated with an error. The quality assessment component 122 may also determine an overall quality metric for the set of automated segmentation masks based on an aggregate similarity metric (e.g., an average similarity metric, an average Dice score, or another aggregate metric) determined for each of the individual segmentation masks relative to the general threshold or their anatomy-specific threshold. In some embodiments, any information determined by the quality assessment component 122 may be included in the quality assessment results data 138.
[0078] At 410, the quality assessment component 408 may also determine whether the quality of the automatic segmentation result is acceptable based on the individual similarity metrics and / or overall quality metrics determined for each of the different automatic segmentation masks and predefined acceptance criteria for the corresponding similarity metrics and / or overall quality metrics. For example, in some embodiments, at 410, the quality assessment component 408 may be configured to simply determine whether the automatic segmentation result is acceptable based on whether the overall quality metric is above (and therefore acceptable) or below (and therefore unacceptable) a defined threshold overall quality metric. In another embodiment, the quality assessment component 408 may be configured to determine that the automatic segmentation result is unacceptable based on any individual segmentation mask having a similarity metric that does not meet its specific anatomy-specific similarity metric threshold.
[0079] In some embodiments, based on determining that the quality is acceptable, at 412, the quality assessment component 122 may report the automatic segmentation result as acceptable. Similarly, based on determining that the quality is unacceptable, at 414, the quality assessment component 122 may report the automatic segmentation result as unacceptable. For example, in some specific implementations, the quality assessment component 122 may generate quality assessment result data 138 including information indicating whether the result is acceptable (e.g., text data, image data, audible data, etc.), and the presentation component 124 may present the quality assessment result data 138 via a suitable electronic output device (e.g., a display, a speaker, etc.). The quality assessment result data 138 may also include information identifying the overall quality metric and the individual similarity metrics determined for each different segmentation mask. The quality assessment result data 138 may also include warning data that identifies any of the individual automatic segmentation masks that are found to have insufficient quality and are therefore outliers (e.g., based on their similarity scores relative to their anatomy-specific thresholds). In some implementations, the presentation component 124 can be configured to present warning data or a warning notification in response to determining that a single automatic segmentation mask is found to be of insufficient quality (e.g., based on its similarity metric being below its defined similarity threshold), where the warning data can identify that the particular automatic segmentation mask may include information regarding the basis on which it is deemed insufficient. In another example, the presentation component 124 can be configured to generate and present (e.g., to a suitable medical professional via a suitable electronic output device) a notification or warning data in response to determining that the automatic segmentation result is unacceptable.
[0080] In some embodiments, the presentation component 124 can also present the segmentation mask and its reconstructed version via a suitable graphical display to facilitate visualization of any detected errors between the respective segmentation masks and review of the errors (e.g., by a suitable medical professional). For example, the presentation component 124 can present an automatic segmentation mask that is determined to be of insufficient quality and therefore an outlier and / or potentially associated with an error adjacent to a visual presentation of its corresponding reconstructed version. In this manner, the presentation component 124 can provide visual information that illustrates physical (e.g., based on size and / or geometry) differences between the respective mask and one or more regions associated with the automatic segmentation mask (determined to be associated with an error), such as Figure 6 exemplified.
[0081] In this regard, Figure 6Example graphical output data 600 that can be included in the quality assessment results data 132 according to one or more embodiments of the disclosed subject matter is illustrated. In this example, the presentation component 124 has presented automatic segmentation masks and reconstruction masks for two organs that were found to be outliers based on their respective Dice scores being lower than their respective organ-specific threshold Dice scores. The graphical output data 600 also includes textual data describing the quality assessment results, including information indicating their outlier classification, their Dice scores, their organ-specific threshold Dice scores, and a basis for describing the geometric errors associated with the automatic segmentation masks.
[0082] In this regard, reference Figures 1 to 6 In some embodiments, in association with comparing the automatic segmentation mask to its reconstructed version, the quality assessment component 122 may also determine information defining any differences between the size and / or geometry of the respective masks (e.g., as determined using one or more geometry-based and / or size-based image object comparison functions employed by the quality assessment component 122 for similarity assessment). For example, the quality assessment component 122 may identify locations of errors, such as areas associated with missing automatic segmentation masks (e.g., undersegmentation), areas corresponding to anatomical structures other than the target structure, areas associated with errors in curvature, geometry, and / or size, and the like. The quality assessment component 122 may also include textual and / or visual information regarding the physical differences determined between the respective segmentation masks (e.g., such as in Figure 6 ). For example, in some embodiments, one or more comparison functions may also generate visual marker data that defines or indicates the location and / or area of the automatic segmentation mask associated with the error (e.g., missing area, auxiliary area, etc.), which may be applied to or overlaid on the automatic segmentation mask in association with presenting the automatic segmentation mask. In this way, due to the ability to visualize the reconstruction and in association with textual and / or visual information about the inaccurate area and / or location of the automatic segmentation error, the quality assessment result data 138 and the error or outlier warnings are easier to interpret, thereby leading to improved human-computer interaction in the medical field.
[0083] In some embodiments, the adjustment component 126 can also adjust or control the use of the automatic segmentation results by one or more other applications 128 based on the quality assessment results. For example, in some implementations, the one or more other applications 128 can include a dose calculation application that automatically calculates dose amounts for IMRT based on the automatic segmentation masks. With these implementations, the adjustment component 126 can be configured to direct the dose calculation application to use the automatic segmentation masks to calculate doses based on a determination that their quality is acceptable, and to prevent the dose calculation application from using the automatic segmentation results for dose calculations based on a determination that their quality is unacceptable.
[0084] Figure 7 A flowchart of an example computer-implemented method 700 for evaluating the output quality of a multi-structure automatic segmentation model according to one or more embodiments of the disclosed subject matter is presented. At 702, the method 700 includes receiving, by a system including a processor (e.g., the computing system 100) (e.g., via the receiving component 110), a segmentation mask generated from medical image data depicting an anatomical region of a subject via one or more segmentation models, wherein each of the segmentation masks depicts a different anatomical structure from a set of different anatomical structures included in the anatomical region. At 704, the method 700 includes generating, by the system (e.g., via the model execution component 120 and the reconstruction model 118), a reconstructed version of the segmentation mask based on applying a multi-channel reconstruction model (e.g., the reconstruction model 118) to the segmentation mask, wherein the reconstructed version corresponds to an optimized version of the segmentation mask. At 706, the method 700 includes determining, by the system (e.g., via the quality assessment component 122), a quality assessment of the segmentation mask based on comparing the segmentation mask to the reconstructed version. At 708, method 700 includes generating, by the system (e.g., via quality assessment component 122), output data regarding the quality assessment. At 710, method 700 includes presenting, by the system (e.g., via presentation component 124), the output data via an electronic output device (e.g., an electronic display, a speaker, etc.).
[0085] Figure 8A high-level flow chart of an example computer-implemented method 800 for generating a multi-channel segmentation mask reconstruction model according to one or more embodiments of the disclosed subject matter is presented. The method 800 includes, at 802, training a multi-channel neural network model (e.g., reconstruction model 118) by a system including a processor (e.g., computing system 100) (e.g., via training component 116) to generate a reconstructed segmentation mask of a corresponding noise-enhanced segmentation mask as included in a corresponding training dataset using a ground-truth segmentation mask for the corresponding noise-enhanced segmentation mask, wherein the ground-truth segmentation mask respectively depicts different anatomical structures (e.g., different types of anatomical structures, different organs, different ROIs, etc.) as extracted from the training medical image data, and wherein the corresponding noise-enhanced segmentation mask includes a noise-enhanced version of the ground-truth segmentation mask. At 804, the method 800 includes, by the system (e.g., via training component 116) generating a trained version of the multi-channel neural network model as a result of the training.
[0086] Figure 9 A flowchart of an example computer-implemented method 900 for evaluating the output quality of a multi-structure automatic segmentation model according to one or more embodiments of the disclosed subject matter is presented. At 902, the method 900 includes: training, by a system including a processor (e.g., the computing system 100) (e.g., via the training component 116), a multi-channel neural network model (e.g., the reconstruction model 118) to generate a reconstructed segmentation mask of a corresponding noise-enhanced segmentation mask as included in a corresponding training dataset using a ground-truth segmentation mask for the corresponding noise-enhanced segmentation mask, wherein the ground-truth segmentation mask respectively depicts different anatomical structures (e.g., different types of anatomical structures, different organs, different ROIs, etc.) as extracted from medical image data, and wherein the corresponding noise-enhanced segmentation mask comprises a noise-enhanced version of the ground-truth segmentation mask. At 904, the method 900 includes: applying, by the system (e.g., via the model execution component 120), the trained version of the multi-channel neural network model to a set of segmentation masks generated from new medical image data via one or more segmentation models, wherein each of the segmentation masks depicts a different anatomical structure of at least some of the different anatomical structures. At 906, method 900 includes generating, by the system, a reconstructed version of the segmentation mask as a result of the applying. At 908, method 900 includes determining, by the system (eg, via quality assessment 122), a quality assessment of the segmentation mask based on comparing the segmentation mask to the reconstructed version.
[0087] Figure 10A flow chart is presented of an example method 1000 for evaluating the output quality of an automatic segmentation model of a single structure in accordance with one or more embodiments of the disclosed subject matter. To this end, the method 1000 corresponds to the method 900, except that the target segmentation model corresponds to a segmentation model configured to automatically segment a single anatomical structure and is a single-channel model of the single anatomical structure, rather than a multi-channel model version, relative to the reconstructed model 118.
[0088] At 1002, method 1000 includes: training, by a system including a processor (e.g., computing system 100) (e.g., via training component 116), a single-channel neural network model (e.g., reconstruction model 118) to generate a reconstructed segmentation mask for a corresponding noise-enhanced segmentation mask using a ground-truth segmentation mask for the corresponding noise-enhanced segmentation mask, wherein the ground-truth segmentation mask respectively depicts the same type of anatomical structure as extracted from medical image data, and wherein the corresponding noise-enhanced segmentation mask comprises a noise-enhanced version of the ground-truth segmentation mask. At 1004, method 1000 includes: applying, by the system (e.g., via model execution component 120), the trained version of the neural network model to a segmentation mask generated from new medical image data via a segmentation model, the segmentation model being configured to segment the same type of anatomical structure from the new medical image data. At 1006, method 1000 includes: generating, by the system, a reconstructed version of the segmentation mask as a result of the applying. At 1008 , method 1000 includes determining, by the system (eg, via quality assessment 122 ), a quality assessment of the segmentation mask based on comparing the segmentation mask to the reconstructed version.
[0089] One or more embodiments may be systems, methods, and / or computer program products at any possible level of technical detail of integration. The computer program product may include a computer-readable storage medium (or multiple media) having computer-readable program instructions thereon for causing a processor to perform aspects of the invention.
[0090] A computer-readable storage medium may be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: a portable computer floppy disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanical encoding device (such as a punch card or a raised structure in a groove on which instructions are recorded), and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be understood as a transient signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., a light pulse passing through a fiber optic cable), or an electrical signal transmitted through a wire.
[0091] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a corresponding computing / processing device, or downloaded to an external computer or external storage device via a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network). The network can include copper transmission cables, optical transmission fibers, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium within the corresponding computing / processing device.
[0092] The computer-readable program instructions for performing the operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, configuration data of an integrated circuit system, or source code or object code written in any combination of one or more programming languages (including object-oriented programming languages such as Smalltalk, C++, etc.), procedural programming languages (such as "C" programming language or similar programming languages), and machine learning programming languages (such as CUDA, Python, Tensorflow, PyTorch, etc.). The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server using suitable processing hardware. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In various embodiments involving machine learning programming instructions, the processing hardware may include one or more graphics processing units (GPUs) and central processing units (CPUs), etc. For example, one or more of the disclosed deep learning models (e.g., segmentation model, reconstruction model 118, and / or a combination thereof) can be written in a suitable machine learning programming language and executed via one or more GPUs, CPUs, or a combination thereof. In some embodiments, electronic circuits including, for example, programmable logic circuits, field programmable gate arrays (FPGAs), or programmable logic arrays (PLAs) can execute computer-readable program instructions by utilizing state information of the computer-readable program instructions to personalize the electronic circuits so as to perform various aspects of the present invention.
[0093] Various aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0094] These computer-readable program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing device create components for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, a programmable data processing device, and / or other device to function in a specific manner, such that the computer-readable storage medium having the instructions stored therein comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0095] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other devices to produce a computer-implemented method, such that the instructions executed on the computer, other programmable apparatus, or other devices implement the functions / actions specified in one or more boxes of the flowchart and / or block diagram.
[0096] The flow charts and block diagrams in the accompanying drawings illustrate the possible specific implementation architecture, function and operation of the system, method and computer program product according to various embodiments of the present invention. In this regard, each frame in the flow chart or block diagram can represent a module, fragment or part of an instruction, which includes one or more executable instructions for realizing a specified logical function. In some alternative specific implementations, the function indicated in the frame may not occur in the order indicated in the figure. For example, depending on the functionality involved, the two frames of the continuous display can actually be executed substantially simultaneously, or the frame can sometimes be executed in the opposite order. It will also be noted that the combination of each frame of the block diagram and / or flow chart illustration and the frame in the block diagram and / or flow chart illustration can be realized by a system based on dedicated hardware that performs a specified function or action or implements a combination of dedicated hardware and computer instructions.
[0097] Combine Figure 11 The systems and methods described below may be embodied in hardware (e.g., a single integrated circuit (IC) chip, multiple ICs, or an application-specific integrated circuit (ASIC)). Furthermore, the order in which some or all method blocks appear in each method should not be considered limiting. Rather, it should be understood that some method blocks can be executed in various orders, not all of which may be explicitly shown herein.
[0098] refer to Figure 11An illustrative example environment 1100 for implementing various aspects of the claimed subject matter includes a computer 1102. The computer 1102 includes a processing unit 1104, a system memory 1106, a codec 1135, and a system bus 1108. The system bus 1108 couples system components, including but not limited to the system memory 1106, to the processing unit 1104. The processing unit 1104 can be any of a variety of available processors. Dual microprocessors, one or more GPUs, and other multi-processor architectures can also be used as the processing unit 1104.
[0099] The system bus 1108 may be any of several types of bus structures including a memory bus or memory controller, a peripheral bus or external bus, or a native bus using any of a variety of available bus architectures including, but not limited to, Industry Standard Architecture (ISA), Micro Channel Architecture (MSA), Extended ISA (EISA), Intelligent Drive Electronics (IDE), VESA Native Bus (VLB), Peripheral Component Interconnect (PCI), card bus, Universal Serial Bus (USB), Advanced Graphics Port (AGP), Personal Computer Memory Card International Association (PCMCIA), FireWire (IEEE 1194), and Small Computer System Interface (SCSI).
[0100] In various embodiments, system memory 1106 includes volatile memory 1110 and non-volatile memory 1112, which may employ one or more of the disclosed memory architectures. A basic input / output system (BIOS), which contains basic routines for transferring information between components within computer 1102, such as during startup, is stored in non-volatile memory 1112. Furthermore, according to the present innovations, codec 1135 may include at least one of an encoder or a decoder, wherein at least one of the encoder or decoder may be comprised of hardware, software, or a combination of hardware and software. Although codec 1135 is depicted as a separate component, codec 1135 may be included within non-volatile memory 1112. By way of example and not limitation, non-volatile memory 1112 may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory, 3D flash memory, or resistive memory (such as resistive random access memory (RRAM)). In at least some embodiments, non-volatile memory 1112 may employ one or more of the disclosed memory devices. Furthermore, non-volatile memory 1112 may be computer memory (e.g., physically integrated with computer 1102 or its motherboard) or removable memory. Examples of suitable removable memory that may be used to implement the disclosed embodiments may include secure digital (SD) cards, compact flash (CF) cards, universal serial bus (USB) memory sticks, and the like. Volatile memory 1110 includes random access memory (RAM) that acts as external cache memory, and in various embodiments, one or more of the disclosed memory devices may also be employed. By way of example and not limitation, RAM may be provided in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), and enhanced SDRAM (ESDRAM).
[0101] The computer 1102 may also include removable / non-removable, volatile / non-volatile computer storage media. Figure 11Illustrated is, for example, a disk storage device 1114. The disk storage device 1114 includes, but is not limited to, devices such as disk drives, solid-state disks (SSDs), flash memory cards, or memory sticks. In addition, the disk storage device 1114 may include a separate storage medium or a storage medium combined with other storage media, including but not limited to an optical disc drive, such as a CD-ROM device (CD-ROM), a CD recordable drive (CD-R drive), a CD rewritable drive (CD-RW drive), or a digital versatile magneto-optical disc ROM drive (DVD-ROM). In order to facilitate connecting the disk storage device 1114 to the system bus 1108, a removable or non-removable interface, such as an interface 1116, is typically used. It should be understood that the disk storage device 1114 can store information related to the user. This type of information can be stored at a server or provided to an application running on a server or user device. In one embodiment, the user can be notified (for example, by an output device 1136) of the type of information stored in the disk storage device 1114 or transmitted to a server or application. The user may be provided with an opportunity to opt-in or opt-out of the collection or sharing of such information by the server or application (eg, via input from input device 1128 ).
[0102] It should be understood that Figure 11 Software is described that acts as an intermediary between a user and the basic computer resources described in the appropriate operating environment 1100. Such software includes an operating system 1118. Operating system 1118, which may be stored on disk storage 1114, is used to control and allocate resources of computer 1102. Applications 1120 utilize operating system 1118's management of resources through program modules 1124 and program data 1126, such as startup / shutdown transaction tables, stored in system memory 1106 or on disk storage 1114. It should be understood that the claimed subject matter can be implemented with various operating systems or combinations of operating systems.
[0103] The user enters commands or information into the computer 1102 through input devices 1128. Input devices 1128 include, but are not limited to, pointing devices such as a mouse, trackball, stylus, touchpad, keyboard, microphone, joystick, gamepad, satellite dish, scanner, TV tuner card, digital camera, digital video camera, and webcam. These and other input devices are connected to the processing unit 1104 via interface ports 1130 through the system bus 1108. Interface ports 1130 include, for example, serial ports, parallel ports, game ports, and universal serial buses (USB). Output devices 1136 use some of the same types of ports as input devices 1128. Thus, for example, a USB port can be used to provide input to the computer 1102 and output information from the computer 1102 to output devices 1136. Output adapters 1134 are provided to illustrate the presence of some output devices 1136 such as monitors, speakers, and printers, as well as other output devices 1136 that require special adapters. By way of illustration and not limitation, output adapters 1134 include video and sound cards that provide a means of connection between output devices 1136 and system bus 1108. It should be noted that other devices or systems of devices provide input and output capabilities, such as remote computer 1138.
[0104] Computer 1102 can operate in a networked environment using logical connections to one or more remote computers, such as remote computer 1138. Remote computer 1138 can be a personal computer, server, router, network PC, workstation, microprocessor-based device, peer device, smartphone, tablet, or other network node, and typically includes many of the elements described with respect to computer 1102. For the sake of simplicity, only memory storage device 1140 is illustrated for remote computer 1138. Remote computer 1138 is logically connected to computer 1102 via network interface 1142, and then connected via communication connection 1144. Network interface 1142 encompasses wired or wireless communication networks, such as local area networks (LANs), wide area networks (WANs), and cellular networks. LAN technologies include fiber distributed data interface (FDDI), copper distributed data interface (CDDI), Ethernet, and token ring, among others. WAN technologies include, but are not limited to, point-to-point links, circuit-switched networks such as Integrated Services Digital Network (ISDN) and its variants, packet-switched networks, and digital subscriber lines (DSL).
[0105] The communication connection 1144 refers to the hardware / software used to connect the network interface 1142 to the bus 1108. Although the communication connection 1144 is shown internal to the computer 1102 for clarity of illustration, the communication connection can also be external to the computer 1102. For exemplary purposes only, the hardware / software required to connect to the network interface 1142 includes internal and external technology, such as modems, including conventional telephone-grade modems, cable modems, and DSL modems, ISDN adapters, as well as wired and wireless Ethernet cards, hubs, and routers.
[0106] Although the subject matter has been described above in the general context of computer-executable instructions of a computer program product running on one and / or multiple computers, it will be appreciated by those skilled in the art that the present disclosure may also be implemented or may be combined with other program modules. Generally speaking, program modules include routines, programs, components, data structures, etc. that perform specific tasks and / or implement specific abstract data types. In addition, it will be appreciated by those skilled in the art that other computer system configurations may be used to practice the computer-implemented methods of the present invention, including single-processor or multi-processor computer systems, small computing devices, mainframe computers, and computers, handheld computing devices (e.g., PDAs, phones), microprocessor-based or programmable consumer or industrial electronic devices, etc. The illustrated aspects may also be practiced in a distributed computing environment, in which tasks are performed by remote processing devices linked by a communication network. However, some (if not all) aspects of the present disclosure may be practiced on stand-alone computers. In a distributed computing environment, program modules may be located in local and remote memory storage devices.
[0107] As used in this application, the terms "component," "system," "platform," and "interface" may refer to and / or include computer-related entities or entities associated with an operating machine having one or more specific functionalities. The entities disclosed herein may be hardware, a combination of hardware and software, software, or software in execution. For example, a component may be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program, and / or a computer. By way of example, both an application running on a server and the server may be components. One or more components may reside within a process and / or execution thread, and a component may be located on a single computer and / or distributed between two or more computers. As another example, the corresponding component may be executed based on various computer-readable media having various data structures stored thereon. Components may communicate via local and / or remote processes (such as based on signals having one or more data packets (e.g., data from a component that interacts with another component in a local system, a distributed system, and / or a network (such as the Internet with other systems))). As another example, a component may be a device having specific functionality provided by mechanical parts operated by electrical or electronic circuitry, which is operated by a software or firmware application executed by a processor. In such cases, the processor may be internal or external to the device and may execute at least a portion of the software or firmware application. As yet another example, a component may be a device that provides specific functionality through electronic components rather than mechanical parts, where the electronic components may include a processor or other means for executing the software or firmware that at least partially imparts the functionality to the electronic components. In one aspect, the component may emulate the electronic component, for example, via a virtual machine within a cloud computing system.
[0108] In addition, the term "or" is intended to mean an inclusive "or" rather than an exclusive "or". That is, unless otherwise specified or clear from the context, "X employs A or B" is intended to mean any natural inclusive permutation. That is, if X employs A; X employs B; or X employs both A and B, then "X employs A or B" is satisfied in any of the foregoing cases. In addition, unless otherwise specified or clear from the context to be directed to a singular form, the articles "a" and "an" used in this specification and the drawings should generally be understood to mean "one or more". As used herein, the terms "example" and / or "exemplary" are used to mean serving as an example, instance, or illustration, and are intended to be non-limiting. For the avoidance of doubt, the subject matter disclosed herein is not limited by such examples. In addition, any aspect or design described herein as "example" and / or "exemplary" is not necessarily to be construed as preferred or advantageous over other aspects or designs, nor is it intended to exclude equivalent exemplary structures and techniques known to those of ordinary skill in the art.
[0109] As used in this specification, the term "processor" may refer to substantially any computational processing unit or device, including but not limited to a single-core processor; a single processor with software multi-threaded execution capability; a multi-core processor; a multi-core processor with software multi-threaded execution capability; a multi-core processor with hardware multi-threading technology; a parallel platform; and a parallel platform with distributed shared memory. Additionally, a processor may refer to an integrated circuit, an application-specific integrated circuit (ASIC), a digital signal processor (DSP), a field-programmable gate array (FPGA), a programmable logic controller (PLC), a complex programmable logic device (CPLD), discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. Furthermore, the processor may utilize nanoscale architectures (such as, but not limited to, molecular and quantum dot-based transistors, switches, and gates) to optimize space usage or enhance the performance of user equipment. The processor may also be implemented as a combination of computational processing units. In this disclosure, terms such as "repository," "storage device," "data repository," "data storage device," "database," and substantially any other information storage component related to the operation and functionality of the component are used to refer to a "memory component," an entity embodied in "memory," or a component that includes memory. It should be understood that the memory and / or memory components described herein may be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. By way of example and not limitation, non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable ROM (EEPROM), flash memory, or non-volatile random access memory (RAM) (e.g., ferroelectric RAM (FeRAM)). For example, volatile memory may include RAM that can act as external cache memory. By way of example and not limitation, RAM can be provided in a variety of forms, such as synchronous RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), Synchlink DRAM (SLDRAM), direct Rambus RAM (DRRAM), direct Rambus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM). Additionally, the disclosed memory components of the systems or computer-implemented methods herein are intended to include, but are not limited to, these and any other suitable types of memory.
[0110] What has been described above only includes examples of systems and computer-implemented methods. Of course, it is not possible to describe every conceivable combination of components or computer-implemented methods for the purposes of describing the present disclosure, but one of ordinary skill in the art will recognize that many other combinations and permutations of the present disclosure are possible. In addition, to the extent that the terms "including," "having," "having," and the like are used in the detailed description, claims, appendices, and drawings, such terms are intended to be inclusive in a manner similar to the term "including," as interpreted when "including" is used as a transitional word in the claims. Descriptions of various embodiments have been given for illustrative purposes, but these descriptions are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to one of ordinary skill in the art without departing from the scope and essence of the described embodiments. The terms used herein are selected to best illustrate the principles of the embodiments, to be superior to practical applications or technical improvements of technologies found on the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A method comprising: receiving, by a system including a processor, segmentation masks generated from medical image data depicting an anatomical region of a subject via one or more segmentation models, wherein each of the segmentation masks depicts a different anatomical structure from a set of different anatomical structures included in the anatomical region; generating, by the system, a reconstructed version of the segmentation mask based on applying a multi-channel reconstruction model to the segmentation mask, wherein the reconstructed version corresponds to an optimized version of the segmentation mask; determining, by the system, a quality assessment of the segmentation mask based on comparing the segmentation mask to the reconstructed version; generating, by the system, output data regarding the quality assessment; as well as The output data is presented by the system via an electronic output device.
2. The method of claim 1 , wherein for each segmentation mask, determining the evaluation comprises: determining, by the system, a similarity measure between the segmentation mask and a reconstructed version of the segmentation mask, wherein the similarity measure represents a quality measure of the segmentation mask as generated via the one or more segmentation models; as well as A determination is made by the system whether the segmentation mask is associated with an error based on whether the similarity metric satisfies a threshold similarity metric.
3. The method of claim 2, wherein the threshold similarity metric varies for the different anatomical structures.
4. The method of claim 2, wherein determining the assessment further comprises: Determining, by the system, whether the segmentation masks collectively meet an acceptable quality criterion based on a total similarity metric determined for the segmentation masks, and wherein the output data indicates whether the segmentation masks collectively meet the acceptable quality criterion.
5. The method according to claim 4, further comprising: Use of the segmentation masks by a clinical application is regulated by the system based on whether the segmentation masks collectively meet the acceptable quality criteria.
6. The method according to claim 2, further comprising: For each segmentation mask, the segmentation mask is associated with the error based on: Warning data is generated by the system indicating that the segmentation mask is associated with the error, wherein the output data includes the warning data, wherein the electronic output device includes a display, and wherein the presenting includes: presenting the warning data via the display in association with presenting the segmentation mask and optionally presenting the reconstructed version of the segmentation mask.
7. The method of claim 2, wherein determining the assessment further comprises: Based on determining that the segmentation mask is associated with the error: Error information regarding differences between a size and / or geometry of the segmentation mask and the reconstructed version is determined by the system based on comparing the segmentation mask with the reconstructed version, and wherein the output data includes the error information.
8. The method of claim 1 , wherein the multi-channel reconstruction model comprises a neural network model, and wherein the method further comprises: The multi-channel reconstruction model is trained by the system, wherein the training comprises: training the multi-channel reconstruction model to generate reconstruction masks of the corresponding noise-enhanced segmentation masks as included in the corresponding training dataset using reference truth segmentation masks for the corresponding noise-enhanced segmentation masks, wherein the reference truth segmentation masks respectively depict the different anatomical structures as extracted from the training medical image data, and wherein the corresponding noise-enhanced segmentation masks comprise noise-enhanced versions of the reference truth segmentation masks.
9. The method according to claim 8, further comprising: The noise enhanced segmentation mask is generated by the system from the ground truth segmentation mask.
10. The method of claim 9, wherein for each ground truth segmentation mask, generating the noise enhanced segmentation mask comprises: An amount of noise data is integrated by the system into the ground truth segmentation mask that is customized based on the size and geometry of the anatomical structure depicted in the ground truth segmentation mask.
11. The method of claim 10, wherein the integrating comprises: The amount of noise data is added to or removed from the ground truth segmentation mask.
12. A system comprising: a memory storing computer-executable components; and a processor that executes the computer-executable components stored in the memory, wherein the computer-executable components include: a receiving component that receives segmentation masks generated from medical image data depicting an anatomical region of a subject via one or more segmentation models, wherein each of the segmentation masks depicts a different anatomical structure from a set of different anatomical structures included in the anatomical region; a model execution component that generates a reconstructed version of the segmentation mask based on applying a multi-channel reconstruction model to the segmentation mask, wherein the reconstructed version corresponds to an optimized version of the segmentation mask; a quality assessment component that determines a quality assessment of the segmentation mask based on comparing the segmentation mask with the reconstructed version and generates output data regarding the quality assessment; and A presentation component presents the output data via an electronic output device.
13. The system of claim 12, wherein for each segmentation mask, the quality assessment component: determining a similarity measure between the segmentation mask and a reconstructed version of the segmentation mask, wherein the similarity measure represents a quality measure of the segmentation mask as generated via the one or more segmentation models; determining whether the segmentation mask is associated with an error based on whether the similarity metric satisfies a threshold similarity metric, and Warning data indicating that the segmentation mask is associated with the error is generated based on determining that the segmentation mask is associated with the error, wherein the output data includes the warning data.
14. The system of claim 13, wherein the threshold similarity metric varies for the different anatomical structures.
15. The system of claim 13, wherein the quality assessment component determines whether the segmentation masks collectively meet an acceptable quality standard based on a total similarity metric determined for the segmentation masks, and wherein the output data indicates whether the segmentation masks collectively meet the acceptable quality standard.
16. The system of claim 15, wherein the computer executable components further comprise: A regulation component regulates usage of the segmentation masks by a clinical application based on whether the segmentation masks collectively meet the acceptable quality criteria.
17. The system of claim 13 , wherein based on determining that the segmentation mask is associated with the error, the quality assessment component determines error information regarding differences between a size and / or geometry of the segmentation mask and the reconstructed version based on comparing the segmentation mask with the reconstructed version, and wherein the presentation component presents the warning data and the error information via an electronic display in association with presenting the segmentation mask and the reconstructed version of the segmentation mask.
18. The system of claim 12, wherein the multi-channel reconstruction model comprises a neural network model, and wherein the computer-executable components further comprise: a training component that trains the multi-channel reconstruction model using an unsupervised machine learning method, wherein the unsupervised machine learning method comprises: training the multi-channel reconstruction model to generate reconstruction masks of the corresponding noise-enhanced segmentation masks as included in the corresponding training dataset using reference truth segmentation masks for the corresponding noise-enhanced segmentation masks, wherein the reference truth segmentation masks respectively depict the different anatomical structures as extracted from the training medical image data, and wherein the corresponding noise-enhanced segmentation masks comprise noise-enhanced versions of the reference truth segmentation masks.
19. The system of claim 12, wherein the computer-executable components further comprise: a noise enhancement component that generates the noise-enhanced segmentation masks from the ground truth segmentation masks, wherein for each ground truth segmentation mask, the noise enhancement component integrates an amount of noise data into the ground truth segmentation mask that is customized based on the size and geometry of the anatomical structure depicted in the ground truth segmentation mask.
20. A non-transitory machine-readable storage medium comprising executable instructions that, when executed by a processor, facilitate performance of operations comprising: receiving segmentation masks generated from medical image data depicting an anatomical region of a subject via one or more segmentation models, wherein each of the segmentation masks depicts a different anatomical structure from a set of different anatomical structures included in the anatomical region; generating a reconstructed version of the segmentation mask based on applying a multi-channel reconstruction model to the segmentation mask, wherein the reconstructed version corresponds to an optimized version of the segmentation mask; determining a quality assessment of the segmentation mask based on comparing the segmentation mask to the reconstructed version; generating output data regarding the quality assessment; as well as The output data is presented via an electronic output device.
Citation Information
Cited By
On-line learning-based sketching model construction method and related equipment
CN121505263A