System and method for carrying out data adaptive single-lens multi-label segmentation by utilizing basic model

The regions of interest in medical images are automatically determined through visual transformers and contrastive similarity metric learning models, which solves the problem of laborious and error-prone positioning in existing technologies and achieves efficient and accurate region of interest segmentation.

CN120707844APending Publication Date: 2025-09-26GE PRECISION HEALTHCARE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510285264.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-03-26
Filing Date
2025-03-11
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing foundational models have a low degree of automation for radiological image localization, making localization laborious and error-prone, increasing clinician fatigue and costs.

Method used

A trained visual transformer model and a contrastive similarity metric learning model are used to automatically determine the similarity between the pixel-level feature vectors in the medical image and the reference pixel-level feature vectors of the template image. The initial segmentation mask is used to mark the region of interest, and a suggestible segmentation model is combined to perform refined segmentation.

Benefits of technology

It realizes the automatic positioning and fine segmentation of the region of interest in medical images, reduces manual intervention, improves the accuracy and efficiency of positioning, and reduces costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120707844A_ABST
    Figure CN120707844A_ABST
Patent Text Reader

Abstract

The invention relates to a system and method for data adaptive single-lens multi-tag segmentation using a base model. A method includes obtaining a medical image and receiving a selection of both a template image and a region of interest within the template image. The method includes inputting both the medical image and the template image into a trained visual transducer model, and outputting both a pixel-level feature vector from the medical image and a reference pixel-level feature vector from the region of interest of the template image from the trained visual transducer model. The method includes inputting both the pixel-level feature vector and the reference pixel-level feature vector into a trained comparative similarity metric learning model, and outputting a pixel similar to a reference pixel from the trained comparative similarity metric learning model. The method includes marking a pixel in the medical image with a segmentation mask, wherein the pixel marked in the medical image corresponds to the region of interest.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] The subject matter disclosed herein relates to medical imaging, and more particularly to systems and methods for data-adaptive single-shot multi-label segmentation using a base model.

[0002] Non-invasive imaging techniques allow images of internal structures or features of a patient / subject to be obtained without performing an invasive procedure on the patient / subject. Specifically, such non-invasive imaging techniques rely on various physical principles (such as differential transmission of X-rays through a target volume, reflection of acoustic waves within a volume, paramagnetism of different tissues and materials within a volume, decomposition of target radionuclides within the body, etc.) to acquire data and construct an image or otherwise represent the observed internal features of the patient / subject.

[0003] During MRI, when a substance, such as human tissue, is subjected to a uniform magnetic field (polarization field B0), the individual magnetic moments of the spins in the tissue attempt to align with the polarization field, but precess around it in a random order at their characteristic Larmor frequency. If the substance or tissue is subjected to a magnetic field (excitation field B1) that lies in the xy plane and is close to the Larmor frequency, the net alignment moment or "longitudinal magnetization" M z can be rotated or "tilted" into the xy plane to produce a net transverse magnetic moment M t After excitation signal B1 terminates, a signal is emitted by the excited spins, and this signal can be received and processed to form an image.

[0004] When these signals are used to generate images, magnetic field gradients (G x , G y and G z Typically, the area to be imaged is scanned in a series of measurement cycles in which these gradient fields are varied according to the particular positioning method being used. The resulting set of received nuclear magnetic resonance (NMR) signals is digitized and processed to reconstruct an image using one of the well-known reconstruction techniques.

[0005] Localization and region of interest segmentation requirements are ubiquitous in different stages of the radiology workflow: planning, guidance, and lesion identification and measurement. However, localization is a laborious and repetitive task. Furthermore, localization increases clinician fatigue, which can lead to inaccuracies. Furthermore, localization increases costs. Grounded models are attractive for automating localization requirements because they demonstrate excellent semantic association capabilities in natural images. However, previous attempts to use off-the-shelf semantic association grounded models for radiology image localization have not been successful. Summary of the Invention

[0006] The following is an overview of certain embodiments disclosed herein. It should be understood that these aspects are provided merely to provide the reader with a brief overview of these specific embodiments, and these aspects are not intended to limit the scope of the present disclosure. In fact, the present disclosure may encompass various aspects that may not be shown below.

[0007] In one embodiment, a computer-implemented method is provided. The computer-implemented method includes obtaining a medical image of a portion of a subject at a processor. The computer-implemented method also includes receiving, at the processor, a selection of both a template image and a region of interest within the template image, wherein the region of interest is labeled in the template image and associated with a label. The computer-implemented method also includes inputting both the medical image and the template image into a trained visual transformer model via the processor. The computer-implemented method even includes outputting both a pixel-level feature vector from the medical image and a reference pixel-level feature vector of the region of interest from the template image from the trained visual transformer model via the processor. The computer-implemented method still includes inputting both the pixel-level feature vector and the reference pixel-level feature vector into a trained contrastive similarity metric learning model via the processor, wherein the trained contrastive similarity metric learning model is configured to automatically determine which pixel-level feature vectors of the pixel-level feature vectors are similar to the reference pixel-level feature vector. The computer-implemented method still includes outputting pixels similar to the reference pixel from the trained contrastive similarity metric learning model via the processor. The computer-implemented method also includes using the initial segmentation mask, via the processor, to mark pixels in the medical image associated with the pixel-level feature vector that is similar to the reference pixel-level feature vector, wherein the marked pixels in the medical image correspond to the region of interest.

[0008] In another embodiment, a system for performing one-time anatomical structure localization is provided. The system includes a memory that encodes a processor-executable routine. The system also includes a processor configured to access the memory and execute the processor-executable routine, wherein the routine, when executed by the processor, causes the processor to perform an action. The action includes obtaining a medical image of a portion of a subject. The action also includes receiving a selection of both a template image and a region of interest within the template image, wherein the region of interest is labeled in the template image and associated with a label. The action also includes inputting both the medical image and the template image into a trained visual transformer model. The action even includes outputting from the trained visual transformer model both a pixel-level feature vector from the medical image and a reference pixel-level feature vector for the region of interest from the template image. The action still further includes inputting both the pixel-level feature vector and the reference pixel-level feature vector into a trained contrastive similarity metric learning model, wherein the trained contrastive similarity metric learning model is configured to automatically determine which of the pixel-level feature vectors are similar to the reference pixel-level feature vector. The actions still further include outputting pixels similar to the reference pixels from the trained contrastive similarity metric learning model. The actions also include labeling pixels in the medical image associated with the pixel-level feature vector that is similar to the reference pixel-level feature vector using an initial segmentation mask, wherein the labeled pixels in the medical image correspond to the region of interest.

[0009] In another embodiment, a non-transitory computer-readable medium is provided. The non-transitory computer-readable medium includes processor executable code that, when executed by a processor, causes the processor to perform an action. The action includes obtaining a medical image of a portion of a subject. The action also includes receiving a selection of a template image and a plurality of regions of interest within the template image, wherein each of the plurality of regions of interest is labeled in the template image and associated with a corresponding label. The action also includes inputting both the medical image and the template image into a trained visual transformer model. The action even includes outputting from the trained visual transformer model both a corresponding pixel-level feature vector from the medical image and a corresponding reference pixel-level feature vector for each of the plurality of regions of interest in the template image. The action still further includes inputting both the pixel-level feature vector and the corresponding reference pixel-level feature vector into a trained contrastive similarity metric learning model, wherein the trained contrastive similarity metric learning model is configured to automatically determine which of the pixel-level feature vectors is similar to each corresponding reference pixel-level feature vector for each region of interest. The action still further includes outputting, from the trained contrastive similarity metric learning model, a corresponding pixel group similar to each corresponding reference pixel in each region of interest. The action also includes respectively labeling, using a corresponding initial segmentation mask, a corresponding pixel group in the medical image associated with each group of corresponding pixel-level feature vectors similar to each corresponding reference pixel-level feature vector in each region of interest, wherein the respectively labeled corresponding pixel groups in the medical image correspond to corresponding regions of interest in the plurality of regions of interest. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] These and other features, aspects, and advantages of the present subject matter will be better understood when the following detailed description is read with reference to the accompanying drawings, in which like characters represent like parts throughout, and in which:

[0011] Figure 1 An embodiment of a magnetic resonance imaging (MRI) system suitable for use with the disclosed technology is illustrated;

[0012] Figure 2 A schematic diagram illustrating the training of a contrastive similarity metric learning model for positioning according to various aspects of the present disclosure;

[0013] Figure 3 A schematic diagram illustrating data-adaptive single-shot segmentation using a basic model according to various aspects of the present disclosure is illustrated;

[0014] Figure 4A flowchart illustrating a method for data-adaptive single-shot segmentation using a base model according to various aspects of the present disclosure is provided;

[0015] Figure 5 A flowchart illustrating a method for data-adaptive single-shot segmentation using a base model (e.g., on multiple medical images) according to aspects of the present disclosure is provided;

[0016] Figure 6 A flowchart illustrating a method for data-adaptive single-shot segmentation using a base model (e.g., on relevant medical images or slices) according to aspects of the present disclosure is provided;

[0017] Figure 7 A flowchart illustrating a method for data-adaptive single-shot multi-label segmentation using a base model (e.g., on relevant medical images or slices) according to various aspects of the present disclosure is provided;

[0018] Figure 8 depicts MR images of a shoulder using different methods to compare region of interest positioning according to aspects of the present disclosure;

[0019] Figure 9 depicts comparing segmented shoulder MR images by using cues from local region of interest locations derived using different methods in accordance with aspects of the present disclosure;

[0020] Figure 10 depicts a table comparing region of interest localization and shoulder segmentation using different methods according to aspects of the present disclosure; and

[0021] Figure 11 Utilization of multi-tag single-shot localization of the entire knee joint volume in accordance with aspects of the present disclosure is illustrated. DETAILED DESCRIPTION

[0022] One or more specific embodiments are described below. In order to provide a concise description of these embodiments, not all features of an actual implementation are necessarily described in the specification. It should be understood that in the development of any such actual implementation, as in any engineering or design project, many implementation-specific decisions must be made to achieve the developer's specific goals, such as complying with system-related and business-related constraints that may vary from implementation to implementation. Furthermore, it should be understood that such development efforts may be complex and time-consuming, but remain a routine task of design, fabrication, and manufacturing for those of ordinary skill having the benefit of this disclosure.

[0023] When introducing elements of various embodiments of the present subject matter, the articles "a," "an," "the," and "said" are intended to indicate that there are one or more elements. The terms "comprising," "including," and "having" are intended to be inclusive and mean that there may be additional elements besides the listed elements. Furthermore, any numerical examples in the following discussion are intended to be non-limiting, and thus the appended numerical values, ranges, and percentages are within the scope of the disclosed embodiments.

[0024] While various aspects of the following discussion are presented in the context of medical imaging, it should be understood that the disclosed technology is not limited to such a medical context. Indeed, examples and explanations are provided in such a medical context solely to facilitate explanation by providing examples of real-world implementations and applications. However, the disclosed technology may also be used in other contexts, such as image reconstruction for non-destructive inspection of manufactured parts or goods (i.e., quality control or quality review applications) and / or non-invasive inspection of packages, boxes, luggage, etc. (i.e., security or screening applications). Generally speaking, the disclosed technology may be used in any imaging or screening context or image processing or photography field in which a set or class of acquired data undergoes a reconstruction process to generate an image or volume.

[0025] The deep learning (DL) methods discussed herein may be based on artificial neural networks and may therefore encompass one or more of the following: deep neural networks, fully interconnected networks, convolutional neural networks (CNNs), unfolded neural networks, perceptrons, codecs, recurrent networks, wavelet filter banks, u-nets, generative adversarial networks (GANs), dense neural networks, or other neural network architectures. Neural networks may include shortcuts, activations, batch normalization layers, and / or other features. These techniques are referred to herein as DL techniques, although the term may also be used with particular reference to the use of deep neural networks, which are neural networks with multiple layers.

[0026] One type of deep learning model is a visual transformer model. A visual transformer model utilizes transformers (e.g., visual transformers) to perform image recognition tasks. Specifically, a visual transformer model decomposes an input image (e.g., a medical image) into blocks, processes these blocks using transformers, and aggregates information for classification or object detection. A visual transformer model utilizes self-attention (i.e., a global operation) because it extracts information from the entire image. This enables the visual transformer model to effectively capture clear semantic dependencies in the image. The visual transformer model achieves similar or better results than other types of deep learning models (e.g., convolutional networks) while requiring far fewer computational resources to train.

[0027] As discussed herein, DL techniques (which may also be referred to as deep machine learning, hierarchical learning, or deep structured learning) are a branch of machine learning techniques that employ mathematical representations of data and artificial neural networks for learning and processing such representations. For example, DL methods can be characterized as using one or more algorithms to extract or model highly abstract concepts of a class of data of interest. This can be accomplished using one or more processing layers, where each layer typically corresponds to a different level of abstraction and may therefore employ or utilize different aspects of the initial data or the output of the previous layer (i.e., a hierarchical or cascaded structure of layers) as the target of the process or algorithm at a given layer. In the context of image processing or reconstruction, this can be characterized as different layers corresponding to different feature levels or resolutions in the data. Generally speaking, the processing of one representation space to the next level representation space can be considered a "stage" of the process. Each stage of the process can be performed by a separate neural network or by different parts of a larger neural network.

[0028] The present disclosure provides systems and methods for data-adaptive single-shot multi-label segmentation using a base model. Specifically, a contrastive learning-based technique is utilized, which allows the task data itself to be used to drive feature similarity without requiring any manual tuning. Furthermore, multiple tasks on the same data can be completed in a single instance, enabling the use of the base model for multi-label single-shot localization and region of interest segmentation with medical imaging data (e.g., three-dimensional (3D) imaging data). Using a visual transformer (e.g., an unsupervised visual transformer) as a backbone, a self-supervised model is trained on a pool of unlabeled data, with the goal of deriving robust feature representations of images that are context-dependent features. The visual transformer architecture enables the derivation of block-level features that can be extended to pixel-level features (via simple post-processing). Furthermore, a contrastive similarity metric learning model is trained on the pixel-level features derived from the visual transformer, aiming to keep similar features as close as possible and dissimilar features as far apart as possible. This is accomplished by creating sample data for the task, augmenting the sample data by simulating expected variations in real-life scenarios for the task, creating positive and negative feature vector pairs for each of a plurality of tasks to account for variability within the feature vectors, and generating a model. The model is applied to any new test data by automatically finding similarities between feature vectors used for localization without utilizing heuristic manual thresholding (e.g., previously utilized with localization attempts utilizing a base model). Specifically, the contrastive similarity metric learning model utilizes a data-driven approach to perform thresholding. The localization output is linked to a suggestable base segmented arbitrary model (SAM) segmentation model with hints automatically selected within the local region to obtain finer segmented regions. In addition, the disclosed systems and methods can use image-level features to automatically select the medical image that is closest to the template image (i.e., the most relevant medical image) to reduce processing time and remove potential false positives that may otherwise be generated in the image.

[0029] The disclosed systems and methods include obtaining a medical image of a portion of a subject. The disclosed systems and methods also include receiving a selection of a template image and a region of interest within the template image, wherein the region of interest is labeled in the template image and associated with a label. The disclosed systems and methods also include inputting both the medical image and the template image into a trained visual transformer model. The disclosed systems and methods even include outputting both a pixel-level feature vector from the medical image and a reference pixel-level feature vector of the region of interest from the template image from the trained visual transformer model. The disclosed systems and methods still further include inputting both the pixel-level feature vector and the reference pixel-level feature vector into a trained contrastive similarity metric learning model, wherein the trained contrastive similarity metric learning model is configured to automatically determine which pixel-level feature vectors of the pixel-level feature vectors are similar to the reference pixel-level feature vector. The disclosed systems and methods still further include outputting pixels similar to the reference pixels from the trained contrastive similarity metric learning model. The disclosed systems and methods also include using the initial segmentation mask to mark pixels in the medical image associated with the pixel-level feature vector that is similar to the reference pixel-level feature vector, wherein the marked pixels in the medical image correspond to the region of interest.

[0030] In certain embodiments, the disclosed systems and methods include labeling the medical image with a refined segmentation mask corresponding to the region of interest using a suggestible segmentation model, wherein the initial segmentation mask serves as an automatic hint for labeling. In certain embodiments, labeling pixels in the medical image associated with the pixel-level feature vector that is similar to the reference pixel-level feature vector includes generating the initial segmentation mask using a connected component analysis of the pixels.

[0031] In certain embodiments, the disclosed systems and methods include obtaining a medical imaging volume of the portion of the subject, wherein the medical imaging volume includes a plurality of medical images, the plurality of medical images including the medical image. In certain embodiments, the disclosed systems and methods include inputting each of the plurality of medical images into the trained visual transformer model. In certain embodiments, the disclosed embodiments also include outputting a corresponding pixel-level feature vector from each of the plurality of medical images from the trained visual transformer model. In certain embodiments, the disclosed systems and methods even include inputting the corresponding pixel-level feature vector from each of the plurality of medical images into the trained contrastive similarity metric learning model. In certain embodiments, the disclosed systems and methods also include outputting a corresponding pixel from each of the plurality of medical images that is similar to a reference pixel from the trained contrastive similarity metric learning model. In certain embodiments, the disclosed systems and methods further include labeling, using a corresponding initial segmentation mask, corresponding pixels in each of the plurality of medical images associated with a corresponding pixel-level feature vector from each of the plurality of medical images that is similar to a reference pixel-level feature vector, wherein the labeled corresponding pixels in each of the plurality of medical images correspond to the region of interest. In certain embodiments, the disclosed systems and methods include labeling, using the suggestable segmentation model, each of the plurality of medical images with a corresponding refined segmentation mask corresponding to the corresponding region of interest, wherein the corresponding initial segmentation mask serves as an automatic hint for labeling.

[0032] In certain embodiments, the disclosed systems and methods include obtaining a medical imaging volume of the portion of the subject, wherein the medical imaging volume includes a plurality of medical images, the plurality of medical images including the medical image. In certain embodiments, the disclosed systems and methods also include inputting each of the plurality of medical images into the trained visual transformer model. In certain embodiments, the disclosed systems and methods also include outputting corresponding pixel-level feature vectors and corresponding image-level features from each of the plurality of medical images from the trained visual transformer model. Determining a group of most relevant medical images from the plurality of medical images. In certain embodiments, the disclosed systems and methods even include inputting the corresponding pixel-level feature vectors from the group of most relevant medical images into the trained contrastive similarity metric learning model. In certain embodiments, the disclosed systems and methods still further include outputting corresponding pixels from the group of most relevant medical images that are similar to reference pixels from the trained contrastive similarity metric learning model. In certain embodiments, the disclosed systems and methods include labeling corresponding pixels in each of the set of most relevant medical images associated with a corresponding pixel-level feature vector similar to a reference pixel-level feature vector from each of the set of most relevant medical images using a corresponding initial segmentation mask, wherein the corresponding pixels labeled in each of the medical images in the set of most relevant medical images correspond to the region of interest. In certain embodiments, the disclosed systems and methods include labeling each of the set of most relevant medical images using a corresponding refined segmentation mask corresponding to the corresponding region of interest using the suggestible segmentation model, wherein the corresponding initial segmentation mask serves as an automatic hint for labeling. In certain embodiments, the disclosed systems and methods include determining the set of most relevant medical images from the plurality of medical images based on the image-level features.

[0033] In certain embodiments, the disclosed systems and methods include receiving the selection of a plurality of regions of interest within the template image, wherein each of the plurality of regions of interest is respectively labeled in the template image and associated with a corresponding label. In certain embodiments, the disclosed systems and methods also include outputting, from the trained visual transformer model, a corresponding reference pixel-level feature vector for each of the plurality of regions of interest from the template image. In certain embodiments, the disclosed systems and methods also include inputting each corresponding reference pixel-level feature vector into a trained contrastive similarity metric learning model, wherein the trained contrastive similarity metric learning model is configured to automatically determine which of the pixel-level feature vectors are similar to each corresponding reference pixel-level feature vector of each region of interest. In certain embodiments, the disclosed systems and methods even include outputting, from the trained contrastive similarity metric learning model via the processor, a corresponding group of pixels that are similar to each corresponding reference pixel in the corresponding reference pixels of each region of interest. In certain embodiments, the disclosed systems and methods include labeling respective groups of pixels in a medical image associated with each of a respective group of pixel-level feature vectors similar to each respective reference pixel-level feature vector for each region of interest using respective initial segmentation masks, wherein the respective labeled respective groups of pixels in the medical image correspond to respective regions of interest among the plurality of regions of interest. In certain embodiments, the disclosed systems and methods include labeling the medical image using respective refined segmentation masks for respective regions, the respective regions corresponding to respective regions of interest among the plurality of regions of interest using a suggestible segmentation model, wherein the respective initial segmentation masks serve as automatic hints for labeling.

[0034] The disclosed technology can be used for localization. Furthermore, the disclosed technology can be used for longitudinal lesion tracking across multiple time points. The disclosed technology can be used with different types of medical images. For example, these images can be obtained from MRI, computed tomography (CT) imaging, or other types of imaging systems. In this disclosure, the technology is described in the context of MRI.

[0035] Considering the above, Figure 1 , a magnetic resonance imaging (MRI) system 100 is schematically illustrated as including a scanner 102, scanner control circuitry 104, and system control circuitry 106. According to embodiments described herein, the MRI system 100 is generally configured to perform MR imaging.

[0036] The system 100 also includes a remote access and storage system or device, such as a picture archiving and communication system (PACS) 108, or other devices, such as teleradiology equipment, that enable on-site or off-site access to data acquired by the system 100. In this way, MR data can be acquired and then processed and evaluated on-site or off-site. Although the MRI system 100 can include any suitable scanner or detector, in the illustrated embodiment, the system 100 includes a whole-body scanner 102 having a housing 120 through which an aperture 122 is formed. An examination table 124 can be moved into the aperture 122 to allow a patient 126 (e.g., a subject) to be positioned therein for imaging selected anatomical structures within the patient's body.

[0037] The scanner 102 includes a series of associated coils for generating a controlled magnetic field for exciting gyromagnetic material within the anatomy of the patient being imaged. Specifically, a primary magnetic coil 128 is provided for generating a primary magnetic field B0 that is generally aligned with the orifice 122. A series of gradient coils 130, 132, and 134 allow for the generation of controlled gradient magnetic fields during an examination sequence for position encoding of certain gyromagnetic nuclei within the patient 126. A radio frequency (RF) coil 136 (e.g., an RF transmit coil) is configured to generate radio frequency pulses for exciting certain gyromagnetic nuclei within the patient. In addition to the coils that may be located locally within the scanner 102, the system 100 also includes a set of receive coils or RF receive coils 138 (e.g., a coil array) configured for placement proximal to (e.g., against) the patient 126. For example, the receive coils 138 may include a cervical / thoracic / lumbar (CTL) coil, a head coil, a single-sided spine coil, and the like. Generally, receive coil 138 is placed near or on top of patient 126 to receive the weak RF signals (relative to the transmit pulses generated by the scanner coils) generated by certain gyromagnetic nuclei in the patient's body when patient 126 returns to his or her relaxed state.

[0038] The various coils of the system 100 are controlled by external circuitry to generate the desired fields and pulses and to read the emissions from the gyromagnetic material in a controlled manner. In the illustrated embodiment, a main power supply 140 provides power to the primary field coil 128 to generate the primary magnetic field Bo. A power input (e.g., power from a utility or grid), a power distribution unit (PDU), a power supply (PS), and a drive circuit 150 can collectively provide power to pulse the gradient field coils 130, 132, and 134. The drive circuit 150 can include amplification and control circuitry for supplying current to the coils as defined by the digitized pulse sequence output by the scanner control circuitry 104.

[0039] Another control circuit 152 is provided for regulating the operation of the RF coil 136. Circuit 152 includes a switching device for alternating between an active mode of operation and a passive mode of operation, in which the RF coil 136 transmits and does not transmit signals, respectively. Circuit 152 also includes amplification circuitry configured to generate RF pulses. Similarly, the receive coil 138 is connected to a switch 154 that switches the receive coil 138 between a receive mode and a non-receive mode. Thus, in the receive mode, the receive coil 138 resonates with RF signals generated by gyromagnetic nuclei released from the patient 126, and in the non-receive mode, they do not resonate with RF energy from the transmit coil (i.e., coil 136) to prevent undesirable operation. Additionally, the receive circuit 156 is configured to receive data detected by the receive coil 138 and may include one or more multiplexing and / or amplification circuits.

[0040] It should be noted that while the scanner 102 and the control / amplification circuitry are illustrated as being coupled by a single line, in practice, multiple such lines may exist. For example, separate lines may be used for control, data communication, power transmission, and the like. Furthermore, appropriate hardware may be provided along each type of line to properly process data and current / voltage. In practice, various filters, digitizers, and processors may be provided between the scanner and either or both of the scanner control circuitry 104 and the system control circuitry 106.

[0041] As shown, the scanner control circuitry 104 includes an interface circuitry 158 that outputs signals for driving the gradient field coils and the RF coils and for receiving data representing magnetic resonance signals generated during an examination sequence. The interface circuitry 158 is coupled to a control and analysis circuitry 160. Based on a defined protocol selected via the system control circuitry 106, the control and analysis circuitry 160 executes commands for driving the circuitry 150 and the circuitry 152.

[0042] The control and analysis circuitry 160 is also used to receive magnetic resonance signals and perform subsequent processing before transmitting the data to the system control circuitry 106. The scanner control circuitry 104 also includes one or more memory circuits 162 that store configuration parameters, pulse sequence descriptions, examination results, etc. during operation.

[0043] Interface circuitry 164 is coupled to control and analysis circuitry 160 for exchanging data between scanner control circuitry 104 and system control circuitry 106. In certain embodiments, control and analysis circuitry 160, while illustrated as a single unit, may include one or more hardware devices. System control circuitry 106 includes interface circuitry 166, which receives data from scanner control circuitry 104 and transmits data and commands back to scanner control circuitry 104. Control and analysis circuitry 168 may include a CPU in a general-purpose or special-purpose computer or workstation. Control and analysis circuitry 168 is coupled to memory circuitry 170 to store programming code for operating MRI system 100 and to store processed image data for later reconstruction, display, and transmission. The programming code may implement one or more algorithms that, when executed by a processor, are configured to perform reconstruction of acquired data as described below. In certain embodiments, memory circuitry 170 may store a visual transformer model used in the techniques described below. In certain embodiments, image reconstruction may occur on a separate computing device having processing and memory circuitry.

[0044] Additional interface circuitry 172 may be provided for exchanging image data, configuration parameters, and the like with external system components, such as the remote access and storage device 108. Finally, the system control and analysis circuitry 168 may be communicatively coupled to various peripheral devices for facilitating operator interface and generating hard copies of reconstructed images. In the illustrated embodiment, these peripheral devices include a printer 174, a monitor 176, and a user interface 178, which may include devices such as a keyboard, a mouse, a touch screen (e.g., integral to the monitor 176), and the like.

[0045] Figure 2 A schematic diagram illustrating the training (e.g., supervised training) of the contrastive similarity metric learning model 218 for localization is shown. A plurality of medical images are obtained. In some embodiments, the plurality of medical images are MR images. In some embodiments, the plurality of medical images may be obtained from other types of imaging (e.g., CT imaging). Each medical image undergoes multiple augmentations (e.g., cropping, transformation, rotation, etc.). This makes the contrastive similarity metric learning model 218 robust to variations in real-life images during training. Figure 2As depicted in , a medical image 220 (representing one of a plurality of medical images) is labeled with regions within a first region (e.g., two regions used to create a positive feature vector) that are selected and labeled (as indicated by reference numeral 222) and regions within a different region (e.g., dissimilar to the first region to create a negative feature vector) that are selected and labeled (as indicated by reference numeral 224). As depicted, the labeling of the medical image 220 is binary. In certain embodiments, the medical image 220 can be labeled with multiple labels. The medical image 220 (along with an enhanced version of the medical image) is input into the trained visual transformer model 180. The trained visual transformer model 180 outputs both block-level features (e.g., block-level feature vectors) (not shown) and image-level features (not shown) from the medical image 220 (and the enhanced version of the medical image). Pixel-level features (e.g., pixel-level feature vectors) 226 are interpolated from the block-level features.

[0046] The pixel-level feature vector 226 is input into the contrastive similarity metric learning model 218. The contrastive similarity metric learning model 218 is trained to make similar pixel-level feature vectors (e.g., positive pairs such as positive pairs 228 on the right side of the dotted line 230) as close as possible (e.g., minimize the distance in the embedding space) and to make dissimilar pixel-level feature vectors (e.g., negative pairs such as negative pairs 232 on the left side of the dotted line 230) as separated as possible (e.g., maximize the distance in the embedding space). The contrastive similarity learning model 218 includes two feed-forward neural networks (FFNs) 234. Positive pairs are given a weight of 1, and negative pairs are given a label of 0. The two feed-forward neural networks 234 have shared weights. The contrastive similarity metric learning model 218 outputs which pixel-level feature vectors are similar and which pixel-level feature vectors are dissimilar.

[0047] In some embodiments, each feedforward neural network 234 has a three-layer network (e.g., with 512, 256, and 128 neurons in the respective layers). In some embodiments, the batch size of the comparative similarity metric learning model 180 is 64. In some embodiments, the learning rate of the comparative similarity metric learning model 180 is 0.01. In some embodiments, the comparative similarity metric learning model 180 can utilize a stochastic optimization technique that implements a per-dimension learning rate method for stochastic gradient descent. The variables of the comparative similarity metric learning model 180 can be different from these variables.

[0048] The contrastive similarity metric learning model 218 as utilized in the present disclosure was trained using 10 medical images and their respective augmentations. The contrastive similarity metric learning model 218 as utilized in the present disclosure was tested using a test set consisting of 5 images with an accuracy of 0.88.

[0049] Figure 3 A schematic diagram illustrating data-adaptive single-shot segmentation using the basic model. Figure 3 The process is depicted for a single task (e.g., localization and segmentation of a single region of interest), but the process can be extended for multiple tasks in a single shot (i.e., localization and segmentation of multiple regions of interest). A template image 236 (e.g., a reference slice) is received or obtained, which includes a selection of a region of interest within the template image (e.g., selected via user input by a user), where the region of interest is marked in the template image 236 with a reference marker (as indicated by reference numeral 238) and associated with a label. The template image 236 includes one or more anatomical landmarks assigned corresponding anatomical labels. The template image 236 is an MR image. The template image 236 is input into the trained visual transformer model 180. The visual transformer model 180 outputs a reference pixel-level feature vector 240 from the region of interest in the template image 236. As depicted, the region of interest is an anatomical landmark. In some embodiments, the region of interest is a lesion.

[0050] Medical imaging data (eg, a medical imaging volume) acquired of a portion (eg, a shoulder) of a subject is obtained. The medical imaging data includes a plurality of slices or medical images. Figure 3 The medical imaging data in is MR imaging data. A medical image 242 (e.g., target slice 1) is input into a trained visual transformer model 180. The trained visual transformer model 180 outputs a pixel-level feature vector 244 from the medical image 242. The pixel-level feature vector 244 is derived from the block-level feature vector via interpolation. In certain embodiments, the trained visual transformer model 180 also outputs image-level features (not shown). The pixel-level feature vector 244 (e.g., all pixel-level features obtained from the medical image 242) and the reference pixel-level feature vector 240 are input into a trained contrastive similarity metric learning model 218, wherein the trained contrastive similarity metric learning model 218 is configured to automatically determine which pixel-level feature vectors in the pixel-level feature vector 244 are similar to the reference pixel-level feature vector 240. The trained contrastive similarity metric learning model 218 outputs pixel-level feature vectors 244 that are similar to the reference pixel-level feature vector 240 and pixel-level feature vectors 244 that are dissimilar to the reference pixel-level feature vector 240 .

[0051] Pixels in the medical image 242 associated with pixel-level feature vectors 244 similar to the reference pixel-level feature vector 240 are labeled using an initial segmentation mask 246, wherein the labeled pixels in the medical image 242 correspond to a region of interest (as selected in the template image 236). In some embodiments, connected component analysis is used to label the pixels to generate the initial segmentation mask 246, as indicated by reference numeral 248. The medical image with the initial segmentation mask 246 is input into a suggestable segmentation model 250. In some embodiments, the suggestable segmentation model 250 is an image segmentation base model or a generalized segmentation refinement model, such as a suggestable base SAM segmentation model configured to refine the segmentation for the region of interest. The suggestable segmentation model 250 outputs the medical image 242 labeled with a more accurate (e.g., refined) segmentation mask 252 corresponding to the region of interest. The initial segmentation mask 246 serves as an automatic hint for labeling.

[0052] In some embodiments, one or more additional medical images 254 (e.g., target slices 254) can be processed in a similar manner as the medical image 254 to locate and segment regions of interest with respective more accurate segmentation masks 256 as depicted in the medical image 254. In some embodiments, the process can be utilized for all medical images in the imaging volume of the portion of the subject. In some embodiments, the process can be performed in its entirety on fewer than all medical images in the imaging volume. Specifically, in some embodiments, the most relevant medical image in the imaging volume (i.e., the image that is closest or most similar to the template image) is processed. In some images, corresponding image-level features can be utilized in the automatic selection of the most relevant medical image in the imaging volume. In some embodiments, data-adaptive single-shot segmentation using a base model can be used to locate and segment multiple different regions of interest in medical imaging data based on multiple and different selections of different regions of interest on the same template image.

[0053] Figure 4 A flow chart of a method 258 for data adaptive single shot segmentation using a base model is illustrated. One or more steps of the method 258 may be performed by Figure 1 The method 258 may be performed by processing circuitry of the magnetic resonance imaging system 100, processing circuitry of another type of imaging system (e.g., a CT imaging system), or processing circuitry of a separate computing device. One or more of the steps of the method 258 may be performed simultaneously or in a manner similar to that described above. Figure 4 The depicted sequence may be performed in a different order. The method 258 may be used for anatomical structure localization, lesion detection, or other types of applications.

[0054] The method 258 includes obtaining a medical image of a portion of a subject (e.g., a target slice from a medical imaging volume) (box 260). The method 258 also includes receiving a selection of both a template image and a region of interest (ROI) (e.g., an anatomical landmark or a lesion) within the template image, wherein the region of interest is labeled in the template image and associated with a label (box 262). The template image includes one or more anatomical landmarks assigned with corresponding anatomical labels. The method 258 also includes inputting both the medical image and the template image (separately) into a trained visual transformer model (box 264). The method 258 further includes outputting from the trained visual transformer model both a pixel-level feature vector from the medical image and a reference pixel-level feature vector for the region of interest from the template image (box 266). Method 258 still further includes inputting both the pixel-level feature vector and the reference pixel-level feature vector into a trained contrastive similarity metric learning model, wherein the trained contrastive similarity metric learning model is configured to automatically determine which of the pixel-level feature vectors are similar to the reference pixel-level feature vector (box 268). Method 258 still further includes outputting pixels similar to the reference pixels from the trained contrastive similarity metric learning model (box 270). Method 258 also includes using an initial segmentation mask to mark pixels in the medical image associated with the pixel-level feature vector that is similar to the reference pixel-level feature vector, wherein the marked pixels in the medical image correspond to the region of interest (box 272). In some embodiments, marking pixels in the medical image associated with the pixel-level feature vector that is similar to the reference pixel-level feature vector includes generating the initial segmentation mask using a connected component analysis of the pixels. The method 258 even further includes utilizing the suggestible segmentation model to label the medical image with a refined segmentation mask corresponding to the region of interest, wherein the initial segmentation mask serves as an automatic hint for labeling (block 274).

[0055] Figure 5 A flow chart illustrating a method 276 for data-adaptive single-shot segmentation using a base model (eg, on multiple medical images or slices) is provided. One or more steps of the method 276 may be performed by Figure 1 The method 276 may be performed by processing circuitry of the magnetic resonance imaging system 100, processing circuitry of another type of imaging system (e.g., a CT imaging system), or processing circuitry of a separate computing device. One or more of the steps of the method 276 may be performed simultaneously or in a manner similar to that described above. Figure 5 The depicted sequence may be performed in a different order. The method 276 may be used for anatomical structure localization, lesion detection, or other types of applications.

[0056] Method 276 includes a medical imaging volume of the portion of the subject, wherein the medical imaging volume includes a plurality of medical images (e.g., slices) (box 278). Method 276 also includes inputting each of the plurality of medical images (e.g., individually) into a trained visual transformer model (box 280). Method 276 also includes outputting corresponding pixel-level feature vectors from each of the plurality of medical images from the trained visual transformer model (e.g., individually) (box 282). Method 276 also includes receiving a selection of both a template image and a region of interest (e.g., an anatomical landmark or lesion) within the template image, wherein the region of interest is labeled in the template image and associated with a label (box 284). The template image includes one or more anatomical landmarks assigned corresponding anatomical labels. Method 276 also inputs the template image into the trained visual transformer model (box 286). The method 276 even further includes outputting a reference pixel-level feature vector for the region of interest from the template image from the trained visual transformer model (box 288). The method 276 still further includes inputting both the corresponding pixel-level feature vector (e.g., for the corresponding medical image) and the reference pixel-level feature vector into a trained contrastive similarity metric learning model, wherein the trained contrastive similarity metric learning model is configured to automatically determine which of the corresponding pixel-level feature vectors are similar to the reference pixel-level feature vector (box 290). The method 276 still further includes outputting pixels similar to the reference pixel from the trained contrastive similarity metric learning model (box 292). The method 276 also includes labeling pixels in the corresponding medical image associated with the pixel-level feature vector that is similar to the reference pixel-level feature vector using the corresponding initial segmentation mask, wherein the labeled pixels in the corresponding medical image correspond to the region of interest (box 294). In some embodiments, labeling pixels in the corresponding medical image associated with the pixel-level feature vector that is similar to the reference pixel-level feature vector includes generating the corresponding initial segmentation mask using a connected component analysis of the pixels. Method 276 further includes labeling the corresponding medical image using a corresponding segmentation mask corresponding to the corresponding region of interest using a suggestible segmentation model, wherein the corresponding initial segmentation mask serves as an automatic hint for labeling (block 296). Blocks 290 through 296 are repeated for each medical image of the medical imaging volume of the portion of the subject.

[0057] Figure 6 A flow chart illustrating a method 298 for data-adaptive single-shot segmentation using a base model (eg, on a relevant medical image or slice) is shown. One or more steps of the method 298 may be performed by Figure 1The method 298 may be performed by processing circuitry of the magnetic resonance imaging system 100, processing circuitry of another type of imaging system (e.g., a CT imaging system), or processing circuitry of a separate computing device. One or more of the steps of the method 298 may be performed simultaneously or in a manner similar to that described above. Figure 6 The depicted sequence may be performed in a different order. The method 298 may be used for anatomical structure localization, lesion detection, or other types of applications.

[0058] Method 298 includes a medical imaging volume of the portion of the subject, wherein the medical imaging volume includes a plurality of medical images (e.g., slices) (box 300). Method 298 also includes inputting each of the plurality of medical images (e.g., individually) into a trained visual transformer model (box 302). Method 298 also includes outputting corresponding pixel-level feature vectors and corresponding image-level features (e.g., image tokens) from the trained visual transformer model (e.g., individually) from each of the plurality of medical images (box 304). Method 298 also includes receiving a selection of both a template image and a region of interest (e.g., an anatomical landmark or lesion) within the template image, wherein the region of interest is labeled in the template image and associated with a label (box 306). The template image includes one or more anatomical landmarks assigned with corresponding anatomical labels. Method 298 also inputs the template image into the trained visual transformer model (box 308). Method 298 even includes outputting a reference pixel-level feature vector and reference image-level features (e.g., reference image tokens) of the region of interest from the template image from the trained visual transformer model (box 310). Method 298 also includes determining one or more (e.g., a group of) most relevant medical images from the plurality of medical images (box 312). The most relevant medical images are those medical images that are most similar to the template image. In some embodiments, determining the most relevant medical images includes comparing the corresponding image-level features of each medical image with the reference image-level features from the template image (e.g., individually). In some embodiments, method 298 may proceed first for the most relevant medical image and then immediately for the next most relevant medical image in the selected related medical images. In some embodiments, method 298 may proceed only for the most relevant medical image. Selecting the most relevant medical image reduces processing time. In addition, selecting the most relevant medical image eliminates potential false positives.

[0059] Method 298 still further includes inputting both the corresponding pixel-level feature vector (e.g., for the corresponding medical image in the selected most relevant medical images) and the reference pixel-level feature vector into a trained contrastive similarity metric learning model, wherein the trained contrastive similarity metric learning model is configured to automatically determine which of the corresponding pixel-level feature vectors are similar to the reference pixel-level feature vector (box 314). Method 298 still further includes outputting pixels similar to the reference pixel from the trained contrastive similarity metric learning model (box 316). Method 298 also includes marking pixels in the corresponding medical image associated with the pixel-level feature vector that is similar to the reference pixel-level feature vector using a corresponding initial segmentation mask, wherein the marked pixels in the corresponding medical image correspond to the region of interest (box 318). In some embodiments, marking pixels in the corresponding medical image associated with the pixel-level feature vector that is similar to the reference pixel-level feature vector includes generating the corresponding initial segmentation mask using a connected component analysis of the pixels. The method 298 further includes labeling the corresponding medical image with a corresponding segmentation mask corresponding to the corresponding region of interest using the suggestible segmentation model, wherein the corresponding initial segmentation mask serves as an automatic hint for labeling (block 319). Blocks 314 to 319 are repeated for each of the one or more related medical images determined for the medical imaging volume of the portion of the subject.

[0060] Figure 7 A flow chart of a method 320 for data adaptive single-shot multi-label segmentation using a base model is illustrated. One or more steps of the method 320 may be performed by Figure 1 The method 320 may be performed by processing circuitry of the magnetic resonance imaging system 100, processing circuitry of another type of imaging system (e.g., a CT imaging system), or processing circuitry of a separate computing device. One or more of the steps of the method 320 may be performed simultaneously or in a manner similar to that described above. Figure 7 The depicted sequence may be performed in a different order. The method 320 may be used for anatomical structure localization, lesion detection, or other types of applications.

[0061] Method 320 includes obtaining a medical image of a portion of a subject (e.g., a target slice from a medical imaging volume) (box 322). Method 320 also includes receiving a selection of a plurality of regions of interest (ROIs) (e.g., anatomical landmarks or lesions) within the template image, wherein each of the plurality of regions of interest is individually labeled in the template image and associated with a corresponding label (box 324). The template image includes one or more anatomical landmarks assigned with corresponding anatomical labels. Method 320 also includes inputting both the medical image and the template image (separately) into a trained visual transformer model (box 326). Method 320 further includes outputting from the trained visual transformer model both pixel-level features from the medical image and a corresponding reference pixel-level feature vector for each of the plurality of regions of interest from the template image (box 328). Method 320 still further includes inputting the pixel-level feature vector and each corresponding reference pixel-level feature vector of each region of interest into the trained contrastive similarity metric learning model, wherein the trained contrastive similarity metric learning model is configured to automatically determine which pixel-level feature vectors of the pixel-level feature vectors are similar to each corresponding reference pixel-level feature vector of each region of interest (box 330). Method 320 still further includes outputting, from the trained contrastive similarity metric learning model, corresponding pixel groups that are similar to each corresponding reference pixel of the corresponding reference pixels of each region of interest (box 332). Method 320 also includes using corresponding initial segmentation masks to respectively label the corresponding pixel groups in the medical image associated with each group of corresponding pixel-level feature vectors that are similar to each corresponding reference pixel-level feature vector of each region of interest, wherein the respectively labeled corresponding pixel groups in the medical image correspond to corresponding regions of interest among the multiple regions of interest (box 334). In some embodiments, labeling the pixels in the medical image associated with the pixel-level feature vector similar to the corresponding reference pixel-level feature vector for each region of interest includes generating the corresponding initial segmentation mask using a connected component analysis of the pixels. The method 320 further includes labeling the medical image using a suggestible segmentation model with corresponding refined segmentation masks for corresponding regions, the corresponding regions corresponding to the corresponding regions of interest in the plurality of regions of interest, wherein the corresponding initial segmentation masks serve as automatic hints for labeling (block 336).

[0062] Figure 8 Depicted is a comparison of MRI images of the shoulder using different methods for locating the region of interest. Figure 8 The left MR image is used to locate the region of interest. Using data-driven methods (i.e., Figure 4 Method 258) in Figure 8 The MR image on the right is used to locate the region of interest. Figure 8 The MR images on the left correspond to Figure 8 MR images on the right. Each row of MR images includes a different slice of the shoulder. Figure 8 As depicted in , there are both false negatives 338 and false positives 340 in ROI localization using heuristic thresholds. There are no false negatives and false positives in ROI localization using a data-driven approach.

[0063] Figure 9 Depicts a comparison of segmented shoulder MR images using cues from local region of interest localization derived using different methods. Figure 9 The MR image on the left is segmented. By using the data-driven approach (i.e., Figure 4 The method 258) in the local region of interest positioning prompts Figure 9 The MR image on the right is segmented. Figure 9 The MR images on the left correspond to Figure 9 MR images on the right. Each row of MR images includes a different slice of the shoulder. Figure 9 As depicted in , segmentation by using cues from local region of interest localization using a heuristic threshold has both false negatives 342 and false positives 344. Segmentation by using cues from local region of interest localization using a data-driven approach has no false negatives and no false positives.

[0064] Figure 10 Table 346 depicts a comparison of region of interest localization and shoulder segmentation using different methods. Specifically, cosine similarity (i.e., manual thresholding or heuristic thresholding methods) is compared with contrast similarity (i.e., data adaptive methods), as described above in Figure 4

[00155] The different methods were evaluated on 15 tri-planar localizer volumes. Localization accuracy was calculated by taking the average of the predicted localizations that fell within the ground truth mask. Segmentation accuracy was analyzed using the average intersection-over-union (i.e., the area of ​​intersection between the predicted segmentation and the ground truth divided by the area of ​​union between the predicted segmentation and the ground truth). As depicted in Table 346, localization using contrast similarity was significantly more accurate than cosine similarity. Additionally, segmentation using contrast similarity was significantly more accurate than cosine similarity, as depicted in Table 346.

[0065] Figure 11 Examples include Figure 7Utilization of multi-label single shot localization of the entire knee joint volume as described in method 320 in . A template image 347 of the knee joint includes three different regions of interest within the template image selected and labeled as indicated by reference numerals 348, 350, and 352. The template image 347 is an MR image. The top row 354 of the MR images depicts the localization of the selected regions of interest in different slices of the knee joint volume. The bottom row 356 of the MR images depicts multiple slices of the knee joint imaging volume. Although Figure 7 The method 320 in FIG. 3 is robust in handling false positives in slices that are far from the template image 347, but this requires additional computation. As described above, image token matching can be used to select (i.e., determine as most relevant) specific slices. This limits processing to the most relevant slices. In the bottom row 356, slices 4-6 would be the most relevant.

[0066] The disclosed subject matter also provides the ability to automatically perform localization and segmentation using a template-based base model and a data-driven feature selection method. Furthermore, the disclosed subject matter also provides the ability to perform multi-label segmentation in a single shot. Furthermore, the disclosed subject matter also provides the ability to utilize a single base model to accomplish multiple tasks.

[0067] The technology presented and claimed herein is referred to and applied to physical and concrete examples of a practical nature that clearly improves upon the state of the art and, therefore, is not abstract, intangible, or purely theoretical. Furthermore, if any claim appended to the end of this specification contains one or more elements designated as "means for [performing] the function of ..." or "steps for [performing] the function of ...," it is intended that such elements be construed under 35 U.S.C. § 112(f). However, for any claim containing elements designated in any other manner, it is not intended that such elements be construed under 35 U.S.C. § 112(f).

[0068] This written description uses examples to disclose the subject matter, including the best mode, and also to enable any person skilled in the art to practice the subject matter, including making and using any devices or systems and performing any incorporated methods. The patentable scope of the subject matter is defined by the claims, and may include other examples that occur to those skilled in the art. Such other examples are intended to be within the scope of the claims if they have structural elements that do not differ from the literal language of the claims, or if they include equivalent structural elements with insignificant differences from the literal language of the claims.

Claims

1. A computer-implemented method, comprising: obtaining, at a processor, a medical image of a portion of a subject; receiving, at the processor, a selection of both a template image and a region of interest within the template image, wherein the region of interest is marked in the template image and associated with a label; inputting both the medical image and the template image into a trained visual transformer model via the processor; outputting from the trained visual transformer model via the processor both a pixel-level feature vector from the medical image and a reference pixel-level feature vector from the region of interest of the template image; inputting, via the processor, both the pixel-level feature vector and the reference pixel-level feature vector into a trained contrastive similarity metric learning model, wherein the trained contrastive similarity metric learning model is configured to automatically determine which of the pixel-level feature vectors are similar to the reference pixel-level feature vector; outputting, via the processor, pixels similar to reference pixels from the trained contrastive similarity metric learning model; as well as Pixels in the medical image associated with the pixel-level feature vector similar to the reference pixel-level feature vector are marked using, via the processor, an initial segmentation mask, wherein the marked pixels in the medical image correspond to the region of interest.

2. The computer-implemented method of claim 1 , further comprising: The medical image is labeled, via the processor, using a hintable segmentation model with a refined segmentation mask corresponding to a region of the region of interest, wherein the initial segmentation mask serves as an automatic hint for labeling.

3. A computer-implemented method according to claim 1, wherein marking pixels in the medical image associated with the pixel-level feature vector that is similar to the reference pixel-level feature vector includes generating the initial segmentation mask using a connected component analysis of the pixels.

4. The computer-implemented method of claim 1 , further comprising: obtaining, at a processor, a medical imaging volume of the portion of the subject, wherein the medical imaging volume comprises a plurality of medical images including the medical image; inputting, via the processor, each of the plurality of medical images into the trained visual transformer model; outputting, via the processor, from the trained visual transformer model a corresponding pixel-level feature vector from each of the plurality of medical images; inputting, via the processor, the corresponding pixel-level feature vector from each of the plurality of medical images into the trained contrastive similarity metric learning model; outputting, via the processor, from the trained comparative similarity metric learning model corresponding pixels from each of the plurality of medical images that are similar to a reference pixel; as well as The processor uses a corresponding initial segmentation mask to mark the corresponding pixels in each of the multiple medical images associated with the corresponding pixel-level feature vector from each of the multiple medical images that is similar to the reference pixel-level feature vector, wherein the corresponding pixels marked in each of the multiple medical images correspond to the region of interest.

5. The computer-implemented method of claim 4 , further comprising: Each medical image of the plurality of medical images is labeled, via the processor, using a hintable segmentation model with a respective refined segmentation mask corresponding to a respective region of the region of interest, wherein the respective initial segmentation mask serves as an automatic hint for labeling.

6. The computer-implemented method of claim 1 , further comprising: obtaining, at a processor, a medical imaging volume of the portion of the subject, wherein the medical imaging volume comprises a plurality of medical images including the medical image; inputting, via the processor, each of the plurality of medical images into the trained visual transformer model; outputting, via the processor, from the trained visual transformer model a respective pixel-level feature vector and a respective image-level feature from each of the plurality of medical images; determining, via the processor, a set of most relevant medical images from the plurality of medical images; inputting, via the processor, the corresponding pixel-level feature vectors from the set of most relevant medical images into the trained contrastive similarity metric learning model; outputting, via the processor, from the trained comparative similarity metric learning model corresponding pixels from the set of most relevant medical images that are similar to reference pixels; as well as The processor uses a corresponding initial segmentation mask to mark the corresponding pixels in each medical image in the set of most relevant medical images associated with the corresponding pixel-level feature vector similar to the reference pixel-level feature vector from each medical image in the set of most relevant medical images, wherein the corresponding pixels marked in each medical image in the set of most relevant medical images correspond to the region of interest.

7. The computer-implemented method of claim 6 , further comprising: Each medical image in the set of most relevant medical images is labeled, via the processor, using a hintable segmentation model with a respective refined segmentation mask corresponding to a respective region of the region of interest, wherein the respective initial segmentation mask serves as an automatic hint for labeling.

8. The computer-implemented method of claim 6, wherein determining the set of most relevant medical images from the plurality of medical images is based on the image-level features.

9. The computer-implemented method of claim 1 , further comprising: receiving, at the processor, the selection of a plurality of regions of interest within the template image, wherein each region of interest in the plurality of regions of interest is respectively labeled in the template image and associated with a corresponding label; outputting, via the processor, from the trained visual transformer model a corresponding reference pixel-level feature vector for each of the plurality of regions of interest from the template image; inputting, via the processor, each corresponding reference pixel-level feature vector into the trained contrastive similarity metric learning model, wherein the trained contrastive similarity metric learning model is configured to automatically determine which of the pixel-level feature vectors are similar to each corresponding reference pixel-level feature vector for each region of interest; outputting, via the processor, from the trained contrastive similarity metric learning model, a corresponding set of pixels similar to each of the corresponding reference pixels of each region of interest; as well as The processor uses the corresponding initial segmentation mask to respectively mark the corresponding pixel groups in the medical image associated with each group of corresponding pixel-level feature vectors similar to each corresponding reference pixel-level feature vector of each region of interest, wherein the corresponding pixel groups respectively marked in the medical image correspond to corresponding regions of interest among the multiple regions of interest.

10. The computer-implemented method of claim 9, further comprising: The medical image is labeled, via the processor, using a hintable segmentation model with respective refined segmentation masks for respective regions, the respective regions corresponding to the respective regions of interest among the plurality of regions of interest, wherein the respective initial segmentation masks serve as automatic hints for labeling.

11. A system, comprising: a memory encoding processor-executable routines; as well as a processor configured to access the memory and to execute the processor-executable routine, wherein the processor-executable routine, when executed by the processor, causes the processor to: obtaining a medical image of a portion of a subject; receiving a selection of both a template image and a region of interest within the template image, wherein the region of interest is marked in the template image and associated with a label; inputting both the medical image and the template image into a trained visual transformer model; outputting from the trained visual transformer model both a pixel-level feature vector from the medical image and a reference pixel-level feature vector of the region of interest from the template image; inputting both the pixel-level feature vector and the reference pixel-level feature vector into a trained contrastive similarity metric learning model, wherein the trained contrastive similarity metric learning model is configured to automatically determine which of the pixel-level feature vectors are similar to the reference pixel-level feature vector; outputting pixels similar to reference pixels from the trained contrastive similarity metric learning model; as well as The pixels in the medical image associated with the pixel-level feature vector similar to the reference pixel-level feature vector are marked using an initial segmentation mask, wherein the marked pixels in the medical image correspond to the region of interest.

12. A system according to claim 11, wherein the processor executable routine, when executed by the processor, further causes the processor to utilize a hintable segmentation model to label the medical image with a refined segmentation mask corresponding to the region of interest, wherein the initial segmentation mask serves as an automatic hint for labeling.

13. The system of claim 11, wherein labeling pixels in the medical image associated with the pixel-level feature vector that are similar to the reference pixel-level feature vector comprises generating the initial segmentation mask using a connected component analysis of the pixels.

14. The system of claim 11 , wherein the processor-executable routine, when executed by the processor, further causes the processor to: obtaining a medical imaging volume of the portion of the subject, wherein the medical imaging volume comprises a plurality of medical images including the medical image; inputting each medical image of the plurality of medical images into the trained visual transformer model; outputting, from the trained visual transformer model, a corresponding pixel-level feature vector from each of the plurality of medical images; inputting the corresponding pixel-level feature vector from each of the plurality of medical images into the trained contrastive similarity metric learning model; outputting, from the trained contrastive similarity metric learning model, corresponding pixels from each of the plurality of medical images that are similar to a reference pixel; as well as The corresponding pixels in each of the multiple medical images associated with the corresponding pixel-level feature vector from each of the multiple medical images that is similar to the reference pixel-level feature vector are marked using the corresponding initial segmentation mask, wherein the corresponding pixels marked in each of the multiple medical images correspond to the region of interest.

15. A system according to claim 14, wherein the processor executable routine, when executed by the processor, further causes the processor to utilize a hintable segmentation model to label each of the plurality of medical images using a corresponding refined segmentation mask corresponding to a corresponding region of the region of interest, wherein the corresponding initial segmentation mask serves as an automatic hint for labeling.

16. The system of claim 11 , wherein the processor-executable routine, when executed by the processor, further causes the processor to: obtaining a medical imaging volume of the portion of the subject, wherein the medical imaging volume comprises a plurality of medical images including the medical image; inputting each medical image of the plurality of medical images into the trained visual transformer model; outputting, from the trained visual transformer model, a corresponding pixel-level feature vector and a corresponding image-level feature from each of the plurality of medical images; determining a set of most relevant medical images from the plurality of medical images; inputting the corresponding pixel-level feature vectors from the set of most relevant medical images into the trained contrastive similarity metric learning model; outputting, from the trained contrastive similarity metric learning model, corresponding pixels from the set of most relevant medical images that are similar to a reference pixel; as well as The corresponding pixels in each of the group of most relevant medical images associated with the corresponding pixel-level feature vector similar to the reference pixel-level feature vector from each of the group of most relevant medical images are marked using the corresponding initial segmentation mask, wherein the corresponding pixels marked in each of the medical images in the group of most relevant medical images correspond to the region of interest.

17. A system according to claim 16, wherein the processor executable routine, when executed by the processor, further causes the processor to utilize a hintable segmentation model to label each medical image in the set of most relevant medical images using a corresponding refined segmentation mask corresponding to a corresponding region of the region of interest, wherein the corresponding initial segmentation mask serves as an automatic hint for labeling.

18. The system of claim 11 , wherein the processor-executable routine, when executed by the processor, further causes the processor to: receiving the selection of a plurality of regions of interest within the template image, wherein each of the plurality of regions of interest is respectively marked in the template image and associated with a corresponding label; outputting, from the trained visual transformer model, a corresponding reference pixel-level feature vector for each of the plurality of regions of interest from the template image; inputting each corresponding reference pixel-level feature vector into the trained contrastive similarity metric learning model, wherein the trained contrastive similarity metric learning model is configured to automatically determine which of the pixel-level feature vectors are similar to each corresponding reference pixel-level feature vector for each region of interest; outputting, from the trained contrastive similarity metric learning model, a corresponding set of pixels similar to each of the corresponding reference pixels of each region of interest; as well as The corresponding pixel groups in the medical image associated with each group of corresponding pixel-level feature vectors similar to each corresponding reference pixel-level feature vector of each region of interest are respectively marked using corresponding initial segmentation masks, wherein the corresponding pixel groups respectively marked in the medical image correspond to corresponding regions of interest among the multiple regions of interest.

19. A system according to claim 18, wherein the processor executable routine, when executed by the processor, further causes the processor to label the medical image using corresponding refined segmentation masks of corresponding regions using a promptable segmentation model via the processor, wherein the corresponding regions respectively correspond to the corresponding regions of interest in the multiple regions of interest, wherein the corresponding initial segmentation masks serve as automatic prompts for labeling.

20. A non-transitory computer-readable medium comprising processor-executable code that, when executed by a processor, causes the processor to: obtaining a medical image of a portion of a subject; receiving a selection of both a template image and a plurality of regions of interest within the template image, wherein each of the plurality of regions of interest is respectively labeled in the template image and associated with a corresponding label; inputting both the medical image and the template image into a trained visual transformer model; outputting from the trained visual transformer model both a corresponding pixel-level feature vector from the medical image and a corresponding reference pixel-level feature vector from each of the plurality of regions of interest in the template image; inputting both the pixel-level feature vectors and the corresponding reference pixel-level feature vectors into a trained contrastive similarity metric learning model, wherein the trained contrastive similarity metric learning model is configured to automatically determine which of the pixel-level feature vectors are similar to each corresponding reference pixel-level feature vector for each region of interest; outputting, from the trained contrastive similarity metric learning model, a corresponding set of pixels similar to each of the corresponding reference pixels of each region of interest; as well as The corresponding pixel groups in the medical image associated with each group of corresponding pixel-level feature vectors similar to each corresponding reference pixel-level feature vector of each region of interest are respectively marked using corresponding initial segmentation masks, wherein the corresponding pixel groups respectively marked in the medical image correspond to corresponding regions of interest among the multiple regions of interest.