Medical image segmentation method and device based on doctor diagnosis decision prompt enhancement
Patent Information
- Application Number
- CN202611247944.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-18
- Publication Date
- 2026-09-25
AI Technical Summary
[0003]本申请提供一种基于医生诊断决策的提示增强医学图像分割方法及装置,解决了在医学图像分割中存在分割区域偏移或不完整的问题
[0013]结合上述第二方面,在一种可能的实现方式中,提示分割单元还用于:提示分割大模型基于医学图像数据及感兴趣区域提示信息,输出与感兴趣区域对应的目标区域概率掩码;对目标区域概率掩码进行阈值化处理,得到与感兴趣区域对应的目标区域二值掩码。
Smart Images

Figure CN122821141A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing, and in particular to a method and apparatus for enhancing medical image segmentation based on physician diagnostic decisions. Background Technology
[0002] Medical image segmentation is a key technology in the field of medical image processing, widely used for the automatic identification and region division of target structures or abnormal areas in medical images. Current technologies often employ deep learning-based segmentation models to directly predict the segmentation region at the pixel level. However, medical images often exhibit significant differences in target morphology, unclear structural boundaries, and inconsistent image quality, making it difficult for segmentation models to accurately determine the spatial extent of the target region, resulting in segmentation region offset or incompleteness. Therefore, a method is urgently needed to address the technical problems of segmentation region offset or incompleteness in existing medical image segmentation techniques. Summary of the Invention
[0003] This application provides a medical image segmentation method and apparatus with enhanced prompts based on physician diagnostic decisions, which solves the problem of segmentation region offset or incompleteness in medical image segmentation.
[0004] To achieve the above objectives, this application adopts the following technical solution: Firstly, a cue-enhanced medical image segmentation method based on physician diagnostic decisions is provided, comprising: acquiring medical image data to be labeled and physician diagnostic decision data; determining region of interest (ROI) cue information based on physician diagnostic decision data; the ROI cue information is used to indicate the spatial range of potential target structures or lesion regions in the medical image data; transmitting the medical image data and the ROI cue information to a large-scale cue segmentation model to obtain a binary mask of the target region corresponding to the ROI; the large-scale cue segmentation model is a large-scale cue segmentation model trained on medical image data; transmitting the medical image data and the target region binary mask through multiple channels to a cue-enhanced segmentation network for segmentation processing to obtain the medical image segmentation result of the target structure or lesion region; the cue-enhanced segmentation network introduces a binary mask as explicit guidance information in the feature fusion stage, and the cue-enhanced segmentation network is a mask-enhanced segmentation network built based on the U-Net architecture.
[0005] In conjunction with the first aspect mentioned above, in one possible implementation, medical image data and region of interest (ROI) cue information are transmitted to a large cue segmentation model to obtain a binary mask of the target region corresponding to the ROI. This includes: the large cue segmentation model outputting a probability mask of the target region corresponding to the ROI based on the medical image data and ROI cue information; and thresholding the probability mask of the target region to obtain a binary mask of the target region corresponding to the ROI.
[0006] In conjunction with the first aspect mentioned above, in one possible implementation, the cue segmentation model outputs a target region probability mask corresponding to the region of interest based on medical image data and region of interest cue information. This includes: encoding the medical image data through an image coding network to obtain medical image features; encoding the region of interest cue information through a cue coding network to obtain cue features; fusing the medical image features and cue features to obtain fused features; and using a mask generation network based on the fused features to obtain a target region probability mask corresponding to the region of interest.
[0007] In conjunction with the first aspect mentioned above, in one possible implementation, the large-scale suggestion segmentation model is a large-scale suggestion segmentation model trained on medical image data. The training process includes: S1: acquiring medical image samples, doctor's diagnostic decision data samples, and corresponding annotation information for the medical image samples; S2: generating corresponding region of interest (ROI) suggestion information based on the doctor's diagnostic decision data samples; S3: inputting the medical image samples and ROI suggestion information into the large-scale suggestion segmentation model for training, and outputting the corresponding target region prediction probability mask; S4: updating the model parameters of the large-scale suggestion segmentation model based on the difference between the target region prediction probability mask and the annotation information; repeating S1 to S4 until the large-scale segmentation model converges.
[0008] In conjunction with the first aspect mentioned above, in one possible implementation, medical image data and a binary mask of the target region are transmitted through multiple channels to a cue-enhancing segmentation network for segmentation processing to obtain a medical image segmentation result for the target structure or lesion region. This includes: constructing multi-channel data by using medical image data as the first input channel and the binary mask of the target region as the second input channel; extracting multi-scale features from the multi-channel data using the encoder of the cue-enhancing segmentation network; generating guiding features based on the binary mask of the target region; transmitting the multi-scale features and guiding features to the decoder of the cue-enhancing segmentation network via skip connections, fusing the guiding features corresponding to the binary mask of the target region to obtain fused features; and performing progressive upsampling and reconstruction on the fused features to output the medical image segmentation result corresponding to the target structure or lesion region.
[0009] In conjunction with the first aspect mentioned above, in one possible implementation, generating guiding features based on a binary mask of the target region includes: performing a scale transformation on the binary mask of the target region to obtain a first binary mask of the target region; the first binary mask of the target region having the same size as the multi-scale feature of the corresponding level; performing feature mapping on the first binary mask of the target region to obtain guiding features; and using the guiding features to characterize the spatial location information of the target region.
[0010] In conjunction with the first aspect mentioned above, in one possible implementation, multi-scale features and guiding features are transmitted to the decoder of the cue enhancement segmentation network via skip connections. The guiding features corresponding to the binary mask of the target region are fused to obtain fused features. This includes: transmitting the multi-scale features output from the corresponding level of the encoder to the corresponding level of the decoder via skip connections; synchronously transmitting the guiding features aligned with the scale of the multi-scale features to the decoder; and performing feature fusion processing on the multi-scale features and guiding features in the decoder to obtain fused features.
[0011] In conjunction with the first aspect mentioned above, in one possible implementation, the large-scale cue segmentation model consists of an image coding layer, a cue coding layer, and a mask generation layer. The image coding layer is used to extract features from medical image data to obtain medical image features. The cue coding layer is used to encode cue information in the region of interest to obtain cue features. The mask generation layer is used to generate a target region mask based on the medical image features and the cue features.
[0012] Secondly, a cue-enhanced medical image segmentation device based on physician diagnostic decisions is provided. The device includes: a data acquisition unit, a cue information generation unit, a cue segmentation unit, and a cue enhancement segmentation unit. The data acquisition unit acquires medical image data to be labeled and physician diagnostic decision data. The cue information generation unit determines region-of-interest (ROI) cue information based on the physician diagnostic decision data. The ROI cue information indicates the spatial range of potential target structures or lesions in the medical image data. The cue segmentation unit transmits the medical image data and the ROI cue information to a large-scale cue segmentation model to obtain a binary mask of the target region corresponding to the ROI. The large-scale cue segmentation model is trained on medical image data. The cue enhancement segmentation unit transmits the medical image data and the target region binary mask through multiple channels to a cue enhancement segmentation network for segmentation processing, obtaining the medical image segmentation result of the target structure or lesion region. The cue enhancement segmentation network introduces a binary mask as explicit guidance information during the feature fusion stage. The cue enhancement segmentation network is a mask enhancement segmentation network built based on the U-Net architecture.
[0013] In conjunction with the second aspect above, in one possible implementation, the prompting segmentation unit is further configured to: output a target region probability mask corresponding to the region of interest based on medical image data and region of interest prompting information; and perform thresholding processing on the target region probability mask to obtain a binary target region mask corresponding to the region of interest.
[0014] This application provides a cue-enhanced medical image segmentation method and apparatus based on physician diagnostic decisions. It generates region-of-interest (ROI) cue information by incorporating physician diagnostic decision data and uses a large-scale cue segmentation model to obtain a binary mask of the target region corresponding to the RIO. This provides explicit spatial guidance constraints for the subsequent segmentation network during the segmentation process, effectively narrowing the segmentation search range and reducing the impact of background interference on the segmentation results. Simultaneously, medical image data and the target region binary mask are input into a cue-enhanced segmentation network built on a U-Net architecture via multiple channels. The binary mask is introduced as explicit guidance information during the feature fusion stage, enabling the network to continuously focus on the target region during encoding and decoding. This improves the boundary integrity and segmentation consistency of target structures or lesions, enhances the accuracy and stability of segmentation results under complex medical imaging conditions, strengthens the segmentation network's ability to focus on the target region, and solves the problem of segmentation region offset or incompleteness in medical image segmentation.
[0015] It should be understood that the descriptions of technical features, technical solutions, beneficial effects, or similar language in this application do not imply that all features and advantages can be achieved in any single embodiment. Rather, it is understood that the description of a feature or beneficial effect means that a specific technical feature, technical solution, or beneficial effect is included in at least one embodiment. Therefore, the descriptions of technical features, technical solutions, or beneficial effects in this specification do not necessarily refer to the same embodiment. Furthermore, the technical features, technical solutions, and beneficial effects described in this embodiment can be combined in any suitable manner. Those skilled in the art will understand that embodiments can be implemented without one or more specific technical features, technical solutions, or beneficial effects of a particular embodiment. In other embodiments, additional technical features and beneficial effects may be identified in specific embodiments that do not embody all embodiments. Attached Figure Description
[0016] Figure 1 A flowchart illustrating a medical image segmentation method based on doctor's diagnostic decision-making enhancement provided in this application embodiment; Figure 2 A flowchart illustrating another medical image segmentation method based on doctor's diagnostic decision-making provided in this application embodiment; Figure 3 A flowchart illustrating another medical image segmentation method based on doctor's diagnostic decision-making provided in this application embodiment; Figure 4 A flowchart illustrating another medical image segmentation method based on doctor's diagnostic decision-making provided in this application embodiment; Figure 5 This is a schematic diagram illustrating the segmentation process and results provided in an embodiment of this application. Figure 6 This is a schematic diagram of a medical image segmentation device that enhances prompts based on doctor's diagnostic decisions, provided as an embodiment of this application. Detailed Implementation
[0017] In the description of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B. The "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. Furthermore, "at least one" means one or more, and "multiple" means two or more. The terms "first," "second," etc., do not limit the quantity or order of execution, and "first," "second," etc., do not necessarily imply differences.
[0018] It should be noted that, in this application, the terms "exemplary" or "for example" are used to indicate that something is being described as an example, illustration, or illustration. Any embodiment or design described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or design solutions. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0019] To address the issues of segmentation region offset or incompleteness in existing medical image segmentation technologies, this application provides a cue-enhanced medical image segmentation method based on physician diagnostic decisions. The method includes: acquiring medical image data to be labeled and physician diagnostic decision data; determining region of interest (ROI) cue information based on the physician diagnostic decision data; the ROI cue information indicates the spatial extent of a potential target structure or lesion region in the medical image data; transmitting the medical image data and the ROI cue information to a large-scale cue segmentation model to obtain a binary mask of the target region corresponding to the ROI; the large-scale cue segmentation model is trained on medical image data; transmitting the medical image data and the target region binary mask through multiple channels to a cue-enhanced segmentation network for segmentation processing to obtain the medical image segmentation result of the target structure or lesion region; the cue-enhanced segmentation network introduces a binary mask as explicit guidance information during the feature fusion stage, and the cue-enhanced segmentation network is a mask-enhanced segmentation network built on the U-Net architecture.
[0020] Figure 1 A flowchart illustrating the medical image segmentation method based on doctor's diagnostic decision-making provided in this application embodiment is shown below. Figure 1 As shown, the method includes: S101. Obtain the medical image data to be labeled and the doctor's diagnostic decision data.
[0021] The medical image data to be labeled indicates the medical image input used for segmentation tasks, which can be two-dimensional slices or three-dimensional volume data, including but not limited to modalities such as CT, MRI, ultrasound or pathological slices. The doctor's diagnostic decision data indicates the positioning or attention information generated by the interpreter during the image reading process, which can be expressed in the form of box selection coordinates, key points, text descriptions or rough outlines.
[0022] In one possible implementation, medical images are preprocessed to meet model input requirements, including resampling to a uniform spatial resolution, intensity normalization, and cropping or padding the images as needed to ensure the efficiency and numerical stability of subsequent model processing. Meanwhile, the doctor's diagnostic decision data is saved in a structured format for the subsequent prompt generation module to read.
[0023] It should be noted that privacy and compliance requirements have been taken into account during the data collection phase, and doctors' diagnostic decision data can be obtained in an anonymized and authorized environment.
[0024] This step ensures that the input data is processable and consistent, providing a stable and repeatable raw data source for subsequent prompt generation and mask inference. This step helps reduce upstream and downstream performance fluctuations caused by differences in image quality and improves the robustness of the entire process.
[0025] S102. Based on the doctor's diagnostic decision data, determine the region of interest prompt information.
[0026] Among them, the Region of Interest Prompt (ROI Prompt) is used to indicate the spatial extent of the target structure or lesion region in medical image data. It can be represented in the form of a rectangular box, a circular region, a sparse point set, or a coarse binary region. This type of prompt is intended to provide spatial prior rather than precise annotation to the subsequent large-scale prompting segmentation model.
[0027] In one possible implementation, the doctor's diagnostic decision data is parsed, mapped to coordinates, and subjected to necessary buffer expansion. Buffer expansion is used to expand single points or narrow annotations into regions with a certain spatial tolerance. Then, the prompts from multiple sources are subjected to consistency checks and merging processes, and the final prompt information is output in a unified format for subsequent use.
[0028] It should be noted that when location information from unstructured sources is converted into spatial cues, semantic parsing and manual or automatic verification should be performed first to ensure the reliability of spatial positioning. Furthermore, the granularity and confidence level of the cues should be marked so that soft or hard cues strategies can be adopted subsequently.
[0029] As an example, if the viewer provides only a single annotation point, a circular prompt area can be generated with that point as the center and a preset radius; if multiple overlapping boxes are provided, they can be combined into a final rectangular or polygonal prompt layer using a union or weighted average method, and the confidence level of the combined prompt can be recorded.
[0030] Based on the above steps, this step structures the doctor's decision-making information into clear spatial cues, thereby effectively constraining the search range and improving the matching probability between the cues and the real target area in the subsequent mask generation stage.
[0031] S103. Transmit the medical image data and region of interest (ROI) hint information to the large-scale hint segmentation model to obtain the binary mask of the target region corresponding to the ROI.
[0032] Among them, the cue segmentation large model refers to a model instance that can receive images and cue information and output pixel-level probability masks. It is pre-trained or fine-tuned on medical image datasets to adapt to the characteristics of image modalities and generates a target region probability map related to the cue region during the inference stage.
[0033] It should be noted that the large-scale cue segmentation model consists of an image encoding layer, a cue encoding layer, and a mask generation layer. The image encoding layer extracts features from the medical image data, obtaining medical image features. It employs a Swin Transformer-based encoder structure, taking standardized medical images as input and outputting image features at four scales (resolutions of 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the input) through four levels of feature extraction. Each feature extraction level includes window attention calculation, residual connections, and layer normalization operations. A global average pooling layer is added at the end of the encoding layer to extract global semantic features of the image. The cue encoding layer encodes cue information from the region of interest, obtaining cue features. The cue encoding layer employs different encoding methods depending on the type of cue information: for spatial geometric cuees such as rectangular boxes and circular regions, the coordinate information of the cue is converted into relative coordinates, input into a multilayer perceptron (MLP), and the output cue features have the same feature dimension as the image encoding layer; for sparse point set cuees, the point set coordinates are converted into a heatmap (Gaussian distribution centered on the points), and the heatmap is converted into cue features through a 1×1 convolutional layer. The mask generation layer generates a target region mask based on medical image features and cue features. The mask generation layer adopts a decoder structure, containing four upsampling modules corresponding to the four scale features of the image encoding layer; skip connections are used to fuse the multi-scale features of the encoding layer with the upsampling features of the decoder; a 1×1 convolutional layer is added at the end of the decoder to convert the number of feature channels to 1, outputting a probability mask of the target region (pixel value range 0-1).
[0034] In one possible implementation, medical image data and region of interest (ROI) cue information are fed into the model as joint input. The model first extracts multi-scale image representations and performs spatial mapping and numericalization of the cue information. Then, in the fusion stage, the image representations and cue mappings are combined to generate a probability mask for the target region. Subsequently, the probability mask is thresholded and morphologically smoothed to obtain a binary mask for the target region.
[0035] Based on the above steps, this step combines structured prompts with image information for mask inference, thereby providing more relevant spatial priors for subsequent segmentation stages and improving the localization accuracy of the overall segmentation process.
[0036] S104. The medical image data and the binary mask of the target region are transmitted to the cue enhancement segmentation network through multiple channels for segmentation processing to obtain the medical image segmentation result of the target structure or lesion region.
[0037] The cue-enhanced segmentation network introduces a binary mask as explicit guidance information during the feature fusion stage. The cue-enhanced segmentation network is a mask-enhanced segmentation network built on the U-Net architecture.
[0038] In one possible implementation, after acquiring the medical image and the binary mask of the target region, the device first uses the binary mask as an additional channel and performs channel-level concatenation with the original image to form a multi-channel input. This multi-channel input is then fed into a cue-enhanced segmentation network to perform layer-by-layer feature extraction and upsampling reconstruction operations. After inference, the network outputs a segmentation map of the target structure or lesion region. Finally, the device formats and restores the size of the segmentation map to make it consistent with the original image, thus obtaining the final segmentation result.
[0039] It should be noted that binary masks can be directly incorporated into the input in the channel dimension as hard constraints, or they can be input in the form of probabilities or soft weights to provide smoother spatial guidance. The specific form to be used can be determined based on the confidence of the output of the cue segmentation large model and the support of the downstream network for soft cues. At the same time, boundary refinement or morphological operations can be used after decoding to further optimize the boundary quality.
[0040] Based on the above steps, this step introduces a binary mask obtained from prompting reasoning as guidance in the feature fusion stage of the segmentation network, which improves the final segmentation results in terms of spatial focus, boundary integrity and consistency.
[0041] This application's embodiments, through a technical solution comprised of data acquisition, prompt generation, prompt-driven segmentation, and mask-enhanced segmentation, achieve a natural connection in information flow between the prompt segmentation large model and the prompt-enhanced segmentation network. During the data input stage, medical image data and decision-related information are utilized simultaneously, providing subsequent steps with clearer priors. Furthermore, by constructing region-of-interest (ROI) prompt information, explicit spatial guidance of the potential target range is achieved. The prompt segmentation large model transforms the prompt information into a structured and directly usable binary mask of the target region, enabling subsequent networks to operate within more precise spatial constraints. Finally, the introduction of a mask as an explicit guiding signal into the segmentation network allows the segmentation process to focus on key regions, reduce background interference, and improve inference stability. This results in more focused feature extraction, clearer region localization, and more consistent output results in the segmentation task, thereby exhibiting higher interpretability and overall segmentation quality in complex image scenarios. This addresses the problems of segmentation region offset or incompleteness in existing technologies for medical image segmentation.
[0042] In one possible implementation of the embodiments of this application, combined with Figure 1 ,like Figure 2 As shown, the above S103 can be implemented through the following S201 and S202, which are explained in detail below: S201. The large-scale segmentation model, based on medical image data and region of interest (ROI) prompts, outputs a probability mask of the target region corresponding to the ROI.
[0043] The target region probability mask indicates the confidence distribution of each pixel or voxel in the image as belonging to the region of interest. The confidence can be a real-valued probability or a normalized score, which reflects the strength of the model's judgment on the probability that each location is the target region.
[0044] In one possible implementation, medical image data and region-of-interest (ROI) cue information are fed into the model as joint input, assembled according to the input format used during model training. The medical image is in tensor format [C,H,W,D], where C is the number of channels, H is the height, W is the width, and D is the depth. The cue information is in tensor format [N,5], where N is the number of cue information and 5 is the cue type identifier. The model's image encoding layer performs four-level feature extraction on the medical image to obtain multi-scale image features. The cue encoding layer encodes the cue information, generating cue features that match the dimensionality of the 1 / 4-scale image features and performs spatial alignment. The fused features are input to the mask generation layer, where multi-scale features are fused through upsampling and skip connections to output a probability mask of the target region. Temperature scaling is performed on the probability mask to adjust the discriminative power of the probability distribution.
[0045] It should be noted that, in order to ensure the availability and stability of the probability mask, the image scale and prompt format should be consistent with those during the training phase during model inference. Temperature scaling or probability calibration steps can be introduced into the model output to adjust the reliability of the probability distribution. In addition, in multimodal or cross-mechanism scenarios, input normalization or domain adaptation techniques can be adopted to reduce performance fluctuations caused by distribution differences.
[0046] Based on the above steps, this step couples the prompt with the image information to produce a continuous spatial confidence distribution. This probability mask is beneficial for using controllable thresholds or soft prompt strategies in subsequent processing, thereby providing a flexible basis for binarization and subsequent guidance while preserving information about uncertain areas.
[0047] S202. Threshold the probability mask of the target region to obtain the binary mask of the target region corresponding to the region of interest.
[0048] Thresholding refers to converting a probability mask into a binary form according to a preset or adaptive threshold rule. The binary mask is used to explicitly indicate the set of locations considered as target regions, so as to serve as explicit spatial guidance information in subsequent segmentation networks.
[0049] In one possible implementation, a fixed threshold or an adaptive threshold strategy based on probability distribution statistics is used to perform binarization. The fixed threshold is, for example, 0.5, or the adaptive threshold can be determined by the mean and variance of the probability mask. After binarization, connected component filtering and morphological opening and closing operations are performed to remove small area noise and fill in internal holes to obtain a more regular mask region.
[0050] As an example, in this embodiment, an adaptive threshold strategy is adopted to calculate the mean and standard deviation of the probability mask, and set the threshold to μ+0.5σ; pixels in the probability mask that are greater than the threshold are set to 1, and pixels that are less than or equal to the threshold are set to 0, thus obtaining an initial binary mask; morphological operations are performed on the initial binary mask, namely, firstly, opening operations are performed using 3×3 structuring elements to remove small-area noise; then, closing operations are performed using 3×3 structuring elements to fill the small holes inside the mask; finally, the target region binary mask is obtained.
[0051] Based on the above steps, this step transforms continuous probability information into explicit binary space constraints. The resulting binary mask is convenient to serve as an explicit guiding input for the downstream segmentation network, thereby providing a practically usable spatial prior for controllable processing of uncertain regions and subsequent feature focusing.
[0052] This application's embodiments, from prompt-driven probabilistic inference to explicit binarization constraints, enable the spatial confidence information output by the prompt-driven segmentation model to be effectively transformed into a binary mask that can be directly used by the downstream segmentation network. While preserving uncertainty information, the stability and interpretability of the mask are improved through a controllable binarization strategy, providing a clear and operable spatial prior for subsequent mask-guided segmentation processing, thereby improving the positioning accuracy and robustness of the overall segmentation process in complex image scenes.
[0053] In one possible implementation of the embodiments of this application, combined with Figure 1 ,like Figure 3 As shown, the training process of the large-scale segmentation model in S103 above can be specifically implemented through the following S1 to S4, which are explained in detail below: S1: Obtain medical image samples, doctor's diagnostic decision data samples, and annotation information corresponding to the medical image samples.
[0054] Among them, medical image samples refer to data instances used to train the large-scale segmentation model, doctor diagnostic decision data samples refer to clinical interpretation information paired with medical image samples, and annotation information indicates the pixel-level reference segmentation or boundary contour corresponding to each medical image sample, which is used as a supervision signal for training and may contain annotation metadata.
[0055] In one possible implementation, a unified preprocessing procedure is first performed on the medical image samples. At the same time, the doctor's diagnostic decision data is structured into machine-readable prompt modalities, such as coordinate point sequences, heat maps, or text labels. The annotation information is then subjected to consistency checks, expert review, and necessary manual corrections to ensure the reliability of the reference annotations.
[0056] As an example, in this embodiment, the collected medical image samples are divided into a training set, a validation set, and a test set in a 7:2:1 ratio; during the division, it is ensured that samples from the same patient do not cross sets to avoid data leakage. A consistency check is performed on the pixel-level reference annotations, calculating the consistency of annotations for the same image by different annotators. If the consistency is <0.8, it is corrected by a senior physician; the corrected annotations are converted into a binary mask format as a supervision signal for training.
[0057] It should be noted that the data sample for doctors' diagnostic decisions has been anonymized and obtained the consent of relevant personnel.
[0058] S2: Generate corresponding region of interest prompts based on doctor's diagnostic decision data samples.
[0059] In one possible implementation, the localization elements in the doctor's diagnostic decision data sample are first parsed, such as the doctor's initial selection of lesions, region highlighting, or layered diagnostic descriptions. Then, this decision information is converted into a cue region aligned with the image spatial coordinates. This cue information is then used as input to Internal SAM2 to generate an initial ROI mask.
[0060] It should be noted that the Internal SAM2 model in this embodiment uses the SAM2-base model structure, with an input image size of 1024×1024. For 3D medical images, a slice-based input is used, with each slice independently generating an initial ROI mask, and then the masks are stacked to form a 3D mask. The structured prompt information is converted into an input format supported by SAM2, with box selection prompts converted into box prompts and key point prompts converted into point prompts. Internal SAM2 is only an exemplary implementation of a large prompt segmentation model, and this application does not limit the specific network structure of the large prompt segmentation model.
[0061] It should be noted that the process of generating prompts must be consistent with the doctor's actual interpretation behavior so that Internal SAM2 can fully learn the doctor's attention patterns, thereby achieving the goal of improving the model's interpretability; at the same time, the prompts must be spatially aligned with the original images to avoid prompt offset affecting the quality of mask generation.
[0062] As an example, in an abdominal MRI task, if a doctor selects a "suspected liver lesion area" in the reading interface, the selection box can be automatically converted into a BOX Prompt, and further converted into a spatial cue signal that SAM2 can recognize for subsequent mask generation.
[0063] S3: Input the medical image samples and region of interest (ROI) hints into the large-scale hint segmentation model for training, and output the corresponding target region prediction probability mask.
[0064] The cue segmentation big model is a joint segmentation system that includes the internal cue big model (Internal SAM2) and the mask enhancement U-Net. It achieves explicit region attention and deep feature extraction through a dual-channel structure of image input and cue input, and outputs a target region prediction mask with continuous probability values.
[0065] In one possible implementation, the system first inputs the region cues generated by the doctor's prompts into InternalSAM2 to obtain a prompt-driven initial ROI mask. Then, the mask and the original medical image are input together as two channels into the mask-enhanced U-Net. The network training and feature learning are completed through mechanisms such as feature concatenation, skip connection mask superposition, and prompt consistency constraints.
[0066] It should be noted that in this step, Internal SAM2 and Mask Enhancement U-Net are connected in series. The cue mask of SAM2 is responsible for providing key region guidance, while U-Net is responsible for performing pixel-level fine segmentation. The two establish an explicit semantic connection through the cue mask channel, which effectively solves the problem of insufficient attention in traditional U-Net.
[0067] As an example, for the liver tumor segmentation task, the doctor's selection prompts generate an initial liver ROI mask after internal SAM2 inference. This mask is used as the second input channel and enters the mask enhancement U-Net together with the original CT image. The network finally outputs a continuous probability map of the liver tumor for subsequent thresholding.
[0068] S4: Based on the difference between the predicted probability mask of the target region and the annotation information, update the model parameters of the large segmentation model; repeat S1 to S4 above until the large segmentation model converges.
[0069] Among them, model updating based on the difference between probability mask and annotation information refers to using the loss function to calculate the deviation between the prediction result and the real annotation, and using backpropagation to adjust the parameters of the large-scale cue segmentation model.
[0070] In one possible implementation, various loss functions can be combined for optimization, such as cross-entropy loss, Dice loss, and boundary loss. During training, learning rate adjustment strategies, batch training strategies, and model validation mechanisms can be used to improve training stability.
[0071] It should be noted that the model training process can determine whether the convergence condition has been met based on validation set performance metrics (such as Dice coefficient, IoU, etc.), and an early stopping strategy can be used during the training phase to avoid overfitting. The optimizer used is the AdamW optimizer, with an initial learning rate of 1e-4 and a weight decay coefficient of 1e-5; the learning rate adjustment strategy uses cosine annealing, with a 10% decay every 10 epochs; the batch size is adjusted according to the hardware configuration. At each layer of the U-Net encoder with mask enhancement, the initial ROI mask is scaled and concatenated channel-by-channel with the image features of the corresponding layer.
[0072] As an example, in the liver tumor segmentation task, the Dice loss is calculated by comparing the predicted probability mask with the manually labeled mask, and gradient updates are performed based on this loss, so that the model can gradually improve the segmentation quality of the tumor boundary.
[0073] This application's embodiments incorporate information of interest from physician diagnostic decision data samples into the medical image segmentation training process. This allows the segmentation model to acquire clinical knowledge cues while learning pixel-level annotation information, thereby improving the model's localization and interpretability of lesion regions or structural targets. Through the joint input of region of interest (ROI) hints and medical image samples, the model can adaptively focus on clinically significant regions, reducing interference from irrelevant features and improving the effectiveness of the training process. Furthermore, a difference feedback mechanism between the target region prediction probability mask and the annotation is introduced during iterative training, enabling the model parameters to continuously converge and optimize segmentation boundary accuracy, thus achieving a more stable, accurate, and generalizable automatic medical image segmentation effect.
[0074] In one possible implementation of the embodiments of this application, combined with Figure 1 ,like Figure 4 As shown, the above S104 can be specifically implemented through the following S401 to S405, which are explained in detail below: S401. Construct multi-channel data by using medical image data as the first input channel and the target region binary mask as the second input channel.
[0075] In one possible implementation, medical images and binary masks are stacked along the channel dimension, allowing the network to simultaneously acquire raw image information and region cues during the input phase.
[0076] It should be noted that the multi-channel structure can enhance the network's sensitivity to regional guidance information and avoid feature dilution caused by indiscriminate processing of the entire map.
[0077] Based on the above steps, explicitly injecting prompts into the input terminal can effectively improve the accuracy of subsequent feature extraction and region localization.
[0078] S402. Enhance the encoder of the segmentation network with prompts to extract multi-scale features from multi-channel data.
[0079] The encoder of the cue enhancement segmentation network includes convolutional layers, residual modules, and a Transformer module, which are used to extract local and global features. The kernel size of the convolutional layer is 3×3, the stride is 2, and the padding is 1. There are 4 residual modules, each containing 2 convolutional layers. The Transformer module has 8 heads and a feature dimension of 256.
[0080] In one possible implementation, the encoder downsamples the multi-channel input layer by layer to form contextual information at different scales.
[0081] It should be noted that multi-scale structures are good at capturing differences in the scale and shape of target structures, thus improving the network's recognition ability in fine-grained region segmentation.
[0082] Based on the above steps, the encoder can output rich image context features for subsequent fusion.
[0083] S403. Generate guiding features based on the binary mask of the target region.
[0084] Among them, the guiding features are used to explicitly highlight the local regions indicated by the binary mask, so that the model can focus on the corresponding structures during the decoding stage.
[0085] In one possible implementation, the binary mask of the target region is scaled to obtain a first target region binary mask; the first target region binary mask has the same size as the multi-scale feature of the corresponding level; the first target region binary mask is subjected to feature mapping to obtain guiding features; the guiding features are used to characterize the spatial location information of the target region.
[0086] As an example, bilinear interpolation is used to scale the binary mask to the same size as the four scale features (1 / 4, 1 / 8, 1 / 16, 1 / 32) output by the encoder, resulting in four scale binary masks of the first target region. For each scale binary mask of the first target region, feature mapping is performed through two 1×1 convolutional layers. The first convolutional layer converts the number of channels to 64, and the second convolutional layer converts the number of channels to be consistent with the number of feature channels of the corresponding level encoder. A ReLU activation function is added after the convolutional layer, and the output is the guiding feature. The guiding feature is used to represent the spatial location information of the target region.
[0087] Based on the above steps, guiding features can become explicit cue information for subsequent fusion. The introduction of guiding features can enhance the region focusing mechanism of the cue-enhancing segmentation network and reduce interference from irrelevant background.
[0088] S404. Multi-scale features and guiding features are transmitted to the decoder of the cue enhancement segmentation network through skip connections, and the guiding features corresponding to the binary mask of the target region are fused to obtain the fused features.
[0089] In one possible implementation, the multi-scale features output from the corresponding level of the encoder are transmitted to the corresponding level of the decoder via skip connections; the guiding features, which are aligned with the scale of the multi-scale features, are synchronously transmitted to the decoder; and the multi-scale features and guiding features are fused in the decoder to obtain the fused features.
[0090] It should be noted that the multi-scale features and guiding features are fused by attention weighting in the decoder, and then the guiding features are converted into a single-channel attention weight map by 1×1 convolution to obtain the fused features.
[0091] Based on the above steps, the model can obtain more localized and structurally clear fusion features during the decoding stage, which helps to improve the segmentation quality of complex target regions.
[0092] S405. Perform stepwise upsampling and reconstruction on the fused features, and output the medical image segmentation results corresponding to the target structure or lesion region.
[0093] The upsampling and reconstruction process is used to restore the low-resolution fusion features to a spatial size consistent with the input image and generate the final segmentation prediction map.
[0094] In one possible implementation, the feature map is amplified in the decoder using bilinear interpolation upsampling combined with convolution, and two convolutional layers are cascaded after each upsampling level to restore local structural details. In the final output layer, a 1×1 convolution is used to compress the number of channels to 1 or the number of classes, and a sigmoid activation function is used to generate a probability map. The output probability map is then processed using the same adaptive thresholding strategy as S103 to generate binary segmentation results. Connectivity analysis is performed on the binary segmentation results; for single-target segmentation, the largest connected component is retained, while for multi-target segmentation, connected components with an area greater than 50 voxels are retained. The size of the segmentation results is restored to the original medical image size using nearest-neighbor interpolation to ensure complete spatial alignment between the segmented region and the original image.
[0095] It should be noted that, in order to improve the continuity of the prediction boundary, a deep supervision mechanism can be introduced in the decoding stage, that is, to generate auxiliary segmentation results in the intermediate layer, so that the model can optimize the structural expression layer by layer during the training process.
[0096] Based on the above steps, this model can generate target region segmentation results with clear structure, continuous boundaries, and that meet the requirements of medical diagnosis.
[0097] As an example, in an embodiment of this application, Figure 5 This is a schematic diagram illustrating the segmentation process and results provided in the embodiments of this application, such as... Figure 5 As shown, the process takes CT images as input. First, based on the doctor's decision information, a bounding box is generated on the image to select suspected lesion areas (i.e., ROI area hints in the form of "rectangular box hints"). This hint is then fed into an internal SAM model trained on a medical dataset (USTCData) to generate the corresponding ROI mask. Then, the original input image and the ROI mask are fed into a "mask-enhanced U-Net network," which fuses multi-scale features using skip connections, and finally outputs a segmented lesion-annotated image (color indicates the segmented and recognized area).
[0098] This application embodiment constructs a multi-channel input by combining the medical image and the target region's binary mask, enabling the cue-enhanced segmentation network to fully utilize the positional information generated by the cue-enhanced segmentation model during the segmentation process. This strengthens the network's ability to focus on the target region during the encoding stage. By explicitly indicating key structures through guided features during the decoding stage, the network can maintain spatial consistency and boundary integrity during multi-scale information fusion. Simultaneously, through the synergistic fusion mechanism of guided features and skip connections, the network can more accurately recover the true shape of the target structure during the reconstruction stage and effectively suppress the interference of background noise and irrelevant regions on the segmentation results. The technical solution of this application embodiment maintains high accuracy while possessing stronger generalization ability and positional sensitivity, which is beneficial for obtaining reliable medical image segmentation results stably under different devices, different body parts, and different image conditions.
[0099] The foregoing mainly describes the solutions of the embodiments of this application from the perspective of device implementation. It is understood that each device, such as a medical image segmentation device for enhancing doctor diagnostic decisions, includes at least one of the hardware structures and software modules corresponding to the execution of each function in order to achieve the above-mentioned functions. Those skilled in the art should readily recognize that, based on the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0100] This application embodiment can divide the medical image segmentation device for doctor diagnostic decision enhancement into functional units according to the above method example. For example, each function can be divided into separate functional units, or two or more functions can be integrated into one processing unit. The integrated unit can be implemented in hardware or as a software functional unit. It should be noted that the unit division in this application embodiment is illustrative and only represents one logical functional division; other division methods may be used in actual implementation.
[0101] When using integrated units, Figure 6 A possible structural schematic diagram of the medical image segmentation device with enhanced prompts based on doctor's diagnostic decision (hereinafter referred to as the medical image segmentation device 60 with enhanced prompts based on doctor's diagnostic decision) involved in the above embodiments is shown. The medical image segmentation device 60 with enhanced prompts based on doctor's diagnostic decision includes a data acquisition unit 601, a prompt information generation unit 602, a prompt segmentation unit 603, and a prompt enhancement segmentation unit 604. Figure 6The schematic diagram shown can be used to illustrate the structure of the medical image segmentation device based on doctor's diagnostic decision-making involved in the above embodiments.
[0102] when Figure 6 The schematic diagram shown illustrates the structure of the medical image segmentation device based on doctor's diagnostic decision-making in the above embodiments. The data acquisition unit 601 acquires the medical image data to be labeled and the doctor's diagnostic decision data; the prompt information generation unit 602 determines the region of interest (ROI) prompt information based on the doctor's diagnostic decision data; the ROI prompt information indicates the spatial range of the medical image data that may contain the target structure or lesion region; the prompt segmentation unit 603 transmits the medical image data and the ROI prompt information to the prompt segmentation large model to obtain a binary mask of the target region corresponding to the ROI; the prompt segmentation large model is a prompt segmentation large model trained on medical image data; the prompt enhancement segmentation unit 604 transmits the medical image data and the target region binary mask through multiple channels to the prompt enhancement segmentation network for segmentation processing to obtain the medical image segmentation result of the target structure or lesion region; the prompt enhancement segmentation network introduces a binary mask as explicit guidance information in the feature fusion stage, and the prompt enhancement segmentation network is a mask enhancement segmentation network built based on the U-Net architecture.
[0103] For example, in one possible implementation, the cue segmentation unit 603 is further configured to: output a target region probability mask corresponding to the region of interest based on medical image data and region of interest cue information; and perform thresholding processing on the target region probability mask to obtain a target region binary mask corresponding to the region of interest.
[0104] The use of the device in the embodiments of this application corresponds to the implementation of the above method. The detailed implementation process can be found in the above method, and will not be repeated here.
[0105] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software programs, implementation can be, in whole or in part, in the form of a computer program product. This computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device containing one or more servers, data centers, etc., that can be integrated with the medium. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state disks (SSDs)).
[0106] Although this application has been described herein in conjunction with various embodiments, those skilled in the art, by reviewing the accompanying drawings, disclosure, and appended claims, will understand and implement other variations of the disclosed embodiments in carrying out the claimed application. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude multiple instances. A single processor or other unit can implement several functions listed in the claims. While different dependent claims may recite certain measures, this does not mean that these measures cannot be combined to produce good results.
[0107] Although this application has been described in conjunction with specific features and embodiments, it is obvious that various modifications and combinations can be made thereto without departing from the spirit and scope of this application. Accordingly, this specification and drawings are merely exemplary illustrations of this application as defined by the appended claims, and are considered to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from the spirit and scope of this application. Thus, if such modifications and modifications of this application fall within the scope of the claims of this application and their equivalents, this application is also intended to include such modifications and modifications.
Claims
1. A medical image segmentation method with cue enhancement based on physician diagnostic decisions, characterized in that, include: Acquire medical image data to be labeled and physician diagnostic decision data; The doctor's diagnostic decision data is positioning information that can be mapped to a medical image spatial coordinate system; Based on the doctor's diagnostic decision data, region of interest (ROI) information is determined; the ROI information is used to indicate the spatial range of the target structure or lesion region contained in the medical image data. The medical image data and the region of interest (ROI) hint information are transmitted to a large-scale hint segmentation model to obtain a binary mask of the target region corresponding to the ROI; the large-scale hint segmentation model is a large-scale hint segmentation model trained on medical image data; The medical image data and the binary mask of the target region are transmitted through multiple channels to a cue-enhanced segmentation network for segmentation processing to obtain the medical image segmentation result of the target structure or lesion region. The cue-enhanced segmentation network introduces the binary mask as explicit guidance information in the feature fusion stage. The cue-enhanced segmentation network is a mask-enhanced segmentation network built on the U-Net architecture.
2. The medical image segmentation method based on doctor's diagnostic decision-making with enhanced prompting as described in claim 1, characterized in that, The step of transmitting the medical image data and the region of interest (ROI) cue information to the cue segmentation model to obtain a binary mask of the target region corresponding to the ROI includes: The large-scale segmentation model, based on medical image data and region of interest (ROI) prompts, outputs a probability mask of the target region corresponding to the ROI. The probability mask of the target region is thresholded to obtain a binary mask of the target region corresponding to the region of interest.
3. The medical image segmentation method based on doctor's diagnostic decision-making with enhanced prompting as described in claim 2, characterized in that, The large-scale segmentation model, based on medical image data and region of interest (ROI) cue information, outputs a target region probability mask corresponding to the ROI, including: The medical image data is encoded using an image coding network to obtain medical image features; The region of interest prompt information is encoded using a prompting encoding network to obtain prompting features; The medical image features and the prompt features are fused to obtain the fused features; Based on the fusion features, a target region probability mask corresponding to the region of interest is obtained through a mask generation network.
4. The medical image segmentation method based on doctor's diagnostic decision-making with enhanced prompting as described in claim 2, characterized in that, The aforementioned large-scale cue segmentation model is a large-scale cue segmentation model trained on medical image data. The training process includes: S1: Obtain medical image samples, doctor's diagnostic decision data samples, and annotation information corresponding to the medical image samples; S2: Generate corresponding region of interest prompt information based on the doctor's diagnostic decision data sample; S3: Input the medical image sample and the region of interest prompt information into the prompt segmentation large model for training, and output the corresponding target region prediction probability mask; S4: Based on the difference between the predicted probability mask of the target region and the annotation information, update the model parameters of the large-scale segmentation model; Repeat steps S1 to S4 until the large segmentation model converges.
5. The medical image segmentation method based on doctor's diagnostic decision-making with enhanced prompting as described in claim 1, characterized in that, The step of transmitting the medical image data and the target region binary mask through multiple channels to a cue-enhanced segmentation network for segmentation processing to obtain the medical image segmentation result of the target structure or lesion region includes: Multi-channel data is constructed by using the medical image data as the first input channel and the target region binary mask as the second input channel. The encoder of the segmentation network is enhanced by prompts to perform multi-scale feature extraction on multi-channel data; Guiding features are generated based on the binary mask of the target region; The multi-scale features and the guiding features are transmitted to the decoder of the cue enhancement segmentation network via skip connections, and the guiding features corresponding to the binary mask of the target region are fused to obtain the fused features; The fused features are upsampled and reconstructed step by step to output the medical image segmentation results corresponding to the target structure or lesion region.
6. The medical image segmentation method based on doctor's diagnostic decision-making with enhanced prompting as described in claim 5, characterized in that, The generation of guiding features based on the binary mask of the target region includes: The target region binary mask is scaled to obtain a first target region binary mask; the first target region binary mask has the same size as the multi-scale feature of the corresponding level; The binary mask of the first target region is subjected to feature mapping processing to obtain guiding features; the guiding features are used to characterize the spatial location information of the target region.
7. The medical image segmentation method based on doctor's diagnostic decision-making with enhanced prompting as described in claim 5, characterized in that, The multi-scale features and the guiding features are transmitted to the decoder of the cue enhancement segmentation network via skip connections, and the guiding features corresponding to the binary mask of the target region are fused to obtain the fused features, including: The multi-scale features output from the corresponding level of the encoder are transmitted to the corresponding level of the decoder via skip connections; The guiding features aligned with the multi-scale feature scale are synchronously transmitted to the decoder; The multi-scale features and the guiding features are fused in the decoder to obtain fused features.
8. The medical image segmentation method based on doctor's diagnostic decision-making with enhanced prompting as described in claim 4, characterized in that, The cue segmentation model consists of an image coding layer, a cue coding layer, and a mask generation layer. The image coding layer is used to extract features from medical image data to obtain medical image features. The prompt encoding layer is used to perform feature encoding on the prompt information of the region of interest to obtain prompt features; The mask generation layer is used to generate a target region mask based on the medical image features and the prompt features.
9. A medical image segmentation device with enhanced prompting based on physician diagnostic decisions, characterized in that, The device includes: a data acquisition unit, a prompt information generation unit, a prompt segmentation unit, and a prompt enhancement segmentation unit; The data acquisition unit is used to acquire medical image data to be labeled and doctor's diagnostic decision data; The prompt information generation unit is used to determine region of interest prompt information based on the doctor's diagnostic decision data; the region of interest prompt information is used to indicate the spatial range of the medical image data that may contain target structures or lesions; The cue segmentation unit is used to transmit the medical image data and the region of interest cue information to the cue segmentation large model to obtain a binary mask of the target region corresponding to the region of interest; the cue segmentation large model is a cue segmentation large model trained on medical image data; The cue-enhancing segmentation unit is used to transmit the medical image data and the target region binary mask through multiple channels to the cue-enhancing segmentation network for segmentation processing, so as to obtain the medical image segmentation result of the target structure or lesion region; the cue-enhancing segmentation network introduces the binary mask as explicit guidance information in the feature fusion stage, and the cue-enhancing segmentation network is a mask enhancement segmentation network built based on the U-Net architecture.
10. The apparatus according to claim 9, characterized in that, The prompt segmentation unit is also used for: The large-scale segmentation model, based on medical image data and region of interest (ROI) prompts, outputs a probability mask of the target region corresponding to the ROI. The probability mask of the target region is thresholded to obtain a binary mask of the target region corresponding to the region of interest.