Brain tumor interactive segmentation method and system based on prompt
Through the interactive segmentation method based on hints, brain images are screened and processed, combined with feature extraction and fusion of two-dimensional and three-dimensional convolutional neural networks, the problem of inaccurate segmentation results when processing complex or boundary blurred images is solved, and higher segmentation accuracy and model stability are achieved.
Patent Information
- Application Number
- CN202510498889.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-04-21
AI Technical Summary
The existing brain tumor segmentation method is not accurate enough when dealing with images with complex tumor morphology or blurred boundaries, and due to data scarcity, diversity and labeling subjectivity, the model cannot be effectively guaranteed in terms of accuracy and stability.
Using an interactive segmentation method based on hints, the slice samples with the most tumor pixels are screened as the prompt image by acquiring and labeling brain images, and brightness enhancement processing is performed, features are extracted in combination with two-dimensional and three-dimensional convolutional neural networks, and prototype features are constructed through feature fusion and average mask pooling calculations, and the final segmentation results of tumor and non-tumor areas are finally obtained through feature matching.
It significantly improves the segmentation effect in complex or boundary blur images, improves the generalization ability of the model and adaptability to different data samples, and enhances the accuracy and reliability of the model when facing data scarcity, diversity and labeling problems.
Smart Images

Figure CN120047471A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image segmentation, and particularly to a prompt-based interactive brain tumor segmentation method and system. Background Art
[0002] Brain tumors are serious diseases caused by the abnormal growth of brain cells, and their treatment and diagnosis have always been one of the challenges in the medical field. With the development of medical imaging technology, the early detection and precise segmentation of brain tumors have become increasingly important. The precise segmentation of brain tumors plays a crucial role in subsequent treatment plan design, tumor monitoring, and patient prognosis evaluation. Currently, deep learning-based brain tumor segmentation methods are developing rapidly and have achieved remarkable progress to a certain extent. However, despite extensive research and application in clinical practice, there are still many technical problems, especially in the processing of complex tumor morphology, blurred boundaries, image noise, and diverse data.
[0003] Currently, the automatic segmentation methods for brain tumors are mainly based on three deep learning architectures: convolutional neural networks (CNNs), generative adversarial networks (GANs), and Transformers. The CNN-based segmentation method uses the method of local image patch classification to transform the image segmentation problem into a classification task of small-scale image patches. This method can implicitly extract features by learning training data and significantly improve the segmentation accuracy. The GAN-based brain tumor segmentation method uses the adversarial mechanism of the generator and discriminator to generate high-quality segmentation results. Transformers perform excellently in processing large-scale image data, can effectively capture long-range dependencies in images, and are gradually applied to brain tumor segmentation tasks, showing good results.
[0004] However, the existing methods do not perform well in dealing with specific difficult samples. Especially when facing images with complex tumor morphology or blurred boundaries, the segmentation results are not accurate enough. Due to the scarcity, diversity, and subjectivity of brain tumor data annotation, the training data is insufficient and of uneven quality, which makes it impossible to effectively guarantee the accuracy and stability of the model. Summary of the Invention
[0005] In order to solve the technical problems that the existing methods do not perform well in dealing with specific difficult samples, especially when facing images with complex tumor morphology or blurred boundaries, the segmentation results are not accurate enough. Due to the scarcity, diversity, and subjectivity of brain tumor data annotation, the training data is insufficient and of uneven quality, which makes it impossible to effectively guarantee the accuracy and stability of the model, the present invention provides a prompt-based interactive brain tumor segmentation method and system.
[0006] The technical solutions provided by the embodiments of the present invention are as follows:
[0007] First aspect:
[0008] A prompt-based interactive brain tumor segmentation method provided by an embodiment of the present invention includes:
[0009] S1: Obtain a brain image containing multiple slice samples;
[0010] S2: Perform tumor annotation on each of the slice samples to form an annotated image;
[0011] S3: Calculate the number of tumor pixels in each of the annotated slice samples, and select the slice sample with the largest number of tumor pixels as the prompt image;
[0012] S4: Perform brightness enhancement processing on the tumor region in the prompt image to obtain an enhanced prompt image;
[0013] S5: Extract two-dimensional features from the enhanced prompt image through a two-dimensional convolutional neural network;
[0014] S6: Extract three-dimensional features from the brain image through a three-dimensional convolutional neural network;
[0015] S7: Concatenate and fuse the two-dimensional features and the three-dimensional features, and restore the spatial resolution of the fused features through the upsampling layer of the three-dimensional convolutional neural network to obtain slice fine segmentation features;
[0016] S8: Extract the annotated image features of the annotated image through the three-dimensional convolutional neural network;
[0017] S9: Perform average mask pooling calculation on the prompt information of the enhanced prompt image and the annotated image features to construct first prototype features of tumor and non-tumor regions;
[0018] S10: Calculate the similarity between the three-dimensional features of the brain image and the first prototype features, and perform feature fusion through average mask pooling to construct second prototype features of the brain image;
[0019] S11: Perform feature matching on the brain image by combining the first prototype features, the second prototype features, and the slice fine segmentation features to obtain the final segmentation results of the tumor and non-tumor regions.
[0020] Second aspect:
[0021] A prompt-based interactive brain tumor segmentation system provided by an embodiment of the present invention includes:
[0022] A processor;
[0023] A memory storing computer-readable instructions, which when executed by the processor, implement the hint-based interactive brain tumor segmentation method as described in the first aspect.
[0024] The third aspect:
[0025] A computer-readable storage medium provided by an embodiment of the present invention, on which a computer program is stored, and when the program is executed by a processor, it implements the hint-based interactive brain tumor segmentation method as described in the first aspect.
[0026] The beneficial effects brought by the technical solutions provided by the embodiments of the present invention at least include:
[0027] (1) In the embodiments of the present invention, the tumor region is highlighted through brightness enhancement, thereby improving the recognizability of segmentation. The features extracted by the two-dimensional convolutional neural network strengthen the local details of the tumor region. At the same time, the fusion of two-dimensional and three-dimensional features and the restoration of spatial resolution make the details of the tumor region clearer. Especially in complex or blurred boundary images, these processes can significantly improve the segmentation effect on difficult samples.
[0028] (2) In the embodiments of the present invention, the slice sample with the most tumor pixels is selected as the hint image to ensure that the most representative sample is selected for training. At the same time, the prototype features are constructed by calculating the average mask pooling, reducing the influence brought by the subjectivity of annotation and enhancing the generalization ability of the model. Feature fusion further improves the adaptability and stability of the model to different data samples, and then improves the accuracy and reliability of the model when facing data scarcity, diversity, and annotation problems. Description of the Drawings
[0029] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0030] Figure 1 It is a schematic flowchart of a hint-based interactive brain tumor segmentation method provided by an embodiment of the present invention;
[0031] Figure 2 It is a schematic structural diagram of a hint-based interactive brain tumor segmentation system provided by an embodiment of the present invention. Detailed Embodiments
[0032] The following describes the technical solutions in the present invention with reference to the drawings.
[0033] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to give examples, illustrations or explanations. Any embodiment or design described as an "example" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Rather, the use of the word "example" is intended to present concepts in a specific manner. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one of the two.
[0034] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, the meaning they express is the same. "(of)", "corresponding", and "corresponding" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, the meaning they express is the same.
[0035] In the embodiments of the present invention, sometimes a subscript such as W 1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meaning they express is the same.
[0036] To make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.
[0037] Refer to the attached Figure 1 illustrates a schematic flow diagram of a prompt-based interactive brain tumor segmentation method provided by an embodiment of the present invention.
[0038] The embodiments of the present invention provide a prompt-based interactive brain tumor segmentation method, which can be implemented by a prompt-based interactive brain tumor segmentation device. The prompt-based interactive brain tumor segmentation device can be a terminal or a server. The processing flow of the prompt-based interactive brain tumor segmentation method can include the following steps:
[0039] S1: Obtain a brain image containing multiple slice samples.
[0040] Among them, a slice sample refers to a two-dimensional image obtained by cutting three-dimensional brain image data (such as MRI or CT scans) along a certain plane. These slices usually appear as a series of parallel and continuous images, and each image represents a cross-section of the brain on this plane. Slice samples are widely used in medical image analysis, especially in the detection and segmentation of brain tumors.
[0041] It should be noted that each slice represents a two-dimensional image of the brain on a certain plane.
[0042] In the present invention, converting brain images into slice samples can more effectively extract local information and improve the accuracy of segmentation and detection.
[0043] S2: Perform tumor annotation on each slice sample to form an annotated image.
[0044] Among them, an annotated image refers to an image in which specific regions (such as tumors, lesions, tissue structures, etc.) are marked on the basis of the original image through manual or automated methods. Annotated images are usually used in supervised learning tasks in medical image analysis and serve as training data to provide guidance for deep learning models so that the models can learn how to identify and segment specific regions.
[0045] In the present invention, accurate annotated images can help the model better understand the morphology, size, and location of the target region. Especially in medical image segmentation tasks such as tumors, precise annotation helps the model provide more refined segmentation results in practical applications.
[0046] S3: Calculate the number of tumor pixels in each annotated slice sample, and select the slice sample with the largest number of tumor pixels as the prompt image.
[0047] Among them, the number of tumor pixels refers to the total number of pixels occupied by the tumor region in medical images (such as MRI, CT scans, etc.). Each medical image usually consists of millions of pixels, and each pixel represents a small region in the image and has a certain gray value or color information. In brain images, the pixels in the tumor region usually exhibit gray values different from those of normal tissues, which enables us to identify the tumor region through image processing techniques.
[0048] In a possible implementation manner, S3 specifically includes:
[0049] Calculate the number of tumor pixels in each annotated slice sample through the following formula, and select the slice sample with the largest number of tumor pixels as the prompt image:
[0050]
[0051] Among them, N i represents the number of tumor pixels in the i-th slice, S i1 represents the number of pixels with a position label of 1 in the i-th slice, S i2 represents the number of pixels with a position label of 2 in the i-th slice, S i4 represents the number of pixels with a label position of 4 in the i-th slice, max represents the maximum value, and i represents the index of the slice.
[0052] Specifically, the screening method is as follows, 、 、 Respectively represent the first The number of pixels labeled 1, 2, and 4 on the slice, For the The total number of pixels of the three types of labels on the slice, find the largest one in slices 0-154 , you can locate the slice The slice with the largest number of tumor pixels was selected.
[0053] In the present invention, by selecting the slices with the largest number of tumor pixels as the prompt image, it is possible to ensure that the most representative samples are selected. These samples usually contain key features of the tumor area and can provide more accurate guidance information for subsequent segmentation tasks, thereby improving the training quality of the model. At the same time, by selecting the slices with the most tumor pixels, the model can adapt to more typical and representative tumor features, which will make the training results more generalizable when facing new or unlabeled images, and improve the stability of segmentation and diagnosis.
[0054] S4: Perform brightness enhancement processing on the tumor area in the prompt image to obtain an enhanced prompt image.
[0055] Among them, brightness enhancement processing is a technology in image preprocessing, which aims to enhance the brightness or contrast of a specific area in the image, so as to make the target area (such as a tumor, lesion or other area of interest) more prominent. This processing method improves the recognizability of the image by changing the gray value (i.e. brightness) of the pixels in the image, especially when the contrast between the target area and the background is low, brightness enhancement can significantly improve the visibility of the target area.
[0056] In a possible implementation, S4 specifically includes:
[0057] The grayscale value of the position labeled 1 in the prompt image is magnified by a first preset multiple, the grayscale value of the position labeled 2 is magnified by a second preset multiple, and the grayscale value of the position labeled 4 is magnified by a third preset multiple to form an enhanced prompt image.
[0058] Optionally, the first preset multiple is 1.5, the second preset multiple is 1.25, and the third preset multiple is 0.7.
[0059] Specifically, then, the first The slice is the "hint" we selected, and its serial number reflects the spatial location information of the tumor, while the pixel value and range reflect the edge information of the tumor. According to the prompt information we just set, the brightness adjustment operation is performed on the tumor area in the sample image with a relatively rich number of pixels. According to the scaling factor, the original image of the four modes is For the pixels corresponding to the labels at positions 1, 2, and 4 in the slice, the gray values are respectively changed to 1.5 times, 1.25 times, and 0.7 times of the original.
[0060] In the present invention, by enhancing the brightness of the pixels at different label positions, the tumor area can be made more prominent. Especially when the contrast between the tumor area and the background is low, the brightness enhancement can significantly improve the visibility of the target area. This helps the automated system to more easily identify the tumor. At the same time, performing different brightness enhancement processes on different labels (such as the areas with labels 1, 2, and 4), for example, magnifying the gray value of label 1 by 1.5 times, magnifying label 2 by 1.25 times, and reducing label 4 to 0.7 times, can help the model better distinguish different parts of the tumor or different types of tumors. In this way, the model can more accurately process the tumor area according to the different labels, thereby improving the segmentation accuracy.
[0061] S5: Extract two-dimensional features from the enhanced prompt image through a two-dimensional convolutional neural network.
[0062] Among them, the two-dimensional convolutional neural network (2D CNN, Convolutional Neural Network) is a deep learning model widely used in the field of image processing, especially in tasks such as image classification, object detection, and image segmentation.
[0063] In a possible implementation manner, S5 includes:
[0064] S501: Perform feature extraction on the enhanced prompt image through the following formula to obtain preliminary features:
[0065]
[0066] Among them, represents the preliminary features, f represents the composite function operation, L represents the total number of feature map layers, represents the output feature map calculation function of the Lth layer in the two-dimensional convolutional neural network, represents the output feature map calculation function of the (L - 1)th layer in the two-dimensional convolutional neural network, represents the output feature map calculation function of the first layer in the two-dimensional convolutional neural network, represents the output feature map calculation function of the lth layer, represents the pooling operation, represents the activation function, represents the output feature map of the (l - 1)th layer in the two-dimensional convolutional neural network, represents the convolutional kernel weight of the (l - 1)th layer in the two-dimensional convolutional neural network, represents the bias term of the lth layer in the two-dimensional convolutional neural network.
[0067] S502: Adopt the attention mechanism to combine the injected prompt information with the preliminary features to obtain two-dimensional features:
[0068]
[0069] wherein, represents the two-dimensional feature, represents the activation function, represents the preliminary feature, represents the prompt information, T represents the transpose operation, represents the dimension of the prompt information.
[0070] Among them, the attention mechanism is a mechanism that mimics the process of human visual attention and is widely used in deep learning, especially in the fields of natural language processing (NLP) and computer vision (CV). Its core idea is to enable the model to dynamically select and focus on certain specific parts of the input when processing the input information, rather than processing all input data equally. This process helps to improve the efficiency and accuracy of the model, especially when dealing with long sequences or complex data.
[0071] Optionally, a recombination operation is performed on the two-dimensional features. This process ensures that it can meet the geometric shape and data format requirements for splicing and fusing with other features according to the shape requirements set by the subsequent splicing conditions.
[0072] In the present invention, by extracting the preliminary features in the enhanced prompt image through a two-dimensional convolutional neural network (2D CNN), spatial structure information (such as edges, textures, shapes, etc.) can be effectively extracted from the image, thereby providing a richer representation for subsequent feature processing. At the same time, the introduction of the attention mechanism enables the model to guide the learning process according to the prompt information (such as the spatial position information of the tumor or the attention to a specific area), thereby improving the accuracy of the model in dealing with specific tasks. For example, in the brain tumor segmentation task, by concentrating the attention on the tumor area, the model can more accurately detect and segment the tumor, reducing background interference and errors.
[0073] Furthermore, the recombination operation enables the two-dimensional features to meet the shape requirements for subsequent feature splicing and fusion. This step ensures the consistency of the features in terms of dimension, shape, and data format, laying a foundation for subsequent feature fusion and avoiding calculation problems caused by mismatched sizes of different features.
[0074] S6: Extract the three-dimensional features of the brain image through a three-dimensional convolutional neural network.
[0075] Among them, the three-dimensional convolutional neural network (3D CNN, 3D Convolutional Neural Network) is a deep learning model widely used in processing data with three-dimensional structures, such as medical images (e.g., MRI, CT scans, etc.), video processing, 3D object recognition, and other fields. Different from the two-dimensional convolutional neural network that processes planar images, the three-dimensional convolutional neural network processes three-dimensional data and can capture the spatial depth information in the data, thereby effectively analyzing complex three-dimensional structures.
[0076] In a possible implementation manner, S6 is specifically:
[0077] Extract the three-dimensional features of the brain image through the following formula:
[0078]
[0079] Among them, represents the three-dimensional feature, f represents the composite function operation, L represents the total number of layers of the feature map, represents the output feature map calculation function of the L-th layer in the three-dimensional convolutional neural network, represents the output feature map calculation function of the (L - 1)-th layer in the three-dimensional convolutional neural network, represents the output feature map calculation function of the first layer in the three-dimensional convolutional neural network, represents the output feature map calculation function of the l-th layer in the three-dimensional convolutional neural network, represents the pooling operation, represents the activation function, represents the output feature map of the (l - 1)-th layer in the three-dimensional convolutional neural network, represents the convolutional kernel weight of the (l - 1)-th layer in the three-dimensional convolutional neural network, represents the bias term of the l-th layer in the three-dimensional convolutional neural network.
[0080] In the present invention, through the three-dimensional convolution operation, the 3D CNN can better recognize and learn the complex spatial structure in the image. For example, tumors usually appear on multiple slices. The three-dimensional convolution can combine the information of these slices to form a more complete description of the tumor area, improving the richness of feature expression. At the same time, in medical image analysis, especially in the analysis of brain images, structures such as tumors and brain tissues often have complex shapes and spatial relationships. By using 3D convolution, the model can learn these complex three-dimensional structures, thereby improving the segmentation accuracy and reducing the cases of mis-segmentation or missed segmentation.
[0081] S7: Concatenate and fuse the two-dimensional features and three-dimensional features, and restore the spatial resolution of the fused features through the upsampling layer of the three-dimensional convolutional neural network to obtain the slice fine segmentation features.
[0082] Among them, the upsampling layer is an operation in the convolutional neural network (CNN) used to increase the spatial resolution of the input data, thereby restoring the size of the image or enhancing the details of the image. It is commonly used in tasks such as image segmentation, generative models (such as generative adversarial networks), and autoencoders, especially when restoring image details or generating high-resolution images.
[0083] Among them, the slice fine segmentation features are features extracted from medical image data, used to describe the specific morphology, position, boundary, etc. of the target area (such as tumors or organs) in each image slice (usually referring to the two-dimensional slice images of MRI or CT scans).
[0084] In a possible implementation, S7 specifically includes:
[0085] S701: Duplicate the two-dimensional features D times in the depth dimension to expand the two-dimensional features to be consistent with the shape of the 3D features.
[0086] In the present invention, by duplicating the two-dimensional features in the depth dimension (S701), the spatial size of the two-dimensional features can be aligned with the three-dimensional features, thereby ensuring the shape consistency of the features during concatenation and fusion. This process avoids computational problems caused by mismatched feature sizes.
[0087] S702: Concatenate and fuse the three-dimensional features and the expanded enhanced features to obtain the fused features:
[0088]
[0089] Among them, d represents the depth dimension index, h represents the height dimension index, w represents the width dimension index, c represents the channel dimension index, D represents the total number of layers in the depth dimension, represents the fused features, represents the concatenation operation along the channel dimension, represents the three-dimensional features, represents the expanded two-dimensional features.
[0090] In the present invention, concatenating and fusing the two-dimensional features and three-dimensional features (S702) enables the model to utilize information from different dimensions simultaneously. Two-dimensional features can usually better capture local details, while three-dimensional features can provide global spatial information. By fusing these two types of features, the model can learn more rich and comprehensive features, improving the recognition ability of complex structures (such as tumors, brain tissues, etc.) in brain images.
[0091] S703: The spatial resolution of the fused features is restored through the upsampling layer of the three-dimensional convolutional neural network to obtain the slice fine segmentation features:
[0092]
[0093] Among them, represents the slice fine segmentation features, represents the function composition executed from right to left, L represents the total number of feature map layers, represents the ordinary 3D convolution, represents the feature splicing connected with, represents the upsampling operation of the 3D transposed convolution, represents the fused features.
[0094] In the present invention, using the upsampling layer of the three-dimensional convolutional neural network can effectively restore the spatial resolution of the fused features, enabling the model to recover more details from the low-resolution feature maps. Especially in the image segmentation task, the detail recovery is crucial for accurately segmenting tumors or other fine structures. Through upsampling, the model can perform fine segmentation of the image at a higher resolution, thereby improving the segmentation accuracy.
[0095] S8: Extract the labeled image features of the labeled image through the three-dimensional convolutional neural network.
[0096] S9: Perform average mask pooling calculation on the prompt information of the enhanced prompt image and the labeled image features to construct the first prototype features of the tumor and non-tumor regions.
[0097] It should be noted that the prompt information is specifically the morphology, size and location of the tumor. The first prototype features are specifically, after completing the labeled image features, further introducing the prompt information to define the prototypes of the tumor and non-tumor regions.
[0098] Among them, the average mask pooling calculation (Average Mask Pooling) is a special pooling method, usually used in deep learning models, especially in tasks such as image segmentation and feature extraction. It is a pooling operation that reduces the spatial dimension by calculating the average value of pixels or features within a specific region while retaining meaningful information.
[0099] In a possible implementation manner, the first prototype features of the tumor and non-tumor regions are specifically:
[0100]
[0101] Among them, represents the first prototype features, represents the average mask pooling calculation, Indicates a prompt message, Indicates the annotation of image features.
[0102] In the present invention, by combining the prompt message with the annotated image features and performing average mask pooling calculation, it can help the model focus on the key features of the tumor and non-tumor regions. The average mask pooling calculation helps to combine this information with the annotated image features, thereby enhancing the model's attention to the tumor region and avoiding interference from background noise or irrelevant regions. At the same time, through the average mask pooling calculation, the features of the tumor and non-tumor regions can be more clearly distinguished, thus helping the model better understand the morphology, size, and location of the tumor. Incorporating the prompt message into the calculation helps to further precisely define the prototypes of the tumor and non-tumor regions, thereby improving the accuracy of segmentation and detection.
[0103] S10: Calculate the similarity between the three-dimensional features of the brain image and the prototype features, and perform feature fusion through average mask pooling to construct the second prototype features of the brain image.
[0104] Among them, similarity is a measurement method for measuring the degree of similarity between two objects or data. In the fields of deep learning, data analysis, information retrieval, image processing, natural language processing, etc., the calculation of similarity is often used to compare the similarity between two data objects (such as images, texts, feature vectors, etc.) in order to make further decisions or perform optimizations.
[0105] It should be noted that the second prototype features are features further constructed on the basis of the first prototype features by combining additional feature information (such as brain image features or other prior knowledge). This process usually involves the comparison, weighting, and fusion of features, with the aim of further refining the learning and recognition capabilities of the model on the basis of the first prototype features.
[0106] In a possible implementation manner, S10 specifically includes:
[0107] S1001: Calculate the similarity between the three-dimensional features of the brain image and the first prototype features through the following formula:
[0108]
[0109] Among them, M represents the similarity between the brain image features and the first prototype features, represents the activation function, represents the cosine similarity calculation, represents the first prototype features, represents the three-dimensional features of the brain image.
[0110] In the present invention, by calculating the similarity between the brain image features and the first prototype features, the model can accurately measure the relationship between the input image and the predefined prototype. This process enables the model to make more precise feature selections based on the similarity, which helps improve the accuracy of image analysis tasks such as tumor segmentation or disease diagnosis.
[0111] S1002: Based on the similarity, perform feature fusion through the average mask pooling calculation method to construct the second prototype features of the brain image:
[0112]
[0113] Wherein, represents the second prototype features, represents the average mask pooling calculation, M represents the similarity between the brain image features and the first prototype features, represents the brain image features.
[0114] In the present invention, through feature fusion by average mask pooling calculation, the model not only combines the brain image features and the first prototype features, but also optimizes the feature fusion process by similarity weighting. This fusion method can better retain important image information while reducing the interference of irrelevant information, making the final second prototype features more representative and distinguishable.
[0115] S11: Combine the first prototype features, the second prototype features and the slice fine segmentation features to perform feature matching on the brain image, and obtain the final segmentation results of the tumor and non-tumor regions.
[0116] In a possible implementation manner, the final segmentation results of the tumor and non-tumor regions are specifically:
[0117]
[0118] Wherein, represents the final segmentation results of the tumor and non-tumor regions, represents the activation function, represents the cosine similarity calculation, represents the concatenation operation along the channel dimension, represents the first prototype features, represents the second prototype features, represents the brain image features.
[0119] In the present invention, by combining multiple feature sources (such as the first and second prototype features and the fine segmentation features) for feature matching, and using similarity calculation and the Softmax activation function, the model can more accurately identify and segment tumor and non-tumor regions. This method not only improves the segmentation accuracy, enhances the discrimination ability of the model, but also optimizes the computational efficiency and enhances the generalization ability of the model, making it perform excellently when processing different types of brain images.
[0120] The beneficial effects brought by the technical solutions provided in the embodiments of the present invention at least include:
[0121] (1) In the embodiments of the present invention, the tumor region is highlighted through brightness enhancement, thereby improving the recognizability of the segmentation. The features extracted by the two-dimensional convolutional neural network strengthen the local details of the tumor region. At the same time, the fusion of two-dimensional and three-dimensional features and the restoration of spatial resolution make the details of the tumor region clearer. Especially in complex or blurred boundary images, these processes can significantly improve the segmentation effect in difficult samples.
[0122] (2) In the embodiments of the present invention, by screening the slice sample with the most tumor pixels as the prompt image, it is ensured that the most representative sample is selected for training. At the same time, the prototype features are constructed by average mask pooling calculation, reducing the influence brought by the subjectivity of annotation and enhancing the generalization ability of the model. Feature fusion further improves the adaptability and stability of the model to different data samples, and thus improves the accuracy and reliability of the model when facing data scarcity, diversity, and annotation problems.
[0123] Refer to the attached Figure 2 illustrates the structural schematic diagram of a prompt-based interactive brain tumor segmentation system provided by the present invention.
[0124] The present invention also provides a prompt-based interactive brain tumor segmentation system 20, which is applied to the above-mentioned prompt-based interactive brain tumor segmentation method, including:
[0125] A processor 201.
[0126] A memory 202, on which computer-readable instructions are stored. When the computer-readable instructions are executed by the processor 201, the prompt-based interactive brain tumor segmentation method as in the method embodiments is implemented.
[0127] The prompt-based interactive brain tumor segmentation system 20 provided by the present invention can execute the above-mentioned prompt-based interactive brain tumor segmentation method and achieve the same or similar technical effects. To avoid repetition, the present invention will not elaborate further.
[0128] The beneficial effects brought by the technical solutions provided in the embodiments of the present invention at least include:
[0129] (1) In the embodiments of the present invention, the tumor region is highlighted through brightness enhancement, thereby improving the recognizability of segmentation. The features extracted by the two-dimensional convolutional neural network strengthen the local details of the tumor region. At the same time, the fusion of two-dimensional and three-dimensional features and the restoration of spatial resolution make the details of the tumor region clearer. Especially in complex or blurred boundary images, these processes can significantly improve the segmentation effect in difficult samples.
[0130] (2) In the embodiments of the present invention, the slice sample with the most tumor pixels is selected as the prompt image to ensure that the most representative sample is selected for training. At the same time, the prototype features are constructed by calculating the average mask pooling, reducing the influence brought by the subjectivity of annotation and enhancing the generalization ability of the model. Feature fusion further improves the adaptability and stability of the model to different data samples, and then improves the accuracy and reliability of the model in the face of data scarcity, diversity and annotation problems.
[0131] It should be understood that the processor in the embodiments of the present invention may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0132] It should also be understood that the memory in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0133] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that contains one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.
[0134] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood by referring to the context before and after.
[0135] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.
[0136] It should be understood that in various embodiments of the present invention, the magnitude of the serial numbers of the above processes does not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0137] Those of ordinary skill in the art will realize that the units and algorithm steps of each example described in connection with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0138] Those skilled in the art can clearly understand that for the sake of convenience and brevity of description, the specific working processes of the devices, apparatuses, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0139] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0140] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0141] In addition, the functional units in each embodiment of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0142] When the above-mentioned function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0143] An embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the method of interactive segmentation of brain tumors based on prompts as described in the method embodiment.
[0144] The computer-readable storage medium provided by the present invention can implement the steps and effects of the method of interactive segmentation of brain tumors based on prompts in the above method embodiment. To avoid repetition, the present invention will not elaborate further.
[0145] The beneficial effects brought by the technical solution provided by the embodiment of the present invention at least include:
[0146] (1) In the embodiment of the present invention, the tumor area is highlighted through brightness enhancement, thereby improving the recognizability of segmentation. The features extracted by the two-dimensional convolutional neural network strengthen the local details of the tumor area. At the same time, the fusion of two-dimensional and three-dimensional features and the restoration of spatial resolution make the details of the tumor area clearer. Especially in complex or blurred boundary images, these processes can significantly improve the segmentation effect in difficult samples.
[0147] (2) In the embodiment of the present invention, the slice sample with the most tumor pixels is selected as the prompt image through screening to ensure that the most representative sample is selected for training. At the same time, the prototype features are constructed by calculating the average mask pooling, reducing the influence brought by the subjectivity of annotation and enhancing the generalization ability of the model. Feature fusion further improves the adaptability and stability of the model to different data samples, and thus improves the accuracy and reliability of the model when facing data scarcity, diversity, and annotation problems.
[0148] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
[0149] The following points need to be explained:
[0150] (1) The attached drawings of the embodiments of the present invention only relate to the structures involved in the embodiments of the present invention, and other structures can refer to the general design.
[0151] (2) For clarity, in the attached drawings used to describe the embodiments of the present invention, the thickness of the layer or region is enlarged or reduced, that is, these drawings are not drawn according to the actual scale. It can be understood that when an element such as a layer, film, region or substrate is referred to as being "on" or "under" another element, the element can be "directly" on or under the other element or there can be an intermediate element.
[0152] (3) Without conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other to obtain new embodiments.
[0153] As above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. The protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A hint-based interactive brain tumor segmentation method, characterized in that: include: S1: Acquire brain images containing multiple slice samples; S2: marking the tumor on each of the slice samples to form a marked image; S3: Calculate the number of tumor pixels in each labeled slice sample, and select the slice sample with the largest number of tumor pixels as the prompt image; S4: performing brightness enhancement processing on the tumor area in the prompt image to obtain an enhanced prompt image; S5: extracting two-dimensional features in the enhanced prompt image through a two-dimensional convolutional neural network; S6: extracting three-dimensional features of the brain image through a three-dimensional convolutional neural network; S7: splicing and fusing the two-dimensional features with the three-dimensional features, and restoring the spatial resolution of the fused features through the upsampling layer of the three-dimensional convolutional neural network to obtain slice fine segmentation features; S8: extracting the annotated image features of the annotated image through the three-dimensional convolutional neural network; S9: performing average mask pooling calculation on the prompt information of the enhanced prompt image and the annotated image features to construct first prototype features of tumor and non-tumor areas; S10: calculating the similarity between the three-dimensional feature of the brain image and the first prototype feature, and performing feature fusion through average mask pooling to construct a second prototype feature of the brain image; S11: performing feature matching on the brain image by combining the first prototype feature, the second prototype feature and the slice fine segmentation feature to obtain a final segmentation result of the tumor and non-tumor areas.
2. The method for interactive brain tumor segmentation based on hints according to claim 1, characterized in that: The S3 specifically includes: The number of tumor pixels in each labeled slice sample is calculated by the following formula, and the slice sample with the largest number of tumor pixels is selected as the prompt image: ; Among them, N i represents the number of tumor pixels in the i-th slice, S i1 represents the number of pixels with the label 1 in the i-th slice position, S i2 represents the number of pixels with position label 2 in the i-th slice, S i4 It indicates the number of pixels with label position 4 in the i-th slice, max indicates the maximum value, and i indicates the index of the slice.
3. The hint-based interactive brain tumor segmentation method according to claim 1, characterized in that: The S4 specifically includes: The grayscale value of the position labeled 1 in the prompt image is magnified by a first preset multiple, the grayscale value of the position labeled 2 is magnified by a second preset multiple, and the grayscale value of the position labeled 4 is magnified by a third preset multiple to form the enhanced prompt image.
4. The hint-based interactive brain tumor segmentation method according to claim 1, characterized in that: The S5 includes: S501: Extract features of the enhanced prompt image using the following formula to obtain preliminary features: ; in, represents preliminary features, f represents composite function operation, L represents the total number of feature map layers, Represents the output feature map calculation function of the Lth layer in a two-dimensional convolutional neural network, Represents the output feature map calculation function of the L-1th layer in the two-dimensional convolutional neural network, Represents the output feature map calculation function of the first layer in a two-dimensional convolutional neural network, Represents the first The calculation function of the layer output feature map, represents the pooling operation, represents the activation function, represents the output feature map of the l-1th layer in the two-dimensional convolutional neural network, represents the convolution kernel weight of the l-1th layer in the two-dimensional convolutional neural network, Represents the bias term of the lth layer in a two-dimensional convolutional neural network; S502: Using the attention mechanism, the injected prompt information is combined with the preliminary features to obtain the two-dimensional features: ; in, represents two-dimensional features, represents the activation function, Represents preliminary features, Indicates prompt information. T represents the transpose operation, Indicates the dimension of the prompt information.
5. The hint-based interactive brain tumor segmentation method according to claim 1, characterized in that: The S6 is specifically: The three-dimensional features of the brain image are extracted by the following formula: ; in, represents three-dimensional features, f represents composite function operation, L represents the total number of feature map layers, Represents the output feature map calculation function of the Lth layer in the three-dimensional convolutional neural network, Represents the output feature map calculation function of the L-1th layer in the three-dimensional convolutional neural network, Represents the output feature map calculation function of the first layer in the three-dimensional convolutional neural network, Represents the first The calculation function of the layer output feature map, represents the pooling operation, represents the activation function, represents the output feature map of the l-1th layer in the three-dimensional convolutional neural network, represents the convolution kernel weight of the l-1th layer in the three-dimensional convolutional neural network, Represents the bias term of the lth layer in a three-dimensional convolutional neural network.
6. The hint-based interactive brain tumor segmentation method according to claim 1, characterized in that: The S7 specifically includes: S701: copying the two-dimensional feature D times in the depth dimension to expand the shape of the two-dimensional feature to be consistent with the shape of the 3D feature; S702: The three-dimensional feature is combined with the expanded enhanced feature to obtain a fused feature: ; Among them, d represents the depth dimension index, h represents the height dimension index, w represents the width dimension index, c represents the channel dimension index, and D represents the total number of layers in the depth dimension. represents the fusion feature, represents the concatenation operation along the channel dimension, Represents three-dimensional features, Represents the expanded two-dimensional features; S703: Restoring the spatial resolution of the fusion feature through the upsampling layer of the three-dimensional convolutional neural network to obtain a slice fine segmentation feature: ; in, represents the slice fine segmentation feature, represents the function compound executed from right to left, L represents the total number of feature map layers, represents ordinary 3D convolution, represents the feature concatenation with the connection, represents the upsampling operation of 3D transposed convolution, Represents fusion features.
7. The hint-based interactive brain tumor segmentation method according to claim 1, characterized in that: The first prototype features of the tumor and non-tumor regions are specifically: ; in, represents the first prototype feature, represents the average mask pooling calculation, Indicates prompt information. Represents the features of the labeled image.
8. The hint-based interactive brain tumor segmentation method according to claim 1, characterized in that: The S10 specifically includes: S1001: Calculate the similarity between the three-dimensional features of the brain image and the first prototype features by the following formula: ; Where M represents the similarity between the brain image features and the first prototype features, represents the activation function, represents the cosine similarity calculation, represents the first prototype feature, Represents the three-dimensional features of brain images; S1002: Based on the similarity, feature fusion is performed by an average mask pooling calculation method to construct a second prototype feature of the brain image: ; in, represents the second prototype feature, represents the average mask pooling calculation, M represents the similarity between the brain image features and the first prototype features, Represents brain image features.
9. The hint-based interactive brain tumor segmentation method according to claim 1, characterized in that: The final segmentation results of the tumor and non-tumor regions are specifically: ; in, Represents the final segmentation results of tumor and non-tumor areas, represents the activation function, represents the cosine similarity calculation, represents the concatenation operation along the channel dimension, represents the first prototype feature, represents the second prototype feature, Represents brain image features.
10. A prompt-based interactive brain tumor segmentation system, characterized in that include: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the prompt-based interactive brain tumor segmentation method according to any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
Pancreatic tumor image segmentation method and system based on reinforcement learning and attention
CN114663431A
Small sample semantic segmentation method based on information interaction enhancement
CN117726809A
SAM-based cross-modal domain generalization medical image segmentation method
CN117808834A
Small sample medical image segmentation method based on multi-level feature guidance
CN118644674A
Pancreatic tumor intelligent analysis method and system based on multi-modal medical image fusion
CN119170255A