SAM2-based brain tissue segmentation method and system
Through the SAM2-based brain tissue segmentation method, the preprocessing and segmentation technology of the SAM2 model are used to solve the difficulties in brain tissue segmentation of different modalities and species, and achieve efficient and accurate brain tissue extraction with cross-modal robustness and consistency between slices.
Patent Information
- Application Number
- CN202510930777.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2025-10-10
AI Technical Summary
Existing technologies lack efficient and generalizable methods for brain tissue segmentation, especially in extracting brains of different modalities and species, which requires a large amount of labeled data and complex parameter adjustments.
A brain tissue segmentation method based on SAM2 is adopted. Through preprocessing, SAM2 model segmentation and Gaussian blur smoothing, combined with cross-modal and cross-species optimization strategies, accurate brain tissue extraction is achieved using a small amount of dedicated training data.
It demonstrates efficient and accurate segmentation results in brain tissue segmentation across different modalities and species without the need for complex parameter adjustment or retraining, and is cross-modal robust and consistent across slices.
Smart Images

Figure CN120765672A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of biomedical technology, and particularly relates to a brain tissue segmentation method and system based on SAM2. BACKGROUND
[0002] Brain extraction is a core preprocessing step in neuroimaging data processing, which aims to accurately segment brain tissue from non-brain tissue (such as skull, dura mater, fat, etc.) in magnetic resonance imaging (MRI), providing a reliable basis for subsequent analysis (such as brain registration, tissue segmentation, cortical thickness measurement, etc.). Even for radiologists, it is a tedious task to separate the brain from the skull, and the accuracy of the results varies from person to person.
[0003] There are mainly two kinds of current mainstream analysis methods. One is based on traditional template / matching technology, such as the BET algorithm of FSL. This kind of method relies on pre-constructed brain atlas or template, and realizes brain tissue segmentation through image registration technology. Its advantage is that it can combine prior anatomical knowledge, but it also has significant defects, such as the registration process involves complex spatial transformation, time-consuming, difficult to meet the needs of large-scale data processing, and the need to manually adjust the registration parameters according to the image quality, poor robustness. The other method is a segmentation method based on deep learning, such as U-net-based segmentation. This method learns complex image features and performs high-precision in specific tasks, such as brain extraction of human brain MRI, but its limitations are also more significant. Such models are usually trained for specific anatomical structures or imaging modes, and are difficult to generalize to unseen image types (such as different species, scanning parameters), in addition, the model relies on a large amount of labeled data, and high-quality labeled data of medical images is scarce, resulting in a sharp drop in performance on new data sets.
[0004] In summary, there is currently a lack of an efficient and highly generalized brain extraction method for brain extraction. SUMMARY
[0005] The purpose of the present application is to provide a brain tissue segmentation method and system based on SAM2, which improves the segmentation efficiency while taking into account the generalization ability.
[0006] In order to achieve the above purpose, the present application provides the following technical solutions:
[0007] In a first aspect, the present application provides a brain tissue segmentation method based on SAM2, comprising the following steps:
[0008] Preprocessing the input brain tissue image;
[0009] Inputting the preprocessed brain tissue image into the SAM2 model for segmentation to obtain a mask image containing brain tissue;
[0010] Smooth the edge of the mask image using Gaussian blur.
[0011] When the brain tissue image is an MRI image, the preprocessing comprises:
[0012] Performing N4 bias field correction on the input three-dimensional MRI image;
[0013] Slicing the three-dimensional MRI image in the sagittal plane, the horizontal plane and the coronal plane respectively to obtain two-dimensional slice data of the three sections, and performing normalization processing on the medical image intensity;
[0014] Converting the normalized single-channel grayscale data into a three-channel RGB format through channel repetition.
[0015] When the brain tissue image is an optical section image, the preprocessing comprises:
[0016] Down-sampling the original section image to a resolution of 1024*1024;
[0017] Removing the artifacts of the down-sampled section image through a three-linear interpolation reconstruction method.
[0018] The SAM2 model is composed of an image encoder, a memory attention module, a mask decoder, a prompt encoder and a memory encoder; the processing of inputting the preprocessed brain tissue image into the SAM2 model for segmentation comprises:
[0019] Training a YOLOv12 model using the obtained two-dimensional slice data to detect the ROI region of interest;
[0020] Selecting a section in the middle position of the section data set, detecting the position of the brain tissue using the trained YOLOv12 model, inputting the detected position information into the prompt encoder of the SAM2 model as BOX prompt information, and then generating a mask image of the section through the mask decoder;
[0021] The image encoder uses the next section and cross-participates in the memory of the target object in the previous section, the mask decoder predicts the segmentation mask of the section, and finally the memory encoder converts the mask and the image embedding vector from the image encoder for predicting the next section; this step is executed in a loop until the last section;
[0022] Returning to the initial mask position, the image encoder uses the previous section and cross-participates in the memory of the target object in the next section, the mask decoder predicts the segmentation mask of the section, and finally the memory encoder converts the mask and the image embedding vector from the image encoder for predicting the previous section; this step is executed in a loop until the first section.
[0023] In a second aspect, the present application provides a brain tissue segmentation system based on SAM2, comprising:
[0024] a preprocessing module configured to preprocess an input brain tissue image;
[0025] a SAM2 model configured to segment the preprocessed brain tissue image to obtain a mask image containing brain tissue;
[0026] a smoothing module configured to smooth edges of the mask image by Gaussian blur.
[0027] In a third aspect, the present application provides a computer program product comprising computer readable instructions, wherein the computer readable instructions, when executed by a processor, implement the steps of the brain tissue segmentation method based on SAM2 of the present application.
[0028] In a fourth aspect, the present application provides a computer readable storage medium comprising computer readable instructions, wherein the computer readable instructions, when executed by a processor, implement the steps of the brain tissue segmentation method based on SAM2 of the present application.
[0029] In a fifth aspect, the present application provides an electronic device comprising: a memory storing program instructions; a processor connected to the memory and executing the program instructions in the memory to implement the steps of the brain tissue segmentation method based on SAM2 of the present application.
[0030] Compared with the prior art, the present application has the following technical advantages:
[0031] The SAM2 model is used to implement cross-modal and cross-species brain tissue segmentation tasks, and has the ability to accurately extract brain tissue with a small amount of special training data. This strong representation capability enables the SAM2 model to accurately perform the target segmentation task without complex parameter adjustment or retraining when facing different modalities (such as MRI, optical sections, and in situ hybridization images) or different species (such as humans, mice, and cynomolgus monkeys).
[0032] The SAM2 model is extended to different brain image modalities, and optimization strategies are proposed for different image spatial resolutions and artifact interference, such as interpolation reconstruction and continuity loss design, to further enhance its cross-modal robustness and slice consistency. Therefore, the brain tissue extraction effect can be stable and generalizable under different data types, species, and imaging conditions.
[0033] The model can adapt to data input under different conditions, and can perform image segmentation without or with only a small amount of image preprocessing steps, complete the brain extraction task more quickly, and achieve results comparable to various expert models such as nnUNetspecialist.
[0034] Other advantages of the present application are described in the following examples. BRIEF DESCRIPTION OF DRAWINGS
[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without any creative effort based on these drawings.
[0036] Figure 1 Flow chart of the brain tissue segmentation method based on SAM2 in embodiment 1.
[0037] Figure 2 Comparison effect diagram of MRI image before and after N4 bias field correction.
[0038] Figure 3 Structure and segmentation process diagram of SAM2 model.
[0039] Figure 4 Comparison diagram of cynomolgus monkey T1 coronal MRI brain extraction before and after extraction.
[0040] Figure 5 Flow chart of the brain tissue segmentation method based on SAM2 in embodiment 2.
[0041] Figure 6 Comparison diagram of mouse brain in situ hybridization before and after trilinear interpolation reconstruction.
[0042] Figure 7 Segmentation result diagram of monkey brain coronal optical section.
[0043] Figure 8 Composition block diagram of the brain tissue segmentation system based on SAM2.
[0044] Figure 9 Composition block diagram of an electronic device. DETAILED DESCRIPTION
[0045] In order to make the objects, technical solutions and advantages of the present application more clear, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.
[0046] Embodiment 1
[0047] The brain tissue segmentation method based on SAM2 provided in the embodiment is mainly aimed at the segmentation of multi-modal (sequence) MRI brain images (sequences including T1, T2, fair, and species including human, macaque and rabbit).
[0048] Referring to Figure 1 , the method comprises the following steps:
[0049] S10, pre-processing the input MRI brain image.
[0050] The original MRI image is a three-dimensional multi-modal MRI (such as T1, T2) sequence, and the original data is a single-channel grayscale matrix. After pre-processing, the format can be accepted by the SAM2 model. Not only does it meet the application requirements of the SAM2 model, but also the information is fully preserved, which can guarantee the segmentation accuracy.
[0051] Specifically, in this step, the pre-processing operation is as follows:
[0052] S101, correcting the input MRI image by N4 bias field, to correct the uneven bright places caused by device differences. As shown in Figure 2 , the left side of the figure is the image before correction, and the right side is the image after correction.
[0053] S102, first decompose the three-dimensional MRI data into two-dimensional slice data, and then normalize the medical image intensity, eliminate the gray scale distribution deviation caused by device differences through dynamic range compression. The formula is:
[0054]
[0055] Where I represents the pixel value of the original image, I norm represents the pixel value of the normalized image, μ 5%-95% represents the mean value of the image pixel value in the 5% to 95% quantile range, and σ 5%-95% represents the standard deviation of the image pixel value in the 5% to 95% quantile range.
[0056] In this step, when the three-dimensional MRI data is decomposed into two-dimensional slice data, the slices are made in the sagittal plane, the horizontal plane and the coronal plane respectively, to obtain two-dimensional slice data of the three sections.
[0057] S103, by means of channel repetition, the normalized single-channel grayscale data is converted into three-channel RGB format, so that morphological matching can be achieved without changing the original semantics.
[0058] S20, inputting the two-dimensional slice data obtained after pre-processing into the SAM2 model for segmentation to obtain the segmentation result.
[0059] As shown in Figure 1 , the two-dimensional slice data obtained from each section is input into the SAM2 model for segmentation, and finally the segmentation images of the three sections are aggregated.
[0060] Referring to Figure 3 The SAM2 model is composed of an image encoder, a memory attention module, a mask decoder, a cue encoder and a memory encoder. The memory encoder can store vector embeddings from the previous slice, fuse with information from the next slice to achieve cross-slice information transmission, and the attention mechanism can ensure the continuity of the anatomical structure.
[0061] The MRI image is preprocessed, and the three-dimensional image data generates a series of two-dimensional slice data sets to complete the segmentation in a streaming manner. Specifically, this step includes the following processing procedures:
[0062] S201, using the obtained two-dimensional slice data to train a YOLOv12 model to detect the ROI region of interest.
[0063] S202, select the slice at the middle position of the slice data set, use the trained YOLOv12 model to detect the position of the brain tissue, input the detected position information into the cue encoder of the SAM2 model as the BOX cue information, and then generate the mask image of the slice through the mask decoder.
[0064] Generally, in the traditional three-view, the image at the middle position occupies the largest object size, and since the segmentation uses the image information of the current slice plus the information of the last segmentation result, selecting the slice at the middle position can maximize the continuity of the slice information and ensure the segmentation effect.
[0065] It is easy to understand that the middle position here does not mean the absolute middle position, but the middle position that can be considered except the first and last positions.
[0066] S203, the image encoder uses the next slice and cross-participates in the memory of the target object in the previous slice, the mask decoder predicts the segmentation mask of the slice, and finally the memory encoder converts the mask and the image embedding vector from the image encoder into a vector for predicting the next slice. Repeat this action until the last slice.
[0067] S204, return to the initial mask position, the image encoder uses the previous slice and cross-participates in the memory of the target object in the next slice, the mask decoder predicts the segmentation mask of the slice, and finally the memory encoder converts the mask and the image embedding vector from the image encoder into a vector for predicting the next slice. Repeat this action until the first slice.
[0068] S30, add the original score probability tensors of the three sections to obtain the final prediction probability of each voxel. Assuming that the probability of predicting brain tissue at voxel (x, y, z) after fusion is P agg (x, y, z), then:
[0069] P agg (x,y,z)=w1P cor (x,y,z)+w2P sag (x,y,z)+w3P axi (x,y,z)
[0070] where w1, w2, w3 are the weights of each slice, which are dynamically assigned by different data sets, that is, a weight matrix is established in advance according to different species and sequences, and the corresponding weight is selected according to the species and imaging method of the input image of the individual. cor (x,y,z), P sag (x,y,z), P axi (x,y,z) are the probabilities of voxel (x, y, z) being predicted as brain tissue in the coronal, sagittal and horizontal planes, respectively, which are obtained by applying the softmax activation function to the original score logits(x, y, z) of each voxel by the SAM2 model, that is,
[0071] P(x,y,z)=softmax(logits(x,y,z))
[0072] S40, using Gaussian blur to smooth the mask image edge, to further improve the continuity and clarity of the segmentation edge.
[0073] Gaussian blur is a common edge smoothing technique that replaces each pixel in the image with a weighted average of its surrounding pixel values to achieve a blur effect. Its formula is:
[0074]
[0075] where I(x, y) is the pixel value of the blurred image at (x, y), I(x', y') is the pixel value of the original image at (x, y), and σ is the variance of the Gaussian distribution.
[0076] The mask image is converted from the probability image. The mask image is actually judged voxel by voxel, and the probability is greater than a certain threshold, which is set to 1, otherwise it is 0. The final effect is that the brain tissue area is white (pixel value is 1) and the non-brain tissue area is black (pixel value is 0).
[0077] S50, reconstruct the smoothed mask image into a three-dimensional segmentation mask, output the standard medical image format NIFTI, to be compatible with ITK-SNAP, 3D Slicer and other medical image processing software.
[0078] The output result after step S40 is a segmentation template (called mask). This template can completely overlap with the original image and clearly indicates the brain tissue and non-brain tissue areas. The user can perform actual brain extraction based on this template. By adjusting the dimensions and voxel spacing of the output mask image to be consistent with the original image, medical image processing software such as ITK-SNAP and 3DSlicer can also view the mask image.
[0079] As an example, the comparison of T1 coronal MRI brain of cynomolgus monkeys before and after extraction is shown in Figure 2. Figure 4 shown. Figure 4 The figure shows the typical segmentation effects of the front, middle and back slices. The top three from left to right are the original images of the front, middle and back slices, and the bottom three are the corresponding extraction result images.
[0080] Example 2
[0081] The SAM2-based brain tissue segmentation method provided in this embodiment is mainly used for the segmentation of brain tissue in optical sections, including block faces and mouse brain in situ hybridization sections. Such images are mostly optical images or microscopic images, generally only one slice image in an anatomical direction, with high resolution and large size.
[0082] See also Figure 5 , the method specifically comprises the following steps:
[0083] S1, preprocessing of optical section images.
[0084] In this step, specifically, you can perform the following operations:
[0085] Downsample the high-resolution original slice image to 1024*1024 resolution;
[0086] Remove artifacts from sliced images after downsampling.
[0087] For some images derived from optical sections but stored in medical image formats (such as NIfTI), missing or discontinuous slices during the stacking process often result in noticeable black streaks along the reconstructed axis, known as tomographic artifacts. These artifacts can severely interfere with subsequent image segmentation tasks. To address this, a trilinear interpolation reconstruction method was employed to compensate for the continuity of image voxels, effectively repairing discontinuous areas and improving the overall image structural integrity and segmentation quality.
[0088] Specifically, first, the original NIfTI image is parsed in three-dimensional direction to identify whether there is abnormal increase or discontinuity of the inter-voxel distance in the Z-axis (i.e. perpendicular to the slice plane), which usually corresponds to slice missing or stacking error. Specifically, the image is sliced along the z-axis, and the proportion of non-zero pixels is counted. When the proportion is too low (lower than a set threshold), it is considered that the slice is missing. Subsequently, a three-linear interpolation algorithm is used to interpolate and complete the missing voxels in the Z-axis direction while keeping the image structure in X and Y directions unchanged. This method uses the gray scale information in the adjacent three directions to construct a linear function in the local space, thereby smoothing and filling the broken area, making the interpolated image more continuous and consistent in space, and effectively removing the black stripe artifacts caused by uneven stacking.
[0089] As an example, the three-linear interpolation reconstructed mouse brain in situ hybridization result image is shown in Figure 6 The upper two images from left to right are the images before reconstruction of the slices in the horizontal plane and the sagittal plane, respectively. The mouse brain has black stripes due to the coronal section. The lower two images are the corresponding reconstructed result images.
[0090] S2, input the slice image after removing the artifacts into the SAM2 model for segmentation.
[0091] Specifically, this step includes the following processing procedures:
[0092] S201, use the obtained two-dimensional slice data to train a YOLOv12 model for detecting the ROI region of interest.
[0093] S202, select the slice in the middle position of the slice data set, use the trained yolo model to detect the position of the brain tissue, input the detected position information into the prompt encoder of the SAM2 model as the BOX prompt information, and then generate the mask image of the slice through the mask decoder.
[0094] S203, the image encoder uses the next slice and cross-participates in the memory of the target object in the previous slice, the mask decoder predicts the segmentation mask of the slice, and finally the memory encoder converts the mask and the image embedding vector from the image encoder for predicting the next slice. Repeat this action until the last slice.
[0095] S204, return to the initial mask position, the image encoder uses the previous slice and cross-participates in the memory of the target object in the next slice, the mask decoder predicts the segmentation mask of the slice, and finally the memory encoder converts the mask and the image embedding vector from the image encoder for predicting the previous slice. Repeat this action until the first slice.
[0096] If the target image is similar to the mouse brain in situ hybridization section, the section segmentation task is also performed on the reconstructed section sagittal plane, horizontal plane, and finally multi-view aggregation is performed. From the perspective of whether the method steps are consistent, the brain tissue extraction method of the mouse brain in situ hybridization section can also be classified in the foregoing embodiment 1.
[0097] S3, using Gaussian blur smoothing mask image edge, to further improve the continuity and clarity of the segmentation edge, and then output the segmentation mask image.
[0098] Gaussian blur is a common edge smoothing technique, which replaces each pixel in the image with a weighted average of its surrounding pixel values to achieve a blur effect. Its formula is:
[0099]
[0100] where I(x, y) is the pixel value of the blurred image at (x, y), I(x', y') is the pixel value of the original image at (x, y), and σ is the variance of the Gaussian distribution.
[0101] As an example, the segmentation result of the monkey brain coronal optical section is shown in Figure 7 . In the figure, the left is the original image before segmentation, and the right is the result image after segmentation.
[0102] The SAM2 model involved in the above two embodiments is the SAM2 (Segment Anything Model 2) base model proposed by Meta, but this model is trained on a large-scale natural image set, which lacks gray-scale, low-contrast, and unclear edge structure medical images, so the model is not sensitive to medical images, and cannot be directly used for brain tissue segmentation, and needs to be fine-tuned to optimize the weight parameters.
[0103] To this end, an image dataset for brain tissue segmentation task is constructed, which contains original image data and its corresponding segmentation mask. The image dataset is a cross-modality (optical / molecular / magnetic resonance imaging) and cross-species composite training set, covering macaque brain coronal section images, mouse brain in situ hybridization sections, and monkey brain three-dimensional T1 weighted MRI image sequences. The SAM2 model is trained on this image dataset, and each batch contains several consecutive sections, so that it can improve its learning ability in learning and make it have the ability to automatically and accurately identify and segment brain tissue under different modalities.
[0104] Specifically, the image encoder and prompt encoder are frozen during the training process, because the encoder is pre-trained based on a large-scale natural image, which already has excellent feature extraction capabilities. Only the mask decoder is fine-tuned to adapt to the new brain tissue segmentation task. The AdamW optimizer is selected, and the base learning rate and batch size are set according to the performance of the training device. Learning rate decay is performed using ReduceLROnPlateau.
[0105] In particular, in order to optimize the learning of the model on the slice structure and segmentation accuracy, a composite loss function is used, which combines four loss functions to simultaneously optimize mask accuracy and spatial continuity. Focal loss is an improved loss function based on cross-entropy, which uses a focal loss function to solve the problem of foreground and background classification imbalance; Dice loss is similar to generalized intersection over union, which is used to enhance the overlap of regions, and dice loss is used to further optimize the segmentation effect; Cross-Entropy (CE) loss provides pixel-level classification supervision; MAE loss is used to optimize the smoothness of the prediction of adjacent slices.
[0106] Unlike other natural image segmentation, the Focal loss is not given too much weight here, because the foreground / background of the medical image dataset is not so extreme, and over-emphasizing Focal will cause gradient instability or underfitting of the background; medical images place more emphasis on the coincidence of structural regions, especially the cerebellum region, so Dice loss is given a higher weight; the continuity between slices is very important, and a higher weight of MAE can avoid "jitter artifacts", so the design ratio of the four losses is 6:8:4:6. That is, the composite loss function is:
[0107] Loss = 6L focal + 8L Dice + 4L MAE + 6L CE
[0108] where Loss is the total loss, L focal is the Focal loss, L Dice is the Dice loss, L MAE is the MAE loss, and L CE is the CE loss.
[0109] The scheme is based on the SAM2 model to realize the cross-modal and cross-species brain tissue segmentation task, and a small amount of special data set is used for training to achieve good segmentation effect. This is different from and better than other models (other models need a large amount of data set for training). The realization of this ability depends on the strong generalization pre-training of the SAM2 model on large-scale and diversified image data sets. The image encoder and prompt encoder have learned universal visual features and prompt response mechanisms. This strong representation capability enables the SAM2 model to accurately perform the target segmentation task without complex parameter adjustment or retraining when facing different modalities (such as MRI, optical sections, and in situ hybridization images) or different species (such as humans, mice, and cynomolgus monkeys).
[0110] The SAM2 is innovatively extended to different brain image modalities, and optimization strategies for different image spatial resolutions and artifact interference are proposed, such as interpolation reconstruction and continuity loss design, which further enhance the cross-modal robustness and slice consistency. Therefore, it can maintain stable and generalizable brain tissue extraction effect under different data types, species and imaging conditions.
[0111] For reference Figure 8 The application also provides a brain tissue segmentation system based on SAM2, which comprises a preprocessing module, a SAM2 model and a smoothing module.
[0112] The preprocessing module is used for preprocessing the input brain tissue image. The preprocessing operation is different for different acquisition methods of brain tissue images, and specific descriptions can be referred to the related descriptions in embodiments 1 and 2.
[0113] The SAM2 model is used for segmenting the preprocessed brain tissue image to obtain a mask image containing brain tissue. The segmentation process and training process of the SAM2 model can be referred to the foregoing related content.
[0114] The smoothing module is used for smoothing the edges of the mask image by Gaussian blur to further improve the continuity and clarity of the segmentation edges.
[0115] The specific execution operations of each module can be referred to the related descriptions in the foregoing method steps, which will not be described here.
[0116] As shown in Figure 9 The present embodiment also provides an electronic device, which can include a processor 41 and a memory 42, wherein the memory 42 is coupled to the processor 41. It is worth noting that this figure is exemplary, and other types of structures can be used to supplement or replace this structure to realize data extraction, report generation, communication or other functions.
[0117] As shown in Figure 9As shown, the electronic device can further include an input unit 43, a display unit 44, and a power supply 45. It is worth noting that the electronic device does not necessarily include all the components shown in Figure 9 In addition, the electronic device can further include components not shown in Figure 9 In addition, the electronic device can further include components not shown in
[0118] The processor 41, which is also sometimes referred to as a controller or operating control, can include a microprocessor or other processor device and / or logic device, which receives input and controls the operation of the various components of the electronic device.
[0119] The memory 42, which can be one or more of a cache, a flash memory, a hard drive, a removable media, a volatile memory, a non-volatile memory, or other suitable device, for example, can store configuration information for the processor 41, instructions for execution by the processor 41, and the like. The processor 41 can execute programs stored in the memory 42 to enable information storage or processing, and the like. In one embodiment, the memory 42 further includes a buffer memory, i.e., a buffer, to store intermediate information.
[0120] The embodiments of the present application also provide a computer program product including computer readable instructions, which, when executed in an electronic device, cause the electronic device to perform the operation steps included in the method of the present application.
[0121] The embodiments of the present application also provide a storage medium having computer readable instructions stored therein, which cause an electronic device to perform the operation steps included in the method of the present application.
[0122] Those skilled in the art can appreciate that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be realized in electronic hardware, computer software, or a combination of both. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been described in general terms in the above description. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. A skilled person can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0123] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or say the part of the prior art that contributes to the present application, or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.
[0124] The above-described embodiments are merely specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications, replacements, and improvements within the technical scope disclosed by the present application, and these modifications, replacements, and improvements should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A brain tissue segmentation method based on SAM2, characterized in that: The following steps are involved: Preprocessing the input brain tissue image; The preprocessed brain tissue image is input into the SAM2 model for segmentation to obtain a mask image containing brain tissue; Gaussian blur is used to smooth the edges of the mask image.
2. The brain tissue segmentation method based on SAM2 according to claim 1, characterized in that: When the brain tissue image is an MRI image, the preprocessing includes: Perform N4 bias field correction on the input 3D MRI image; The three-dimensional MRI image is sliced in the sagittal, horizontal, and coronal planes to obtain two-dimensional slice data of the three sections, and the medical image intensity is normalized; The normalized single-channel grayscale data is converted into three-channel RGB format by channel duplication.
3. The brain tissue segmentation method based on SAM2 according to claim 2, characterized in that: Before using Gaussian blur to smooth the edge of the mask image, the method further includes the steps of: adding the original score probability tensors of the three slices, and obtaining the probability of each voxel as the final prediction; After using Gaussian blur to smooth the edge of the mask image, the method further includes the step of reconstructing the smoothed mask image into a three-dimensional segmentation mask image.
4. The brain tissue segmentation method based on SAM2 according to claim 1, characterized in that: When the brain tissue image is an optical section image, the preprocessing includes: Downsample the original slice image to 1024*1024 resolution; The artifacts of the downsampled slice images were removed by trilinear interpolation reconstruction method.
5. The brain tissue segmentation method based on SAM2 according to claim 1, characterized in that: The SAM2 model is composed of an image encoder, a memory attention module, a mask decoder, a prompt encoder, and a memory encoder; the process of inputting the preprocessed brain tissue image into the SAM2 model for segmentation includes: Use the obtained 2D slice data to train the YOLOv12 model to detect the ROI region; A slice in the middle of the slice dataset is selected, and the location of the brain tissue is detected using the trained YOLOv12 model. The detected location information is input into the prompt encoder of the SAM2 model as the box prompt information, and then the mask image of the slice is generated through the mask decoder; The image encoder uses the next slice and cross-references the memory of the target object in the previous slice. The mask decoder predicts the segmentation mask for the slice. Finally, the memory encoder converts the mask and the image embedding vector from the image encoder to predict the next slice. This step is repeated until the last slice is obtained. Return to the initial mask position, the image encoder uses the previous slice and cross-references the memory of the target object in the next slice, the mask decoder predicts the segmentation mask of the slice, and finally the memory encoder converts the mask and the image embedding vector from the image encoder to predict the previous slice; this step is repeated until the first slice.
6. The brain tissue segmentation method based on SAM2 according to claim 5, characterized in that: During the training process of the SAM2 model, the composite loss function used is Loss=6L focal +8L Dice +4L MAE +6L CE , Among them, Loss is the total loss, L focal is the Focal loss, L Dice is the Dice loss, L MAE is the MAE loss, L CE is the CE loss.
7. A brain tissue segmentation system based on SAM2, characterized in that: include: A preprocessing module, used for preprocessing the input brain tissue image; The SAM2 model is used to segment the preprocessed brain tissue image and obtain a mask image containing brain tissue; A smoothing module is used to smooth the edge of the mask image by Gaussian blurring.
8. A computer program product comprising computer-readable instructions, characterized in that: When the computer-readable instructions are executed by a processor, the steps of the brain tissue segmentation method based on SAM2 are implemented.
9. A computer-readable storage medium comprising computer-readable instructions, characterized in that: When the computer-readable instructions are executed by a processor, the steps of the brain tissue segmentation method based on SAM2 are implemented.
10. An electronic device, characterized in that: include: Memory, which stores program instructions; A processor is connected to the memory and executes program instructions in the memory to implement the steps of the brain tissue segmentation method based on SAM2 according to any one of claims 1 to 6.