Medical image segmentation method, system and equipment based on twin network and multi-view fusion, and medium
Through the medical image segmentation method based on twin network and multi-view fusion, the problems of undersegment and oversegment in cardiac magnetic resonance image segmentation are solved, and efficient and accurate segmentation effect is achieved, the robustness and applicability of the model are enhanced, and the clinical application value is important.
Patent Information
- Application Number
- CN202510108920.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art is difficult to accurately segment the left ventricle, right ventricle and myocardium in cardiac magnetic resonance images, and there are problems of undersegment and oversegment, and the traditional methods are complex and time-consuming, making it difficult to achieve fully automatic segmentation.
Using a medical image segmentation method based on twin network and multi-view fusion, rich feature information is extracted through multi-view fusion of original slice image encoder, region constraint slice encoder, high-frequency information encoder and low-frequency information encoder, and region constraint is used to reduce interference from background noise.
It improves the accuracy and efficiency of cardiac magnetic resonance image segmentation, reduces unnecessary calculations, enhances the robustness and applicability of the model, and is especially suitable for medical image segmentation tasks, and has important clinical application value.
Smart Images

Figure CN120047474A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of medical image segmentation, and particularly relates to a medical image segmentation method, system, device and medium based on a siamese network and multi-view fusion. Background Art
[0002] Cardiovascular Magnetic Resonance (CMR) images are a method for diagnosing cardiovascular diseases. It is an imaging technique that reconstructs the internal structure image of an object based on the intensity and distribution of signals through nuclear magnetic resonance imaging, and has good soft tissue contrast and safety. Dividing the heart into anatomically significant parts and marking the positions of the left ventricle, right ventricle, and myocardium in the image stacks of the short-axis and long-axis sections of the heart can non-invasively qualitatively and quantitatively evaluate the structure and function of the heart. However, the target segmentation region in CMR images has an unbalanced distribution with the surrounding background pixels, and there are motion artifacts and noises caused by cardiac dynamics, resulting in unclear boundaries between the foreground and background regions after segmentation. There are fuzzy regions in the boundaries and edge information between the left ventricle, myocardium, and right ventricle, and there are tissues with intensity values similar to those of the myocardium, which will lead to under-segmentation and over-segmentation at the image boundaries. Traditional image segmentation methods such as region growing, level set, and clustering are too complex, time-consuming, and laborious, and the structure of the heart itself is complex with variable shapes and sizes, making it a very challenging task to fully automatically segment CMR images. Therefore, there is an urgent need for a CMR image segmentation method that can accurately solve the above problems. Summary of the Invention
[0003] The purpose of the present invention is to provide a medical image segmentation method, system, device and medium based on a siamese network and multi-view fusion to solve the problems existing in the above-mentioned prior art.
[0004] To achieve the above purpose, the present invention provides a medical image segmentation method based on a siamese network and multi-view fusion, including:
[0005] Obtaining sequence slice data of cardiovascular magnetic resonance images;
[0006] Inputting the sequence slice data into a medical image segmentation model for image segmentation to obtain segmentation results corresponding to each slice data; wherein, the segmentation results include segmentation masks of the left ventricle, right ventricle, and myocardium; the medical image segmentation model is constructed based on an encoder-decoder network, the encoder includes a plurality of encoder branches with the same structure and non-shared weights, the encoder branches include an original slice image encoding branch, a region-constrained slice encoder branch, a high-frequency information encoding branch, and a low-frequency information encoding branch, and the decoder includes a plurality of decoding blocks with bottleneck structures.
[0007] Optionally, the training process of the medical image segmentation model specifically includes:
[0008] Obtain training data, where the training data includes training slice data and corresponding segmentation results;
[0009] Construct an initial medical image segmentation model, input the training data into the initial medical image segmentation model for image segmentation, and take the minimum loss between the initial training result after image segmentation and the segmentation result corresponding to the training slice data as the goal for training to obtain the trained medical image segmentation model.
[0010] Optionally, the processing process of the medical image segmentation model specifically includes:
[0011] If the current slice data is the first image of the sequence slice data, input the current slice data into the original slice image encoding branch, high-frequency information encoding branch, and low-frequency information encoding branch for feature extraction to obtain original features, high-frequency features, and low-frequency features; perform weighted summation on the same-level output features output by each branch to obtain the fusion features at each level, input the fusion features at each level into the decoder for upsampling and feature extraction, and use a linear layer to restore the output of the decoder to the original image scale to output the segmentation result of the current slice data;
[0012] If the current slice data is not the first image of the sequence slice data, input the current slice data into the original slice image encoding branch, high-frequency information encoding branch, and low-frequency information encoding branch for feature extraction to obtain original features, high-frequency features, and low-frequency features; input the current slice data and the previous slice data into the region-constrained slice encoder branch for feature extraction to obtain region-constrained features; perform weighted summation on the same-level output features output by each branch to obtain the fusion features at each level, input the fusion features at each level into the decoder for upsampling and feature extraction, and use a linear layer to restore the output of the decoder to the original image scale to output the segmentation result of the current slice data.
[0013] Optionally, the processing process of inputting the current slice data into the original slice image encoding branch specifically includes:
[0014] Input the current slice data into the original slice image encoding branch, and perform feature extraction on the original slice image through a convolutional neural network to obtain original features.
[0015] Optionally, the processing process of inputting the current slice data into the high-frequency information encoding branch specifically includes:
[0016] Perform wavelet decomposition on the current slice data to obtain several components including low-frequency information and high-frequency information;
[0017] Superimpose the components containing high-frequency information to obtain a component containing all high-frequency signals, and restore the superimposed component to the original scale through the bilinear interpolation algorithm to obtain a high-frequency information image;
[0018] Extract features from the high-frequency information image through a convolutional neural network to obtain high-frequency features.
[0019] Optionally, the process of inputting the current slice data into the low-frequency information encoding branch specifically includes:
[0020] Perform wavelet decomposition on the current slice data to obtain several components containing low-frequency information and high-frequency information;
[0021] Extract the component containing only low-frequency information, convert the extracted component to the spatio-temporal domain based on the inverse wavelet transform, and restore the component converted to the spatio-temporal domain to the original scale through the bilinear interpolation algorithm to obtain a low-frequency information image;
[0022] Extract features from the low-frequency information image through a convolutional neural network to obtain low-frequency features.
[0023] Optionally, the process of inputting the current slice data and the previous slice data into the region-constrained slice encoder branch for feature extraction specifically includes:
[0024] Adjust the current slice data to a preset size based on the bilinear interpolation algorithm;
[0025] Input the previous slice data and the current slice data adjusted to the preset size into the previous slice image branch and the current slice branch respectively for feature extraction, use the output of the previous slice branch as a convolutional kernel, and perform a convolutional operation on the output of the current slice branch to obtain a cross-correlation operation result;
[0026] Output the result after the cross-correlation operation to a convolutional layer for feature extraction to obtain a class positive and negative confidence output branch and a position and size offset output branch;
[0027] Obtain the highest-confidence anchor box based on the class positive and negative confidence output branch, extract the offset corresponding to the highest-confidence anchor box based on the position and size offset output branch to obtain a target position rectangle;
[0028] Crop the current slice data based on the target position rectangle, and fill the part outside the rectangle with 0 to obtain a region-constrained slice image;
[0029] Extract features from the region-constrained slice image through a convolutional neural network to obtain region-constrained features.
[0030] A medical image segmentation system based on a siamese network and multi-view fusion includes:
[0031] A data acquisition module for acquiring sequence slice data of cardiac magnetic resonance images;
[0032] An image segmentation module for inputting the sequence slice data into a medical image segmentation model for image segmentation to obtain segmentation results corresponding to each slice of data; wherein, the segmentation results include segmentation masks of the left ventricle, right ventricle, and myocardium; the medical image segmentation model is constructed based on an encoder-decoder network, the encoder includes a plurality of encoder branches with the same structure and non-shared weights, the encoder branches include an original slice image encoding branch, a region-constrained slice encoder branch, a high-frequency information encoding branch, and a low-frequency information encoding branch, and the decoder includes a plurality of decoding blocks with bottleneck structures.
[0033] An electronic device includes a memory and a processor, the memory is used for storing a computer program, and the processor runs the computer program to enable the electronic device to execute the described medical image segmentation method based on a siamese network and multi-view fusion.
[0034] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the described medical image segmentation method based on a siamese network and multi-view fusion.
[0035] The technical effects of the present invention are as follows:
[0036] Through the multi-view fusion of the original slice image encoder, region-constrained slice encoder, high-frequency information encoder, and low-frequency information encoder, the present invention can extract rich feature information from different angles, thereby improving the accuracy of segmentation. The present invention uses the target position rectangular box predicted by the siamese network to perform region constraint on the current slice image, limits the segmentation range, reduces the interference of background noise, and improves the segmentation accuracy of the target region. At the same time, using the segmentation result of the previous slice and the image information of the current slice, the target position rectangular box is quickly predicted through the siamese network, reducing unnecessary calculations and improving the segmentation efficiency. At the same time, the present application extracts high-frequency and low-frequency information through wavelet transform. The high-frequency information contains edge information, which helps to accurately segment the boundary of the target; the low-frequency information contains stable structural information, which helps to handle the inconsistency of the image data distribution.
[0037] Through technologies such as multi-view fusion, region constraint, high-frequency and low-frequency information extraction, and siamese network, the present invention realizes the efficient and accurate segmentation of cardiac magnetic resonance images. The combination of these technologies not only improves the accuracy and efficiency of segmentation, but also enhances the robustness and applicability of the model, is particularly suitable for medical image segmentation tasks, and has important clinical application value. Description of the Drawings
[0038] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0039] The drawings forming a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:
[0040] Figure 1 It is a schematic diagram of the medical image segmentation model structure in the embodiments of the present invention;
[0041] Figure 2 It is a schematic diagram of the visual comparison of the segmentation results of the proposed RCSiamCANet and other models on the ACDC dataset in the embodiments of the present invention;
[0042] Figure 3 It is a flowchart of the image segmentation implementation in the embodiments of the present invention. Detailed implementation manners
[0043] Now, various exemplary implementation manners of the present invention will be described in detail. This detailed description should not be considered as a limitation to the present invention, but rather as a more detailed description of certain aspects, characteristics, and implementation schemes of the present invention.
[0044] It should be understood that the terms described in the present invention are only used to describe specific implementation manners and are not used to limit the present invention. Additionally, for the numerical ranges in the present invention, it should be understood that each intermediate value between the upper and lower limits of the range is also specifically disclosed. Each intermediate value within any stated value or stated range, as well as each smaller range between any other stated value or intermediate value within the stated range, is also included in the present invention. The upper and lower limits of these smaller ranges can be independently included or excluded from the range.
[0045] Without departing from the scope or spirit of the present invention, various improvements and changes can be made to the specific implementation manners of the description of the present invention, which are obvious to those skilled in the art. Other implementation manners obtained from the description of the present invention are obvious to those skilled in the art. The description and embodiments of this application are only exemplary.
[0046] Regarding the terms "comprising", "including", "having", "containing", etc. used herein, they are all open-ended terms, meaning including but not limited to.
[0047] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0048] As Figure 1 - Figure 3 shown, in this embodiment, a medical image segmentation method based on a Siamese network and multi-view fusion is provided, including: obtaining sequence slice data of cardiac magnetic resonance images; inputting the sequence slice data into a medical image segmentation model for image segmentation to obtain segmentation results corresponding to each slice data; wherein, the segmentation results include segmentation masks of the left ventricle, right ventricle, and myocardium; the medical image segmentation model is constructed based on an encoder-decoder network, the encoder includes a plurality of encoder branches with the same structure and non-shared weights, the encoder branches include an original slice image encoding branch, a region-constrained slice encoder branch, a high-frequency information encoding branch, and a low-frequency information encoding branch, and the decoder includes a plurality of decoding blocks with a bottleneck structure.
[0049] Through the multi-view fusion of the original slice image encoder, region-constrained slice encoder, high-frequency information encoder, and low-frequency information encoder in this embodiment, rich feature information can be extracted from different angles, thereby improving the accuracy of segmentation. In this embodiment, the target position rectangular box predicted by the Siamese network is used to perform region constraint on the current slice image, restricting the segmentation range, reducing the interference of background noise, and improving the segmentation accuracy of the target region. At the same time, using the segmentation result of the previous slice and the image information of the current slice, the target position rectangular box is quickly predicted through the Siamese network, reducing unnecessary calculations and improving the segmentation efficiency. At the same time, in the present application, high-frequency and low-frequency information are extracted through wavelet transform. The high-frequency information contains edge information, which helps to accurately segment the boundary of the target; the low-frequency information contains stable structural information, which helps to handle the inconsistency of the image data distribution.
[0050] Through the combination of the multi-view fusion encoder and the region-constrained slice encoder in this embodiment, as well as the use of technologies such as wavelet transform and Siamese network, the efficient segmentation of cardiac magnetic resonance images is realized, which can automatically extract the features in the images, has stronger anti-interference ability against noise, intensity inhomogeneity, shape changes, etc., and is more suitable for processing complex cardiac MR images.
[0051] The network of this embodiment adopts an encoder-decoder architecture. The encoder part is a multi-view fusion coding module composed of an original slice image encoder branch, a region-constrained slice encoder branch, a high-frequency information encoder branch, and a low-frequency information encoder branch.
[0052] The original slice image encoder branch uses a CNN to extract features from the original slice image, obtaining the most complete feature information. The region-constrained slice encoder branch extracts features from the target segmentation rectangular region predicted by the Siamese network, and the regions of non-interest will not be feature-extracted, thereby establishing a segmentation region constraint.
[0053] The high-frequency information encoder branch will learn the high-frequency edge information of the heart image obtained by wavelet transform. Similarly, the low-frequency information encoder branch will learn the low-frequency information with stable signals in the heart image.
[0054] The decoder part upsamples and extracts features from the image through 4 decoder blocks, and finally restores it to the original image scale through a linear layer. The network structure is as Figure 1 shown.
[0055] The encoder consists of 4 branches with the same structure but non-shared weights. Each encoder branch contains 4 encoding blocks, and each encoding block is a bottleneck structure composed of 3 convolutional layers and a Relu activation function stacked. The convolutional kernel sizes are 1x1, 3x3, and 1x1 respectively. The input size of the magnetic resonance image is 1x224x224. The outputs of each layer of the 4 encoder branches are 128x112x112, 256x56x56, 512x28x28, and 1024x14x14 respectively. Finally, the outputs of the same level of the 4 encoder branches are summed according to the weights to obtain 3 total outputs. The formula is as follows:
[0056] e i =e si ·w s +e hi ·w h +e li ·w l +e pi ·w p
[0057] Among them, e i (i = 1, 2, 3, 4) is the total output, e si is the output of the original slice image encoder branch, e hi is the output of the high-frequency information encoder branch, e li is the output of the low-frequency information encoder branch, e pi is the output of the region-constrained slice encoder branch, w s 、w h 、w l and w p are all learnable parameters. It should be noted that if the slice is the first slice and there is no previous slice, it only contains the image encoder branch, the high-frequency information encoder branch, and the low-frequency information encoder branch.
[0058] In the decoder stage, there is also a decoding block containing 4 bottleneck structures. After the output of each decoding block, the original image scale is gradually restored through the bilinear interpolation algorithm. By concatenating each output with the 3 outputs of the encoder, the k-class segmentation result output of size kx224x224 is restored from 256x28x28, 128x56x56, and 64x112x112.
[0059] Processing of high-frequency and low-frequency information: The original image is decomposed by horizontal and vertical filtering along the x-axis and y-axis directions into 4 vectors, namely LL (low-frequency signals in both directions), HL (high-frequency signal in the horizontal direction and low-frequency signal in the vertical direction), LH (high-frequency signal in the vertical direction and low-frequency signal in the horizontal direction), and HH (high-frequency signals in both directions). For cardiac magnetic resonance images, segmentation of the left ventricle, right ventricle, and myocardium is required. This embodiment notes that the edges of these three targets are usually very clear. Therefore, this embodiment can use edge information to guide feature extraction and result segmentation, where the high-frequency information in the frequency domain contains edge information. By superimposing HL, LH, and HH, all components containing high-frequency signals are obtained. At the same time, due to differences between devices in the acquisition of cardiac magnetic resonance images, the distribution of image data is not completely consistent, while the low-frequency information in the frequency domain is relatively stable. Therefore, this embodiment extracts the component containing only the low-frequency information LL. This embodiment uses the inverse wavelet transform to reverse the signal components from the frequency domain back to the spatio-temporal domain. At this time, due to wavelet decomposition, the scale is reduced to half of the original image size, that is, 1x112x112, and then the two components are restored to the original image size scale, that is, 1x224x224, and at this time, this embodiment obtains images of high-frequency and low-frequency information.
[0060] Region-constrained slice coding branch: This embodiment introduces a siamese network. This embodiment observes that there are regular changes in the cardiac structure and tissues between adjacent slices. By capturing the regularities of the feature changes between adjacent slices to capture this spatial relationship, the network can better learn the cardiac structure and tissues and faster fit the data distribution, thereby finally improving the segmentation accuracy. This embodiment uses the segmented result of the previous slice and inputs it into the siamese network together with the current slice. This embodiment notes that the targets for cardiac magnetic resonance image segmentation, the left ventricle, right ventricle, and myocardium, are usually distributed together, and there will be no scattered small segmented blocks around the final segmentation mask. Therefore, this embodiment can obtain the rectangular frame of the target position of the segmentation result of the current slice, thus obtaining a region-constrained prediction of the current target position based on the spatial relationship. Because in cardiac magnetic resonance images, due to factors such as the long acquisition time during data acquisition and the movement of patients and equipment, the boundaries of target segmentation and regions of no interest will be blurred and the data distribution will be unclear. Through the region-constrained prediction, the segmentation range can be restricted and the segmentation accuracy can be improved.
[0061] The input of the region-constrained slice encoder is the segmentation result of the target position rectangular box obtained by tracking the current slice image and the previous slice image through a Siamese network. Specifically, the size of the previous slice image is 1x127x127, and the size of the current slice image is bilinearly interpolated to 1x255x255 and then respectively input into a 50-layer ResNet network, with the output sizes being 256x6x6 and 256x22x22 respectively. At this time, the four feature maps of the two branches are cross-input into four 50-layer ResNet networks in pairs to obtain four outputs. Cross-correlation operation is performed on one output of the previous slice image branch and one output of the current slice image branch, and the same is done for the other two outputs. The cross-correlation operation uses the output of the previous slice image branch as the convolutional kernel to perform slicing processing on the output of the current slice image branch. The formula is as follows:
[0062]
[0063] The result after cross-correlation is output to the convolutional layer for further feature extraction, obtaining two outputs, namely the positive and negative confidence output branch of 17x17x2k for categories and the position size offset output branch of 17x17x4k. Calculate the bounding box of the target position by using the anchor box with the highest confidence and the position size offset, and crop the target position rectangular box of the current slice image, filling the part outside the rectangular box with 0 to obtain the input of the region-constrained slice encoder.
[0064] As shown in Table 1, the ACDC dataset is clinical data obtained from the University Hospital of Dijon, France. A series of short-axis slices of the heart are obtained using two Siemens MRI scanners with different magnetic intensities. The slice thickness is 5mm - 10mm (usually 5mm), and the spatial resolution is from 1.34 - 1.68 mm^2 / pixel.
[0065] Table 1 ACDC Dataset
[0066]
[0067]
[0068] Table 2 Performance of Each Model on Four Key Quantitative Indicators
[0069] Dice mJC mASD mHD Average Myo RV LV FCN 87.65 86.57 83.51 92.87 0.78 0.59 3.06 UNet 88.31 85.97 85.03 93.94 0.79 0.56 2.84 TransUNet 89.56 86.42 87.54 94.71 0.81 0.54 2.79 SwinUNet 87.56 82.97 85.69 94.03 0.79 0.42 1.92 EMCAD 91.28 88.81 89.26 95.77 0.84 0.28 1.28 RCSiamCANet 92.35 91.44 86.82 93.15 0.86 0.25 1.20
[0070] Through comparative analysis, this embodiment shows the performance of these models on four key quantitative indicators in Table 2, especially the specific values of the average Dice indicators for the left ventricle, myocardium, and right ventricle. At the same time Figure 2Shows the visual comparison of the segmentation results of the proposed RCSiamCANet and other models on the ACDC dataset. Among them, each row is a case slice in the test set, and each column is the original image, the ground truth label, and the visual results of different models. Green represents the right ventricle, blue represents the myocardium, and red represents the left ventricle. From the visual results of this embodiment, it can be seen that RCSiamCANet performs very well both in terms of the segmentation edges and the target segmentation regions. EMCAD, UNet, and FCN have problems of under-segmentation and segmentation redundancy in the segmentation results. The segmentation results of SwinUNet show that the segmentation edges are not smooth, and TransUNet also has the problem of non-smooth segmentation edges with sporadic redundant segmentation small pieces around;
[0071] Comparison of segmentation effects on the ACDC dataset: RCSiamCANet achieved the best performance in terms of average Dice, average JC, average ASD, and average HD. Specifically, RCSiamCANet reached 92.35% in the mDice coefficient, which was improved by 4.7%, 4.04%, 2.79%, 4.79, and 1.12% compared with FCN, UNet, TransUNet, SwinUNet, and EMCAD respectively, and was particularly prominent in the segmentation of the myocardium.
[0072] This solution realizes the efficient and accurate segmentation of cardiac magnetic resonance images through technologies such as multi-view fusion, region constraint, high-frequency and low-frequency information extraction, and siamese network. The combination of these technologies not only improves the accuracy and efficiency of segmentation, but also enhances the robustness and applicability of the model, and is particularly suitable for medical image segmentation tasks, with important clinical application value.
[0073] Implementable, this embodiment also provides a medical image segmentation system based on a siamese network and multi-view fusion, including:
[0074] A data acquisition module for acquiring sequence slice data of cardiac magnetic resonance images;
[0075] An image segmentation module for inputting the sequence slice data into a medical image segmentation model for image segmentation to obtain the segmentation results corresponding to each slice data; wherein, the segmentation results include segmentation masks of the left ventricle, right ventricle, and myocardium; the medical image segmentation model is constructed based on an encoder-decoder network, the encoder includes a plurality of encoder branches with the same structure and non-shared weights, the encoder branches include an original slice image encoding branch, a region constraint slice encoder branch, a high-frequency information encoding branch, and a low-frequency information encoding branch, and the decoder includes a plurality of decoding blocks with bottleneck structures.
[0076] Implementable, this embodiment also provides an electronic device, including a memory and a processor. The memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the described medical image segmentation method based on a twin network and multi-view fusion.
[0077] Implementable, this embodiment also provides a computer-readable storage medium that stores a computer program. When the computer program is executed by a processor, it implements the described medical image segmentation method based on a twin network and multi-view fusion.
[0078] As described above, only the preferred specific implementation manners of this application are provided, but the protection scope of this application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in this application should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. A medical image segmentation method based on Siamese network and multi-view fusion, characterized in that: include: Acquiring serial slice data of cardiac magnetic resonance images; The serial slice data is input into a medical image segmentation model for image segmentation to obtain a segmentation result corresponding to each slice data; wherein the segmentation result includes a segmentation mask of the left ventricle, the right ventricle and the myocardium; the medical image segmentation model is constructed based on an encoder-decoder network, the encoder includes a plurality of encoder branches with the same structure and non-shared weights, the encoder branches include an original slice image encoding branch, a regional constraint slice encoder branch, a high-frequency information encoding branch and a low-frequency information encoding branch, and the decoder includes a plurality of decoding blocks with bottleneck structures.
2. According to claim 1, a medical image segmentation method based on Siamese network and multi-view fusion is characterized in that: The training process of the medical image segmentation model specifically includes: Acquire training data, wherein the training data includes training slice data and corresponding segmentation results; An initial medical image segmentation model is constructed, the training data is input into the initial medical image segmentation model for image segmentation, and training is performed with the goal of minimizing the loss between the initial training result after image segmentation and the segmentation result corresponding to the training slice data to obtain a trained medical image segmentation model.
3. According to claim 1, a medical image segmentation method based on Siamese network and multi-view fusion is characterized in that: The processing process of the medical image segmentation model specifically includes: If the current slice data is the first image of the sequence slice data, the current slice data is input into the original slice image encoding branch, the high-frequency information encoding branch and the low-frequency information encoding branch for feature extraction to obtain original features, high-frequency features and low-frequency features; the output features of the same level output by each branch are weighted summed to obtain fused features of each level, and the fused features of each level are input into the decoder for upsampling and feature extraction, and the output of the decoder is restored to the original image scale by using a linear layer, and the segmentation result of the current slice data is output; If the current slice data is not the first image of the sequence slice data, the current slice data is input into the original slice image encoding branch, the high-frequency information encoding branch and the low-frequency information encoding branch for feature extraction to obtain the original features, high-frequency features and low-frequency features; the current slice data and the previous slice data are input into the regional constraint slice encoder branch for feature extraction to obtain the regional constraint features; the output features of the same level output by each branch are weighted summed to obtain the fusion features of each level, the fusion features of each level are input into the decoder for upsampling and feature extraction, and the output of the decoder is restored to the original image scale using a linear layer to output the segmentation result of the current slice data.
4. The medical image segmentation method based on Siamese network and multi-view fusion according to claim 3, characterized in that: The process of inputting the current slice data into the original slice image encoding branch specifically includes: The current slice data is input into the original slice image encoding branch, and the original slice image is subjected to feature extraction through a convolutional neural network to obtain the original features.
5. The medical image segmentation method based on Siamese network and multi-view fusion according to claim 3, characterized in that: The process of inputting the current slice data into the high-frequency information encoding branch specifically includes: Perform wavelet decomposition on the current slice data to obtain several components containing low-frequency information and high-frequency information; The components containing high-frequency information are superimposed to obtain components containing all high-frequency signals, and the superimposed components are restored to the original scale by a bilinear interpolation algorithm to obtain a high-frequency information image; The high-frequency information image is extracted through convolutional neural network to obtain high-frequency features.
6. The medical image segmentation method based on Siamese network and multi-view fusion according to claim 3, characterized in that: The process of inputting the current slice data into the low-frequency information encoding branch specifically includes: Perform wavelet decomposition on the current slice data to obtain several components containing low-frequency information and high-frequency information; Extract the components containing only low-frequency information, convert the extracted components into the time-space domain based on the inverse wavelet transform, and restore the components converted into the time-space domain to the original scale through the bilinear interpolation algorithm to obtain a low-frequency information image; The low-frequency information image is extracted through convolutional neural network to obtain low-frequency features.
7. The medical image segmentation method based on Siamese network and multi-view fusion according to claim 3, characterized in that: The step of inputting the current slice data and the previous slice data into the region constraint slice encoder branch for feature extraction specifically includes: Adjust the current slice data to the preset size based on the bilinear interpolation algorithm; The previous slice data and the current slice data adjusted to a preset size are respectively input into the previous slice image branch and the current slice branch for feature extraction, and the output of the previous slice branch is used as the convolution kernel to perform a convolution operation on the output of the current slice branch to obtain a cross-correlation operation result; The result after the cross-correlation operation is output to the convolution layer to extract features, and the positive and negative confidence output branches of the category and the position size offset output branches are obtained; Based on the category positive and negative confidence output branch, the highest confidence anchor frame is obtained, and based on the position size offset output branch, the offset corresponding to the highest confidence anchor frame is extracted to obtain a target position rectangular frame; The current slice data is cropped based on the target position rectangular frame, and the part outside the rectangular frame is filled with 0 to obtain a region-constrained slice image; The feature extraction of the region constraint slice image is performed through a convolutional neural network to obtain the region constraint feature.
8. A medical image segmentation system based on Siamese network and multi-view fusion, characterized in that: include: A data acquisition module, used for acquiring serial slice data of cardiac magnetic resonance images; An image segmentation module is used to input the serial slice data into a medical image segmentation model for image segmentation, and obtain a segmentation result corresponding to each slice data; wherein the segmentation result includes a segmentation mask of the left ventricle, the right ventricle and the myocardium; the medical image segmentation model is constructed based on an encoder-decoder network, the encoder includes a plurality of encoder branches with the same structure and non-shared weights, the encoder branches include an original slice image encoding branch, a regional constraint slice encoder branch, a high-frequency information encoding branch and a low-frequency information encoding branch, and the decoder includes a plurality of decoding blocks with bottleneck structures.
9. An electronic device, characterized in that: It includes a memory and a processor, the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to perform a medical image segmentation method based on a twin network and multi-view fusion according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that: It stores a computer program, which, when executed by a processor, implements a medical image segmentation method based on a twin network and multi-view fusion as described in any one of claims 1 to 7.
Citation Information
Cited By
AR heart digital twinning method and system for individualized medical treatment
CN120977595A