Medical image sequence semi-supervised segmentation model construction method and application
By designing a semi-supervised segmentation model for medical image sequences and utilizing a spatiotemporal memory module and a dual cross-attention module, we solved the problem of insufficient utilization of spatiotemporal correlation features between frames in existing methods, and improved the segmentation accuracy of medical image sequences, especially the ability to quantify cardiac physiological indicators in cardiac magnetic resonance imaging.
Patent Information
- Application Number
- CN202511091803.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-05
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-08-05
AI Technical Summary
Existing semi-supervised learning methods cannot effectively utilize the spatiotemporal correlation features between frames when processing dynamic medical image sequences, resulting in a decrease in segmentation performance. In particular, it is difficult to accurately quantify cardiac physiological indicators in fields such as cardiac magnetic resonance imaging.
A semi-supervised segmentation model for medical image sequences was designed, which includes a first segmentation network, a feature projection module, a spatiotemporal memory module, and a dual cross-attention perception module. Through feature query-matching and feature fusion, the segmentation accuracy of the model for medical image sequences was improved.
Through the design of the spatiotemporal memory module and the dual cross-attention module, the segmentation accuracy of medical image sequences has been significantly improved. Especially under the condition of few labels, it can better mine the spatiotemporal correlation features between frames and improve the segmentation performance of the model.
Smart Images

Figure CN120598982A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of medical image processing, and more specifically, relates to a method for constructing a semi-supervised segmentation model for medical image sequences and its application. Background Art
[0002] Medical image sequences (such as magnetic resonance imaging and CT scans) contain rich temporal information, making accurate segmentation crucial. For example, cardiac magnetic resonance imaging (CMRI), with its high temporal and spatial resolution and lack of ionizing radiation, is the gold standard for assessing cardiac function. To automatically quantify cardiac physiological parameters (including left ventricular myocardial mass, cardiac volumes, and ejection fraction) using CMRI, accurate segmentation of anatomical structures in CMRI images is essential. Deep learning algorithms, when trained on large labeled datasets, achieve excellent segmentation results for medical image sequences. However, in real-world clinical scenarios, collecting large amounts of well-labeled data is time-consuming and labor-intensive due to various limitations. Furthermore, obtaining sufficiently large unlabeled datasets can be challenging due to healthcare privacy regulations.
[0003] In recent years, semi-supervised learning methods have demonstrated promising performance in addressing these challenges. These methods can effectively leverage unlabeled data, even with limited labeled data, thereby improving the model's generalization capabilities. These methods can generally be categorized into two categories: pseudo-label-based methods and consistency regularization-based methods. In pseudo-label-based models, predictions from labeled and unlabeled images are used as pseudo-labels and combined with true labels to retrain the model. Consistency regularization-based methods, on the other hand, encourage the model to maintain consistent predictions despite varying perturbations of the input data.
[0004] The above methods are all static segmentation methods, which have good segmentation effects for single-frame medical images, but the segmentation performance is insufficient when processing dynamic medical image sequences. This is mainly manifested in that the above methods are unable to handle the spatiotemporal connections between frames, and ignore the spatiotemporal correlation characteristics between frames, which will lead to a decrease in the segmentation performance of medical image sequences. For example, in movie cardiac MRI, the existing semi-supervised methods generally segment two specific frames of images. If the segmentation of each frame of movie cardiac MRI images can be achieved, segmentation samples can be provided for clinical measurement of dynamic physiological indicators such as filling fraction. Therefore, there is an urgent need for a semi-supervised segmentation method for medical image sequences to improve the segmentation performance of medical image sequences. Summary of the Invention
[0005] In response to the above-mentioned defects or improvement needs of the prior art, the present invention provides a method for constructing a semi-supervised segmentation model for medical image sequences and its application, the purpose of which is to improve the segmentation accuracy of medical image sequences.
[0006] To achieve the above-mentioned object, the present invention provides a method for constructing a semi-supervised segmentation model for medical image sequences, comprising: training a semi-supervised segmentation network for medical image sequences using a training sample set to obtain a trained semi-supervised segmentation model for medical image sequences; wherein the training sample is a medical image sequence, and the medical image sequence contains at least one frame of manually annotated image. GT ; The medical image sequence semi-supervised segmentation network includes: The first segmentation network is used to segment the medical image sequence input frame by frame to obtain the segmentation probability map of each frame of medical image Segmentation probability map and the corresponding segmentation labels ; Feature projection module, used to transform the segmentation probability map The image after splicing with the corresponding medical image is downsampled, and the downsampled features corresponding to each frame of medical image are used as memory keys ; The second segmentation network includes an encoder and a decoder, wherein the encoder is used to perform feature encoding on the medical image sequence input in the form of a sequence to obtain a feature ; The spatiotemporal memory module is used to store the memory key value and query value features Perform feature query-matching to generate semantically enhanced features of the medical image sequence ; Wherein, the query value feature The encoder output and the memory key value Intermediate features with consistent resolution; Channel and spatial attention modules are used to integrate the features After performing channel and space perception respectively, obtain time compression features and spatial compression features ;in, and The fused features With the characteristics After inputting into the decoder, the segmentation probability map is obtained and segmentation labels ; The training process includes: calculating the segmentation labels and The consistency loss and the segmentation probability map and The segmentation probability maps with labels are respectively associated with their corresponding manually labeled images supervised loss; adding the consistency loss and the supervised loss as the total loss to train the medical image sequence semi-supervised segmentation network to obtain a trained medical image sequence semi-supervised segmentation model.
[0007] Furthermore, the spatiotemporal memory module includes: Channel sensing unit, used to memorize the key value Compress the time feature in the time dimension to obtain the time feature value ; A spatial perception unit for detecting the query value characteristics Perform feature compression in the spatial dimension to obtain spatial feature values ; Similarity matrix calculation unit, used to calculate the memory key value and the query value feature Similarity matrix , ;in, represents the two-norm operator; Time and space feature reading unit, used to read features containing spatial information and read the special ;in, , ; Convolution unit, used to and The features after channel splicing are convolved to obtain the features .
[0008] Furthermore, the medical image sequence semi-supervised segmentation network also includes a dual cross-attention perception module; The features The way to obtain is: Will and After being input into the dual cross attention perception module for bidirectional feature fusion, the temporal fine-grained feature T@S and the spatial fine-grained feature S@T are generated respectively, and T@S and S@T are convolutionally fused to obtain the feature ; Among them, T@S and S@T are calculated as follows: ; ; in, Represents cross attention; query features for Features after the linear layer, time key value and spatial key values for Features after convolutional layers; query features for Features after the linear layer, time key value and spatial key values for Features after the convolutional layer respectively; Represents the dimension of the feature, Represents convolution.
[0009] Furthermore, the channel and spatial attention module includes a channel perception unit and a spatial perception unit; The channel perception unit is used to Perform time feature compression on the time dimension to generate the time compression feature ; The spatial perception unit is used to Perform feature compression in the spatial dimension to generate the spatial compression feature .
[0010] Furthermore, the time compression feature and the spatial compression feature The calculation method is: ; ; in, represents a multilayer perceptron, represents the sigmoid activation function, represents average pooling, stands for max pooling.
[0011] Furthermore, the segmentation probability map output by the first segmentation network is and the corresponding segmentation labels for: ; in, For the first segmentation network, is the input medical image sequence, are the trainable parameters of the first segmentation network; The segmentation probability map output by the second segmentation network and segmentation labels for: ; in, For the second segmentation network, are the trainable parameters of the second segmentation network.
[0012] The present invention also provides a semi-supervised segmentation method for medical image sequences, comprising: The unlabeled medical image or medical image sequence to be segmented is input into the second segmentation network of the medical image sequence semi-supervised segmentation model constructed by any of the above-mentioned methods for constructing a medical image sequence semi-supervised segmentation model to obtain a segmentation label.
[0013] The present invention also provides an electronic device comprising a computer-readable storage medium and a processor; The computer-readable storage medium is used to store executable instructions; The processor is used to read the executable instructions stored in the computer-readable storage medium to execute any one of the above-mentioned methods for constructing a semi-supervised segmentation model for a medical image sequence, or to execute the above-mentioned semi-supervised segmentation method for a medical image sequence.
[0014] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for constructing a semi-supervised segmentation model for a medical image sequence as described in any one of the above items, or implements the semi-supervised segmentation method for a medical image sequence as described above.
[0015] The present invention also provides a computer program product, comprising a computer program, which, when run on a computer, enables the computer to execute any of the above-mentioned methods for constructing a semi-supervised segmentation model for a medical image sequence, or to execute the above-mentioned semi-supervised segmentation method for a medical image sequence.
[0016] In general, the above technical solutions conceived by the present invention can achieve the following beneficial effects: (1) The present invention designs two sub-segmentation networks to learn the features of medical image sequences in different forms, which helps the network learn differentiated knowledge to update network parameters and avoids the problem of learning homogenized information using the same segmentation network, thereby improving the segmentation accuracy of the model. The segmentation probability map output by the first segmentation network and the corresponding pathological image are spliced at the channel and then downsampled to obtain the features as memory keys. , the encoder output of the second segmentation network and the memory key Features with consistent resolution are used as query value features , based on memory key value and query value features , using the constructed spatiotemporal memory module to perform feature query-matching, from the perspective of time and space Matching semantic information fully exploits the spatiotemporal correlation features between frames, providing semantically enhanced features for the decoder , which can alleviate the problem of network performance degradation under few-label training conditions and improve the segmentation accuracy of the model.
[0017] (2) As a preferred option, the designed spatiotemporal memory module is based on the memory key value and query value features Calculate the L2 similarity function (similarity matrix ), the similarity matrix It reflects the similarity between the features of the first segmentation network and the second segmentation network, based on the memory key value Time feature value after time feature compression in the time dimension , query value features Spatial eigenvalues after feature compression in spatial dimensions and similarity matrix , specifically reads temporal dimension features and spatial dimension features to provide semantically enhanced information for the decoder.
[0018] (3) Furthermore, through the designed dual cross-attention perception module, temporal fine-grained features and spatial fine-grained features are obtained and bidirectionally fused. The temporal and spatial features in the sequence images are interactively fused to obtain stronger feature expression capabilities, and the spatiotemporal connections between frames are explicitly mined, further improving the segmentation performance of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 Flowchart of the semi-supervised segmentation method for medical image sequences in an embodiment of the present invention.
[0020] Figure 2 Schematic diagram of the spatiotemporal memory module structure in an embodiment of the present invention.
[0021] Figure 3 Schematic diagram of the structure of the dual cross-attention perception module in an embodiment of the present invention.
[0022] Figure 4 This is a visualization result diagram of the segmentation probability map of the method in the embodiment of the present invention and other segmentation methods on the public dataset ACDC.
[0023] Figure 5 This is a visualization diagram of the prediction results of the method in the embodiment of the present invention and other segmentation methods on the public dataset CAMUS. DETAILED DESCRIPTION
[0024] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only intended to illustrate the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below may be combined with each other as long as they do not conflict with each other.
[0025] In the present invention, the terms "first", "second", etc. in the present invention and the accompanying drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.
[0026] Example 1 like Figure 1 As shown, the embodiment of the present invention provides a method for constructing a semi-supervised segmentation model for medical image sequences based on deep learning, which mainly includes: The training sample set is used to train the semi-supervised segmentation network of the medical image sequence to obtain a trained semi-supervised segmentation model of the medical image sequence; wherein the training sample in the training sample set is each case data, each case data is a time series image with several frames, at least one frame of image is an image with artificial label , and the rest are unlabeled images. The label is usually an organ mask. For example, for a cardiac MRI sequence, the label is a heart mask.
[0027] The semi-supervised segmentation network for medical image sequences includes: a first segmentation network, a second segmentation network, a feature projection module, a spatiotemporal memory module, a channel and spatial attention module.
[0028] The specific training process includes: S1. Send the time series images frame by frame into the first segmentation network to obtain the segmentation probability map of each frame of pathological image and the corresponding segmentation labels , i Indicates the first i Frame pathological image; segmentation probability map The corresponding i The pathological images are spliced at the channel and then sent to the feature projection module for downsampling. The downsampled features corresponding to each pathological image are used as memory keys. Among them, the first segmentation network is a spatial feature segmentation network, that is, it does not extract the temporal dimension information of the sample and only processes the spatial features of a single frame image.
[0029] S2, the time series image is sent to the encoder of the second segmentation network in the form of a sequence. The encoder includes multiple convolutional layers, which are connected to the memory key value. The features obtained by the convolutional layer with consistent resolution are used as query value features , query value feature and memorized key values Together they are sent to the direct spatiotemporal memory module for feature query-matching, from the perspective of time and space. Match semantic information to generate semantic enhancement features of the time series image ; Among them, the second segmentation network is a time-series-aware segmentation network.
[0030] S3, features output from the last layer of the encoder After inputting into the channel and spatial attention modules for channel and spatial perception respectively, the time compression features are obtained. and spatial compression features .
[0031] S4. Time compression features and spatial compression features The fused features and semantic enhancement features The images are sent to the decoder of the second segmentation network and upsampled to restore the original image resolution to generate the segmentation probability map of the time series image. and segmentation labels .
[0032] S5. Segmentation Label and segmentation labels Calculate consistency loss and segmentation probability map The labeled part and the corresponding manual label calculate the supervised loss, segmentation probability map The labeled part and the corresponding manual label are calculated with supervised loss; among them, the segmentation probability map Segmentation probability map of each frame of pathological image Composition, segmentation tags The segmentation labels of each frame of pathological images composition.
[0033] S6. Add the consistency loss and the supervised loss as the total loss function, perform loss training on the sequence semi-supervised segmentation network, and update the network parameters. After the loss converges or reaches the preset training rounds, a trained medical image sequence semi-supervised segmentation model is obtained.
[0034] As a further design of the present invention, the medical image sequence semi-supervised segmentation network also includes a dual cross-attention perception module.
[0035] Time compression feature and spatial compression features The fused features It is the feature output after the dual cross attention perception module.
[0036] The dual cross-attention perception module is used to compress the time features and spatial compression features Perform bidirectional feature fusion to generate temporal fine-grained features T@S and spatial fine-grained features S@T respectively, and fuse the two for convolutional feature fusion to obtain the fused dual cross attention features .
[0037] Preferably, in S1 or S2, before each frame of the pathological image or time series image is input into the first segmentation network or the second segmentation network, the image is also preprocessed, and the preprocessing includes cropping to the same size, and also includes performing the same scaling, translation, rotation and flipping on the cropped labeled image and the unlabeled image.
[0038] In S1, the first segmentation network segments the sequence image as follows: ; in, is the segmentation probability map of each frame of pathological image The segmentation probability map of the time series image is composed of is the first segmentation network, is the input time series image, are the trainable parameters of the first segmentation network, is the segmentation label of each frame of pathological image The segmentation labels of the time series images.
[0039] like Figure 2 As shown, in S2, the spatiotemporal memory module includes a channel perception unit, a spatial perception unit, a similarity matrix calculation unit, a temporal and spatial feature reading unit, and a convolution unit.
[0040] The channel perception unit is used to memorize key values Compress the time feature in the time dimension to obtain the time feature value .
[0041] The spatial perception unit is used to characterize the query value Perform feature compression in the spatial dimension to obtain spatial feature values .
[0042] Similarity matrix calculation unit is used to calculate memory key value and query value features Similarity matrix , calculated as: ; in, represents the two-norm operator, symbol Represents a multiplication operation.
[0043] The time and space feature reading unit is used to read the features containing spatial information and read characteristics containing time information : ; ; The convolution unit is used to and Convolution is performed on the features after channel splicing to obtain semantic enhancement features : ; Among them, Conv stands for convolution.
[0044] The feature output by the second layer of the encoder of the second segmentation network in the embodiment of the present invention is used as the query value feature .
[0045] In S3, the channel and spatial attention module includes a channel perception unit and a spatial perception unit.
[0046] Features output by the last layer of the encoder Input to the channel perception unit, compress the time feature in the time dimension, and generate time compression features , the calculation formula is: ; in, represents a multilayer perceptron, represents the sigmoid activation function, represents average pooling, stands for max pooling.
[0047] Features output by the last layer of the encoder Input to the spatial perception unit, perform feature compression in the spatial dimension, and generate spatial compression features , the calculation formula is: ; in, Represents convolution.
[0048] In S4, the time compression feature and spatial compression features The two are sent to the dual cross attention perception module to perform bidirectional feature fusion, respectively generating temporal fine-grained features T@S and spatial fine-grained features S@T, and the two are fused by convolutional features to obtain the fused dual cross attention features. .
[0049] As a derivative version of self-attention, cross attention shows many advantages in utilizing the relationship between different features and effectively fusing these unique information. In this embodiment of the present invention, a dual cross attention mechanism is proposed for matching and fusion. and , to dig and The potential for information to complement each other.
[0050] like Figure 3 As shown, the space compression feature As feature one, the time compression feature As feature 2. In the embodiment of the present invention, the spatial compression feature The features after the linear layer are used as query features , the time compression feature The features after the convolution layer are used as time keys and spatial key values , perform cross attention operation to obtain the spatial fine-grained feature S@T. The calculation formula is: ; in, Indicates cross-attention; Represents the dimension of the feature, T Indicates transposition. At the same time, in the softmax output layer of the cross attention and The residual connection is used between the original time information ( ) fusion.
[0051] Time compression features The features after the linear layer are used as query features , compress the space feature The features after the convolution layer are used as time keys and spatial key values , perform cross attention operation to obtain the temporal fine-grained feature T@S. The calculation formula is: ; Then, T@S is fused with S@T using the following formula: ; in, express Activation function.
[0052] In S4, segmentation probability map and segmentation labels The calculation method is: ; in, is the second segmentation network, are the trainable parameters of the second segmentation network.
[0053] In S5, and Calculate the consistency loss function ,at the same time and The labeled part and the manual label Calculating supervised loss function ; The consistency loss function is added to the full supervision loss function to get the total loss function As shown below: ; in, For the i The segmentation probability map of the frame has a label. In the embodiment of the present invention, the first pathological image in the time series image is selected as the artificially labeled image. , so in this formula i =1; is the total number of frames of pathological images in the time series image. In the embodiment of the present invention, and Both ; is the cross entropy loss, Dice loss.
[0054] Example 2 An embodiment of the present invention provides a semi-supervised segmentation method for a medical image sequence, comprising: Inputting an unlabeled medical image or medical image sequence to be segmented into a trained segmentation network to obtain segmentation labels, wherein the trained segmentation network is the second segmentation network in the semi-supervised segmentation model for medical image sequences constructed by the method for constructing a semi-supervised segmentation model for medical image sequences in Example 1.
[0055] For related solutions, please refer to the corresponding description in Example 1 and will not be repeated here.
[0056] To verify the effectiveness of the method of the present invention, in an embodiment of the present invention, case data were obtained from the publicly available cardiac cine magnetic resonance image dataset ACDC (hereinafter referred to as the ACDC dataset) and the CAMUS dataset. The ACDC dataset contained data from 150 cases, and the CAMUS dataset contained data from 500 cases. Each case contained 10-20 frames, with only the first frame being a labeled image and the rest being unlabeled. The case data in the ACDC and CAMUS datasets were divided into training and validation sets in a ratio of 8:2.
[0057] In the embodiment of the present invention, the framework used by both the first segmentation network and the second segmentation network is the Unet framework, and the network backbone is ResNet50. After the input image is sent to the first segmentation network, it will be downsampled 4 times and upsampled 4 times to generate a segmentation probability map.
[0058] The preprocessing process includes cropping the labeled and unlabeled images in the sample to the same size (in this embodiment, the size is 224×224×20 in width×height×depth), and also includes performing the same scaling, translation, rotation, and flipping on the cropped labeled and unlabeled images.
[0059] The semi-supervised segmentation model for medical image sequences constructed based on the above embodiment was compared with other existing advanced methods on two public datasets. The quantitative comparison was evaluated using the Dice similarity coefficient (DSC), the average symmetric surface distance (ASSD), and the Haussdorf distance (HD). To verify the effectiveness of this method, it was compared with nine commonly used advanced methods. The experimental results are shown in Tables 1 and 2, as well as Figure 4 and Figure 5As shown; among them, 9 commonly used advanced methods are URPC (Uncertainty rectified pyramid consistency), CPS (Cross Pseudosupervision), SLC-Net (Shape-aware and local context constraints Network), MLRP (Mutual learning with reliable pseudo label network), TarVis (Unified videosegmentation network), Xmem (Atkinson-Shiffrin memory segmentation) model), MemSAM (Spatial temporal memory segment anything model), PKEcho (Proxy- andkernel-based semi-supervised network), SSCF (Spatiotemporal semantic calibration and fusion network). In Tables 1 and 2, RVC indicates that the test image is a right ventricular image, Myo indicates that the test image is a myocardial image, LVC indicates that the test image is a left ventricular image, and LA indicates that the test image is a right atrial image; TransUnet (LB) indicates the lower limit of performance, and TransUnet (UB) indicates the upper limit of performance. Method 1 indicates the method proposed in this embodiment of the present invention, and Method 2 indicates the method in this embodiment of the present invention after removing the spatiotemporal memory module. Figure 4 Neutralization Figure 5 In the figure, T represents different times. ; .
[0060] Table 1 and Table 2 respectively list the DSC values and ASSD values corresponding to the present invention and other segmentation methods. Table 1 shows the results of the present invention on the ACDC dataset, with 10 cases as training samples and 50 cases as test samples. Among the 10 training samples, each training sample contains 10 frames of images, of which only the first frame of the image is manually labeled. The black bold part indicates the best performance under this evaluation indicator. Table 2 shows the results of the present invention on the CAMUS dataset, with 100 cases as training samples and 400 cases as test samples. The black bold part indicates the best performance under this evaluation indicator. It can be seen from the table that the performance of the proposed method exceeds other semi-supervised segmentation methods and is very close to the upper bound of fully supervised segmentation.
[0061] Example 3 An embodiment of the present invention provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the method in the above-mentioned embodiment 1 or embodiment 2 when executing the computer program.
[0062] The relevant technical solutions are the same as above and will not be repeated here.
[0063] Example 4 An embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the method in the above-mentioned embodiment 1 or embodiment 2 are implemented.
[0064] Specifically, the memory may include a high-speed random access memory and may also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage device.
[0065] The relevant technical solutions are the same as above and will not be repeated here.
[0066] Example 5 An embodiment of the present application provides a computer program product, including a computer program. When the computer program is run on a computer, the computer executes the steps of the method in the above-mentioned embodiment 1 or embodiment 2.
[0067] The relevant technical solutions are the same as above and will not be repeated here.
[0068] It will be easily understood by those skilled in the art that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for constructing a semi-supervised segmentation model for medical image sequences, characterized in that: include: The semi-supervised segmentation network of the medical image sequence is trained using a training sample set to obtain a trained semi-supervised segmentation model of the medical image sequence; wherein the training sample is a medical image sequence, and the medical image sequence contains at least one frame of manually annotated image GT ; The medical image sequence semi-supervised segmentation network includes: The first segmentation network is used to segment the medical image sequence input frame by frame to obtain the segmentation probability map of each frame of medical image Segmentation probability map and the corresponding segmentation labels ; Feature projection module, used to transform the segmentation probability map The image after splicing with the corresponding medical image is downsampled, and the downsampled features corresponding to each frame of medical image are used as memory keys ; The second segmentation network includes an encoder and a decoder, wherein the encoder is used to perform feature encoding on the medical image sequence input in the form of a sequence to obtain a feature ; The spatiotemporal memory module is used to store the memory key value and query value features Perform feature query-matching to generate semantically enhanced features of the medical image sequence ; Wherein, the query value feature The encoder output and the memory key value Intermediate features with consistent resolution; Channel and spatial attention modules are used to integrate the features After performing channel and space perception respectively, obtain time compression features and spatial compression features ;in, and The fused features With the characteristics After inputting into the decoder, the segmentation probability map is obtained and segmentation labels ; The training process includes: calculating the segmentation labels and The consistency loss and the segmentation probability map and The segmentation probability maps with labels are respectively associated with their corresponding manually labeled images supervised loss; adding the consistency loss and the supervised loss as the total loss to train the medical image sequence semi-supervised segmentation network to obtain a trained medical image sequence semi-supervised segmentation model.
2. The method for constructing a semi-supervised segmentation model for medical image sequences according to claim 1, wherein: The spatiotemporal memory module comprises: Channel sensing unit, used to memorize the key value Compress the time feature in the time dimension to obtain the time feature value ; A spatial perception unit for detecting the query value characteristics Perform feature compression in the spatial dimension to obtain spatial feature values ; Similarity matrix calculation unit, used to calculate the memory key value and the query value characteristic Similarity matrix , ;in, represents the two-norm operator; Time and space feature reading unit, used to read features containing spatial information and read the special ;in, , ; Convolution unit, used to and The features after channel splicing are convolved to obtain the features .
3. The method for constructing a semi-supervised segmentation model for medical image sequences according to claim 1 or 2, wherein: The medical image sequence semi-supervised segmentation network also includes a dual cross-attention perception module; The features The way to obtain is: Will and After being input into the dual cross attention perception module for bidirectional feature fusion, the temporal fine-grained feature T@S and the spatial fine-grained feature S@T are generated respectively, and T@S and S@T are convolutionally fused to obtain the feature ; Among them, T@S and S@T are calculated as follows: in, Represents cross attention; query features for Features after the linear layer, time key value and spatial key values for Features after convolutional layers; query features for Features after the linear layer, time key value and spatial key values for Features after the convolutional layer respectively; Represents the dimension of the feature, Represents convolution.
4. The method for constructing a semi-supervised segmentation model for medical image sequences according to claim 3, wherein: The channel and spatial attention module includes a channel perception unit and a spatial perception unit; The channel perception unit is used to Perform time feature compression on the time dimension to generate the time compression feature ; The spatial perception unit is used to Perform feature compression in the spatial dimension to generate the spatial compression feature .
5. The method for constructing a semi-supervised segmentation model for medical image sequences according to claim 4, wherein: The time compression feature and the spatial compression feature The calculation method is: in, represents a multilayer perceptron, represents the sigmoid activation function, represents average pooling, stands for max pooling.
6. The method for constructing a semi-supervised segmentation model for medical image sequences according to claim 1, wherein: The segmentation probability map output by the first segmentation network and the corresponding segmentation labels for: in, For the first segmentation network, is the input medical image sequence, are the trainable parameters of the first segmentation network; The segmentation probability map output by the second segmentation network and segmentation labels for: in, For the second segmentation network, are the trainable parameters of the second segmentation network.
7. A semi-supervised segmentation method for medical image sequences, characterized in that: include: The unlabeled medical image or medical image sequence to be segmented is input into the second segmentation network of the medical image sequence semi-supervised segmentation model constructed by the method for constructing a medical image sequence semi-supervised segmentation model according to any one of claims 1 to 6 to obtain a segmentation label.
8. An electronic device, characterized in that: comprising a computer-readable storage medium and a processor; The computer-readable storage medium is used to store executable instructions; The processor is used to read the executable instructions stored in the computer-readable storage medium to execute the medical image sequence semi-supervised segmentation model construction method described in any one of claims 1 to 6, or to execute the medical image sequence semi-supervised segmentation method described in claim 7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method for constructing a semi-supervised segmentation model for a medical image sequence as described in any one of claims 1 to 6 is implemented, or the method for semi-supervised segmentation of a medical image sequence as described in claim 7 is implemented.
10. A computer program product, characterized in that It includes a computer program, which, when running on a computer, enables the computer to execute the method for constructing a semi-supervised segmentation model for a medical image sequence according to any one of claims 1 to 6, or execute the semi-supervised segmentation method for a medical image sequence according to claim 7.
Citation Information
Patent Citations
Self-adaptive semi-supervised image segmentation method and system based on uncertainty knowledge domain
CN114549842A
Global feature enhanced semi-supervised video target segmentation method and system
CN117876931A
Semi-supervised medical image segmentation method of mutual pseudo supervised edge perception double CNN
CN118297976A
Semi-supervised medical image segmentation method based on cross-image learning and shape fusion
CN120072213A
Three-dimensional medical image segmentation method
CN120219754A
Cited By
A breast cancer medical image segmentation method fusing expert knowledge and entropy optimization
CN122574402B