Semi-supervised medical image segmentation method and device based on global geometric prior
Through the global geometric prior semi-supervised medical image segmentation method, the image data carrying labels and uncarried labels are used, combined with data perturbation and multiple loss functions optimization models, the problem of deep learning algorithm dependence on labeled data is solved, and the accuracy and robustness of medical image segmentation is improved.
Patent Information
- Application Number
- CN202510274633.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-07-18
AI Technical Summary
Existing deep learning algorithms require a large amount of labeled data in medical image segmentation. The noise information in the unlabeled data limits the model effect, and the receptive field of the convolutional neural network is limited, making it difficult to effectively capture global geometric features.
The semi-supervised medical image segmentation method of global geometric priors is adopted. By obtaining medical images carrying labels and not carrying labels, performing data perturbation processing, inputting the semi-supervised learning model, constructing the target loss function, integrating supervision loss, consistency loss and geometric moment consistency loss, and optimizing model parameters.
The segmentation performance and generalization ability of the model under complex noise conditions can be improved, and the global geometric features of medical images can be better captured and segmentation accuracy and robustness can be improved.
Smart Images

Figure CN120339602A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep learning, and more particularly, to a semi-supervised medical image segmentation method and device based on global geometric prior knowledge. Background Art
[0002] Medical image segmentation is crucial for computer-aided diagnosis. However, deep learning algorithms usually require a large amount of labeled data. Semi-supervised learning can make up for this deficiency by simultaneously using labeled and unlabeled data, which not only improves the model performance but also enhances the efficiency of medical image training. However, due to the inherently complex background of medical images, the information extracted from unlabeled samples often contains noise, which to a certain extent limits the effectiveness of general models.
[0003] In medical data, the same organ usually exhibits a consistent anatomical structure, and these structures are closely related to geometric features because anatomical structures are essentially defined by specific geometric attributes (such as shape, size, position, and orientation). These geometric feature extraction methods can effectively improve the segmentation performance, but usually rely on a limited variety of geometric features. At the same time, due to the inherent limitations of convolutional neural networks, their receptive fields are restricted by the size of the convolutional kernel and mainly capture local patterns. Therefore, integrating more types of global geometric features is crucial for enabling the model to learn more discriminative and semantically meaningful representations from a large amount of unlabeled data. Summary of the Invention
[0004] The purpose of the present invention is to provide a semi-supervised medical image segmentation method and device based on global geometric prior knowledge to improve the above problems. To achieve the above purpose, the technical solutions adopted by the present invention are as follows:
[0005] In a first aspect, the present application provides a semi-supervised medical image segmentation method based on global geometric prior knowledge, including:
[0006] Obtaining first information and second information, where the first information is a medical image carrying a label, and the second information is a medical image not carrying a label;
[0007] Performing data perturbation processing on the second information to obtain third information;
[0008] Inputting the first information, second information, and third information into a preset semi-supervised learning model, and constructing an objective loss function based on a supervision loss, a consistency loss, and a geometric moment consistency loss. When the calculated value of the objective loss function meets a set condition, stop optimizing the model parameters to obtain a semi-supervised image segmentation model, where the geometric moment consistency loss is used to measure the degree of consistency of the model for the same image region in terms of symmetry, shape, and orientation information;
[0009] Input the medical image to be segmented into the semi-supervised image segmentation model to obtain a predicted segmentation result.
[0010] In a second aspect, the present application also provides a semi-supervised medical image segmentation device based on global geometric prior, including:
[0011] An acquisition unit, configured to acquire first information and second information, where the first information is a medical image with labels, and the second information is a medical image without labels;
[0012] A perturbation unit, configured to perform data perturbation processing on the second information to obtain third information;
[0013] An optimization unit, configured to input the first information, second information, and third information into a preset semi-supervised learning model, and construct an objective loss function based on a supervision loss, a consistency loss, and a geometric moment consistency loss. When the calculated value of the objective loss function meets a set condition, stop optimizing the model parameters to obtain a semi-supervised image segmentation model, where the geometric moment consistency loss is used to measure the degree of consistency of the model for the same image region in terms of symmetry, shape, and direction information;
[0014] A first input unit, configured to input the medical image to be segmented into the semi-supervised image segmentation model to obtain a predicted segmentation result.
[0015] In a third aspect, the present application also provides a semi-supervised medical image segmentation device based on global geometric prior, including:
[0016] A memory, configured to store a computer program;
[0017] A processor, configured to implement the steps of the semi-supervised medical image segmentation method based on global geometric prior when executing the computer program.
[0018] In a fourth aspect, the present application also provides a readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above-mentioned semi-supervised medical image segmentation method based on global geometric prior are implemented.
[0019] The beneficial effects of the present invention are as follows:
[0020] By inputting the image with labels and the unlabeled image with perturbations into a preset semi-supervised learning model and constructing an objective function, the objective loss function integrates a supervision loss, a consistency loss, and a geometric moment consistency loss that can accurately measure the degree of consistency of the model for the same image region in terms of symmetry, shape, and direction information, and can comprehensively and deeply optimize the model parameters to obtain a semi-supervised image segmentation model with better performance.
[0021] Other features and advantages of the present invention will be described in the subsequent specification, and in part will be obvious from the specification, or can be understood by implementing the embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention, and therefore should not be regarded as a limitation of the scope. For those of ordinary skill in the art, without creative efforts, other relevant drawings can also be obtained based on these drawings.
[0023] Figure 1 Schematic flow diagram of the semi-supervised medical image segmentation method based on global geometric prior described in the embodiments of the present invention;
[0024] Figure 2 Schematic framework diagram of the semi-supervised medical image segmentation based on global geometric prior described in the embodiments of the present invention;
[0025] Figure 3 Schematic flow diagram of the multi-view geometric moment attention mechanism described in the embodiments of the present invention;
[0026] Figure 4 Schematic flow diagram of the global information interaction described in the embodiments of the present invention;
[0027] Figure 5 Schematic four-way geometric rearrangement of the information interaction Mamba (IIM) described in the embodiments of the present invention;
[0028] Figure 6 Schematic structural diagram of the semi-supervised medical image segmentation device based on global geometric prior described in the embodiments of the present invention;
[0029] Figure 7 Schematic structural diagram of the semi-supervised medical image segmentation device based on global geometric prior described in the embodiments of the present invention.
[0030] Reference numerals in the figure: 10, acquisition unit; 20, perturbation unit; 30, optimization unit; 40, first input unit; 800, semi-supervised medical image segmentation device based on global geometric prior; 801, processor; 802, memory; 803, multimedia component; 804, I / O interface; 805, communication component. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0031] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. The components of the embodiments of the present invention usually described and illustrated in the accompanying drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0032] It should be noted that like reference numerals and letters denote like items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of the present invention, the terms "first", "second", etc. are only used for distinguishing descriptions and cannot be construed as indicating or implying relative importance.
[0033] Embodiment 1:
[0034] This embodiment provides a semi-supervised medical image segmentation method based on global geometric prior.
[0035] Medical image datasets usually contain a small labeled subset and a large unlabeled subset. The labeled data is not sufficient to represent all the diversities, resulting in distribution differences. In the prior art, a copy-paste method is adopted to spread the semantic information of the labeled data to the unlabeled data. There are also other methods that assume that regions with similar semantics in the labeled and unlabeled data from the same distribution can share labels. Such as cosine similarity, complete correlation matrix, and graph-based methods are all used to identify and transfer label information. However, these methods mainly focus on local similarity matching and may not be able to effectively capture the overall data distribution, which is crucial for alleviating the bias between scarce labeled data and abundant unlabeled data.
[0036] Inspired by the above, this application proposes a novel semi-supervised medical image segmentation framework (GIGP) based on global geometric priors from a global perspective. First, the global information interaction Mamba module can integrate global context information with linear computational complexity, thus better aligning the shared features between labeled data and unlabeled data. Second, geometric moments can capture various global geometric features of samples, such as shape, symmetry, and directionality, which helps to characterize the anatomical structure of samples. At the same time, the normalized geometric moments provide scale invariance and retain information during the multi-scale fusion process, reducing the information loss caused by upsampling or downsampling. Third, the same organ naturally exhibits dynamic changes and geometric variations between different individuals or over time. The sine wave transform can effectively simulate these changes, such as the periodic shape changes of the atrium during the heartbeat and the morphological differences of the pancreas between different individuals.
[0037] See Figure 1 , which shows that the method includes step S10, step S20, step S30, and step S40.
[0038] Step S10. Obtain the first information and the second information. The first information is a medical image with a label, and the second information is a medical image without a label;
[0039] Specifically, considering that deep learning-based image segmentation algorithms require a large amount of labeled data for training to obtain good segmentation performance, and medical images usually have the characteristics of complex backgrounds and blurred boundaries, which greatly increases the difficulty of labeling. Therefore, it is unrealistic to obtain a large amount of labeled medical data. Therefore, semi-supervised learning can be used to try to fully utilize labeled and unlabeled data to improve the performance of the model.
[0040] Step S20. Perform data perturbation processing on the second information to obtain the third information;
[0041] Specifically, the data perturbation processing includes noise addition processing and image distortion processing. The noise addition processing can simulate the noise interference situation that may be encountered during image acquisition in the real scene by introducing random noise into the original medical image samples, greatly expanding the diversity of training data, enhancing the adaptability of the model to medical images in different noise environments, and improving the robustness of the model in segmenting medical images under complex noise conditions. The image distortion processing changes the geometric features such as the shape and angle of the medical image samples, providing rich deformed image samples for the model, prompting the model to learn the features of the image in different deformed states, enhancing the model's recognition ability for complex shape changes of medical images, enabling it to accurately segment even when facing various image deformations that may occur in practical applications, and improving the accuracy and generalization ability of the model's segmentation results.
[0042] Step S20 specifically includes steps S21 to S25:
[0043] Step S21. Construct a distorted coordinate based on a preset amplitude and frequency to obtain a distorted grid coordinate;
[0044] Step S22. Map the second information into the distorted grid coordinate based on a three-dimensional interpolation algorithm to obtain the second information after image distortion processing;
[0045] Specifically, performing a sine transform on the shape of the input data to distort its shape can highlight its geometric features, thereby enhancing the adaptability of the model to changes in the data shape. Since the performance difference between using sine or cosine for image distortion is small, we choose to use a sine transform to distort the input data. The steps of Global Geometric Perturbation Consistency (GGPC) are as follows:
[0046] First, we use the following formula to generate a wave grid to distort the three-dimensional coordinates and generate a distorted grid with wave characteristics:
[0047]
[0048] where G(x, y, z) is the distorted grid coordinate; A is the amplitude; f is the frequency.
[0049] By making the distortion subtle enough, the model can capture the tiny deformations in the input data.
[0050] Second, use three-dimensional interpolation technology to map n unlabeled samples to the distorted grid, expressed as:
[0051]
[0052] where, is the unlabeled sample after mapping; F grid is the function used to perform the interpolation; G(x, y, z) is the distorted grid coordinate.
[0053] By effectively performing geometric transformation on three-dimensional data, it provides a richer feature representation for subsequent deep learning models.
[0054] Step S23. Generate a noise matrix based on a random number generator for a preset noise frequency, type, and distribution characteristics, and the noise matrix has the same size as the image in the second information;
[0055] Step S24. Combine the noise matrix with the second information to obtain the second information after noise addition processing;
[0056] Step S25. Use the second information after image distortion processing and the second information after noise addition processing as the third information;
[0057] Specifically, a noise matrix consistent with the image size in the second information is carefully constructed using a random number generator according to the preset noise frequency, type, and distribution characteristics. Subsequently, this noise matrix is cleverly combined with the second information to complete the noise addition process for the second information. The combination of noise addition and image distortion provides more comprehensive and diverse data for subsequent model training, comprehensively assisting in optimizing the model performance.
[0058] Step S30. Input the first information, the second information, and the third information into a preset semi-supervised learning model, and construct an objective loss function based on the supervised loss, the consistency loss, and the geometric moment consistency loss. When the calculated value of the objective loss function meets the set conditions, stop optimizing the model parameters to obtain a semi-supervised image segmentation model. The geometric moment consistency loss is used to measure the degree of consistency of the model for the same image region in terms of symmetry, shape, and orientation information.
[0059] Specifically, the semi-supervised algorithm framework of this application is as Figure 2 shown. We adopt the Mean Teacher (MT) model as the backbone network. In the forward propagation process, each mini-batch contains m medical images with labels, that is, labeled samples and n medical images without labels, that is, unlabeled samples
[0060] Specifically, step S30 specifically includes steps S31 to S36:
[0061] Step S31. Input the first information and the second information into the student network to obtain multiple first feature maps, labeled image probability maps, and unlabeled image probability maps.
[0062] Specifically, the student network processes the labeled samples and unlabeled samples to obtain multiple first feature maps, labeled data probability maps and unlabeled data probability maps
[0063] Step S32. Input the second information into the teacher network to obtain multiple second feature maps and perturbation probability maps.
[0064] Specifically, the teacher network processes the unlabeled samples with added noise and their global geometric perturbation samples to obtain multiple second feature maps, as well as the teacher network noise perturbation probability map and the teacher network geometric perturbation probability map Among them, the noise perturbation probability map corresponds to the samples processed by the noise addition process, and the geometric perturbation probability map corresponds to the samples processed by the image distortion process.
[0065] Step S33. Construct a geometric moment consistency loss function based on the first feature map and the second feature map;
[0066] Specifically, both the student network and the teacher network in this application consist of multiple encoding layers, an information interaction module, and multiple decoding layers. As Figure 2 shown, the student network includes five encoding layers. Among the five decoding layers from left to right, the first encoding layer is a single convolutional layer, and the second to fifth encoding layers are each composed of a downsampling layer and multiple convolutional layers. After the processing of the first four decoding layers, a multi-view geometric attention mechanism (MVMA) is added, that is, using geometric moments to impose multi-view constraints to provide comprehensive guidance for various geometric features. This attention mechanism weights these features, enabling the model to precisely focus on and extract key information in the image. The information interaction module is the Mamba (GIIM) module, which can jointly learn the features of labeled data and unlabeled data at corresponding spatial positions. By globally transferring the features learned from the labeled data to the unlabeled data, it can more effectively capture shared characteristics, thereby narrowing the distribution difference between them and improving the overall learning performance. The teacher network also includes five encoding layers, and the composition of the encoding layers is the same as that of the student network, but no geometric attention mechanism is added between the encoding layers.
[0067] The student network also contains four decoding layers, and each of the four decoding layers is composed of an upsampling layer and multiple convolutional layers. The decoding layers of the teacher network are exactly the same as those of the student network.
[0068] Specifically, step S33 specifically includes steps S331 to S336:
[0069] Step S331. Input the first information and the second information into the student network, and process them successively through multiple encoding layers, the information interaction module, and multiple decoding layers to obtain multiple encoded feature maps, a first interaction feature map, and multiple first decoded feature maps respectively;
[0070] Step S332. Input the third information into the teacher network, and process it through all the encoding layers and the information interaction module to obtain a second interaction feature map;
[0071] Specifically, the first information carrying tags and the second information without tags are input into the student network together. First, they enter multiple encoding layers. After operations such as convolution, sampling, and geometric attention mechanism in the encoding layers, the image resolution is gradually reduced, and at the same time, features from low-level to high-level, from simple to complex are extracted, generating multiple encoded feature maps. These encoded feature maps contain rich semantic information of the image and are the basis for subsequent processing. Then, it enters the information interaction module, which promotes the communication and fusion between different features, further integrates the image information, and generates the first interaction feature map. Finally, it flows through multiple decoding layers. The decoding layers are opposite to the encoding layers. By means of deconvolution, upsampling, etc., the image resolution is gradually restored, and the abstract feature information is transformed into specific feature representations, obtaining multiple first decoded feature maps, which are prepared for the generation of subsequent segmentation results.
[0072] After performing noise perturbation and geometric perturbation on the medical image without tags, the obtained third information is input into the teacher network. The teacher network first passes through all encoding layers, also extracts image features and reduces the image resolution, then passes through the information interaction module to integrate the features, and finally outputs the second interaction feature map.
[0073] Specifically, step S332 specifically includes steps S3321 to S3325:
[0074] Step S3321. After passing the third information through all encoding layers, an initial feature map is obtained;
[0075] Step S3322. Perform spatial convolution processing on the initial feature map to obtain the third feature map;
[0076] Specifically, a geometric moment attention mechanism GMAM is added to the encoding layer. GMAM constructs a multi-view second-order geometric moment attention mechanism (MVMA) and a multi-scale second-order normalized geometric moment consistency (MSGC). As Figure 3 shown, the second-order geometric moment can help the model capture the symmetry, shape, and orientation information of the samples. MVMA can extract global geometric features from three directions of the three-dimensional features, namely the length, width, and height directions, to help the model comprehensively understand the three-dimensional space structure. MSGC utilizes scale invariance to achieve the complementarity of deep and shallow geometric features.
[0077] Step S3323. Perform normalization processing on the third feature map, and perform four-way geometric rearrangement on the normalized third feature map to obtain the fourth feature map. The four-way geometric rearrangement includes rearranging the forward direction, reverse direction, inter-channel direction, and inter-sample direction of the feature map;
[0078] Step S3324. Perform fusion processing on the fourth feature map and the third feature map to obtain the fifth feature map;
[0079] Step S3325. Normalize and perform fully connected processing on the fifth feature map in sequence, and fuse the processed fifth feature map and the fourth feature map to obtain the second interaction feature map;
[0080] Specifically, the present application introduces global information interaction Mamba in the information interaction module, enabling the geometric information of the labeled data to be effectively transmitted to the unlabeled data, so that the unlabeled data can utilize the geometric information in the labeled data to improve the accuracy of feature extraction. Through this information propagation from labeled to unlabeled data, during the learning process of unlabeled data, the global geometric information crucial for medical image segmentation can be naturally prioritized, thereby significantly enhancing the performance and generalization ability of the model in medical image segmentation tasks.
[0081] As Figure 4 shown, the specific calculation process of the information interaction module is as follows:
[0082]
[0083] Specifically, F 0 is the initial feature map; F 1 is the third feature map; F 2 is the fourth feature map; F out is the second interaction feature map; LayerNorm() is normalization; MLP() is the fully connected processing multi-layer perceptron, i.e., fully connected; IIM() is the information interaction Mamba.
[0084] The Mamba theory is to achieve the interaction between labeled and unlabeled features in the channel dimension, effectively encoding and learning their shared representation.
[0085] As Figure 5 shown, IIM consists of four directional components: the forward direction f, the reverse direction r, the inter-channel direction c, and the inter-sample direction t. The output of IIM is defined as follows:
[0086] IIM(F) = Mamba(F f ) + Mamba(F r ) + λ1Mamba(F c ) + λ2Mamba(F t )
[0087] where λ1 and λ2 are the corresponding weight coefficients; IIM() is the information interaction Mamba; F f is the forward direction; F r is the reverse direction; F c is the inter-channel direction; F t is the inter-sample direction; Mamba() is the Mamba layer, which is used to model the global information in the sequence.
[0088] The relevant steps corresponding to the student network and the teacher network are exactly the same and will not be elaborated here.
[0089] Step S333. The second interactive feature map is decoded through all decoding layers to obtain a second decoded feature map;
[0090] Step S334. Construct a first loss function based on multiple encoded feature maps, the first interactive feature map, and the second interactive feature map;
[0091] Specifically, step S334 specifically includes steps S3341 to S3347:
[0092] Step S3341. Calculate the centroid position in the encoded feature map to obtain the image centroid;
[0093] Step S3342. Calculate the second-order geometric moment corresponding to the encoded feature map based on the image centroid;
[0094] Step S3343. Normalize the second-order geometric moment to obtain a first normalized geometric moment;
[0095] Step S3344. Based on the above steps of calculating the image centroid, second-order geometric moment, and normalizing, process the first interactive feature map and the second interactive feature map respectively, and sequentially obtain a second normalized geometric moment and a third normalized geometric moment;
[0096] Step S3345. Calculate the difference between each first normalized geometric moment and the third normalized geometric moment, and calculate the product of the difference and a preset weight coefficient to obtain multiple first products;
[0097] Step S3346. Calculate the difference between the second normalized geometric moment and the third normalized geometric moment, and calculate the product of the difference and a preset weight coefficient to obtain a second product;
[0098] Step S3347. Calculate the sum of multiple first products and the second product to obtain a first loss function;
[0099] Specifically, after MVMA is applied to each encoder layer for encoding, the output of the k-th layer encoder is where C k is the category output by the k-th layer encoder; H k is the height output by the k-th layer encoder; W k is the width output by the k-th layer encoder; L k is the length output by the k-th layer encoder. And calculate the second-order geometric moment in different directions, as Figure 2 shown:
[0100]
[0101] Among them, is the geometric moment of order p+q+r in the height direction of the encoded feature map output by the k-th layer encoder; P k (x,y,z) is the class prediction value corresponding to the position (x,y,z) in the encoded feature map output by the k-th layer encoder; (x0,y0,z0) is the centroid; p+q+r = 2.
[0102] The calculation methods of the second-order geometric moments corresponding to the width direction and the length direction are exactly the same as those of the above height direction. The second-order geometric moment in the height direction The second-order geometric moment in the width direction and the second-order geometric moment in the length direction contain global geometric information of different dimensions. Construct a spatial attention map:
[0103]
[0104] Among them, is the geometric moment of order p+q+r in the i direction of the encoded feature map output by the k-th layer encoder; P k (x,y,z) is the class prediction value corresponding to the position (x,y,z) in the encoded feature map output by the k-th layer encoder; λ P is the coefficient optimized by neural network learning; z k (x,y,z) is the spatial attention map; σ() is the activation function; Conv() is the 1×1 convolution function.
[0105] MSGC is applied to each resolution layer of the student network, as well as the final layers of the encoder and decoder in the teacher network. Its normalized geometric moment can be expressed as:
[0106]
[0107] Among them, is the normalized geometric moment of the encoded feature map output by the k-th layer encoder in the student network; is the normalized geometric moment of the encoded feature map output by the k-th layer decoder in the student network; is the normalized geometric moment of the encoded feature map corresponding to the noise perturbation in the teacher network; is the normalized geometric moment of the decoded feature map corresponding to the noise perturbation in the teacher network; is the normalized geometric moment of the encoded feature map corresponding to the geometric perturbation in the teacher network; is the normalized geometric moment of the decoded feature map corresponding to the geometric perturbation in the teacher network; (x0,y0,z0) is the centroid; is the encoder or decoder feature that is converted into a single channel after channel alignment processing in the student network; After channel alignment processing in the teacher network, the noise perturbation is converted into single-channel encoder or decoder features; After channel alignment processing in the teacher network, the geometric perturbation is converted into single-channel encoder or decoder features; σ x is the normalization scale in the x direction; H k is the height of the output of the k-th layer encoder.
[0108] σ y and σ z are calculated in the same way as σ x and will not be elaborated here.
[0109] The calculation formula of the first loss function is:
[0110]
[0111] where L1 is the first loss function; α k and β k are weight coefficients; is the normalized geometric moment of the encoded feature map output by the k-th layer encoder in the student network; is the normalized geometric moment of the encoded feature map corresponding to the noise perturbation in the teacher network; is the normalized geometric moment of the encoded feature map corresponding to the geometric perturbation in the teacher network.
[0112] Step S335. Construct a second loss function based on multiple first decoded feature maps and second decoded feature maps;
[0113] Specifically, each decoding layer in the student network will output corresponding first decoded feature maps, and the teacher network will obtain second decoded feature maps only after all decoding layers. The construction of the second loss function is the same as the construction method of the above first loss function, and the calculation of the normalized geometric moment is required, which will not be elaborated here.
[0114] The corresponding second function calculation formula is:
[0115]
[0116] where L2 is the second loss function; α k and β k are weight coefficients; is the normalized geometric moment of the decoded feature map output by the k-th layer decoder in the student network; is the normalized geometric moment of the decoded feature map corresponding to the noise perturbation in the teacher network; is the normalized geometric moment of the decoded feature map corresponding to the geometric perturbation in the teacher network.
[0117] Step S336. Obtain the geometric moment consistency loss function based on the first loss function and the second loss function;
[0118] Specifically, the geometric moment consistency loss function is defined as follows:
[0119]
[0120] where L gmc is the geometric moment consistency loss function; α k and β k are weight coefficients; is the normalized geometric moment of the encoded feature map output by the k-th layer encoder in the student network; is the normalized geometric moment of the decoded feature map output by the k-th layer decoder in the student network; is the normalized geometric moment of the encoded feature map corresponding to the noise perturbation in the teacher network; is the normalized geometric moment of the decoded feature map corresponding to the noise perturbation in the teacher network; is the normalized geometric moment of the encoded feature map corresponding to the geometric perturbation in the teacher network; is the normalized geometric moment of the decoded feature map corresponding to the geometric perturbation in the teacher network.
[0121] Step S34. Construct a supervision loss function based on the labeled image probability map and the true label probability map;
[0122] Specifically, according to the labeled data probability map and the true label supervise the labeled data, and construct the supervision loss function as:
[0123] L s = D(P l s , Y l ) + α E E(P l s , Y l )
[0124] where D() is the loss function constructed by the Dice coefficient; E() is the cross-entropy loss function; α E is the balance coefficient; P l s is the labeled data probability map; Y l is the true label.
[0125] Step S35. Construct a consistency loss function based on the unlabeled image samples and the perturbation probability map;
[0126] Specifically, according to the unlabeled data probability map the teacher network noise perturbation probability map and the geometric perturbation probability map of the teacher network Construct the consistency loss function as follows:
[0127]
[0128] where MSE() is the mean square error function; L c is the consistency loss function; is the probability map of unlabeled data; is the noise perturbation probability map of the teacher network; is the geometric perturbation probability map of the teacher network.
[0129] Step S36. Based on preset weight coefficients, construct the target loss function from the consistency loss function, the supervised loss function, and the consistency loss function;
[0130] Specifically, the target loss function of the GIGP framework is defined as:
[0131] L = L s + γ1L c + γ2L gmc
[0132] where γ1 and γ2 are weight coefficients; L is the target loss function; L s is the supervised loss function; L c is the consistency loss function; L gmc is the geometric moment consistency loss function.
[0133] Step S40. Input the medical image to be segmented into the semi-supervised image segmentation model to obtain the predicted segmentation result;
[0134] Specifically, input the target medical image into the trained semi-supervised image segmentation model, and the segmentation result of the target medical image will be obtained, that is, a predicted label will be assigned to each pixel position.
[0135] Embodiment 2:
[0136] As Figure 6 shown, this embodiment provides a semi-supervised medical image segmentation device based on global geometric prior, and the device includes:
[0137] An acquisition unit 10, configured to acquire first information and second information, where the first information is a medical image carrying labels, and the second information is a medical image not carrying labels;
[0138] A perturbation unit 20, configured to perform data perturbation processing on the second information to obtain third information;
[0139] The optimization unit 30 is configured to input the first information, the second information, and the third information into a preset semi-supervised learning model, and construct an objective loss function based on a supervision loss, a consistency loss, and a geometric moment consistency loss. After the calculated value of the objective loss function meets the set conditions, the optimization of the model parameters is stopped to obtain a semi-supervised image segmentation model. The geometric moment consistency loss is used to measure the degree of consistency of the model for the same image region in terms of symmetry, shape, and orientation information;
[0140] The first input unit 40 is configured to input the medical image to be segmented into the semi-supervised image segmentation model to obtain a predicted segmentation result.
[0141] In a specific embodiment disclosed in the present application, the perturbation unit 20 includes:
[0142] The distortion unit is configured to construct distorted coordinates based on a preset amplitude and frequency to obtain distorted grid coordinates;
[0143] The mapping unit is configured to map the second information into the distorted grid coordinates based on a three-dimensional interpolation algorithm to obtain the second information after image distortion processing;
[0144] The generation unit is configured to generate a noise matrix based on a random number generator for a preset noise frequency, type, and distribution characteristics. The noise matrix has the same size as the image in the second information;
[0145] The combination unit is configured to combine the noise matrix with the second information to obtain the second information after noise addition processing;
[0146] The as unit is configured to use the second information after image distortion processing and the second information after noise addition processing as the third information.
[0147] In a specific embodiment disclosed in the present application, the optimization unit 30 includes:
[0148] The second input unit is configured to input the first information and the second information into the student network to obtain a plurality of first feature maps, an annotated image probability map, and an unannotated image probability map;
[0149] The third input unit is configured to input the second information into the teacher network to obtain a plurality of second feature maps and a perturbation probability map;
[0150] The first construction unit is configured to construct a geometric moment consistency loss function based on the first feature map and the second feature map;
[0151] The second construction unit is configured to construct a supervision loss function based on the annotated image probability map and the true label probability map;
[0152] The third construction unit is configured to construct a consistency loss function based on the unannotated image samples and the perturbation probability map;
[0153] The fourth construction unit is used to construct a target loss function based on a consistency loss function, a supervision loss function, and a consistency loss function according to a preset weight coefficient.
[0154] In a specific implementation manner disclosed in the present application, the first construction unit includes:
[0155] The fourth input unit is used to input the first information and the second information into the student network, and after passing through multiple encoding layers, an information interaction module, and multiple decoding layers for processing, respectively obtain multiple encoded feature maps, a first interaction feature map, and multiple first decoded feature maps;
[0156] The fifth input unit is used to input the third information into the teacher network, and after passing through all the encoding layers and the information interaction module for processing, obtain a second interaction feature map;
[0157] The decoding unit is used to perform decoding processing on the second interaction feature map through all the decoding layers to obtain a second decoded feature map;
[0158] The fifth construction unit is used to construct a first loss function based on multiple encoded feature maps, a first interaction feature map, and a second interaction feature map;
[0159] The sixth construction unit is used to construct a second loss function based on multiple first decoded feature maps and a second decoded feature map;
[0160] The first obtaining unit is used to obtain a geometric moment consistency loss function based on the first loss function and the second loss function.
[0161] In a specific implementation manner disclosed in the present application, the fifth construction unit includes:
[0162] The first calculation unit is used to calculate the centroid position in the encoded feature map to obtain the image centroid;
[0163] The encoding unit is used to calculate the second-order geometric moment corresponding to the encoded feature map based on the image centroid;
[0164] The first normalization unit is used to perform normalization processing on the second-order geometric moment to obtain a first normalized geometric moment;
[0165] The second obtaining unit is used to process the first interaction feature map and the second interaction feature map respectively based on the steps of calculating the image centroid, the second-order geometric moment, and normalizing, and sequentially obtain a second normalized geometric moment and a third normalized geometric moment respectively;
[0166] The second calculation unit is used to calculate the difference between each first normalized geometric moment and the third normalized geometric moment, and calculate the product of the difference and a preset weight coefficient to obtain multiple first products;
[0167] A third calculation unit, configured to calculate the difference between the second normalized geometric moment and the third normalized geometric moment, and calculate the product of the difference and a preset weight coefficient to obtain a second product;
[0168] A fourth calculation unit, configured to calculate the sum of a plurality of first products and the second product to obtain a first loss function.
[0169] In a specific implementation manner disclosed in this application, the fifth input unit includes:
[0170] A third obtaining unit, configured to obtain an initial feature map after the third information passes through all encoding layers;
[0171] A convolution unit, configured to perform spatial convolution processing on the initial feature map to obtain a third feature map;
[0172] A second normalization unit, configured to perform normalization processing on the third feature map, and perform four-way geometric rearrangement on the normalized third feature map to obtain a fourth feature map, where the four-way geometric rearrangement includes rearranging the forward direction, reverse direction, inter-channel direction, and inter-sample direction of the feature map;
[0173] A fusion unit, configured to perform fusion processing on the fourth feature map and the third feature map to obtain a fifth feature map;
[0174] A fully connected unit, configured to perform normalization and fully connected processing on the fifth feature map in sequence, and perform fusion processing on the processed fifth feature map and the fourth feature map to obtain a second interaction feature map.
[0175] It should be noted that regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0176] Embodiment 3:
[0177] Corresponding to the above method embodiment, a semi-supervised medical image segmentation device based on global geometric prior is further provided in this embodiment. A semi-supervised medical image segmentation device based on global geometric prior described below can be correspondingly referred to the semi-supervised medical image segmentation method based on global geometric prior described above.
[0178] Figure 7 It is a block diagram of a semi-supervised medical image segmentation device 800 based on global geometric prior shown according to an exemplary embodiment. As Figure 7As shown, the semi-supervised medical image segmentation device 800 based on global geometric prior may include: a processor 801 and a memory 802. The semi-supervised medical image segmentation device 800 based on global geometric prior may also include one or more of a multimedia component 803, an I / O interface 804, and a communication component 805.
[0179] Among them, the processor 801 is used to control the overall operation of the semi-supervised medical image segmentation device 800 based on global geometric prior to complete all or part of the steps in the above-mentioned semi-supervised medical image segmentation method based on global geometric prior. The memory 802 is used to store various types of data to support the operation of the semi-supervised medical image segmentation device 800 based on global geometric prior. These data may include, for example, instructions for any application or method operating on the semi-supervised medical image segmentation device 800 based on global geometric prior, as well as application-related data, such as contact data, received and sent messages, pictures, audio, video, and so on. The memory 802 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc. The multimedia component 803 may include a screen and an audio component. The screen may be, for example, a touch screen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone, and the microphone is used to receive external audio signals. The received audio signals may be further stored in the memory 802 or sent through the communication component 805. The audio component further includes at least one speaker for outputting audio signals. The I / O interface 804 provides an interface between the processor 801 and other interface modules, and the above-mentioned other interface modules may be a keyboard, a mouse, buttons, etc. These buttons may be virtual buttons or physical buttons. The communication component 805 is used for the semi-supervised medical image segmentation device 800 based on global geometric prior to communicate with other devices in a wired or wireless manner. Wireless communication, such as Wi-Fi, Bluetooth, near field communication (NFC), 2G, 3G, or 4G, or a combination of one or more of them. Accordingly, the communication component 805 may include: a Wi-Fi module, a Bluetooth module, and an NFC module.
[0180] In an exemplary embodiment, the semi-supervised medical image segmentation device 800 based on global geometric prior can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components, and is used to execute the above-mentioned semi-supervised medical image segmentation method based on global geometric prior.
[0181] In another exemplary embodiment, there is also provided a computer-readable storage medium including program instructions. When the program instructions are executed by a processor, the steps of the above-mentioned semi-supervised medical image segmentation method based on global geometric prior are implemented. For example, the computer-readable storage medium can be the above-mentioned memory 802 including program instructions, and the above program instructions can be executed by the processor 801 of the semi-supervised medical image segmentation device 800 based on global geometric prior to complete the above-mentioned semi-supervised medical image segmentation method based on global geometric prior.
[0182] Embodiment 4:
[0183] Corresponding to the above method embodiment, in this embodiment, there is also provided a readable storage medium. A readable storage medium described below can be correspondingly referred to with a semi-supervised medical image segmentation method based on global geometric prior described above.
[0184] A readable storage medium stores a computer program. When the computer program is executed by a processor, the steps of the semi-supervised medical image segmentation method based on global geometric prior in the above method embodiment are implemented.
[0185] The readable storage medium can specifically be various readable storage media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disc that can store program codes.
[0186] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
[0187] As mentioned above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or replacements, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A semi-supervised medical image segmentation method based on global geometric prior, characterized in that Including: Obtain first information and second information, where the first information is a medical image with a label, and the second information is a medical image without a label; Perform data perturbation processing on the second information to obtain third information; Input the first information, second information, and third information into a preset semi-supervised learning model, and construct an objective loss function based on a supervision loss, a consistency loss, and a geometric moment consistency loss. After the calculated value of the objective loss function meets the set conditions, stop optimizing the model parameters to obtain a semi-supervised image segmentation model. The geometric moment consistency loss is used to measure the degree of consistency of the model for the same image region in terms of symmetry, shape, and direction information; Input the medical image to be segmented into the semi-supervised image segmentation model to obtain a predicted segmentation result.
2. The semi-supervised medical image segmentation method based on global geometric prior according to claim 1, wherein , The data perturbation processing includes noise addition processing and image distortion processing. Performing data perturbation processing on the second information to obtain third information includes: Construct distorted coordinates based on a preset amplitude and frequency to obtain distorted grid coordinates; Map the second information to the distorted grid coordinates based on a three-dimensional interpolation algorithm to obtain the second information after image distortion processing; Generate a noise matrix based on a random number generator for a preset noise frequency, type, and distribution characteristics. The noise matrix has the same image size as the second information; Combine the noise matrix with the second information to obtain the second information after noise addition processing; Use the second information after image distortion processing and the second information after noise addition processing as the third information.
3. The semi-supervised medical image segmentation method based on global geometric prior according to claim 2, wherein , The semi-supervised learning model includes a student network and a teacher network. Inputting the first information, second information, and third information into a preset semi-supervised learning model and constructing an objective loss function based on a supervision loss, a consistency loss, and a geometric moment consistency loss includes: Input the first information and the second information into the student network to obtain multiple first feature maps, an annotated image probability map, and an unannotated image probability map; Input the second information into the teacher network to obtain multiple second feature maps and a perturbed probability map; Construct a geometric moment consistency loss function based on the first feature map and the second feature map; Construct a supervision loss function based on the annotated image probability map and the true label probability map; Construct a consistency loss function based on the unannotated image samples and the perturbed probability map; Construct the objective loss function by combining the consistency loss function, the supervision loss function, and the consistency loss function based on preset weight coefficients.
4. The semi-supervised medical image segmentation method based on global geometric prior according to claim 3, characterized in that , Both the student network and the teacher network include multiple encoding layers, an information interaction module, and multiple decoding layers. Constructing a geometric moment consistency loss function based on the first feature map and the second feature map includes: Input the first information and the second information into the student network, and process them sequentially through multiple encoding layers, the information interaction module, and multiple decoding layers to obtain multiple encoded feature maps, a first interaction feature map, and multiple first decoded feature maps; Input the third information into the teacher network, and process it through all the encoding layers and the information interaction module to obtain a second interaction feature map; The second interactive feature map is decoded through all decoding layers to obtain a second decoded feature map; Construct a first loss function based on multiple encoded feature maps, a first interactive feature map, and a second interactive feature map; Construct a second loss function based on multiple first decoded feature maps and a second decoded feature map; Obtain the geometric moment consistency loss function based on the first loss function and the second loss function.
5. The semi-supervised medical image segmentation method based on global geometric prior according to claim 4, wherein ,Constructing a first loss function based on multiple encoded feature maps and the first interactive feature map, and the second interactive feature map includes: Calculate the centroid position in the encoded feature map to obtain an image centroid; Calculate the second-order geometric moment corresponding to the encoded feature map based on the image centroid; Perform normalization processing on the second-order geometric moment to obtain a first normalized geometric moment; Based on the above steps of calculating the image centroid, second-order geometric moment, and normalization, process the first interactive feature map and the second interactive feature map respectively, and sequentially obtain a second normalized geometric moment and a third normalized geometric moment; Calculate the difference between each first normalized geometric moment and the third normalized geometric moment, and calculate the product of the difference and a preset weight coefficient to obtain multiple first products; Calculate the difference between the second normalized geometric moment and the third normalized geometric moment, and calculate the product of the difference and a preset weight coefficient to obtain a second product; Calculate the sum of multiple first products and second products to obtain the first loss function.
6. A semi-supervised medical image segmentation device based on global geometric prior, characterized in that Including: An acquisition unit for acquiring first information and second information, where the first information is a medical image with a label, and the second information is a medical image without a label; A perturbation unit for performing data perturbation processing on the second information to obtain third information; An optimization unit for inputting the first information, second information, and third information into a preset semi-supervised learning model, and constructing an objective loss function based on a supervision loss, a consistency loss, and a geometric moment consistency loss. When the calculated value of the objective loss function meets a set condition, stop optimizing the model parameters to obtain a semi-supervised image segmentation model. The geometric moment consistency loss is used to measure the consistency degree of the model for the same image region in terms of symmetry, shape, and direction information; A first input unit for inputting a medical image to be segmented into the semi-supervised image segmentation model to obtain a predicted segmentation result.
7. The semi-supervised medical image segmentation device based on global geometric prior according to claim 6, characterized in that, The perturbation unit includes: A distortion unit for constructing distorted coordinates based on a preset amplitude and frequency to obtain distorted grid coordinates; A mapping unit for mapping the second information into the distorted grid coordinates based on a three-dimensional interpolation algorithm to obtain second information after image distortion processing; A generation unit for generating a noise matrix based on a random number generator for a preset noise frequency, type, and distribution characteristics, where the noise matrix has the same size as the image in the second information; A combination unit for combining the noise matrix with the second information to obtain second information after noise addition processing; An as unit for using the second information after image distortion processing and the second information after noise addition processing as the third information.
8. The semi-supervised medical image segmentation device based on global geometric prior according to claim 7, characterized in that, The optimization unit includes: A second input unit for inputting the first information and the second information into the student network to obtain a plurality of first feature maps, an annotated image probability map, and an unannotated image probability map; A third input unit for inputting the second information into the teacher network to obtain a plurality of second feature maps and a perturbation probability map; A first construction unit for constructing a geometric moment consistency loss function based on the first feature map and the second feature map; A second construction unit for constructing a supervision loss function based on the annotated image probability map and the true label probability map; A third construction unit for constructing a consistency loss function based on the unannotated image samples and the perturbation probability map; A fourth construction unit for constructing the target loss function based on the consistency loss function, the supervision loss function, and the consistency loss function according to a preset weight coefficient; 9. The semi-supervised medical image segmentation device based on global geometric prior according to claim 8, wherein The first construction unit includes: A fourth input unit for inputting the first information and the second information into the student network, and successively processing them through a plurality of encoding layers, an information interaction module, and a plurality of decoding layers to respectively obtain a plurality of encoded feature maps, a first interaction feature map, and a plurality of first decoded feature maps; A fifth input unit for inputting the third information into the teacher network and processing it through all the encoding layers and the information interaction module to obtain a second interaction feature map; A decoding unit for decoding the second interaction feature map through all the decoding layers to obtain a second decoded feature map; A fifth construction unit for constructing a first loss function based on the plurality of encoded feature maps, the first interaction feature map, and the second interaction feature map; A sixth construction unit for constructing a second loss function based on the plurality of first decoded feature maps and the second decoded feature map; A first obtaining unit for obtaining the geometric moment consistency loss function based on the first loss function and the second loss function; 10. The semi-supervised medical image segmentation device based on global geometric prior according to claim 9, characterized in that, The fifth construction unit includes: A first calculation unit for calculating the centroid position in the encoded feature map to obtain an image centroid; An encoding unit for calculating the second-order geometric moment corresponding to the encoded feature map based on the image centroid; A first normalization unit for normalizing the second-order geometric moment to obtain a first normalized geometric moment; A second obtaining unit for processing the first interaction feature map and the second interaction feature map respectively according to the steps of calculating the image centroid, the second-order geometric moment, and normalizing them, and successively obtaining a second normalized geometric moment and a third normalized geometric moment; A second calculation unit for calculating the difference between each first normalized geometric moment and the third normalized geometric moment, and calculating the product of the difference and a preset weight coefficient to obtain a plurality of first products; A third calculation unit for calculating the difference between the second normalized geometric moment and the third normalized geometric moment, and calculating the product of the difference and a preset weight coefficient to obtain a second product; A fourth calculation unit for calculating the sum of the plurality of first products and the second product to obtain the first loss function.
Citation Information
Cited By
End-to-end point cloud denoising model construction method and system
CN121328618A