A small-sample segmentation method for medical images based on transduction inference

Through the medical image small sample segmentation method based on transduction reasoning, the transduction reasoning module is used to fuse the information of the support set and the query set to extract prototype features, solving the problem of dependence on a large amount of data in the existing technology, achieving high-precision medical image segmentation, and reducing R&D costs.

CN115937232BActive Publication Date: 2025-05-27XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211738213.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-31
Publication Date
2025-05-27
Estimated Expiration
2042-12-31

AI Technical Summary

Technical Problem

The existing medical image segmentation method based on deep learning requires a large amount of data for training, and it is difficult to effectively utilize the information of the query set itself, ignoring the distribution difference between the support set and the query set, resulting in limited segmentation effect.

Method used

The medical image small sample segmentation method based on transduction reasoning is used to extract image features through neural networks, and the transduction reasoning module is used to fuse the information of the support set and query set, extract prototype features, and segment the unlabeled query set.

Benefits of technology

It improves the accuracy of medical image segmentation, and can accurately segment based on a small number of labeled samples when the neural network is not learning for specific tasks, reducing dependence on massive data and reducing R&D costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115937232B_ABST
    Figure CN115937232B_ABST
Patent Text Reader

Abstract

The present invention discloses a small-sample segmentation method for medical images based on transduction inference, comprising the following steps: (1) Using a neural network as a feature extraction network to extract feature maps of two-dimensional medical images or three-dimensional image slices; (2) Based on transduction inference, performing fusion analysis on the support set data and the query set data, and respectively extracting prototype feature vectors corresponding to the foreground class and prototype feature vectors corresponding to the background class for segmenting the query set; (3) For each feature in the feature map of the query set, calculating its similarity with the prototype features of each class respectively, and then by comparing the magnitudes of the similarities of each class, determining the class corresponding to each feature in the query set, that is, its class is the class corresponding to the prototype feature with the maximum similarity. The present invention can improve the segmentation effect and maximize the accuracy of the segmentation result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of artificial intelligence, computer vision and medical image analysis, and particularly relates to a medical image small-sample segmentation method based on transduction reasoning. Background Art

[0002] With the development of medical imaging technologies, especially the emergence of imaging means such as CT and MRI, it has greatly promoted the progress of medical diagnosis and treatment means, and also greatly promoted the development of medical image automated analysis algorithms. Among them, computer vision algorithms based on deep learning perform outstandingly in tasks related to medical image classification, detection, segmentation, etc., and can provide assistance to doctors or even achieve automated diagnosis algorithms, significantly improving the modern medical level. However, the common medical image segmentation method based on deep learning, "U-Net: Convolutional Networks for Biomedical Image Segmentation", usually requires a large amount of data for learning to segment lesions or organs within a limited range, which makes it difficult for U-Net to be applied in actual scenarios. At the same time, the imaging data itself is restricted by privacy, ethics, etc., and annotation requires professional physicians to spend a lot of time for annotation. These two restrictions make the research and promotion of traditional medical image segmentation algorithms face great difficulties.

[0003] As the inventor understands, for the problem that current medical image segmentation algorithms usually require a large amount of data for training, although researchers such as Ouyang have proposed medical image segmentation methods based on small samples, such as "Self-supervision with Superpixels: Training Few-Shot Medical Image Segmentation Without Annotation" as solutions, the existing methods only extract task-critical information from the support set for the segmentation task, and cannot effectively utilize the information of the query set itself, while ignoring the distribution difference between the support set and the query set, resulting in limited segmentation effects and requiring more support data. Summary of the Invention

[0004] In order to overcome the deficiencies of the above-mentioned prior art, the purpose of the present invention is to provide a medical image small-sample segmentation method based on transduction reasoning. By analyzing the data features of the support set and the query set through transduction reasoning, more effective prototype features are extracted to segment the areas involved in the unlearned segmentation tasks, which can improve the segmentation effect and maximize the accuracy of the segmentation result.

[0005] In order to achieve the above purpose, the technical solution adopted by the present invention is:

[0006] A small-sample medical image segmentation method based on transduction inference, where the support set data and the query set data are used as inputs. The support set is image data with segmentation annotations, and the query set is image data without segmentation. The segmentation method includes the following steps;

[0007] (1) Use a neural network as a feature extraction network to extract the feature map z of the two-dimensional medical image or three-dimensional image slice x;

[0008] (2) Based on transduction inference, perform fusion analysis on the support set data and the query set data, and respectively extract the prototype feature vectors corresponding to the foreground classes and the prototype feature vectors corresponding to the background classes for segmenting the query set;

[0009] (2a) Adopt an annotation information embedding network, and use the segmentation annotation y corresponding to the image as an input to construct the annotation information v of the image;

[0010] (2b) Use an information fusion network, and use the feature map z corresponding to the image and the annotation information v constructed through the image annotation as inputs to fuse these two kinds of information to form the fused information u;

[0011] (2c) Through the transduction inference module, take the fused information u of the support set and the query set S and u Q as inputs, and perform fusion and learning through the transduction inference module, and then complete the input fused information u of the query set Q and generate new fused information of the query set

[0012] (2d) Adopt a solving network to process the fused information output by the transduction inference module to decode the potential query set annotation information v therein Q ;

[0013] (2e) Use a prototype selection network to fuse and analyze the feature map z and the annotation information v of the support set and the query set, and select the prototype features z of each category for segmenting the query set from them P ;

[0014] (3) For each feature in the feature map z Q of the query set, calculate its similarity with the prototype features of each category respectively, and then determine the category corresponding to each feature in the query set by comparing the sizes of the similarities s i of each category, that is, its category is the category corresponding to the prototype feature with the largest similarity.

[0015] In the above steps, in the definitions of various symbols, the superscript S represents the data information corresponding to the support set, and the superscript Q represents the data information corresponding to the query set. For example, if z represents a feature map, then z Q represents the feature map of the query set, and z S b represents the feature map of the support set.

[0016] The feature extraction network F mentioned in step (1) feat is composed of multiple layers of neural or multi-layer convolutional neural networks. The input is a single-channel two-dimensional image x, and the output is a high-dimensional feature

[0017] The annotation information embedding network T mentioned in step (2a) anno The specific formula is as follows.

[0018]

[0019] Among them, y is the network input, representing the segmentation annotation corresponding to the image; v is the network output, representing the annotation information constructed by the annotation information embedding network;

[0020] reshape means to combine the two dimensions of length and width in the input data into one dimension, and reshape′ means to split the dimension combined by length and width in the input data into two dimensions of length and width;

[0021] E anno (·) represents a phrase embedding network. By means of a sliding window, the annotation information represented by each pixel in the annotation y of the image and its adjacent area is converted into a representation form based on feature vectors;

[0022] F anno (·) is a multi-layer convolutional neural network, which is used to further extract the features hidden in the vectorized annotation and mine more potential information;

[0023] G anno (·) represents a multi-layer fully connected neural network, which is used to fuse the annotation information from a global perspective;

[0024] The annotation information of the support set image is generated from the annotation corresponding to the support set image. For the query set data without annotation information, blank is used as the input to generate annotation information. Here, blank is a pseudo-annotation that represents neither the foreground nor the background.

[0025] The information fusion network T in step (2b) fuse The calculation formula is as follows.

[0026] u = T fuse (z, v) = Trans(encode = v, decode = z)

[0027] Among them, Trans is a Transformer neural network including an encoder and a decoder. The input of the encoder is the annotation information v, the input of the decoder is the feature map z corresponding to the image, and the output of the network is the result of the fusion of the feature map of the image and the annotation information, that is, the fusion information u.

[0028] The expression and calculation steps of the transduction inference module in step (2c) are as follows;

[0029]

[0030] First, use the operation of splicing along the length and width dimensions to splice the support set fusion feature u S and the query set fusion feature u Q to obtain At the same time, in the same way, splice the annotation information v of the support set S and the blank annotation information of the query set to obtain

[0031] Then, use the Transformer neural network (Trans) including an encoder and a decoder to perform transduction inference on the spliced fusion information u and annotation information v to complete the annotation information v in the query set Q , where the input of the encoder is the input of the decoder is and the output is the updated support set and query set fusion information;

[0032] Finally, through the method of masking, extract the part corresponding to the query set in the updated fusion information .

[0033] The calculation process of the solution network in step (2d) is as follows:

[0034] First, use the solution network to solve the updated query set fusion information , where the solution network T resolve (·) is a multi-layer convolutional neural network;

[0035] Then, optionally, bring the result v Q back into the calculation formula of the transduction inference module again and replace to perform the calculation to obtain a new and repeat the above process until the stop condition is reached, that is, the number of loops reaches the upper limit or the query set fusion information no longer updates;

[0036] After one or more transduction inferences and solutions, the query set annotation information v solved by the solution networkQ , as the final output result of the solution network.

[0037] The prototype selection network expression of the step (2e) is z P = T proto (z S , v S , z Q , v Q ), and its calculation process is as follows.

[0038] First, calculate the scores i' corresponding to all query set and support set features:

[0039]

[0040] Among them, Trans is a Transformer neural network including an encoder and a decoder. The input of the encoder is the feature map z, and the input of the decoder is the annotation information v. F i is a multi-layer fully connected neural network, and the output result i is the score of each feature;

[0041] Then, normalize the scores of all features through the softmax function so that the sum of all scores is 1

[0042] i = softmax(i')

[0043] Finally, according to the scores, sort and select the N features with the largest scores as the prototype features

[0044]

[0045] Among them, top(i, N) represents the index values corresponding to the N vectors with the largest scores in i.

[0046] In the step (3), when there are multiple prototype features for segmentation in a certain category, a normalization method is used for fusion, or for the segmentation score s of this type:

[0047] s = sum(s Q ·softmax(s Q ))

[0048] At the same time, when comparing similarities to judge pixel categories, a bias-softmax function is used for normalization to determine the category c corresponding to the pixel:

[0049] c = argmax(bias_softmax(s Q ; b))

[0050] = argmax(softmax([s 0 , s 1 + b1 ,…,s n +b n ))

[0051] where b = [b 1 , b 2 ,…, b n is the bias, which is used as a hyperparameter that can be artificially adjusted during the segmentation process to obtain more optimized results; s Q = [s 0 , s 1 , s 2 ,…, s n represents the similarity between the features in the query set feature map and the prototype features corresponding to n categories.

[0052] Advantages of the present invention:

[0053] First, based on the few-shot segmentation method, the present invention segments medical images, which can achieve accurate segmentation of unlabeled samples (i.e., the query set) according to a small number of labeled samples (i.e., the support set) when the neural network has not learned for a specific task, making full use of the characteristic that the same segmentation region of different individuals in medical images has relatively small differences. This avoids the disadvantages of traditional deep learning-based medical image segmentation algorithms that need to first learn relevant segmentation tasks through a large amount of data before they can complete related tasks. The acquisition of a large amount of medical image data usually requires a lot of manpower and material resources for image acquisition, data annotation, etc., and is subject to privacy restrictions and ethical constraints, resulting in high R & D costs and long cycles for related products. Therefore, the present invention can greatly reduce the cost and threshold of R & D of deep learning-based medical image segmentation algorithms.

[0054] Second, in the few-shot segmentation algorithm of the present invention, transduction inference is used to analyze the support set and the query set, and prototype features are extracted for segmenting the query set. Compared with traditional few-shot methods, it effectively integrates the information of the support set and the query set, deeply mines the query set data, effectively utilizes the unlabeled query set data, and avoids segmentation errors caused by distribution differences between the support set and the query set, improving the segmentation effect.

[0055] Third, in the few-shot segmentation algorithm of the present invention, multiple prototype features generated by transduction inference are used to segment the query set. In a differential way, appropriate prototype features are selected for different regions in the foreground and background, and pixels are classified according to the similarity, thereby improving the overall segmentation accuracy. At the same time, the bias-normalized exponential method is adopted to reconcile the magnitude relationship of the similarity between pixels and prototype features of different categories. In practical applications, users can set the bias according to specific situations to maximize the accuracy of the segmentation result. Description of the Drawings

[0056] Figure 1 It is the overall framework of a small-sample segmentation algorithm for medical images based on transduction inference.

[0057] Figure 2 It is a schematic diagram of the DenseNet and U-Net networks.

[0058] Figure 3 It is a schematic diagram of the process for extracting prototype features based on transduction inference.

[0059] Figure 4 It is a schematic diagram of the annotation information embedding network.

[0060] Figure 5 It is a schematic diagram of the information fusion network.

[0061] Figure 6 It is a schematic diagram of the transduction inference module.

[0062] Figure 7 It is a schematic diagram of the prototype selection network.

[0063] Figure 8 It is a schematic diagram of prototype segmentation. Detailed implementation manners

[0064] The present invention will be further described in detail below with reference to the accompanying drawings.

[0065] Refer to the attached Figure 1 A method for small-sample segmentation of medical images based on transduction inference according to the present invention generally includes: ① feature map extraction, ② prototype feature extraction based on transduction inference, and ③ prototype segmentation algorithm.

[0066] Step 1: Use a variety of neural networks including the Dense-Net and U-Net shown in the attached Figure 2 as the feature extraction network to extract the feature map z of the two-dimensional medical image or three-dimensional image slice x. These neural networks use neural networks such as convolutional neural networks, fully connected neural networks, or Transformer networks, and use the two-dimensional medical image or three-dimensional slice x as the input. The network usually includes multiple downsampling modules based on average pooling or max pooling, and an upsampling module is added behind these networks to keep the size of the feature map. So when the input image is x ∈ R H×W , the output feature map is z ∈ R C×H′×W′ , where C is the number of channels. In this embodiment, the U-Net shown in the attached Figure 2 is used as the feature extraction network, and its output number of channels is 128. H′ and W′ are usually H / 4 and W / 4.

[0067] Generally speaking, when the size of the feature map output by the network is smaller than the size required by the transductive inference module, interpolation algorithms are used to upsample the feature map, or a transposed convolution module is used to enlarge the size of the feature map, or the downsampling module in the feature extraction network is deleted.

[0068] In this embodiment, for the convenience of explanation, the number of channels of the feature map is set to 128, and the length and width are set to 256, which is the same as the original input image. The upsample module is used to transform the size of the network output.

[0069] Generally, there is no limit to the number of images contained in the support set and the query set. In this embodiment, only one two-dimensional medical image is included as the support set and one two-dimensional medical image is included as the query set.

[0070] Step 2: The process of analyzing and fusing the support set and query set data based on transductive inference and extracting the prototype features for query set segmentation is as shown in the appendix Figure 3 shown. In this embodiment, the method of the present invention is described by a binary classification task. The prototype features include foreground prototype features and background prototype features The prototype features can contain a single vector or multiple vectors. Therefore, the vector shape size of the prototype features is z P ∈R N×C , where N represents the number of vectors. The specific calculation process of the transductive inference module includes:

[0071] Step 2a: As shown in the appendix Figure 3 shown, the first step is to construct the annotation information v by using the annotation information embedding network T anno , with the corresponding annotation information y of the image as the input. The shape size of the annotation information is usually y ∈ R (C-1)×H×W where C represents the number of segmentation categories. In this embodiment, the segmentation categories include foreground and background, so C = 2. Then the shape size of the output annotation information is v ∈ R C′×H×W , where C′ is the number of feature channels corresponding to the generated annotation information. In this embodiment, C′ = 128. Its calculation formula is as follows:

[0072]

[0073] Specifically, as shown in the appendix Figure 4 shown, first, through the phrase embedding network E anno (·), the annotation y ∈ R of the image 1 ×256×256, in a windowed manner, extract the surrounding 7×7 region corresponding to each pixel's annotation, and through the method of embedding encoding, encode the 1×7×7 matrix into a 128-dimensional vector. For a 256×256 annotation, a high-dimensional matrix of 128×256×256 will be obtained.

[0074] Then, use the multi-layer convolutional network F anno (·) to further extract features from the extracted features. In this embodiment, F anno consists of two cascaded ResNet basic modules, and the ReLU function is used as the activation function. The size of the input data is 128×256×256, and the size of the output data remains 128×256×256.

[0075] Finally, through the reshape operation, convert the input data from 128×256×256 to 128×65536, and use the two-layer fully connected neural network G anno to further extract features. The number of neurons in the hidden layer of the fully connected network is 2048. And the output data is converted from the 128×65536 size to 128×256×256 by the reshape' operation and used as the annotation information v.

[0076] For the support set with annotations, the annotation information v S is generated by the above steps. However, for the query set without annotations, its corresponding empty table annotation information is first generated by E anno to generate blank embedding information that does not belong to any category, and then through F anno and G anno process the blank embedding information to obtain the empty table annotation information

[0077] Step 2b: As shown in the appendix Figure 3 , after converting the image annotation y into the annotation information v, it is necessary to use the annotation information v and the feature map z of the image as the input, and through the information fusion network T fuse for fusion to further process through transduction inference. The calculation formula of the information fusion network T fuse is as follows:

[0078] u = T fuse (z, v) = Trans(encode = v, decode = z)

[0079] where as Figure 5As shown, Trans is a Transformer network including an encoder and a decoder. The input of the encoder is the annotation information v, and the input of the decoder is the feature map z. Here, both the annotation information and the feature map are three-dimensional matrices. To be the input of the Transformer network, first, the two dimensions of length and width in the annotation information v and the feature map z are fused. In this embodiment, the annotation information and the feature map with the shape of 128×256×256 are fused in the length and width dimensions to obtain a matrix with the shape of 128×65536 as the input. Then, in order to retain the position information of the annotation information v and the feature map z before dimension fusion, a position information encoder is used to add position information to them.

[0080] Next, the annotation information is encoded by the encoder and fused with the feature map passing through the decoder. The encoding includes a multi-head attention mechanism module, a feed-forward neural network, an addition operator, and a regularization operator; the decoder includes the same components. Different from the encoder, the decoder includes two multi-head attention mechanism modules. The first module is the same as the attention mechanism module of the encoder, taking the annotation information v or the feature map z as its three inputs of "query", "key", and "value". The second module takes the output of the encoder as the two inputs of "key" and "value", and the "query" input is the output of the previous layer attention module.

[0081] Finally, the output of the decoder is processed by a multi-layer perceptron (MLP) to obtain the fused information u.

[0082] Step 2c: As shown in the appendix Figure 3 After obtaining the fused information u, use the transduction inference module T trans to fuse the support set information with the annotation of the query set. The calculation formula is as follows:

[0083]

[0084] where, as Figure 6 shown, Trans is a Transformer network including multiple layers of encoders and decoders. For the encoder, first, the support set fused information u S and the query set fused information u Q are fused in the two dimensions of length and width. In this embodiment, the fused information with the size of 128×256×256 is transformed into a matrix with the size of 128×65536; then the fused information of the support set and the query set is concatenated in the length and width dimensions to obtain a matrix with the size of 128×131072 as the input of the encoder.

[0085] For the decoder, in the same way as the encoder, the annotation information v of the support set S and the blank annotation The information is fused. In this embodiment, the annotation information with a size of 128×256×256 is converted into a matrix with a size of 128×65536. Then, the two annotation information are concatenated in the length and width dimensions to obtain a matrix with a size of 128×131072 as the input of the decoder.

[0086] In the Transformers network of the transduction inference module, the structures of the encoder and the decoder are the same as those of the Transformers structure used in the fusion network T Figure 5 in the appendix fuse And similarly, in this embodiment, the output of the last encoder (3) will be used as the input of the "key" and "value" in the input of the other three decoders.

[0087] Finally, the result output by the decoder (3) is subjected to further feature fusion through a multi-layer perceptron as the output to update the fused information. The updated fused information here contains the information of the support set and the query set, and the updated query set fusion information needs to be updated by means of a mask.

[0088] Step 2d: As shown in the appendix Figure 3 , after obtaining the updated query set fusion information , use the resolution network T resolve to extract the hidden query set annotation information v Q in the updated fusion information, and its calculation formula is as follows:

[0089]

[0090] First, use the resolution network to resolve the updated query set fusion information v Q , where the resolution network T resolve (·) is a multi-layer convolutional neural network. In this embodiment, ResNetBlock is used as the multi-layer convolutional network for the resolution operation.

[0091] Optionally, the resolved query set annotation information v Q is brought back to the transduction inference module in step 2c again to replace and perform transduction inference again to obtain a new The above steps can be repeated until the stop condition is met. In this embodiment, the stop condition includes two:

[0092] ① The similarity between the currently resolved query set annotation information and the query set annotation information resolved in the previous iteration is greater than the threshold t resolve =0.9:

[0093]

[0094] ② When the number of iterations is more than 10, stop the iteration.

[0095] Finally, after single or multiple transduction inferences and calculations, the final query set annotation information v calculated by the calculation network Q , as the final output result of the settlement network.

[0096] Step 2e: As shown in the appendix Figure 3 , after obtaining the updated query set annotation information v Q , use the prototype selection network T proto to select appropriate prototype features z from the support set and query set feature maps P , for the segmentation task of the query set. The calculation process of the prototype selection network is as follows:

[0097] As shown in the appendix Figure 7 , first, calculate the scores corresponding to the features in all support set and query set feature maps:

[0098]

[0099] Among them, Trans is a Transformer neural network including an encoder and a decoder. The input of the encoder is the feature maps of the support set and the query set. The feature maps first merge the two channels of length and width, and are converted from a matrix with a shape size of C×H×W to C×HW, and the query set and the support set are concatenated; the input of the decoder is the annotation information v. Similarly, the two dimensions of length and width of the annotation information are merged, and then the query set and the support set annotation information are concatenated. In this embodiment, both the feature map and the annotation information are matrices with a size of 128×256×256, which are converted into matrices with a size of 128×65536, and then the query set and the support set are concatenated to obtain a matrix with a size of 128×131072 as the input. The output of the Transformer network is still 128×131072 in size.

[0100] Then, F i is a multi-layer fully connected neural network used to reduce the information output by the Transformer network and regress it to the original score i' for each feature. In this embodiment, F i is a two-layer fully connected network. The number of neurons in the first layer is 512, the number of neurons in the second layer is 1, and no activation function is added. In addition, the shape size of the original score is i'∈R 131072 , corresponding to 131072 features here.

[0101] Next, the normalized exponential function is used to process the original scores so that the sum of the scores corresponding to all features is 1:

[0102] i = softmax(i′)

[0103] Finally, the normalized scores are sorted from high to low, and the top N prototype features with the highest scores are selected:

[0104]

[0105] In this embodiment, for the foreground and background, 10 prototype features are respectively extracted for the segmentation task. For the foreground and background, two independent fully connected networks F i+ and F i- are respectively instantiated to score the features corresponding to the foreground and background. The inputs used by these two networks are outputs of the same Transformer network. In addition, for the binary classification segmentation task, the top N and bottom N vectors with the highest scores can be respectively selected as prototype features.

[0106] Step 3, as shown in the appendix Figure 1 After the prototype features are selected by transduction inference, the query set is segmented by prototype segmentation. As shown in the appendix Figure 8 As shown, the segmentation process is as follows: First, for each feature Q in the query set feature map z and each prototype feature the similarity s ij is calculated. As shown in the appendix Figure 8 In this embodiment, the cosine trigonometric function is used as the measurement method for the similarity of two vectors:

[0107] Then, for the query set feature map z Q for each pixel, the similarities with multiple prototype features z P of the same category are fused in a normalized weight manner to obtain the overall similarity of the current pixel to a single category:

[0108] s = sum(s Q ·softmax(s Q ))

[0109] For example, for a specific category, there are prototype features 1 and 2, and the query set feature map z QCalculate the similarity through the cosine trigonometric function and obtain two similarity maps. Then, in each similarity map, normalize the similarity of each pixel to the corresponding two prototype features through the normalization-exponential function as the weight, multiply it by the similarity itself, and then add the results corresponding to the two prototype features as the similarity of the pixel;

[0110] Finally, perform normalization using the bias softmax function, compare the magnitude relationship between the similarities of different categories corresponding to the features, and take the category with the largest similarity as the category corresponding to the feature, that is, the corresponding category in the image. In this embodiment, the categories include foreground and background. Therefore, perform normalization for the foreground and background:

[0111] c = argmax(bias_softmax(s Q ; b)) = argmax(softmax([s 0 , s 1 +

[0112] b 1 , …, s n + b n )) where b = [b 1 , b 2 , …, b n is the bias, which is used as a hyperparameter that can be manually adjusted during the segmentation process to obtain a more optimized result; s Q = [s 0 , s 1 , s 2 , …, s n represents the similarity of the features in the query set feature map to the prototype features corresponding to n categories.

Claims

1. A small-sample segmentation method for medical images based on transduction inference, characterized in that, where the support set data and the query set data are used as inputs, the support set is image data with segmentation annotations, and the query set is image data to be segmented. The segmentation method includes the following steps; (1) Use a neural network as a feature extraction network to extract the feature map z of the two-dimensional medical image or three-dimensional image slice x; (2) Based on transduction inference, perform fusion analysis on the support set data and the query set data, and respectively extract the prototype feature vectors corresponding to the foreground class and the prototype feature vectors corresponding to the background class for segmenting the query set from them; (2a) Adopt an annotation information embedding network, and use the segmentation annotation y corresponding to the image as an input to construct the annotation information v of the image; (2b) Use an information fusion network, and use the feature map z corresponding to the image and the annotation information v constructed through the image annotation as inputs to fuse these two kinds of information to form the fusion information u; (2c) Through the transduction inference module, fuse the information u of the support set and the query set S with u Q as input, and perform fusion and learning through the transduction inference module, and then complete the fused information u of the input query set Q and generate new fused information of the query set (2d) Use a solution network to process the fused information output by the transduction inference module to decode the potential query set annotation information v therein Q ; (2e) Use the prototype selection network to fuse and analyze the feature maps z and annotation information v of the support set and the query set, and select the prototype features z of each category for segmenting the query set from them. P ; (3) For each feature in the feature map z of the query set Q calculate its similarity with the prototype features of each category respectively, and then determine the category corresponding to each feature in the query set by comparing the sizes of the similarities s i of each category, that is, its category is the category corresponding to the prototype feature with the largest similarity.

2. The small-sample segmentation method for medical images based on transduction inference according to claim 1, characterized in that, The feature extraction network F described in the step (1) feat is composed of multiple layers of neural or multi-layer convolutional neural networks, with a single-channel two-dimensional image x as the input and high-dimensional features as the output 3. The small-sample segmentation method for medical images based on transduction inference according to claim 1, characterized in that, The annotation information embedding network T described in the step (2a) anno The specific formula is as follows v = T anno (y) = reshape’(G anno (reshape(F anno (E anno (y))))) where, y is the network input, representing the segmentation annotation corresponding to the image; v is the network output, representing the annotation information constructed by the annotation information embedding network; reshape means merging the length and width dimensions in the input data into one dimension, and reshape′ means splitting the dimension merged by the length and width in the input data into the length and width dimensions; E anno (·) represents a phrase embedding network, which, by means of a sliding window, converts the annotation information represented by each pixel and its adjacent area in the annotation y of the image into a representation form based on feature vectors; F anno (·) is a multi-layer convolutional neural network, which is used to further extract the features implicit in the vectorized representation of the annotation and mine more potential information; G anno (·) represents a multi-layer fully connected neural network for fusing annotation information globally; The annotation information of the support set image is generated from the annotation corresponding to the support set image. For the query set data without annotation information, blank is used as the input to generate the annotation information, where blank is a pseudo-annotation that represents neither the foreground nor the background.

4. The small-sample segmentation method for medical images based on transduction inference according to claim 1, characterized in that, The information fusion network T in the step (2b) fuse has the following calculation formula u = T fuse (z, v) = Trans(encode = v, decode = z) where, Trans is a Transformer neural network including an encoder and a decoder. The input of the encoder is the annotation information v, the input of the decoder is the feature map z corresponding to the image, and the output of the network is the result of the fusion of the feature map of the image and the annotation information, that is, the fusion information u.

5. The small-sample segmentation method for medical images based on transduction inference according to claim 1, characterized in that, The expression and calculation steps of the transduction inference module in step (2c) are as follows; First, perform an operation of splicing along the length and width dimensions Fuse the support set feature u S With the query set feature u Q Perform splicing to obtain At the same time, in the same way, the annotation information v of the support set S With the blank annotation information of the query set Perform splicing to obtain Then, a Transformer neural network (Trans) including an encoder and a decoder is used to perform transduction inference on the spliced fusion information u and the annotation information v to complete the annotation information v in the query set Q , where the input of the encoder is The input of the decoder is The output is the updated support set and the query set fusion information; Finally, in the way of masking, the part corresponding to the query set in the updated fusion information is extracted.

6. The small-sample segmentation method for medical images based on transduction inference according to claim 1, characterized in that, The calculation process of the solution network in step (2d) is as follows: First, use the solution network to fuse information for the updated query set for solution, where the solution network T resolve (·) is a multi-layer convolutional neural network; Then, optionally, the result v Q is brought into the calculation formula of the transduction reasoning module again and replaces for calculation to obtain a new and the above process is repeated until the stopping condition is reached, that is, the number of loops reaches the upper limit or the fusion information of the query set no longer updates; After one or more transduction inferences and calculations, the query set annotation information v calculated by the calculation network Q is used as the final output result of the calculation network.

7. The small-sample segmentation method for medical images based on transduction inference according to claim 1, characterized in that, The prototype selection network expression of the step (2e) is z P = T proto (z S , v S , z Q , v Q ), and its calculation process is as follows First, calculate the scores i′ corresponding to all query set and support set features: Among them, Trans is a Transformer neural network including an encoder and a decoder. The input of the encoder is the feature map z, and the input of the decoder is the annotation information v. F i is a multi-layer fully connected neural network, and the output result i is the score of each feature; Then, normalize the scores of all features through the softmax function so that the sum of all scores is 1 i = softmax(i′) Finally, according to the scores, sort and select the N features with the largest scores as the prototype features Among them, top(i, N) represents the index values corresponding to the N vectors with the largest scores in i.

8. A small-sample segmentation method for medical images based on transduction inference according to claim 1, characterized in that in step (3), when there are multiple prototype features for segmentation in a certain category, a normalization method is used for fusion, or for the segmentation score s of this type: s = sum(s Q · softmax(s Q )) At the same time, when comparing similarities to determine pixel categories, a bias-normalized exponential function is used for normalization to determine the category c corresponding to the pixel: c = argmax(bias_softmax(s Q ; b)) = argmax(softmax([s 0 , s 1 + b 1 , …, s n + b n )) where b = [b 1 , b 2 , …, b n is the bias, which is used as a hyperparameter that can be manually adjusted during the segmentation process to obtain more optimized results; s Q = [s 0 , s 1 , s 2 , …, s n represents the similarity between the features in the query set feature map and the prototype features corresponding to n classes.

Citation Information

Patent Citations

  • Multi-organ segmentation method based on self-supervised feature small sample learning

    CN113706487A

  • Device for detecting an edge using segmentation information and method thereof

    US20220262006A1