A method, device, storage medium and electronic device for segmenting medical images

By using the combination method of encoder, measurement network and memory network in medical image processing, the similarity characteristics of images are extracted and learned, and the accuracy of medical image segmentation under small sample data is solved, and the efficient image segmentation effect is achieved.

CN115393369BActive Publication Date: 2025-06-13BOE TECHNOLOGY GROUP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211055254.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-31
Publication Date
2025-06-13
Estimated Expiration
2042-08-31

AI Technical Summary

Technical Problem

Under the condition of small amount of labeled data, how to accurately segment medical images is a problem that needs to be solved urgently.

Method used

Feature extraction of support graphs and query graphs is performed through the encoder, and the metric network and memory network are used to learn based on similarity features, forward and backward feature maps are output, and upsampled through the decoder to achieve accurate segmentation of medical images.

Benefits of technology

This method improves the segmentation accuracy of medical images under small sample data, enhances the information interaction ability between adjacent images, and realizes accurate segmentation of medical images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115393369B_ABST
    Figure CN115393369B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method, apparatus, storage medium, and electronic device for segmenting medical images. The method includes: extracting features of M support graphs and M query graphs in N dimensions respectively; determining M first feature vectors, and determining M similarity matrices based on the M first feature vectors and the M second type of feature vector matrices in the Nth dimension; outputting M forward feature map matrices and M backward feature map matrices; respectively splicing the M forward feature map matrices and the M backward feature map matrices and performing upsampling processing in N dimensions to obtain segmentation maps corresponding to the M query graphs. The present disclosure inputs the support graphs and query graphs into the encoder simultaneously for feature extraction, and uses similarity features to improve the segmentation prediction effect. Further, a memory module is combined to learn the similarity features between adjacent images, enhance the information interaction ability between adjacent images, and thus achieve accurate segmentation of medical images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of medical image processing, and in particular, to a method, apparatus, storage medium, and electronic device for segmenting medical images. Background Art

[0002] With the development of computer vision technology, computer vision technology is used to assist in lesion localization based on medical images (i.e., semantic segmentation of medical images) to reduce the workload of doctors and improve efficiency. For semantic segmentation of medical images, semantic segmentation methods based on deep learning are usually adopted.

[0003] Generally, deep learning requires a large amount of labeled data. However, due to the high sensitivity of data in the medical industry, the amount of labeled data is very limited. Therefore, how to accurately perform semantic segmentation on medical images under the condition of less labeled data is an urgent problem to be solved. Summary of the Invention

[0004] The purpose of the embodiments of the present disclosure is to provide a method, apparatus, storage medium, and electronic device for segmenting medical images, so as to solve the problem of how to accurately perform semantic segmentation on medical images under the condition of less labeled data in the prior art.

[0005] The embodiments of the present disclosure adopt the following technical solutions: A method for segmenting medical images, including: extracting features of M support graphs and M query graphs in N dimensions respectively through an encoder to obtain a first type of feature vector matrix of each support graph in N dimensions and a second type of feature vector matrix of each query graph in N dimensions, where the M query graphs are M consecutive medical images to be processed; determining M first feature vectors based on the labeled graph corresponding to each support graph and the first type of feature vector matrix of each support graph in the Nth dimension through a metric network, and determining M similarity matrices based on the M first feature vectors and the M second type of feature vector matrices in the Nth dimension; learning the M similarity matrices through a memory network, and outputting M forward feature map matrices and M backward feature map matrices; splicing the M forward feature map matrices and the M backward feature map matrices respectively through a decoder and performing upsampling processing in N dimensions to obtain segmentation maps corresponding to the M query graphs.

[0006] The present disclosure also provides a medical image segmentation device, including: an encoding module, configured to perform feature extraction on M support graphs and M query graphs respectively in N dimensions to obtain a first type of feature vector matrix of each support graph in N dimensions, and a second type of feature vector matrix of each query graph in N dimensions, where the M query graphs are M consecutive medical images to be processed; a metric module, configured to determine M first feature vectors based on the labeled graph corresponding to each support graph and the first type of feature vector matrix of each support graph in the Nth dimension, and determine M similarity matrices based on the M first feature vectors and the M second type of feature vector matrices in the Nth dimension; a memory module, configured to learn the M similarity matrices and output M forward feature map matrices and M backward feature map matrices; a decoding module, configured to perform upsampling processing on the concatenated M forward feature map matrices and M backward feature map matrices respectively in N dimensions to obtain segmentation maps corresponding to the M query graphs.

[0007] The present disclosure also provides a storage medium storing a computer program, which when executed by a processor implements the steps of the above-mentioned medical image segmentation method.

[0008] The present disclosure also provides an electronic device, including at least a memory and a processor, where the memory stores a computer program, and the processor implements the steps of the above-mentioned medical image segmentation method when executing the computer program on the memory.

[0009] The beneficial effects of the embodiments of the present disclosure are as follows: The support graphs and query graphs are simultaneously input into the encoder for feature extraction, and the segmentation prediction effect is improved by using similarity features. Further, a memory module is combined to learn the similarity features between adjacent images, enhance the information interaction ability between adjacent images, and thus achieve accurate segmentation of medical images. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0011] Figure 1 It is a flowchart of the medical image segmentation method according to the first embodiment of the present disclosure;

[0012] Figure 2 It is a schematic diagram of the architecture of the segmentation network according to the first embodiment of the present disclosure;

[0013] Figure 3 It is a schematic diagram of the training process of the segmentation network in the first embodiment of the present disclosure;

[0014] Figure 4 It is a schematic diagram of the structure of the medical image segmentation device in the second embodiment of the present disclosure;

[0015] Figure 5 It is another schematic diagram of the structure of the medical image segmentation device in the second embodiment of the present disclosure;

[0016] Figure 6 It is a schematic diagram of the structure of the electronic device in the fourth embodiment of the present disclosure. Detailed implementation manners

[0017] Reference is made herein to the accompanying drawings to describe various aspects and features of the present disclosure.

[0018] It should be understood that various modifications can be made to the embodiments claimed herein. Accordingly, the above description should not be construed as limiting, but merely as exemplifications of embodiments. Those skilled in the art will envision other modifications within the scope and spirit of the present disclosure.

[0019] The accompanying drawings, which are included in and constitute a part of this specification, illustrate embodiments of the present disclosure and, together with the general description of the present disclosure given above and the detailed description of the embodiments given below, serve to explain the principles of the present disclosure.

[0020] These and other features of the present disclosure will become apparent from the following description of the preferred forms of the embodiments given by way of non-limiting example with reference to the accompanying drawings.

[0021] It should also be understood that, although the present disclosure has been described with reference to some specific examples, those skilled in the art can surely implement many other equivalent forms of the present disclosure, which have the features as claimed and thus are all within the protection scope defined thereby.

[0022] When combined with the accompanying drawings, the above and other aspects, features and advantages of the present disclosure will become more apparent in view of the following detailed description.

[0023] Specific embodiments of the present disclosure are described hereinafter with reference to the accompanying drawings; however, it should be understood that the claimed embodiments are merely examples of the present disclosure and can be implemented in various ways. Well-known and / or repetitive functions and structures are not described in detail to avoid obscuring the present disclosure with unnecessary or redundant details. Therefore, the specific structural and functional details claimed herein are not intended to be limiting, but merely as a basis and representative basis for the claims to teach those skilled in the art to use the present disclosure in substantially any suitable detailed structure in a variety of ways.

[0024] This specification may use the phrase "in one embodiment", "in another embodiment", "in yet another embodiment", or "in other embodiments", all of which may refer to one or more of the same or different embodiments according to the present disclosure.

[0025] To solve the above problems existing in the prior art, the first embodiment of the present disclosure provides a method for segmenting medical images, which uses a segmentation network formed by a Unet architecture and a memory learning network to implement small-sample medical image segmentation. Figure 1 The flowchart of the segmentation method provided in this embodiment is shown, which mainly includes steps S10 to S40:

[0026] S10, respectively extract features of M support images and M query images in N dimensions through an encoder, so as to obtain a first type of feature vector matrix of each support image in N dimensions, and a second type of feature vector matrix of each query image in N dimensions.

[0027] In this embodiment, the support image is a medical image that has completed target segmentation, and the query image is a medical image to be processed that currently needs to be segmented. Generally, the target to be segmented in the embodiment is an organ, tissue, or lesion site. The support image is used as a reference for segmentation to achieve accurate segmentation of the target in the query image. In addition, in this embodiment, M query images can be segmented simultaneously, and these M query images are M consecutive medical images to be processed. For example, when taking CT images of a patient's head, continuous shooting will be carried out in a specific order during one shooting process to form slice images of different positions of the patient's head; when target recognition of the above CT images is required, they can be input into the encoder in sequence to achieve more accurate segmentation while segmenting a single query image and combining the associations between multiple query images.

[0028] Figure 2 The schematic diagram of the architecture of the segmentation network in this embodiment is shown. As Figure 2 shown, the encoder 10 mainly includes M encoding units 101, and each encoding unit 101 corresponds to the input of one support image (I s ) and one query image (I q ). Multiple query images are respectively input into the first encoding unit, the second encoding unit until the Mth encoding unit according to the image order, and there may be no association between multiple support images, as long as they are images of the same part as the query image and have achieved accurate target segmentation.

[0029] For each coding unit 101, it is actually the encoder of the Unet architecture, and the specific layers include convolutional layers, activation layers, downsampling processing, etc. For example, an input image with a size of 512*512 can be processed to form a feature vector matrix of 64*64*256. For the two inputs of the support graph and the query graph, the coding unit 101 outputs two different types of feature vector matrices in N dimensions. Among them, the first type of feature vector matrix corresponds to the feature vector matrix of the support graph in N dimensions, and the second type of feature vector matrix corresponds to the feature vector matrix of the query graph in N dimensions. In actual implementation, the number of convolutional layers, downsampling layers, etc. in the coding unit 101 can be adjusted. Generally, the number of convolutional layers does not exceed 3, and the number of downsampling layers does not exceed 5.

[0030] S20. Through the metric network, M first feature vectors are determined based on the labeled graph corresponding to each support graph and the first type of feature vector matrix of each support graph in the Nth dimension, and M similarity matrices are determined based on the M first feature vectors and the M second type of feature vector matrices in the Nth dimension.

[0031] For each support graph, there is a corresponding labeled graph (mask) to mark the position of the target segmented in the support graph. During the segmentation process of the query graph, the support graph and its corresponding mask can be used as reference samples for query graph segmentation to improve the segmentation accuracy. In this embodiment, for the first type of feature vector matrix output by each coding unit 101, the metric network downsamples the mask of the corresponding support graph to the corresponding dimension (i.e., the Nth dimension) and then extracts feature vectors from the first type of feature vector matrix of the support graph in the Nth dimension to obtain the first feature vector, so as to use the first feature vector to represent the relationship between the mask and the support graph. Subsequently, the first feature vector is calculated for similarity with each feature vector in the second type of feature vector matrix in the Nth dimension one by one, and then the similarity matrix is obtained to represent the correlation between the features of the query graph and the features of the support graph.

[0032] Reference Figure 2 For the metric network 20, it also includes M metric units 201 and is connected to the M coding units 101 in the encoder 1 one by one. That is, the feature vector matrix output by the first coding unit is input into the first metric unit, and the content of the labeled graph of the corresponding support graph is also input into the first metric unit. In addition, this embodiment mainly uses cosine similarity for measurement. For a mask with a size of 512*512, after average pooling, it is reduced by 8 times. After feature extraction from the first type of feature vector matrix, a feature vector of 1*1*256 is obtained, and finally a similarity matrix of 64*64*1 is obtained.

[0033] S30. Learn the M similarity matrices through the memory network, and output M forward feature map matrices and M backward feature map matrices.

[0034] Since there are certain dependencies among multiple query graphs, for any one query graph, the feature extraction situation of other adjacent query graphs can also be used as the segmentation basis for the current query graph. Therefore, in this embodiment, a memory network is added to transfer the features learned from the previous image to the next image, and the features between adjacent layer images are used to further improve the final segmentation effect. Specifically, the memory network in this embodiment at least includes a bidirectional long short-term memory network (LSTM, Long Short-Term Memory), which can simultaneously achieve bidirectional knowledge interaction in the forward and backward directions.

[0035] It should be noted that LSTM is usually used to process one-dimensional constants, while the input of the bidirectional LSTM in this embodiment is a two-dimensional vector. Therefore, the parameters of the LSTM need to be adjusted accordingly as follows:

[0036] i t = σ(x t W xi + h t-1 W hi + b i ) (1)

[0037] f t = σ(x t W xf + h t-1 W hf + b f ) (2)

[0038] c t = c t-1 ⊙ f t + i t ⊙ tanh(x t W xc + h t-1 W hc + b c ) (3)

[0039] o t = σ(x t W xo + h t-1 W ho + b o ) (4)

[0040] h t = o t ⊙ tanh(c t ) (5)

[0041] Among them, σ() and tanh() are activation functions, and i t is the input gate, f t is the forget gate, o t is the output gate, b i and b f and b c and b o are biases, x t is the input, c t is the cell activation state, h t is the hidden state, W is the weight, t is the coordinate of the nth slice, and ⊙ represents the dot product.

[0042] In actual implementation, referring to Figure 2 , the memory network 30 also includes M memory units 301. Each memory unit 301 corresponds to a metric unit 201. For a single memory unit 301, it outputs a forward feature map matrix and a backward feature map matrix based on the similarity matrix output by the corresponding metric unit 201 and the knowledge interaction between the M memory units 301 during forward and backward propagation. For the M memory units 301, a total of M forward feature map matrices and M backward feature map matrices are output.

[0043] It should be noted that due to the large computational amount of the two-dimensional input of the LSTM and the possible forgetting problem of the LSTM itself, the number of memory units included in its network should not be too large. The number of units in other modules and the number of support graphs and query graphs input can be the same as the number of memory units. That is, it is appropriate to take an integer value of M within 3 to 9. Generally, the value of M can be selected as 4 or 5.

[0044] S40. The M forward feature map matrices and M backward feature map matrices are respectively spliced and then subjected to upsampling processing in N dimensions to obtain the segmentation maps corresponding to the M query graphs.

[0045] The decoder is the decoder corresponding to the encoder in the Unet architecture, mainly used to restore the feature vector matrix to the original size image in the first dimension through upsampling processing, and perform object recognition and segmentation based on the pixel categories in the original size image. Specifically, for the forward feature map matrix and the backward feature map matrix output by a single memory unit 301, the decoder splices the two, and based on the feature vector matrices of different dimensions output by the encoder, through operations such as convolution and upsampling, finally obtains the output prediction.

[0046] Referring to Figure 2, the decoder 40 includes M decoding units 401, which are connected to the M memory units 301 in one-to-one correspondence and are also connected to the corresponding encoding units 101 at the same time. For example, the first decoding unit 401 is connected to the first encoding unit 101, and the Mth decoding unit 401 is connected to the Mth encoding unit 101. When decoding in the decoding unit 401, first splice the input forward feature map matrix and backward feature map matrix to obtain a spliced matrix, then splice the second type of feature vector matrix of the corresponding query graph output by the encoding unit 101 in N dimensions with this spliced matrix, and then perform upsampling processing in N dimensions to correspondingly obtain a feature map in the first dimension. Finally, perform target segmentation based on the softmax layer, which can integrate the same-level features output by the encoder during the decoding process to improve the segmentation prediction effect.

[0047] In some embodiments, in order to improve the segmentation effect of the segmentation network, before actually segmenting the query graph, the encoder 10, metric network 20, memory network 30, and decoder 40 in the segmentation network can also be trained using sample graphs. The sample graphs are medical images that have already completed target segmentation, and each sample graph has its corresponding labeled graph for marking the location of the target and segmenting the image.

[0048] During the training process, taking M = 5 as an example, the segmentation network can be trained by inputting 5 consecutive sample graphs and the corresponding number of support graphs, and multiple repeated trainings can be performed until the segmentation effect is optimal. It should be noted that there can be a certain overlap between the 5 samples input each time to enhance the feature correlation between images. For example, the first time, the 1st, 2nd, 3rd, 4th, and 5th sample graphs are input, and the second time, the 4th, 5th, 6th, 7th, and 8th sample graphs are input, etc.

[0049] Figure 3 shows a schematic flowchart of training the segmentation network. As Figure 3 shown, its basic process is similar to Figure 1 that is, using the support graph and sample graph as inputs, enabling the segmentation network to output the segmentation graphs corresponding to each sample graph, and Figure 1In contrast, during the training process, the segmentation map of the sample map output by the segmentation network and the corresponding labeled map of the known sample map are further compared to determine the loss function, and the parameter values of the encoder 10, the metric network 20, the memory network 30, and the decoder 40 are optimized and adjusted based on the value of the loss function. After the parameters of each module are optimized, the same sample map or different sample maps are input into the segmentation network to re-output the corresponding segmentation map, and the loss function is recalculated until the loss function is minimized to complete the training of the segmentation network. In actual use, a corresponding number of support maps and query maps are input into the trained segmentation network, and the segmentation map corresponding to the query map can be obtained, realizing the accurate segmentation of medical images under few-shot learning.

[0050] Specifically, the loss function can be expressed by the following formula:

[0051]

[0052] where y n,c is the nth pixel label of the cth class, and p n,c is the predicted value of the nth pixel of the cth class.

[0053] In this embodiment, the support map and the query map are simultaneously input into the encoder for feature extraction, and the segmentation prediction effect is improved by using the similarity features. Further, the memory network is combined to learn the similarity features between adjacent images, enhance the information interaction ability between adjacent images, and thus achieve the accurate segmentation of medical images.

[0054] The second embodiment of the present disclosure provides a medical image segmentation device, which can be installed in any electronic device with processing functions to realize the segmentation of each target in the medical image. Figure 4The structural schematic diagram of the processing device in this embodiment is shown, which mainly includes an encoding module 100, a metric module 200, a memory module 300, and a decoding module 400. Among them, the encoding module 100 is used to extract features of M support graphs and M query graphs in N dimensions respectively, so as to obtain a first type of feature vector matrix of each support graph in N dimensions and a second type of feature vector matrix of each query graph in N dimensions. The M query graphs are M consecutive medical images to be processed; the metric module 200 is used to determine M first feature vectors based on the labeled graph corresponding to each support graph and the first type of feature vector matrix of each support graph in the Nth dimension, and determine M similarity matrices based on the M first feature vectors and the M second type of feature vector matrices in the Nth dimension; the memory module 300 is used to learn the M similarity matrices and output M forward feature map matrices and M backward feature map matrices; the decoding module 400 is used to perform upsampling processing in N dimensions after splicing the M forward feature map matrices and the M backward feature map matrices respectively, so as to obtain the segmentation maps corresponding to the M query graphs.

[0055] In some embodiments, M is an integer between 3 and 9; the memory module at least includes a bidirectional long short-term memory network.

[0056] In some embodiments, the decoding module 400 is specifically used to splice the M forward feature map matrices and the M backward feature map matrices respectively through the decoding module to obtain M splicing matrices; splice the second type of feature vector matrices of the M query graphs in N dimensions with the M splicing matrices respectively and then perform upsampling processing in N dimensions to obtain M feature maps in the first dimension; perform segmentation processing on the M feature maps in the first dimension to obtain the segmentation maps corresponding to the M query graphs.

[0057] In some embodiments, the segmentation device further includes an optimization module 500, such as Figure 5As shown, the optimization module 500 mainly trains the encoding module, the metric module, the memory module, and the decoding module through M support graphs and M sample graphs. Specifically, the optimization module 500 extracts features of the M support graphs and the M sample graphs in N dimensions through the encoding module to obtain a first type of feature vector matrix of each support graph in N dimensions and a second type of feature vector matrix of each sample graph in N dimensions, where the M sample graphs are consecutive medical images; the metric module determines M first feature vectors based on the labeled graph corresponding to each support graph and the first type of feature vector matrix of each support graph in the Nth dimension, and determines M similarity matrices based on the M first feature vectors and the M second type of feature vector matrices in the Nth dimension; the memory module learns the M similarity matrices and outputs M forward feature map matrices and M backward feature map matrices; the decoding module performs upsampling processing on the concatenated M forward feature map matrices and M backward feature map matrices in N dimensions to obtain segmentation maps corresponding to the M sample graphs; the loss function is determined according to the segmentation maps corresponding to the M sample graphs and the labeled graphs corresponding to the M sample graphs, and the parameters of the encoding module, the metric module, the memory module, and the decoding module are adjusted based on the loss function.

[0058] In some embodiments, the optimization module also re-determines the loss function according to the optimized encoding module, metric module, memory module, and decoding module until the value of the loss function is minimized.

[0059] In this embodiment, the support graph and the query graph are simultaneously input into the encoding module for feature extraction, and the segmentation prediction effect is improved by using the similarity features. Further, the memory module is combined to learn the similarity features between adjacent images, enhance the information interaction ability between adjacent images, and thus achieve accurate segmentation of medical images.

[0060] The third embodiment of the present disclosure provides a storage medium, which can be installed in any electronic device with processing capabilities. Specifically, it is a computer-readable medium storing a computer program, which when executed by a processor implements the method provided in any embodiment of the present disclosure, including the following steps S31 to S34:

[0061] S31, extract features of the M support graphs and the M query graphs in N dimensions through an encoder to obtain a first type of feature vector matrix of each support graph in N dimensions and a second type of feature vector matrix of each query graph in N dimensions, where the M query graphs are consecutive M medical images to be processed;

[0062] S32. Based on each labeled graph corresponding to the support graph and the first type of feature vector matrix of each support graph in the Nth dimension, determine M first feature vectors through the metric network, and determine M similarity matrices based on the M first feature vectors and the M second type of feature vector matrices in the Nth dimension;

[0063] S33. Learn the M similarity matrices through the memory network, and output M forward feature map matrices and M backward feature map matrices;

[0064] S34. Through the decoder, splice the M forward feature map matrices and the M backward feature map matrices respectively and perform upsampling processing in N dimensions to obtain the segmentation maps corresponding to the M query graphs.

[0065] Specifically, M is an integer between 3 and 9; the memory network includes at least a bidirectional long short-term memory network.

[0066] When the computer program is executed by the processor to perform upsampling processing in N dimensions after splicing the M forward feature map matrices and the M backward feature map matrices respectively through the decoder, it is specifically executed by the processor as follows: splice the M forward feature map matrices and the M backward feature map matrices respectively through the decoder to obtain M splicing matrices; splice the second type of feature vector matrices of the M query graphs in N dimensions with the M splicing matrices respectively and perform upsampling processing in N dimensions to obtain M feature maps in the first dimension; perform segmentation processing on the M feature maps in the first dimension to obtain the segmentation maps corresponding to the M query graphs.

[0067] The computer program is further executed by the processor as follows: train the encoder, metric network, memory network, and decoder through the M support graphs and the M sample graphs.

[0068] When the computer program is executed by the processor to train the encoder, metric network, memory network, and decoder using M support graphs and M sample graphs, the following steps are specifically executed by the processor: Extract features of the M support graphs and M sample graphs in N dimensions through the encoder to obtain the first type of feature vector matrix of each support graph in N dimensions and the second type of feature vector matrix of each sample graph in N dimensions. The M sample graphs are consecutive medical images; Determine M first feature vectors through the metric network based on the labeled graph corresponding to each support graph and the first type of feature vector matrix of each support graph in the Nth dimension, and determine M similarity matrices based on the M first feature vectors and the M second type of feature vector matrices in the Nth dimension; Learn the M similarity matrices through the memory network and output M forward feature map matrices and M backward feature map matrices; Concatenate the M forward feature map matrices and the M backward feature map matrices respectively through the decoder and perform upsampling processing in N dimensions to obtain the segmentation maps corresponding to the M sample graphs; Determine the loss function based on the segmentation maps corresponding to the M sample graphs and the labeled graphs corresponding to the M sample graphs, and adjust the parameters of the encoder, metric network, memory network, and decoder based on the loss function.

[0069] After the computer program is executed by the processor to adjust the parameters of the encoder, metric network, memory network, and decoder based on the loss function, the following steps are further executed by the processor: Re-determine the loss function according to the optimized encoder, metric network, memory network, and decoder until the value of the loss function is minimized.

[0070] In this embodiment, the support graph and the query graph are input into the encoder for feature extraction at the same time, and the segmentation prediction effect is improved by using the similarity features. Further, the memory network is combined to learn the similarity features between adjacent images, enhance the information interaction ability between adjacent images, and thus achieve accurate segmentation of medical images.

[0071] The fourth embodiment of the present disclosure provides an electronic device. This electronic device can be directly connected to a medical imaging device to directly access the medical images output by the imaging device, and can also be connected to a display device to directly output the processed images on the display device. Its structural schematic diagram is as Figure 6 shown, and it includes at least a memory 601 and a processor 602. A computer program is stored on the memory 601, and when the processor 602 executes the computer program on the memory 601, it implements the method provided in any embodiment of the present disclosure. Exemplarily, the steps of the computer program of the electronic device are as follows S41 to S44:

[0072] S41. The encoder extracts features of M support graphs and M query graphs in N dimensions respectively to obtain a first type of feature vector matrix of each support graph in N dimensions and a second type of feature vector matrix of each query graph in N dimensions, where the M query graphs are M consecutive medical images to be processed;

[0073] S42. Based on the labeled graph corresponding to each support graph and the first type of feature vector matrix of each support graph in the Nth dimension, the metric network determines M first feature vectors, and based on the M first feature vectors and the M second type of feature vector matrices in the Nth dimension, it determines M similarity matrices;

[0074] S43. The memory network learns the M similarity matrices and outputs M forward feature map matrices and M backward feature map matrices;

[0075] S44. The decoder concatenates the M forward feature map matrices and the M backward feature map matrices respectively and then performs upsampling processing in N dimensions to obtain the segmentation maps corresponding to the M query graphs.

[0076] Specifically, M is an integer between 3 and 9; the memory network includes at least a bidirectional long short-term memory network.

[0077] When the processor executes the computer program stored in the memory to perform upsampling processing in N dimensions after concatenating the M forward feature map matrices and the M backward feature map matrices respectively through the decoder, it specifically executes the following computer program: concatenating the M forward feature map matrices and the M backward feature map matrices respectively through the decoder to obtain M concatenated matrices; concatenating the second type of feature vector matrices of the M query graphs in N dimensions with the M concatenated matrices respectively and then performing upsampling processing in N dimensions to obtain M feature maps of the first dimension; performing segmentation processing on the M feature maps of the first dimension to obtain the segmentation maps corresponding to the M query graphs.

[0078] The processor also executes the following computer program stored in the memory: training the encoder, the metric network, the memory network, and the decoder with the M support graphs and the M sample graphs.

[0079] When the processor executes the computer program for training the encoder, metric network, memory network, and decoder using M support graphs and M sample graphs stored in the memory, it specifically performs the following steps: Extract features of the M support graphs and M sample graphs in N dimensions through the encoder to obtain the first type of feature vector matrix of each support graph in N dimensions and the second type of feature vector matrix of each sample graph in N dimensions. The M sample graphs are consecutive medical images. Determine M first feature vectors through the metric network based on the labeled graph corresponding to each support graph and the first type of feature vector matrix of each support graph in the Nth dimension, and determine M similarity matrices based on the M first feature vectors and the M second type of feature vector matrices in the Nth dimension. Learn the M similarity matrices through the memory network and output M forward feature map matrices and M backward feature map matrices. Concatenate the M forward feature map matrices and M backward feature map matrices respectively through the decoder and perform upsampling processing in N dimensions to obtain the segmentation maps corresponding to the M sample graphs. Determine the loss function based on the segmentation maps corresponding to the M sample graphs and the labeled graphs corresponding to the M sample graphs, and adjust the parameters of the encoder, metric network, memory network, and decoder based on the loss function.

[0080] After the processor adjusts the parameters of the encoder, metric network, memory network, and decoder based on the loss function, it also executes the following computer program: Re-determine the loss function according to the optimized encoder, metric network, memory network, and decoder until the value of the loss function is minimized.

[0081] In this embodiment, the support graph and the query graph are simultaneously input into the encoder for feature extraction, and the segmentation prediction effect is improved by using the similarity features. Further, the memory network is combined to learn the similarity features between adjacent images, enhance the information interaction ability between adjacent images, and thus achieve accurate segmentation of medical images.

[0082] The above has described multiple embodiments of the present disclosure in detail, but the present disclosure is not limited to these specific embodiments. Those skilled in the art can make various variations and modifications to the embodiments based on the concept of the present disclosure, and these variations and modifications should fall within the scope of protection required by the present disclosure.

Claims

1. A method for segmenting medical images, characterized in that, comprising: Performing feature extraction on M support graphs and M query graphs respectively in N dimensions through an encoder to obtain a first type of feature vector matrix of each support graph in N dimensions, and a second type of feature vector matrix of each query graph in N dimensions, where the M query graphs are M consecutive medical images to be processed; Determining M first feature vectors through a metric network based on the labeled graph corresponding to each support graph and the first type of feature vector matrix of each support graph in the Nth dimension, and determining M similarity matrices based on the M first feature vectors and the M second type of feature vector matrices in the Nth dimension; Learning the M similarity matrices through a memory network, and outputting M forward feature map matrices and M backward feature map matrices; Performing upsampling processing on N dimensions after splicing the M forward feature map matrices and the M backward feature map matrices respectively through a decoder to obtain segmentation graphs corresponding to the M query graphs.

2. The segmentation method according to claim 1, characterized in that, M is an integer between 3 and 9.

3. The segmentation method according to claim 1, characterized in that, The memory network at least includes a bidirectional long short-term memory network.

4. The segmentation method according to claim 1, characterized in that, Performing upsampling processing on N dimensions after splicing the M forward feature map matrices and the M backward feature map matrices respectively through a decoder, including: Splicing the M forward feature map matrices and the M backward feature map matrices respectively through a decoder to obtain M splicing matrices; Splicing the second type of feature vector matrices of the M query graphs in N dimensions with the M splicing matrices respectively and then performing upsampling processing on N dimensions to obtain M feature maps of the first dimension; Performing segmentation processing on the M feature maps of the first dimension to obtain segmentation graphs corresponding to the M query graphs.

5. The segmentation method according to claim 1, characterized in that, further comprising: Training the encoder, the metric network, the memory network and the decoder through M support graphs and M sample graphs.

6. The segmentation method according to claim 5, characterized in that, Training the encoder, the metric network, the memory network and the decoder through M support graphs and M sample graphs, including: Performing feature extraction on M support graphs and M sample graphs respectively in N dimensions through an encoder to obtain a first type of feature vector matrix of each support graph in N dimensions, and a third type of feature vector matrix of each sample graph in N dimensions, where the M sample graphs are consecutive medical images; Determining M first feature vectors through a metric network based on the labeled graph corresponding to each support graph and the first type of feature vector matrix of each support graph in the Nth dimension, and determining M similarity matrices based on the M first feature vectors and the M third type of feature vector matrices in the Nth dimension; Learn the M similarity matrices through a memory network, and output M forward feature map matrices and M backward feature map matrices; Through a decoder, the M forward feature map matrices and the M backward feature map matrices are respectively concatenated and then subjected to upsampling processing in N dimensions to obtain the segmentation maps corresponding to the M sample maps; Determine a loss function according to the segmentation maps corresponding to the M sample maps and the labeled maps corresponding to the M sample maps, and adjust the parameters of the encoder, the metric network, the memory network, and the decoder based on the loss function.

7. The segmentation method according to claim 6, wherein, after adjusting the parameters of the encoder, the metric network, the memory network, and the decoder based on the loss function, it further includes: Re-determine the loss function according to the optimized encoder, metric network, memory network, and decoding module until the value of the loss function is minimized.

8. A medical image segmentation device, wherein, it includes: An encoding module for respectively extracting features of M support maps and M query maps in N dimensions to obtain a first type of feature vector matrix of each support map in N dimensions, and a second type of feature vector matrix of each query map in N dimensions, where the M query maps are M consecutive medical images to be processed; A metric module for determining M first feature vectors based on the labeled map corresponding to each support map and the first type of feature vector matrix of each support map in the Nth dimension, and determining M similarity matrices based on the M first feature vectors and the M second type of feature vector matrices in the Nth dimension; A memory module for learning the M similarity matrices and outputting M forward feature map matrices and M backward feature map matrices; A decoding module for respectively concatenating the M forward feature map matrices and the M backward feature map matrices and then performing upsampling processing in N dimensions to obtain the segmentation maps corresponding to the M query maps.

9. A storage medium storing a computer program, wherein, When the computer program is executed by a processor, it implements the steps of the medical image segmentation method according to any one of claims 1 to 7.

10. An electronic device, at least including a memory and a processor, with a computer program stored on the memory, wherein, When the processor executes the computer program on the memory, it implements the steps of the medical image segmentation method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method and device for detecting position of target object in image, electronic equipment and storage medium

    CN112581522A

  • Method, device and medium for segmenting medical image

    CN113947681A