Image segmentation method, device and storage medium
By using deep learning methods to segment multimodal pathological images, and utilizing feature fusion and loss function training of neural networks, the problem of insufficient accuracy in multimodal pathological image segmentation in existing technologies is solved, and the accuracy of segmentation is improved.
Patent Information
- Application Number
- CN202210153314.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-18
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2042-02-18
AI Technical Summary
Existing pathological image segmentation methods mainly focus on RGB images, making it difficult to effectively utilize complementary information in multimodal pathological images, resulting in insufficient segmentation accuracy.
Deep learning methods are employed to segment multimodal pathological images using neural networks. This includes a first-modality image processing part, a second-modality image processing part, and a feature fusion part. Features are extracted and fused separately, and the network is trained using cross-entropy loss function, Dice loss function, and multimodal interaction loss function to generate the final segmentation result.
It improves the accuracy of multimodal pathological image segmentation, especially for pixels where the prediction results differ significantly under different modalities, thus enhancing the overall segmentation performance of the neural network.
Smart Images

Figure CN116664472B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates generally to image processing, and more specifically to methods, apparatus and storage media for image segmentation of multimodal pathological images. Background Technology
[0002] With the rapid development of artificial intelligence, deep learning technology has been applied to the field of pathological image analysis. AI can, for example, automatically locate cancerous and non-cancerous areas in pathological images and statistically analyze relevant indicators, providing a basis for pathological analysis and treatment. This can effectively reduce the workload of pathologists and improve their efficiency.
[0003] Many deep learning segmentation methods based on RGB pathological images have been proposed. For example, by analyzing RGB pathological images of lymph nodes, the location of cancerous cells can be automatically located, and the extent of cancer cell metastasis can be classified.
[0004] On the other hand, with the development of hyperspectral image acquisition technology, pathological studies of hyperspectral images have also begun. For example, Yun B, Wang Y, Chen J, and others published the first hyperspectral pathological image database for cholangiocarcinoma and proposed using a spectral transformer to perform image segmentation of hyperspectral pathological images based on this database. For details, please refer to the paper "SpecTr: Spectral Transformer for Hyperspectral Pathology Image Segmentation" [J]. arXiv preprint arXiv:2103.03604,2021. Summary of the Invention
[0005] This disclosure proposes a technique for image segmentation of multimodal pathological images using deep learning. In this document, multimodal pathological images may include, but are not limited to, RGB images, hyperspectral images, depth images, etc.
[0006] According to one aspect of this disclosure, a computer-implemented method for performing image segmentation on multimodal pathological images using a neural network is provided. The multimodal pathological images include at least a first modality image and a second modality image. The first modality image is an image of an object obtained in a first modality, and the second modality image is an image of the object obtained in a second modality. The neural network includes at least a first modality image processing part, a second modality image processing part, and a feature fusion part. The method includes: extracting a first feature from the first modality image by the first modality image processing part, and extracting a second feature from the second modality image by the second modality image processing part; and fusing the second feature with the first feature by the feature fusion part to generate a feature provided to the first modality image processing part. The first modality image processing part performs image segmentation on the first modality image based on the first fusion feature, and the second modality image processing part performs image segmentation on the second modality image based on the second fusion feature, wherein the segmentation prediction results of the first modality image processing part and the second modality image processing part are merged as the final segmentation result; the first modality image processing part is trained based on the first modality training image and the first loss function, and the second modality image processing part is trained based on the second modality training image and the second loss function; and image segmentation is performed on the multimodal pathological image to be segmented using the trained neural network.
[0007] According to another aspect of this disclosure, an apparatus is provided for performing image segmentation on multimodal pathological images using a neural network, wherein the multimodal pathological images include at least a first modality image and a second modality image, the first modality image being an image of an object obtained in a first modality, and the second modality image being an image of the object obtained in a second modality. The apparatus includes: a memory storing computer program instructions; and one or more processors, the processors implementing at least the following portions of the neural network by executing the computer program instructions: a first modality image processing portion configured to extract a first feature from the first modality image and perform image segmentation on the first modality image based on a first fusion feature; and a second modality image processing portion configured to perform image segmentation on the first modality image. The second modality image extracts a second feature, and performs image segmentation on the second modality image based on the second fused feature; and a feature fusion part is configured to fuse the second feature based on the first feature to generate the first fused feature provided to the first modality image processing part, and fuse the first feature based on the second feature to generate the second fused feature provided to the second modality image processing part, wherein the segmentation prediction results of the first modality image processing part and the second modality image processing part are merged as the final segmentation result, wherein the first modality image processing part is trained based on the first modality training image and the first loss function, and the second modality image processing part is trained based on the second modality training image and the second loss function.
[0008] According to another aspect of this disclosure, a storage medium storing a computer program is provided, which, when executed by a computer, causes the computer to perform the above-described method. Attached Figure Description
[0009] The accompanying drawings illustrate embodiments of the present disclosure and provide a further understanding of the disclosure. The above and other objects, features, and advantages of the present disclosure can be more readily understood by referring to the following description of various embodiments in conjunction with the accompanying drawings, in which:
[0010] Figure 1 The architecture of a neural network according to a first embodiment of the present disclosure is illustrated schematically.
[0011] Figure 2 schematically shown Figure 1 The structure of each image processing component.
[0012] Figure 3 The feature fusion according to this disclosure is illustrated schematically.
[0013] Figure 4A neural network according to a second embodiment of the present disclosure is illustrated schematically.
[0014] Figure 5 A neural network according to a third embodiment of the present disclosure is illustrated schematically.
[0015] Figure 6 A flowchart of an image segmentation method according to a first embodiment of the present disclosure is shown.
[0016] Figure 7 A flowchart of an image segmentation method according to a second embodiment of the present disclosure is shown.
[0017] Figure 8 A flowchart of an image segmentation method according to a third embodiment of the present disclosure is shown.
[0018] Figure 9 An exemplary configuration block diagram of computer hardware implementing the present disclosure is shown. Detailed Implementation
[0019] In the following, embodiments according to this disclosure will be described in detail with reference to the accompanying drawings. In the drawings, the same or similar elements will be indicated by the same or similar reference numerals. Furthermore, detailed descriptions of known technologies and configurations incorporated herein will be omitted where such omissions may obscure the subject matter of this disclosure.
[0020] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of this disclosure. Unless the context clearly indicates otherwise, singular expressions also include plural forms. Furthermore, the terms “comprising,” “including,” and “having” as used herein are intended to indicate the presence of the described features, entities, operations, and / or components, but do not exclude the presence or addition of one or more other features, entities, operations, and / or components.
[0021] In the following description, numerous specific details are set forth to provide a full understanding of this disclosure. However, this disclosure may also be practiced without some or all of these specific details. The accompanying drawings show only components closely related to the embodiments according to this disclosure, while other details less relevant to this disclosure are omitted.
[0022] Figure 1 The architecture of a neural network according to a first embodiment of this disclosure is schematically illustrated, the neural network being used to perform image segmentation on multimodal pathological images. For example... Figure 1As shown, the neural network includes multiple image processing parts 101-10M and a feature fusion part 110. Specifically, the first modality image processing part 101 extracts a first feature f1 from the first modality image, the second modality image processing part 102 extracts a second feature f2 from the second modality image, and so on, the Mth modality image processing part 10M extracts the Mth feature f1 from the Mth modality image. M The extracted features are input to the feature fusion section 110. The feature fusion section 110 fuses the various features to generate a first fused feature f. 1-f Second fusion feature f 2-f ...Mth fusion feature f M-f And each of these is provided to its respective image processing unit (101-10M). The following will combine... Figure 3 The specific description is the feature fusion performed by the feature fusion section 110.
[0023] Figure 2 schematically shown Figure 1 The specific structure of each image processing component. For example... Figure 2 As shown, each image processing unit 101-10M has an encoder-decoder structure. The encoder includes multiple cascaded downsampling convolutional layers, and the decoder includes multiple cascaded upsampling convolutional layers. The output of the last convolutional layer of the encoder is input to the first convolutional layer of the decoder. It should be noted that, although in Figure 2 The diagram shows that both the encoder and decoder include three convolutional layers, but this disclosure is not limited to this, and other numbers of convolutional layers are also possible.
[0024] also, Figure 1 The first feature f1, the second feature f2, ..., the Mth feature f M The first fused feature f is output from the last convolutional layer of the encoder in the corresponding image processing part. 1-f Second fusion feature f 2-f ...Mth fusion feature f M-f The first-stage convolutional layer of the decoder is fed into the corresponding image processing part. Figure 2 (Not shown in the image).
[0025] Figure 3 The feature integration according to this disclosure is illustrated schematically. Figure 1 In the feature fusion section 110 shown, the first feature f1, the second feature f2, ..., the Mth feature f1 are first fused together. M Add elements together As shown in the following mathematical expression (1):
[0026]
[0027] Then, the resulting M1 is input into the convolutional layer F with the ReLU activation function. θ1 In the middle, convolutional layer F θ1 The output can be represented as Then, the resulting M2 is input into two cascaded convolutional layers F with ReLU activation functions. θ2 and F θ3 In the middle, convolutional layer F θ3 The output can be expressed as M3 = F θ3 (F θ2 (M2)).
[0028] Therefore, the first fusion feature f for the first modality image can be calculated according to the following mathematical formula (2). 1-f :
[0029]
[0030] Figure 3 Only the first fusion feature f is shown in the image. 1-f The calculation of the second fusion feature f for the second modality image can be performed using a similar method. 2-f And the Mth fusion feature f for the Mth modality image M-f As shown in the following mathematical expressions (3) and (4):
[0031]
[0032]
[0033] When training the neural network according to the first embodiment, the first modality image processing section 101 to the Mth modality image processing section 10M are respectively input as training images. In particular, the first modality image to the Mth modality image input in one iteration are images obtained for the same object in different modalities.
[0034] Furthermore, the first modality image processing part 101 is trained based on the first modality training image and the first loss function L1, the second modality image processing part 102 is trained based on the second modality training image and the second loss function L2, and so on, based on the Mth modality training image and the Mth loss function L... M To train the Mth modality image processing part 10M.
[0035] Loss function L1-L M Each of these can include a pixel-level cross-entropy loss function and an image-level Dice loss function. Taking the first loss function L1 as an example, its cross-entropy loss function can be expressed as the following mathematical formula (5):
[0036]
[0037] Among them, y i This represents the truth value of the i-th pixel. For example, a truth value of 0 indicates that the i-th pixel is located in a non-cancerous region, while a truth value of 1 indicates that the i-th pixel is located in a cancerous region. σ represents the confidence level of the prediction result of the first modality image processing part 101 for the i-th pixel, N represents the number of pixels, and σ is the pixel-level sigmoid function.
[0038] The Dice loss function included in the first loss function L1 can be expressed as the following mathematical formula (6):
[0039]
[0040] Where y represents the true value of a pixel. This indicates the confidence level of the prediction result for a pixel.
[0041] Based on the above, the first loss function L1 can be expressed as the following mathematical formula (7):
[0042] L1=αL bce +βL dice -(7)
[0043] Where α and β represent weights.
[0044] Similarly, the second loss function L2 to the Mth loss function L can be calculated for other image processing components. M Each of them.
[0045] In the first embodiment, the overall loss function used to train the neural network can be expressed as the following mathematical formula (8):
[0046] L = L1 + L2 + ... + L M -(8)
[0047] Once training is complete, the trained neural network is used to perform image segmentation on multimodal pathological images used in real-world applications. These multimodal pathological images include multiple images of the same object obtained in different modalities, such as RGB images, hyperspectral images, and depth images. Specifically, each trained image processing unit 101-10M performs image segmentation on the corresponding modality of the image, and then the segmentation prediction results from each image processing unit are merged to generate the final segmentation result. As shown in the following mathematical expression (9):
[0048]
[0049] in, These represent the predicted probabilities of each pixel in the corresponding image by each image processing unit 101-10M. As an example, the merging method is addition.
[0050] Figure 4 A neural network according to a second embodiment of the present disclosure is schematically illustrated, differing from the neural network of the first embodiment in that it further includes a segmentation result fusion portion 410. (Refer to...) Figure 4 The segmentation result fusion section 410 fuses the image segmentation results performed by each image processing section 101-10M on the corresponding modality image to generate a fused segmentation result. As an example, the segmentation prediction results of each image processing section 101-10M can be multiplied element-wise to generate the fused segmentation result. However, this disclosure is not limited thereto, and those skilled in the art can readily employ other known fusion methods.
[0051] In training Figure 4 In the process of the neural network shown, in addition to using the first to Mth loss functions described in the first embodiment, a fusion loss function L is also used. fusion To train the segmentation result fusion part 410. This is done using the first loss function L1 to the Mth loss function L... M Similarly, the fusion loss function L fusion This also includes the cross-entropy loss function and the Dice loss function, as shown in the following mathematical formulas (10) and (11), respectively:
[0052]
[0053]
[0054] Where y represents the true value of a pixel. This indicates the confidence level of the prediction result for a pixel.
[0055] In the second embodiment, the overall loss function used to train the neural network can be expressed as the following mathematical formula (12):
[0056] L = L1 + L2 + ... + L M +L fusion -(12)
[0057] Once training is complete, the trained neural network is used to perform image segmentation on multimodal pathological images in practical applications. Specifically, each trained image processing unit 101-10M performs image segmentation for its corresponding modality, and the segmentation result fusion unit 410 fuses the segmentation prediction results from each image processing unit 101-10M. In this case, the predicted probabilities of each image processing unit are... The predicted probability of fusion generated by the fusion part 410 with the segmentation result Merge the results to obtain the final split. As shown in the following mathematical expression (13):
[0058]
[0059] Figure 5 The diagram schematically illustrates a neural network according to a third embodiment of the present disclosure, which differs from the neural network of the second embodiment in that a multimodal interaction loss function L is further used during the training process. interaction .
[0060] The interaction loss function can be expressed as the following mathematical formula (14):
[0061]
[0062] Where Norm is the normalization factor, N represents the number of pixels, M represents the number of modalities, and l and m each correspond to one of the M modalities. This represents the prediction result based on the l-th modality image. This represents the prediction result based on the m-th modality image. From mathematical formula (14), it can be seen that the interaction loss function is based on the difference between the segmentation prediction results under the two modalities. By applying the interaction loss function during training, the prediction results under various modalities can be made closer to each other, which can improve the accuracy of segmentation prediction for pixels whose prediction results under different modalities differ significantly.
[0063] In the third embodiment, the overall loss function used to train the neural network can be expressed as the following mathematical formula (15):
[0064] L = L1 + L2 + ... + L M +L fusion +γL interaction -(15)
[0065] Where γ represents the weight.
[0066] Once training is complete, the trained neural network is used to perform image segmentation on multimodal pathological images in practical applications. The final segmentation result is the same as described in the second embodiment, as shown in mathematical formula (13).
[0067] Figure 6 A flowchart illustrating a method for performing image segmentation on multimodal pathological images according to a first embodiment of this disclosure is shown. Figure 6As shown, in step S610, the first modality image processing unit 101 to the Mth modality image processing unit 10M respectively extract the first feature f1 to the Mth feature f1 from the first modality image to the Mth modality image. M .
[0068] In step S620, the feature fusion part 110 processes each feature f1-f M Perform fusion to generate the first fusion feature f 1-f up to the Mth fusion feature f M-f And provide them to the first modal image processing section 101 to the Mth modal image processing section 10M respectively.
[0069] In step S630, the first modality image processing unit 101 to the Mth modality image processing unit 10M respectively base their processing on the first fusion feature f. 1-f up to the Mth fusion feature f M-f Image segmentation is performed on the image of the corresponding modality. The segmentation prediction results of each image processing part (101-10M) can be merged to obtain the final segmentation result.
[0070] In step S640, the first modality image processing part 101 to the Mth modality image processing part 10M are trained using the training image set. The loss function shown in mathematical formula (8) is used in the training.
[0071] In step S650, image segmentation is performed on the pathological images in the actual application using the trained first modality image processing unit 101 to the Mth modality image processing unit 10M.
[0072] Figure 7 A flowchart of an image segmentation method according to a second embodiment of the present disclosure is shown. Figure 7 As shown, in step S710, the first modality image processing unit 101 to the Mth modality image processing unit 10M respectively extract the first feature f1 to the Mth feature f1 from the first modality image to the Mth modality image. M .
[0073] In step S720, the feature fusion part 110 processes each feature f1-f M Perform fusion to generate the first fusion feature f 1-f up to the Mth fusion feature f M-f And provide them to the first modal image processing section 101 to the Mth modal image processing section 10M respectively.
[0074] In step S730, the first modality image processing unit 101 to the Mth modality image processing unit 10M respectively base their processing on the first fusion feature f. 1-f up to the Mth fusion feature f M-fImage segmentation is performed on the image of the corresponding modality.
[0075] In step S740, the segmentation prediction results of each image processing part 101-10M are fused by the segmentation result fusion part 410 to generate a fused segmentation result. In this case, the segmentation results of each image processing part 101-10M and the fused segmentation result generated by the segmentation result fusion part 410 can be merged as the final segmentation result.
[0076] In step S750, the first modality image processing part 101 to the Mth modality image processing part 10M and the segmentation result fusion part 410 are trained. The loss function shown in mathematical formula (12) is used in the training.
[0077] In step S760, image segmentation is performed on the actual pathological image using the trained first modality image processing part 101 to the Mth modality image processing part 10M and the segmentation result fusion part 410.
[0078] Figure 8 A flowchart of an image segmentation method according to a third embodiment of this disclosure is shown. Figure 8 As shown, in step S810, the first modality image processing unit 101 to the Mth modality image processing unit 10M respectively extract the first feature f1 to the Mth feature f1 from the first modality image to the Mth modality image. M .
[0079] In step S820, the feature fusion part 110 processes each feature f1-f M Perform fusion to generate the first fusion feature f 1-f up to the Mth fusion feature f M-f And provide them to the first modal image processing section 101 to the Mth modal image processing section 10M respectively.
[0080] In step S830, the first modality image processing unit 101 to the Mth modality image processing unit 10M respectively base their processing on the first fusion feature f. 1-f up to the Mth fusion feature f M-f Image segmentation is performed on the image of the corresponding modality.
[0081] In step S840, the segmentation prediction results of each image processing part 101-10M are fused by the segmentation result fusion part 410 to generate a fused segmentation result. In this case, the segmentation results of each image processing part 101-10M and the fused segmentation result generated by the segmentation result fusion part 410 can be merged as the final segmentation result.
[0082] In step S850, the first modality image processing part 101 to the Mth modality image processing part 10M and the segmentation result fusion part 410 are trained. The loss function shown in mathematical formula (15) is used in the training.
[0083] In step S860, image segmentation is performed on the actual pathological image using the trained first modality image processing part 101 to the Mth modality image processing part 10M and the segmentation result fusion part 410.
[0084] The technology of this disclosure has been described above in conjunction with various embodiments. In this disclosure, by fusing features extracted from pathological images of different modalities, complementary information contained in images of different modalities (e.g., RGB images, hyperspectral images, and depth images) can be utilized to improve segmentation accuracy. Furthermore, by fusing segmentation prediction results for images of various modalities, the overall segmentation performance of the neural network can be further improved. Moreover, by using a multimodal interaction loss function during neural network training, the accuracy of segmentation prediction can be improved for pixels where prediction results differ significantly across modalities.
[0085] This disclosure can be applied to automatic image segmentation of various types of pathological images, including but not limited to: breast cancer pathological images, gastric cancer pathological images, bile duct cancer pathological images, thyroid cancer pathological images, etc.
[0086] The methods described in the above embodiments can be implemented by software, hardware, or a combination of software and hardware. Programs included in the software can be stored beforehand in a storage medium located internally or externally to the device. As an example, during execution, these programs are written to random access memory (RAM) and executed by a processor (e.g., a CPU) to implement the various processes described herein.
[0087] Figure 9 An example configuration block diagram of computer hardware is shown, which is an example of an apparatus for performing image segmentation on multimodal pathological images according to the present disclosure.
[0088] like Figure 9 As shown, in computer 900, central processing unit (CPU) 901, read-only memory (ROM) 902 and random access memory (RAM) 903 are connected to each other via bus 904.
[0089] The input / output interface 905 is further connected to the bus 904. The input / output interface 905 is connected to the following components: an input unit 906 formed by a keyboard, mouse, microphone, etc.; an output unit 907 formed by a display, speaker, etc.; a storage unit 908 formed by a hard disk, non-volatile memory, etc.; a communication unit 909 formed by a network interface card (such as a local area network (LAN) card, modem, etc.); and a driver 910 for driving a mobile medium 911, such as a disk, optical disk, magneto-optical disk, or semiconductor memory.
[0090] In a computer with the above structure, the CPU 901 loads the program stored in the storage unit 908 into the RAM 903 via the input / output interface 905 and the bus 904, and executes the program to perform the method described above.
[0091] The program to be executed by the computer (CPU 901) can be recorded on a portable medium 911, which is formed as a packaging medium, such as a magnetic disk (including a floppy disk), an optical disk (including a compact optical disk-read-only memory (CD-ROM)), a digital multifunction optical disk (DVD), etc.), a magneto-optical disk, or a semiconductor memory. Furthermore, the program to be executed by the computer (CPU 901) can also be provided via wired or wireless transmission media such as a local area network, the Internet, or digital satellite broadcasting.
[0092] When the removable medium 911 is installed in the driver 910, the program can be installed in the storage unit 908 via the input / output interface 905. Alternatively, the program can be received by the communication unit 909 via a wired or wireless transmission medium and installed in the storage unit 908. Alternatively, the program can be pre-installed in the ROM 902 or the storage unit 908.
[0093] A program executed by a computer may be a program that performs processing in the order described in this specification, or it may be a program that performs processing in parallel or when needed (such as when invoked).
[0094] The units or devices described herein are for logical purposes only and do not strictly correspond to physical devices or entities. For example, the function of each unit described herein may be implemented by multiple physical entities, or the function of multiple units described herein may be implemented by a single physical entity. Furthermore, the features, components, elements, steps, etc., described in one embodiment are not limited to that embodiment, but can also be applied to other embodiments, such as replacing specific features, components, elements, steps, etc., in other embodiments, or in combination with them.
[0095] The scope of this invention is not limited to the specific embodiments described herein. Those skilled in the art will understand that various modifications or variations can be made to the embodiments described herein, depending on design requirements and other factors, without departing from the principles and spirit of the invention. The scope of this invention is defined by the appended claims and their equivalents.
[0096] Postscript
[0097] (1). A computer-implemented method for performing image segmentation on multimodal pathological images using a neural network, wherein the multimodal pathological images include at least a first modality image and a second modality image, the first modality image being an image of an object obtained in a first modality, the second modality image being an image of the object obtained in a second modality, the neural network including at least a first modality image processing part, a second modality image processing part, and a feature fusion part, the method comprising:
[0098] The first modal image processing unit extracts a first feature from the first modal image, and the second modal image processing unit extracts a second feature from the second modal image;
[0099] The feature fusion part fuses the second feature based on the first feature to generate a first fused feature provided to the first modal image processing part, and fuses the first feature based on the second feature to generate a second fused feature provided to the second modal image processing part;
[0100] The first modality image processing part performs image segmentation on the first modality image based on the first fusion feature, and the second modality image processing part performs image segmentation on the second modality image based on the second fusion feature, wherein the segmentation prediction results of the first modality image processing part and the second modality image processing part are merged as the final segmentation result;
[0101] The first modality image processing part is trained based on the first modality training image and the first loss function, and the second modality image processing part is trained based on the second modality training image and the second loss function; and
[0102] Image segmentation is performed on multimodal pathological images to be segmented using a trained neural network.
[0103] (2). According to the method described in (1), wherein the neural network further includes a segmentation result fusion part, and the method further includes:
[0104] The segmentation prediction results of the first modality image processing part and the second modality image processing part are fused by the segmentation result fusion part to generate a fused segmentation result, wherein the segmentation prediction results of the first modality image processing part and the second modality image processing part, as well as the fused segmentation result, are combined to form the final segmentation result; and
[0105] The fusion part of the segmentation result is trained based on the fusion loss function.
[0106] (3). According to the method of (2), each of the first loss function, the second loss function and the fusion loss function includes a cross-entropy loss function and a Dice loss function.
[0107] (4). The method according to (1) or (2) further includes: training the neural network based on an interaction loss function, the interaction loss function being based on the difference between the segmentation prediction results of the first modality image processing part and the segmentation prediction results of the second modality image processing part.
[0108] (5). The method according to (1) further includes: the feature fusion portion,
[0109] Add the first feature and the second feature element by element;
[0110] The features obtained after addition are input into one or more convolutional layers;
[0111] The features output from the convolutional layer are element-wise added to the first feature to generate the first fused feature; and
[0112] The features output from the convolutional layer are added element-wise to the second feature to generate the second fused feature.
[0113] (6). The method according to (2) further includes: multiplying the segmentation prediction result of the first modality image processing part and the segmentation prediction result of the second modality image processing part element-wise by the segmentation result fusion part to generate the fused segmentation result.
[0114] (7) The method according to (1), wherein each of the first modality image processing part and the second modality image processing part has an encoder-decoder structure, the encoder comprising a plurality of downsampling convolutional layers, and the decoder comprising a plurality of upsampling convolutional layers.
[0115] Wherein, the first feature is output from the encoder of the first modality image processing section, and the first fused feature is input to the decoder of the first modality image processing section.
[0116] The second feature is output from the encoder of the second modality image processing section, and the second fused feature is input to the decoder of the second modality image processing section.
[0117] (8). The method according to (1), wherein the multimodal pathological image further includes a third modality image obtained for the object in a third modality, the neural network further includes a third modality image processing part, and the method further includes:
[0118] The third modality image processing unit extracts a third feature from the third modality image;
[0119] The feature fusion component fuses the second feature and the third feature based on the first feature to generate the first fused feature, fuses the first feature and the third feature based on the second feature to generate the second fused feature, and fuses the first feature and the second feature based on the third feature to generate the third fused feature.
[0120] The third modality image processing part performs image segmentation on the third modality image based on the third fusion feature, wherein the segmentation prediction results of the first modality image processing part, the second modality image processing part, and the third modality image processing part are merged to form the final segmentation result;
[0121] The third-modality image processing component is trained based on the third-modality image training set and the third loss function; and
[0122] Image segmentation is performed on multimodal pathological images to be segmented using a trained neural network.
[0123] (9) An apparatus for performing image segmentation on a multimodal pathological image using a neural network, wherein the multimodal pathological image includes at least a first modal image and a second modal image, the first modal image being an image of an object obtained in a first modality, and the second modal image being an image of the object obtained in a second modality, the apparatus comprising:
[0124] A memory that stores computer program instructions; and
[0125] One or more processors configured to implement at least the following portions of the neural network by executing the computer program instructions:
[0126] The first modality image processing section is configured to extract a first feature from the first modality image and perform image segmentation on the first modality image based on the first fusion feature;
[0127] The second modality image processing section is configured to extract a second feature from the second modality image and perform image segmentation on the second modality image based on the second fused feature; and
[0128] The feature fusion section is configured to fuse the second feature based on the first feature to generate the first fused feature provided to the first modality image processing section, and to fuse the first feature based on the second feature to generate the second fused feature provided to the second modality image processing section.
[0129] In this process, the segmentation prediction results from the first modality image processing part and the second modality image processing part are combined to form the final segmentation result.
[0130] Specifically, the first modality image processing part is trained based on the first modality training image and the first loss function, and the second modality image processing part is trained based on the second modality training image and the second loss function.
[0131] (10) The apparatus according to (9), wherein the processor is further configured to implement a segmentation result fusion portion of the neural network, the segmentation result fusion portion being configured to fuse the segmentation prediction result of the first modality image processing portion with the segmentation prediction result of the second modality image processing portion to generate a fused segmentation result, wherein the segmentation prediction results of the first modality image processing portion and the second modality image processing portion and the fused segmentation result are combined as a final segmentation result.
[0132] The fusion part of the segmentation result is trained based on the fusion loss function.
[0133] (11). The apparatus according to (10), wherein each of the first loss function, the second loss function and the fusion loss function includes a cross-entropy loss function and a Dice loss function.
[0134] (12). The apparatus according to (9) or (10), wherein the neural network is trained based on an interactive loss function, the interactive loss function being based on the difference between the segmentation prediction results of the first modality image processing part and the segmentation prediction results of the second modality image processing part.
[0135] (13). The apparatus according to (9), wherein the feature fusion portion is configured as follows:
[0136] Add the first feature and the second feature element by element;
[0137] The features obtained after addition are input into one or more convolutional layers;
[0138] The features output from the convolutional layer are element-wise added to the first feature to generate the first fused feature; and
[0139] The features output from the convolutional layer are added element-wise to the second feature to generate the second fused feature.
[0140] (14). According to the apparatus of (10), wherein the segmentation result fusion portion is configured to multiply the segmentation prediction result of the first modality image processing portion and the segmentation prediction result of the second modality image processing portion element-wise to generate the fused segmentation result.
[0141] (15) The apparatus according to (9), wherein each of the first modality image processing portion and the second modality image processing portion has an encoder-decoder structure, the encoder comprising a plurality of downsampling convolutional layers, and the decoder comprising a plurality of upsampling convolutional layers.
[0142] Wherein, the first feature is output from the encoder of the first modality image processing section, and the first fused feature is input to the decoder of the first modality image processing section.
[0143] The second feature is output from the encoder of the second modality image processing section, and the second fused feature is input to the decoder of the second modality image processing section.
[0144] (16) The apparatus according to (9), wherein the multimodal pathological image further includes a third modality image obtained for the object in a third modality, and the processor is further configured to implement a third modality image processing portion of the neural network, the third modality image processing portion being configured to extract a third feature from the third modality image.
[0145] The feature fusion section is configured to fuse the second feature and the third feature based on the first feature to generate the first fused feature, fuse the first feature and the third feature based on the second feature to generate the second fused feature, and fuse the first feature and the second feature based on the third feature to generate the third fused feature.
[0146] The third modality image processing part is configured to perform image segmentation on the third modality image based on the third fusion feature, wherein the segmentation prediction results of the first modality image processing part, the second modality image processing part, and the third modality image processing part are merged as the final segmentation result.
[0147] The third-modality image processing part is trained based on the third-modality image training set and the third loss function.
[0148] In this process, a trained neural network is used to perform image segmentation on the multimodal pathological images to be segmented.
[0149] (17). A storage medium storing a computer program, which, when executed by a computer, causes the computer to perform an image segmentation method for multimodal pathological images according to (1)-(8).
Claims
1. A computer-implemented method for performing image segmentation on multimodal pathological images using a neural network, wherein the multimodal pathological images include at least a first modality image and a second modality image, the first modality image being an image of an object obtained in a first modality, the second modality image being an image of the object obtained in a second modality, the neural network including at least a first modality image processing part, a second modality image processing part, and a feature fusion part, the method comprising: The first modal image processing unit extracts a first feature from the first modal image, and the second modal image processing unit extracts a second feature from the second modal image; The feature fusion part fuses the second feature based on the first feature to generate a first fused feature provided to the first modal image processing part, and fuses the first feature based on the second feature to generate a second fused feature provided to the second modal image processing part; The first modality image processing part performs image segmentation on the first modality image based on the first fusion feature, and the second modality image processing part performs image segmentation on the second modality image based on the second fusion feature, wherein the segmentation prediction results of the first modality image processing part and the second modality image processing part are merged as the final segmentation result; The first modality image processing part is trained based on the first modality training image and the first loss function, and the second modality image processing part is trained based on the second modality training image and the second loss function; and Image segmentation is performed on multimodal pathological images to be segmented using a trained neural network.
2. The method according to claim 1, wherein, The neural network also includes a segmentation result fusion part. The method further includes: The segmentation prediction results of the first modality image processing part and the second modality image processing part are fused by the segmentation result fusion part to generate a fused segmentation result, wherein the segmentation prediction results of the first modality image processing part and the second modality image processing part, as well as the fused segmentation result, are combined to form the final segmentation result; and The fusion part of the segmentation result is trained based on the fusion loss function.
3. The method according to claim 2, wherein, Each of the first loss function, the second loss function, and the fusion loss function includes a cross-entropy loss function and a Dice loss function.
4. The method according to claim 1 or 2, further comprising: The neural network is trained based on an interactive loss function, which is based on the difference between the segmentation prediction results of the first modality image processing part and the segmentation prediction results of the second modality image processing part.
5. The method according to claim 1, further comprising: From the feature fusion portion, Add the first feature and the second feature element by element; The features obtained after addition are input into one or more convolutional layers; The features output from the convolutional layer are added element-wise to the first feature to generate the first fused feature; as well as The features output from the convolutional layer are added element-wise to the second feature to generate the second fused feature.
6. The method according to claim 2, further comprising: The fusion segmentation result is generated by element-wise multiplying the segmentation prediction result of the first modality image processing part with the segmentation prediction result of the second modality image processing part.
7. The method according to claim 1, wherein, Each of the first modality image processing part and the second modality image processing part has an encoder-decoder structure, wherein the encoder includes multiple downsampling convolutional layers and the decoder includes multiple upsampling convolutional layers. Wherein, the first feature is output from the encoder of the first modality image processing section, and the first fused feature is input to the decoder of the first modality image processing section. The second feature is output from the encoder of the second modality image processing section, and the second fused feature is input to the decoder of the second modality image processing section.
8. The method according to claim 1, wherein, The multimodal pathological image also includes a third-modal image obtained for the object in the third modality, and the neural network further includes a third-modal image processing part. The method further includes: The third modality image processing unit extracts a third feature from the third modality image; The feature fusion component fuses the second feature and the third feature based on the first feature to generate the first fused feature, fuses the first feature and the third feature based on the second feature to generate the second fused feature, and fuses the first feature and the second feature based on the third feature to generate the third fused feature. The third modality image processing part performs image segmentation on the third modality image based on the third fusion feature, wherein the segmentation prediction results of the first modality image processing part, the second modality image processing part, and the third modality image processing part are merged to form the final segmentation result; The third-modality image processing component is trained based on the third-modality image training set and the third loss function; and Image segmentation is performed on multimodal pathological images to be segmented using a trained neural network.
9. An apparatus for performing image segmentation on a multimodal pathological image using a neural network, wherein the multimodal pathological image includes at least a first modality image and a second modality image, the first modality image being an image of an object obtained in a first modality, and the second modality image being an image of the object obtained in a second modality, the apparatus comprising: A memory that stores computer program instructions; as well as One or more processors configured to implement at least the following portions of the neural network by executing the computer program instructions: The first modality image processing part is configured to extract a first feature from the first modality image and perform image segmentation on the first modality image based on the first fusion feature; The second modality image processing section is configured to extract a second feature from the second modality image and perform image segmentation on the second modality image based on the second fusion feature; as well as The feature fusion section is configured to fuse the second feature based on the first feature to generate the first fused feature provided to the first modality image processing section, and to fuse the first feature based on the second feature to generate the second fused feature provided to the second modality image processing section. In this process, the segmentation prediction results from the first modality image processing part and the second modality image processing part are combined to form the final segmentation result. Specifically, the first modality image processing part is trained based on the first modality training image and the first loss function, and the second modality image processing part is trained based on the second modality training image and the second loss function.
10. A storage medium storing a computer program, which, when executed by a computer, causes the computer to perform an image segmentation method for multimodal pathological images according to any one of claims 1-8.
Citation Information
Patent Citations
Image segmentation method and image segmentation device
CN110009598A
Multimodal three-dimensional medical image fusion method and system, and electronic device
WO2021022752A1