A Multi-View Multi-Modal Image Segmentation Method and System
By generating convolution kernels with different views and modalities for image segmentation, the problem of failing to fully consider multi-view multi-modal features in the prior art is solved, and higher segmentation accuracy and computing efficiency are achieved.
Patent Information
- Application Number
- CN202210044339.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-14
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-01-14
AI Technical Summary
The existing multi-view and multi-modal medical image segmentation methods fail to fully consider the multi-view multi-modal features of the image, resulting in large errors in the segmentation results.
A multi-view multi-modal image segmentation method is adopted, by obtaining multi-modal multi-view images, encoding and generating a first feature map, decoding and generating a second feature map, and using prior information to generate different convolution kernels for segmentation operations, fully considering the multi-view multi-modal features of the image.
It improves the accuracy of image segmentation, can help doctors more accurately detect the probability of disease, is highly practical and scalable, and reduces network parameters, improves computing efficiency and speed.
Smart Images

Figure CN114549411B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and particularly relates to a multi-view multi-modal image segmentation method and system. Background Art
[0002] The statements in this part only provide background technical information related to the present invention and do not necessarily constitute prior art.
[0003] Medical image segmentation is an important research branch of medical image processing, which is currently a hot topic in medical image processing and has important clinical diagnostic and therapeutic value. Medical image segmentation is the process of dividing a medical image into several non-overlapping "connected" regions according to certain similarity features of the medical image (such as brightness, color texture, etc.). The relevant features show consistency or similarity in the same region, while showing obvious differences in different regions. That is to say, there is a certain discontinuity characteristic of pixels at the region boundary. It is a complex and crucial step in the field of medical image processing and analysis. Its purpose is to segment out the parts with certain special meanings in the medical image, extract relevant features, provide a reliable basis for clinical diagnosis and treatment and pathological research, and assist doctors to make more accurate diagnoses.
[0004] Currently, many methods have been proposed for multi-view and multi-modal image segmentation, such as segmentation models based on thresholds, deformation models, neural networks, etc. Among them, multi-view and multi-modal images include mammary gland images, and its two views are cranio-caudal (CC) and mediolateral oblique (MLO) respectively. There are two different modalities under each view, namely low energy (LE) and high energy (HE) two energy images. As far as the inventor knows, most of the neural network-based image segmentation models do not take into account the characteristics of multi-view and multi-modal images. For example, a fixed neural network is trained by network models such as FCN network and UNet network for image segmentation methods. Due to the multi-view characteristics of images, such methods often have certain limitations when applied to image segmentation with multi-view characteristics. On the one hand, these methods regard all pictures as a kind of image and do not consider the characteristics of multi-view and multi-modal images; on the other hand, images with different views and modalities have different characteristics, and there should be a large error in the results obtained by using a fixed convolution kernel to convolve images with different views and modalities. Summary of the Invention
[0005] In order to solve the technical problems existing in the above background art, the present invention provides a multi-view multi-modal image segmentation method and system, which can utilize the multi-view multi-modal characteristics of images to improve the accuracy of segmentation.
[0006] In order to achieve the above object, the present invention adopts the following technical solutions:
[0007] The first aspect of the present invention provides a multi-view multi-modal image segmentation method, which includes:
[0008] Obtaining a set of multi-modal multi-view images;
[0009] Encoding each multi-modal multi-view image to generate a first feature map;
[0010] Decoding the first feature map to generate a second feature map;
[0011] Selecting any view and modality information in a set of multi-modal multi-view images as prior information and encoding it into a one-hot vector;
[0012] Performing global pooling operation on the first feature map to obtain an aggregated feature and concatenating it with the one-hot vector;
[0013] Generating a convolution kernel for segmenting the corresponding view and modality based on the concatenated vector and the convolution kernel generator;
[0014] Performing a convolution operation on the second feature map using the convolution kernel to obtain the segmentation result of the corresponding view and modality image.
[0015] The second aspect of the present invention provides a multi-view multi-modal image segmentation system, which includes:
[0016] An image acquisition module for obtaining a set of multi-modal multi-view images;
[0017] A first feature map generation module for encoding each multi-modal multi-view image to generate a first feature map;
[0018] A second feature map generation module for decoding the first feature map to generate a second feature map;
[0019] A prior information encoding module for selecting any view and modality information in a set of multi-modal multi-view images as prior information and encoding it into a one-hot vector;
[0020] A vector concatenation module for performing global pooling operation on the first feature map to obtain an aggregated feature and concatenating it with the one-hot vector;
[0021] A convolution kernel generation module for generating a convolution kernel for segmenting the corresponding view and modality based on the concatenated vector and the convolution kernel generator;
[0022] An image segmentation module for performing a convolution operation on the second feature map using the convolution kernel to obtain the segmentation result of the corresponding view and modality image.
[0023] The third aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps in the multi-view multi-modal image segmentation method as described above are implemented.
[0024] The fourth aspect of the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps in the multi-view multi-modal image segmentation method as described above are implemented.
[0025] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0026] (1) In terms of the image segmentation effect, the present invention first proposes a technique for segmenting multi-view multi-modal breast images using a single network. It generates different convolution kernels according to different views and modalities as prior information for the parameters of the final convolution operation, fully considering the multi-modal and multi-view characteristics of breast images, and improving the accuracy of image segmentation.
[0027] (2) In terms of practicability and scalability, the multi-view multi-modal breast image segmentation method of the present invention improves the accuracy of image segmentation for different views and different modalities, and can more accurately help doctors detect the probability of corresponding diseases based on images. Therefore, it has strong practicability. At the same time, this method can also be used in other multi-modal and multi-view medical images, so it also has good scalability.
[0028] (3) In terms of operation efficiency and speed, the method of the present invention based on encoding and decoding feature extraction realizes the segmentation of breast images with different modalities and different views using only a single network, reducing the parameters of the network. Therefore, the operation speed is faster and the efficiency is higher.
[0029] The advantages of the additional aspects of the present invention will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] The accompanying drawings forming a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.
[0031] Figure 1 is a flowchart of the multi-view multi-modal image segmentation method according to an embodiment of the present invention;
[0032] Figure 2 is a schematic diagram of the multi-view multi-modal image segmentation method according to an embodiment of the present invention;
[0033] FIG. 3(a) is the first type of breast image according to an embodiment of the present invention;
[0034] Figure 3(b) is the second type of breast image according to the embodiment of the present invention;
[0035] Figure 3(c) is the third type of multi-modal breast image according to the embodiment of the present invention;
[0036] Figure 3(d) is the fourth type of breast image according to the embodiment of the present invention;
[0037] Figure 4(a) is the tumor segmentation result of the first type of breast image according to the embodiment of the present invention;
[0038] Figure 4(b) is the tumor segmentation result of the second type of breast image according to the embodiment of the present invention;
[0039] Figure 4(c) is the tumor segmentation result of the third type of breast image according to the embodiment of the present invention;
[0040] Figure 4(d) is the tumor segmentation result of the fourth type of breast image according to the embodiment of the present invention;
[0041] Figure 5 It is a schematic structural diagram of a multi-view multi-modal image segmentation system according to the embodiment of the present invention. Detailed implementation manners
[0042] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0043] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0044] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they specify the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0045] Embodiment 1
[0046] As Figure 1 shown, this embodiment provides a multi-view multi-modal image segmentation method, which specifically includes the following steps:
[0047] S101: Obtain a set of multi-modal multi-view images.
[0048] In this embodiment, the multi-modal multi-view images take breast images as an example.
[0049] It should be noted here that the multi-view multi-modal image segmentation method provided in this embodiment is applicable to other images in addition to breast images, such as cardiac images and so on.
[0050] Breast images include two views: craniocaudal (CC) and mediolateral oblique (MLO). For each radiographic position, two different modal images, low energy (LE) and high energy (HE), can be obtained, resulting in a total of four types of breast images as Figures 3(a) - 3(d) shown.
[0051] These four different types of images are respectively represented as
[0052] where represents the i-th type of breast image containing the n i data set. Y ij represents the ground truth label of the j-th breast image of the i-th type, and X ij is the corresponding predicted segmentation result.
[0053] This embodiment realizes the segmentation of multi-modal multi-view breast images using the same network. Therefore, for this network architecture, only the parameter θ needs to be continuously updated during the training process to minimize the loss function, that is:
[0054]
[0055] Regarding the segmentation of the breast and the tumor as a binary segmentation task, binary cross-entropy loss and DiceLoss are jointly used as the objective for each task. Therefore, the loss function can be described as:
[0056]
[0057] where p i and y i represent the predicted value and the ground truth value of the i-th pixel respectively. V represents the number of all pixels, and ∈ is added as a smoothing factor.
[0058] S102: Encode each multi-modal multi-view image to generate a first feature map.
[0059] In step S102, an encoder is used to encode each multi-modal multi-view image to generate a first feature map.
[0060] Specifically, convolutional and pooling downsampling operations are sequentially performed on the multi-modal multi-view image to obtain the first feature map for feature extraction.
[0061] As Figure 2 shown in process 1, the feature map F for feature extraction is obtained through a series of convolutional and pooling downsampling operations of the encoder ijAfter each convolutional layer, group normalization and the ReLU activation function are applied. Downsampling is performed using a convolutional layer with a stride of 2. The number of convolutional kernels is set to 32 in the first layer and doubled after each downsampling step. For each given input image X ij ,
[0062] F ij = f E (X ij ; θ E )
[0063] where F ij represents the first feature map generated by the encoder, f E represents the entire encoder architecture, and θ E represents the parameters of the entire encoder.
[0064] S103: Decode the first feature map to generate a second feature map.
[0065] In step S103, the decoder is used to decode the first feature map to generate a second feature map.
[0066] Specifically, in the decoder, the first feature map is continuously upsampled and added to the low-level feature maps of the encoder, and then refined through a residual network to obtain a second feature map with all the semantic information of a set of pictures.
[0067] As shown in process 6 of Figure 2 , the feature map F ij generated by the encoder is continuously upsampled and added to the low-level feature maps of the encoder, and then refined through a residual network to obtain a feature map M ij with all the semantic information of a set of pictures:
[0068] M ij = f D (F ij ; θ D )
[0069] where f D represents the decoder architecture and θ D represents the parameters of the entire decoder.
[0070] Finally, the number of channels of M ij is 8. The feature map M ij generated by this process has the semantic information of four pictures and is not restricted by task encoding.
[0071] Among them, the above encoder and decoder constitute the UNet network model architecture.
[0072] S104: Select any view and modality information from a set of multimodal multi-view images as prior information and encode it into a one-hot vector.
[0073] Which type of view and modality of images to be segmented is an important prior information. Therefore, a method similar to the attention mechanism is used to encode the prior information into a one-hot vector.
[0074] For example: Encode the low-energy (LE) image in the cranio-caudal (CC) view into a one-hot vector: e1 = (1, 0, 0, 0) T ; Similarly, encode the high-energy (HE) in the cranio-caudal (CC) view into a one-hot vector: e2 = (0, 1, 0, 0) T ; Encode the low-energy (LE) in the mediolateral oblique (MLO) view into a one-hot vector: e3 = (0, 0, 1, 0) T ; Encode the high-energy (HE) in the mediolateral oblique (MLO) view into a one-hot vector: e4 = (0, 0, 1, 0) T .
[0075] S105: Perform global pooling operation on the first feature map to obtain an aggregated feature and concatenate it with the one-hot vector.
[0076] As Figure 2 shown in Process 2, perform global pooling operation (GAP) on the feature map F generated by the encoder ij to obtain a vector with aggregated features. Then concatenate this vector with the encoded vector into the same vector. Then this vector will have both the feature information of the input image and the view and modality information of the breast image.
[0077] S106: Based on the concatenated vector and the convolutional kernel generator, generate convolutional kernels for segmenting the corresponding views and modalities.
[0078] As Figure 2 shown in Process 3, use the concatenated vector in Step 4 as the input of the convolutional kernel generator ψ(·) to generate convolutional kernels for segmenting specific views and modalities.
[0079] ω ij = ψ(GAP(F ij ) || T ij ; θ ψ )
[0080] where GAP represents the global pooling operation on the first feature map, and θ ψ represents all the parameters of this convolutional kernel generator. || represents the concatenation operation. T ij represents the one-hot vector encoded in the S104 process. ωij Represents the convolution kernel generated by the convolution kernel generator.
[0081] S107: Perform a convolution operation on the second feature map using the convolution kernel to obtain the segmentation results for the corresponding views and modality images.
[0082] Use the generated convolution kernel to perform a convolution operation on the feature map M generated by the decoder. ij As shown in processes 4 and 5 in Figure 2 , the feature map M ij is a feature map with semantic information of four types of pictures and not restricted by task encoding. Performing a convolution operation on it with the convolution kernel generated by the vector with encoding information can then segment the breast pictures of specific views and modalities, achieving the goal of segmenting breast images of different views and modalities by one network. The convolution kernel parameters are divided into three layers, with eight channels in the first two layers and two channels in the last layer. The two channels are used to segment the breast and tumors:
[0083] P ij = ((M ij * ω ij1 ) * ω ij2 ) * ω ij3
[0084] where * represents convolution, and P ij represents the prediction of the breast and tumors. The depth is set to 3, the width of the first two convolution layers is set to 8, and the width of the last convolution layer is set to 1, which is used to segment the tumors of specific views and modules. The results are as shown in Figures 4(a) - 4(d) .
[0085] Embodiment 2
[0086] As shown in Figure 5 , this embodiment provides a multi-view multi-modal image segmentation system, which specifically includes the following modules:
[0087] (1) Image acquisition module, which is used to acquire a set of multi-modal multi-view images.
[0088] (2) First feature map generation module, which is used to encode each multi-modal multi-view image to generate the first feature map.
[0089] Specifically, in the first feature map generation module, use the encoder to encode each multi-modal multi-view image to generate the first feature map.
[0090] In the encoder, perform convolution and pooling downsampling operations on the multi-modal multi-view image in sequence to obtain the first feature map for feature extraction.
[0091] (3) The second feature map generation module is used to decode the first feature map to generate a second feature map.
[0092] Specifically, in the second feature map generation module, a decoder is used to decode the first feature map to generate a second feature map.
[0093] In the decoder, the first feature map is continuously subjected to upsampling operations and added to the low-level feature maps of the encoder, and then refined through a residual network to obtain a second feature map with all semantic information of a set of pictures.
[0094] (4) The prior information encoding module is used to select any view and modal information from a set of multi-modal multi-view images as prior information and encode it into a one-hot vector.
[0095] (5) The vector concatenation module is used to perform global pooling operations on the first feature map to obtain an aggregated feature and concatenate it with the one-hot vector.
[0096] (6) The convolution kernel generation module is used to generate convolution kernels for segmenting corresponding views and modalities based on the concatenated vector and a convolution kernel generator.
[0097] (7) The image segmentation module is used to perform convolution operations on the second feature map using the convolution kernels to obtain segmentation results of corresponding view and modal images.
[0098] It should be noted here that each module in this embodiment corresponds one by one to each step in Embodiment 1, and its specific implementation process is the same, so it will not be repeated here.
[0099] Embodiment 3
[0100] This embodiment provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps in the multi-view multi-modal image segmentation method as described above.
[0101] Embodiment 4
[0102] This embodiment provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in the multi-view multi-modal image segmentation method as described above.
[0103] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device generate means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0104] The foregoing is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A multi-view multi-modal image segmentation method, characterized in that, Including: Obtain a set of multi-modal multi-view images; The multi-modal multi-view images are breast images, which include two views of craniocaudal and mediolateral oblique, and each imaging modality can obtain two different modality images of low energy and high energy, resulting in a total of four types of breast images; Four different types of images are respectively represented as Among them represents the n i th type of breast image in the dataset, and Y ij represents the true label of the jth breast image in the ith type, and X ij is the corresponding predicted segmentation result; Regarding the segmentation of the breast and tumor as a binary segmentation task, jointly use binary cross-entropy loss and Dice Loss as the objective for each task, and the loss function is: where p i and y i represent the predicted value and the true value of the i-th pixel respectively, V represents the total number of all pixels, and ∈ is added as a smoothing factor; Encode each multi-modal multi-view image to generate a first feature map; Decode the first feature map to generate a second feature map; Use the decoder to decode the first feature map to generate a second feature map; In the decoder, continuously perform upsampling operations on the first feature map and add it to the low-level feature map of the encoder, and then refine it through a residual network to obtain a second feature map with all semantic information of a set of pictures; Select any view and modality information in a set of multi-modal multi-view images as prior information and encode it as a one-hot vector; Encode the low-energy image in the cranio-caudal view as a one-hot vector: e1 = (1, 0, 0, 0) T ; Encode the high-energy in the cranio-caudal view as a one-hot vector: e2 = (0, 1, 0, 0) T ; Encode the low-energy in the medial-lateral oblique view as a one-hot vector: e3 = (0, 0, 1, 0) T ; Encode the high-energy in the medial-lateral oblique view as a one-hot vector: e2 = (0, 0, 0, 1) T ; Perform global pooling operation on the first feature map to obtain an aggregated feature, and concatenate it with the one-hot vector; Based on the concatenated vector and the convolutional kernel generator, generate convolutional kernels for segmenting the corresponding view and modality; The convolutional kernel parameters are divided into three layers, with eight channels in the first two layers and two channels in the last layer. The two features are used to segment the breast and tumor; Use the convolutional kernel to perform convolution operation on the second feature map to obtain the segmentation result of the corresponding view and modality image.
2. The multi-view multi-modal image segmentation method according to claim 1, characterized in that, Use the encoder to encode each multi-modal multi-view image to generate a first feature map.
3. The multi-view multi-modal image segmentation method according to claim 2, characterized in that, In the encoder, sequentially perform convolution and pooling downsampling operations on the multi-modal multi-view image to obtain the first feature map for feature extraction.
4. A multi-view multi-modal image segmentation system, characterized in that, Including: An image acquisition module for obtaining a set of multi-modal multi-view images; The multi-modal multi-view images are breast images, which include two views of craniocaudal and mediolateral oblique, and each imaging modality can obtain two different modality images of low energy and high energy, resulting in a total of four types of breast images; Four different types of images are respectively represented as Among them represents the $n$ i th type of breast image in the dataset, and $Y$ ij represents the true label of the $j$th breast image in the $i$th type, and $X$ ij is the corresponding predicted segmentation result; Regarding the segmentation of the breast and tumor as a binary segmentation task, jointly use binary cross-entropy loss and Dice Loss as the objective for each task, and the loss function is: where p i and y i represent the predicted value and the true value of the i-th pixel respectively, V represents the total number of all pixels, and ∈ is added as a smoothing factor; A first feature map generation module for encoding each multi-modal multi-view image to generate a first feature map; A second feature map generation module for decoding the first feature map to generate a second feature map; Use the decoder to decode the first feature map to generate a second feature map; In the decoder, continuously perform upsampling operations on the first feature map and add it to the low-level feature map of the encoder, and then refine it through a residual network to obtain a second feature map with all semantic information of a set of pictures; A prior information encoding module for selecting any view and modality information in a set of multi-modal multi-view images as prior information and encoding it as a one-hot vector; Encode the low-energy image in the cranio-caudal view as a one-hot vector: e1 = (1, 0, 0, 0) T ; Encode the high-energy in the cranio-caudal view as a one-hot vector: e2 = (0, 1, 0, 0) T ; Encode the low-energy in the medial-lateral oblique view as a one-hot vector: e3 = (0, 0, 1, 0) T ; Encode the high-energy in the medial-lateral oblique view as a one-hot vector: e2 = (0, 0, 0, 1) T ; A vector concatenation module for performing global pooling operation on the first feature map to obtain an aggregated feature and concatenating it with the one-hot vector; A convolutional kernel generation module for generating convolutional kernels for segmenting the corresponding view and modality based on the concatenated vector and the convolutional kernel generator; The convolution kernel parameters are divided into three layers, with eight channels in the first two layers and two channels in the last layer. Two features are used to segment the breast and the tumor; An image segmentation module, which is used to perform a convolution operation on the second feature map by using the convolution kernel to obtain the segmentation results of the corresponding views and modality images.
5. The multi-view multi-modal image segmentation system according to claim 4, wherein, In the first feature map generation module, an encoder is used to encode each multi-modal multi-view image to generate a first feature map.
6. A computer-readable storage medium, on which a computer program is stored, wherein, When the program is executed by a processor, it implements the steps in the multi-view multi-modal image segmentation method according to any one of claims 1-3.
7. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, When the processor executes the program, it implements the steps in the multi-view multi-modal image segmentation method according to any one of claims 1-3.
Citation Information
Patent Citations
Division and identification method and device based on dense network image
CN108062754A
Medical image segmentation method based on task dynamic learning partial marker
CN113112454A