Multi-modal image segmentation model training method and multi-modal image segmentation method
By introducing a dual-stream fusion module into the multimodal image segmentation model, the image segmentation accuracy problem caused by the single-stream structure in the prior art is solved, and higher image feature extraction and segmentation accuracy are achieved.
Patent Information
- Application Number
- CN202510065374.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-13
AI Technical Summary
The existing multimodal fusion image segmentation method uses a single-stream structure, which may cause distortions where image information overlaps, making it difficult to accurately capture the relevant information between multiple modal images, affecting the accuracy of segmented images.
The training method of multimodal image segmentation model is adopted, including obtaining a multimodal sample image set, inputting it into the initial unbiased multimodal fusion segmentation network, feature extraction and fusion are performed through feature extraction module and dual-stream fusion module, and training the network based on the loss function value until the training cutoff condition is met.
Through the feature fusion of the dual-stream fusion module, the accuracy of image feature extraction is improved, the accuracy of image segmentation is enhanced, and the distortion problem of single-stream structure at the overlap of image information is solved.
Smart Images

Figure CN119992251A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a training method for a multimodal image segmentation model and a multimodal image segmentation method. Background Art
[0002] Since a single image contains limited information, it is very difficult to perform tasks such as target detection and image segmentation based on a single image, so multimodal image fusion technology came into being. Multimodal image fusion refers to the joint processing of images from different imaging methods or different devices to obtain a fused image containing multiple information, thereby integrating the advantages of different modal images and extracting more comprehensive information.
[0003] Currently, most of the segmentation methods for multimodal fusion images adopt a single-stream structure. Using a single-stream structure to perform feature extraction and other processing on the fused image may cause distortion where the image information overlaps, and it is difficult to accurately capture the relevant information between multiple modal images, thereby affecting the accuracy of the segmented image. Summary of the invention
[0004] In view of this, the present application is committed to providing a training method for a multimodal image segmentation model and a multimodal image segmentation method to improve the accuracy of image segmentation.
[0005] In a first aspect, the present application provides a method for training a multimodal image segmentation model, the method comprising:
[0006] Acquire a set of multimodal sample images of a target object, wherein the set of multimodal sample images includes sample images of multiple modalities and pre-annotated real label regions in the sample images;
[0007] Inputting the multimodal sample image set into an initial unbiased multimodal fusion segmentation network to obtain an image segmentation result of the target object, wherein the image segmentation result includes a predicted label area;
[0008] Determine a loss function value based on the true label area and the predicted label area;
[0009] Based on the loss function value, the initial unbiased multimodal fusion segmentation network is trained until a training cutoff condition is met to obtain a trained unbiased multimodal fusion segmentation network;
[0010] Among them, the initial unbiased multimodal fusion segmentation network includes an image fusion function module, and the image fusion function module includes a feature extraction module and a dual-stream fusion module. The feature extraction module is used to extract the image features of the sample image, and the dual-stream fusion module is used to perform feature fusion based on the union and intersection of the multiple image features.
[0011] In a second aspect, the present application provides a multimodal image segmentation method, the method comprising:
[0012] Acquire a multimodal image of the object to be processed;
[0013] The multimodal image is input into an unbiased multimodal fusion segmentation network to obtain an image segmentation result of the object to be processed, wherein the unbiased multimodal fusion segmentation network is trained based on the training method of the multimodal image segmentation model described in any one of the implementation modes of the first aspect above.
[0014] In a third aspect, the present application provides a training device for a multimodal image segmentation model, the device comprising:
[0015] A first acquisition unit is used to acquire a multimodal sample image set of a target object, wherein the multimodal sample image set includes sample images of multiple modalities and real label areas pre-annotated in the sample images;
[0016] A first image segmentation unit, used for inputting the multimodal sample image set into an initial unbiased multimodal fusion segmentation network to obtain an image segmentation result of the target object, wherein the image segmentation result includes a predicted label area;
[0017] A loss function determining unit, configured to determine a loss function value based on the true label region and the predicted label region;
[0018] A training unit is used to train the initial unbiased multimodal fusion segmentation network based on the loss function value until the training cutoff condition is met, thereby obtaining a trained unbiased multimodal fusion segmentation network; wherein the initial unbiased multimodal fusion segmentation network includes an image fusion function module, the image fusion function module includes a feature extraction module and a dual-stream fusion module, the feature extraction module is used to extract the image features of the sample image, and the dual-stream fusion module is used to perform feature fusion based on the union and intersection of the multiple image features.
[0019] In a fourth aspect, the present application provides a multimodal image segmentation device, the device comprising:
[0020] A second acquisition unit, used to acquire a multimodal image of the object to be processed;
[0021] The second image segmentation unit is used to input the multimodal image into an unbiased multimodal fusion segmentation network to obtain an image segmentation result of the object to be processed, wherein the unbiased multimodal fusion segmentation network is trained based on the training method of the multimodal image segmentation model described in any one of the implementation methods of the first aspect above.
[0022] In a fifth aspect, the present application provides an electronic device, the device comprising: a memory and a processor;
[0023] The memory is used to store relevant program codes;
[0024] The processor is used to call the program code to execute the training method of the multimodal image segmentation model described in any one of the implementations of the first aspect or the multimodal image segmentation method described in any one of the implementations of the second aspect.
[0025] In a sixth aspect, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium is used to store a computer program, and the computer program is used to execute the training method of the multimodal image segmentation model described in any one of the implementations of the first aspect or the multimodal image segmentation method described in any one of the implementations of the second aspect.
[0026] In the seventh aspect, the present application provides a computer program product, which includes a computer program / instructions, and when the computer program / instructions are executed by a processor, the training method of the multimodal image segmentation model described in any one of the implementations of the first aspect or the multimodal image segmentation method described in any one of the implementations of the second aspect is implemented.
[0027] In the above implementation of the present application, in order to train the model for image segmentation, it is first necessary to obtain a set of multimodal sample images of the target object. Among them, the multimodal sample image set includes sample images of multiple modes and real label areas pre-marked in the sample images. That is, the real label area in the sample image is used as the training target, and the real label area represents the label type corresponding to the area in the sample image. The multimodal sample image set is input into the initial unbiased multimodal fusion segmentation network to obtain the image segmentation result of the target object. Among them, the image segmentation result includes the predicted label area, that is, the label type corresponding to the area in the sample image predicted by the initial unbiased multimodal fusion segmentation network. Then, the loss function value is determined based on the real label area and the predicted label area, which represents the error between the real label area and the predicted label area. The initial unbiased multimodal fusion segmentation network is trained based on the loss function value until the training cutoff condition is met to obtain a trained unbiased multimodal fusion segmentation network. Among them, the initial unbiased multimodal fusion segmentation network includes an image fusion function module, and the image fusion function module includes a feature extraction module and a dual-stream fusion module. The feature extraction module is used to extract image features of the sample image, and the dual-stream fusion module is used to perform feature fusion based on the union and intersection of multiple image features. Through the method provided by the present application, since the initial unbiased multimodal fusion segmentation network includes a dual-stream fusion module, feature fusion can be performed using the union and intersection of multiple image features, which can not only ensure the integrity of image features, but also consider the common characteristics between image features, thereby improving the accuracy of image feature extraction, thereby improving the accuracy of image segmentation. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. It is obvious that the drawings described below are only some embodiments provided in the present application, and a person skilled in the art can also obtain other drawings based on these drawings.
[0029] Figure 1 A flowchart of a method for training a multimodal image segmentation model provided in an embodiment of the present application.
[0030] Figure 2 A schematic diagram of the structure of an initial unbiased multimodal fusion segmentation network provided in an embodiment of the present application.
[0031] Figure 3 A schematic diagram of the structure of an image fusion function module provided in an embodiment of the present application.
[0032] Figure 4 A schematic diagram of the structure of a dual-stream cross-selection conversion model provided in an embodiment of the present application.
[0033] Figure 5 A schematic diagram of the structure of a dual-stream fusion module provided in an embodiment of the present application.
[0034] Figure 6 A schematic diagram of the structure of an image segmentation function module provided in an embodiment of the present application.
[0035] Figure 7 A schematic diagram of an image segmentation result provided in an embodiment of the present application.
[0036] Figure 8 A flowchart of a multimodal image segmentation method provided in an embodiment of the present application.
[0037] Fig. 9 A schematic diagram of a training device for a multimodal image segmentation model provided in an embodiment of the present application.
[0038] Fig.10 A schematic diagram of a multimodal image segmentation device provided in an embodiment of the present application.
[0039] Fig.11 A schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0040] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. The described embodiments are only exemplary implementation methods of the present application, not all implementation methods. Those skilled in the art can combine the embodiments of the present application to obtain other embodiments without creative work, and these embodiments are also within the scope of protection of the present application.
[0041] Since a single image contains limited information, it is very difficult to perform tasks such as target detection and image segmentation based on a single image, so multimodal image fusion technology came into being. Multimodal image fusion refers to the joint processing of images from different imaging methods or different devices to obtain a fused image containing multiple information, thereby integrating the advantages of different modal images and extracting more comprehensive information.
[0042] Taking the medical application field as an example, image analysis is an indispensable part of disease diagnosis, treatment and scientific research. However, single-modality medical images often fail to provide comprehensive and accurate information, so it is often necessary to fuse medical images of multiple modalities and perform image analysis based on the fused images. Multimodal images in the medical field can include computed tomography (CT), magnetic resonance imaging (MRI), positron emission tomography (PET), etc.
[0043] Currently, most of the segmentation methods for multimodal fusion images adopt a single-stream structure. Using a single-stream structure to perform feature extraction and other processing on the fused image may cause distortion where the image information overlaps, and it is difficult to accurately capture the relevant information between multiple modal images, thereby affecting the accuracy of the segmented image.
[0044] Based on this, an embodiment of the present application provides a training method for a multimodal image segmentation model to improve the accuracy of image segmentation. In specific implementation, it is first necessary to obtain a set of multimodal sample images of the target object. Among them, the multimodal sample image set includes sample images of multiple modes and real label areas pre-marked in the sample images. That is, the real label area in the sample image is used as the training target, and the real label area represents the label type corresponding to the area in the sample image. The multimodal sample image set is input into the initial unbiased multimodal fusion segmentation network to obtain the image segmentation result of the target object. Among them, the image segmentation result includes a predicted label area, that is, the label type corresponding to the area in the sample image predicted by the initial unbiased multimodal fusion segmentation network. Then, the loss function value is determined based on the real label area and the predicted label area, indicating the error between the real label area and the predicted label area. The initial unbiased multimodal fusion segmentation network is trained based on the loss function value until the training cutoff condition is met to obtain a trained unbiased multimodal fusion segmentation network. Among them, the initial unbiased multimodal fusion segmentation network includes an image fusion function module, and the image fusion function module includes a feature extraction module and a dual-stream fusion module. The feature extraction module is used to extract image features of the sample image, and the dual-stream fusion module is used to perform feature fusion based on the union and intersection of multiple image features. Through the method provided by the present application, since the initial unbiased multimodal fusion segmentation network includes a dual-stream fusion module, feature fusion can be performed using the union and intersection of multiple image features, which can not only ensure the integrity of image features, but also consider the common characteristics between image features, thereby improving the accuracy of image feature extraction, thereby improving the accuracy of image segmentation.
[0045] In order to facilitate understanding of the technical solution provided by the embodiments of the present application, a detailed introduction will be given below in conjunction with the drawings in the embodiments.
[0046] See also Figure 1 As shown, it is a flowchart of a training method of a multimodal image segmentation model provided in an embodiment of the present application.
[0047] The method may be performed by a model training device. Optionally, the model training device may be a terminal device or a server.
[0048] The method may include the following steps:
[0049] S101: Acquire a set of multimodal sample images of a target object.
[0050] In order to train the model, it is necessary to obtain training samples. Among them, multimodal images can represent images taken from different imaging methods or different devices for the target object. For example, in the medical field, multimodal images can represent CT, MRI, PET, etc., and the target object can represent an organ including a lesion. In the field of autonomous driving, multimodal images can represent RGB images taken by cameras, point cloud images taken by radars, etc.
[0051] The multimodal sample image set includes sample images of multiple modalities, and a true label region is pre-marked in each sample image, and the true label region represents the label type corresponding to different regions in the sample image, wherein the label type may be one or more. That is, the sample image may include one or more categories of true label regions.
[0052] In a possible implementation, different numbers may be used to represent labels of different categories, so that the category of each true label area may be determined.
[0053] S102: Input the multimodal sample image set into the initial unbiased multimodal fusion segmentation network to obtain the image segmentation result of the target object.
[0054] Among them, the image segmentation result can be understood as an image with the same size as the sample image, and the image segmentation result includes a predicted label area, which represents the label category corresponding to different areas in the sample image predicted and output by the initial unbiased multimodal fusion segmentation network. That is, the image segmentation result can include one or more categories of predicted label areas. The initial unbiased multimodal fusion segmentation network represents the initial model for fusion segmentation of a multimodal sample image set. By using the multimodal sample image set to train the initial unbiased multimodal fusion segmentation network, it can accurately perform image segmentation.
[0055] In a possible implementation, the initial unbiased multimodal fusion segmentation network includes an image fusion function module, which includes a feature extraction module and a dual-stream fusion module (DFM). Among them, the feature extraction module is used to extract the image features of the sample image, and the dual-stream fusion module is used to perform feature fusion based on the union and intersection of multiple image features, so as to perform image segmentation based on the fused image features later. Among them, feature fusion based on the union of image features of multiple sample images can ensure the integrity of multimodal image features and extract richer image feature information. Feature fusion based on the intersection of image features of multiple sample images can extract the common points between sample images of different modalities, consider the consistency between multiple image features, so that image features can be extracted more accurately. It should be noted that the specific structure of the initial unbiased multimodal fusion segmentation network can be referred to in the subsequent embodiments, and will not be introduced in detail here.
[0056] S103: Determine a loss function value based on the true label area and the predicted label area.
[0057] After the initial unbiased multimodal fusion segmentation network outputs the predicted label area, since each sample image includes a pre-annotated true label area as the sample true value, the loss function value can be determined based on the true label area and the predicted label area, that is, the error between the true label area and the predicted label area. When the loss function value is larger, it indicates that the error between the true label area and the predicted label area is also larger, and the accuracy of the initial unbiased multimodal fusion segmentation network is lower.
[0058] S104: Training the initial unbiased multimodal fusion segmentation network based on the loss function value until a training cutoff condition is met, thereby obtaining a trained unbiased multimodal fusion segmentation network.
[0059] Since the loss function value can reflect the accuracy of the initial unbiased multimodal fusion segmentation network, when the loss function value is greater than the preset threshold, it indicates that the accuracy of the initial unbiased multimodal fusion segmentation network does not meet the requirements, and the parameters of the initial unbiased multimodal fusion segmentation network need to be readjusted for training. That is, after the parameters of the initial unbiased multimodal fusion segmentation network are readjusted, the process of inputting the multimodal sample image set into the initial unbiased multimodal fusion segmentation network after adjusting the parameters and the subsequent process of determining the loss function value is continued, and the above process is iteratively executed until the iterative training is stopped when the training cutoff condition is met, and the trained unbiased multimodal fusion segmentation network (UMF-SegNet) is obtained.
[0060] In a possible implementation, the training cutoff condition may be that the loss function value is less than a preset threshold, or the number of iterative training reaches a preset number, and the iterative training process can be stopped when either of the two conditions is met. That is, after iteratively executing the input of the multimodal sample image set into the initial unbiased multimodal fusion segmentation network after adjusting the parameters, the image segmentation result of the target object is re-obtained, and the image segmentation result includes the predicted label area. The loss function value is re-determined based on the predicted label area and the true label area. When the loss function value is less than the preset threshold, or the cumulative number of iterative training reaches the preset number, the training process is stopped to obtain the trained unbiased multimodal fusion segmentation network.
[0061] Through the method provided in the above embodiment, since the initial unbiased multimodal fusion segmentation network includes a dual-stream fusion module, the union and intersection of multiple image features can be used for feature fusion, which can not only ensure the integrity of the image features, but also consider the common characteristics between the image features, thereby improving the accuracy of image feature extraction, thereby improving the accuracy of image segmentation.
[0062] In order to more clearly understand the technical solutions provided by the embodiments of the present application, the specific implementation process and beneficial effects of the above steps will be specifically introduced in combination with the embodiments below.
[0063] The following first introduces the specific implementation process of step S102 "inputting the multimodal sample image set into the initial unbiased multimodal fusion segmentation network to obtain the image segmentation result of the target object".
[0064] According to the above embodiments, the initial unbiased multimodal fusion segmentation network includes an image fusion function module, which includes a feature extraction module and a dual-stream fusion module. In a possible implementation, the initial unbiased multimodal fusion segmentation network also includes an image segmentation function module. Figure 2 As shown, Figure 2 A schematic diagram of the structure of an initial unbiased multimodal fusion segmentation network provided in an embodiment of the present application. Figure 2 It can be seen that the initial unbiased multimodal fusion segmentation network 200 includes an image fusion function module 201 and an image segmentation function module 202. The image fusion function module 201 includes a feature extraction module 2011 and a dual-stream fusion module 2012.
[0065] In specific implementation, step S102 can be implemented in the following manner:
[0066] A1: Input the multimodal sample image set into the feature extraction module to obtain the image features of multiple sample images.
[0067] Among them, the feature extraction module is used to extract image features of sample images.
[0068] A2: Determine the union and intersection of image features of multiple sample images.
[0069] In order to ensure the integrity of multimodal image features, the union of image features of multiple sample images can be determined to extract richer image feature information. In order to extract the common points between sample images of different modalities, considering the consistency between multiple image features, the intersection of image features of multiple sample images can be determined for feature fusion, so that image features can be extracted more accurately.
[0070] A3: Input the union and intersection of image features into the dual-stream fusion module for feature fusion to obtain fused image features.
[0071] Among them, the dual-stream fusion module is used to perform feature fusion based on the union and intersection of multiple image features to obtain fused image features.
[0072] A4: Input the fused image features into the image segmentation function module to obtain the image segmentation result.
[0073] Among them, the image segmentation result includes the predicted label area.
[0074] Regarding step A1 "inputting the multimodal sample image set into the feature extraction module to obtain image features of multiple sample images", since the multiple sample images in the multimodal sample image set may have different scales and formats, before processing the sample images, the sample images can be first input into the encoder for unified encoding processing to obtain sample images of a unified format and size, and then subsequent feature extraction and other operations can be performed.
[0075] In a possible implementation, the image fusion function module may include multiple feature extraction modules, so that each sample image corresponds to a feature extraction module. That is, each feature extraction module is used to extract image features of a sample image. Optionally, an encoder may be added before each feature extraction module to encode the sample image. The feature extraction module then extracts features from the encoded sample image.
[0076] Each feature extraction module may include multiple feature extractors, each of which is used to extract image features of different scales of the sample image. Specifically, for any sample image in the multimodal sample image set, the sample image is input into multiple feature extractors to obtain multiple image features of the sample image, each image feature corresponds to a feature extractor one by one, that is, each feature extractor is used to extract an image feature of the sample image.
[0077] In a possible implementation, the feature extraction module may include four feature extractors, which are respectively used to extract four image features of the sample image. For example, the first feature extractor extracts the first image feature of the sample image, and the resolution of the image feature is 1 / 2 of the resolution of the sample image. Then the first image feature can be input into the second feature extractor, and the second image feature of the sample image can be output, and the resolution of the image feature is 1 / 4 of the resolution of the sample image. The second image feature is input into the third feature extractor, and the third image feature of the sample image is output, and the resolution of the image feature is 1 / 8 of the resolution of the sample image. The third image feature is input into the fourth feature extractor, and the fourth image feature of the sample image is output, and the resolution of the image feature is 1 / 16 of the resolution of the sample image.
[0078] It should be noted that the number of feature extractors and the resolution of extracted image features provided in the above embodiment are only exemplary descriptions and are not limited to the above implementations. The number of feature extractors and the resolution of extracted image features can be adjusted in combination with actual application scenarios. Based on this, the implementation principle of the image fusion function module will be introduced below in combination with a specific application scenario.
[0079] See also Figure 3 , which is a structural diagram of an image fusion function module provided in an embodiment of the present application.
[0080] according to Figure 3 It can be seen that the image fusion function module 201 includes a feature extraction module 2011 and a dual-stream fusion module 2012, and each sample image corresponds to a feature extraction module 2011 and a dual-stream fusion module 2012. Among them, the feature extraction module 2011 includes an encoder and four feature extractors, which are used to extract four image features of the sample image. The feature extractor includes a Swin Transformer, a DCFormer, and a regularization layer. The dual-stream fusion module 2012 is used for feature fusion to obtain fused image features.
[0081] In a possible implementation, the feature extractor includes a shifted window (Swin) transformer and a dual-stream cross-selective transformer (DCFormer), and the DCFormer includes a dual-stream multimodal cross-attention (DMCSA) module and a multilayer perceptron (MLP). The following will introduce the implementation principle of the feature extractor in combination with its structure.
[0082] In the specific implementation, the sample image is input into Swin Transformer to obtain the first image features. Swin Transformer introduces a sliding window mechanism, which allows the attention mechanism to work in a local range through the sliding window, while maintaining the connection across windows to reduce the complexity of the calculation. Swin Transformer also adopts a hierarchical structure design, gradually increasing the receptive field through the downsampling layer to extract features of different scales. By introducing relative position encoding, the contextual relationship can be further enhanced, making the extracted first image features more accurate.
[0083] After obtaining the image features corresponding to the multiple sample images, the second image features corresponding to the sample images adjacent to the current sample image can be obtained. Then the first image features and the second image features are input into the DMCSA module to obtain the attention feature matrix. Among them, the DMCSA module is mainly used to find the cross-position relationship between different image features, and select the highly correlated parts of the multimodal image features for attention calculation. Finally, the attention feature matrix is input into the MLP for feature extraction to obtain the image features.
[0084] In one possible implementation, the feature extractor may also include a regularization layer to prevent overfitting during training. The regularization layer may be connected after the DCFormer, and after the attention feature matrix is input to the MLP in the DCFormer, the MLP may output the initial image features. The initial image features are then input to the regularization layer, and finally the image features are output. The implementation principle of DCFormer will be introduced in conjunction with an application scenario.
[0085] See also Figure 4 , which is a structural schematic diagram of a dual-stream cross-selection conversion model provided in an embodiment of the present application.
[0086] In this application scenario, DCFormer includes convolution kernel, regularization layer, DMCSA and MLP. and the second image feature After that, we can use the convolution kernel Perform convolution operation and combine the image features and Input to the regularization layer for processing, and output the processed first image feature I m Similarly, Perform the same process to obtain the processed second image feature I m+1 Then I m and I m+1Input into DMCSA to get the attention feature matrix. Then the attention feature matrix can be input into the regularization layer for processing to get the processed attention feature matrix. The processed attention feature matrix is input into MLP and jump-connected with the input of MLP to get the initial image matrix output by DCFormer
[0087] In one possible implementation, the DMCSA module can obtain the attention feature matrix in the following way:
[0088] B1: Divide the first image feature into multiple non-overlapping regions to obtain a third image feature corresponding to each region, and divide the second image feature into multiple non-overlapping regions to obtain a fourth image feature corresponding to each region.
[0089] For example, the matrix size of the first image feature (second image feature) is expressed as HxW, and the first image feature (second image feature) can be divided into SxS regions to obtain the third image feature (fourth image feature) corresponding to each region. The size can be expressed as HW / S 2 .
[0090] B2: The third image feature and the fourth image feature are processed based on the attention mechanism to obtain the first query matrix, the first key matrix, the first value matrix corresponding to the third image feature and the second query matrix, the second key matrix, and the second value matrix corresponding to the fourth image feature.
[0091] Among them, the attention mechanism mainly assigns different weights to different parts of the input data (third image features, fourth image features), and pays more attention to processing the more critical information in the input data, that is, assigning parts with higher weights, which can simplify the calculation complexity and improve the accuracy of feature extraction. The attention mechanism mainly generates representations of query matrix, key matrix and value matrix through linear transformation of input data. Among them, linear transformation is achieved by multiplying the input data with three weight matrices respectively. These weight matrices can be obtained in the model training process of applying the attention mechanism. The following will be explained in conjunction with a specific application scenario.
[0092] In this application scenario, I m Represents the first image feature, with I m+1 Denotes the second image feature as the input of DMCSA. DMCSA first divides the two image features into S×S non-overlapping regions to obtain the third image feature and the fourth image feature Thus I m and I m+1 Reshape To represent. Using the attention mechanism, the third image feature and the fourth image feature Processing is performed to obtain the third image feature The corresponding first query matrix Q m , the first bond matrix K m , the first value matrix V m , and the fourth image feature The corresponding second query matrix Q m+1 , the second bond matrix K m+1 , the second value matrix V m+1 .in, and They represent the corresponding weight matrices respectively. and They represent the corresponding weight matrices respectively.
[0093] B3: Determine an adjacency matrix based on the first query matrix and the second key matrix.
[0094] For the first image feature, each region obtained by division (the third image feature) corresponds to a set of first query matrices, first key matrices, and first value matrices, so the first query mean matrix can be determined based on the first query matrices corresponding to the multiple regions. That is, the average values of the multiple first query matrices are calculated to obtain the first query mean matrix. Similarly, for the second image feature, each region obtained by division (the fourth image feature) corresponds to a set of second query matrices, second key matrices, and second value matrices, so the second key mean matrix can be determined based on the second key matrices corresponding to the multiple regions. That is, the average values of the multiple second key matrices are calculated to obtain the second key mean matrix.
[0095] Then, the product of the first query mean matrix and the transpose of the second key mean matrix is calculated to determine the adjacency matrix, which can represent the degree of semantic association between two regions in different modalities. represents the first query mean matrix, Represents the second key mean matrix, with A r represents the adjacency matrix, then we have
[0096] B4: Use the TopK algorithm to process the adjacency matrix and obtain k indexes.
[0097] The goal of the TopK algorithm is to find the first k largest elements from each row of the adjacency matrix as k indexes to obtain In r =topk(A r ), ensuring that Inr The i-th row contains the k indexes of the regions most relevant to the i-th region. The main principle of the TopK algorithm includes first building a heap with the first k elements in the data set (each row of the adjacency matrix). If the top k largest elements are required, a small heap is built. Then the remaining multiple elements are compared with the top element of the heap in turn. If it is greater than the top element of the heap, the top element of the heap is replaced. After comparing the remaining multiple elements with the top element of the heap in turn, the remaining k elements in the heap are the top k largest elements. It should be noted that the value of k can be set in combination with the actual application scenario, and the embodiments of the present application do not make specific limitations on this.
[0098] B5: Processing the second key matrix based on the k indexes to obtain a key matrix, and processing the second value matrix based on the k indexes to obtain a value matrix.
[0099] That is, the rows of the second key matrix are screened using k indexes to obtain the key matrix The gather operation in PyTorch is an index collection operation that collects elements from the input tensor according to the given index. The second value matrix is screened using k indexes to obtain the value matrix
[0100] B6: Use the attention function to process the first query matrix, key matrix and value matrix to obtain the regional attention feature matrix.
[0101] The attention function can represent the query matrix as a weighted sum of the value matrix, so that each first query matrix can only process the most semantically relevant key-value pairs. The attention function calculation can be expressed as
[0102] In a possible implementation, local context enhancement features can also be introduced to obtain a regional attention feature matrix. Specifically, the first query matrix, the key matrix, and the value matrix are calculated using the attention function to obtain a weighted feature matrix. The second value matrix is convolved to obtain a convolution value matrix. Based on the weighted feature matrix and the convolution value matrix, the regional attention feature matrix is determined.
[0103] It should be noted that the embodiment of the present application does not limit the size of the convolution kernel of the convolution operation. For example, a convolution kernel of 5x5 can be set to perform the convolution operation. Then the regional attention feature matrix O m It can be expressed as
[0104] B7: Combine the regional attention feature matrices corresponding to the multiple regions to obtain an attention feature matrix, wherein the attention matrix has the same scale as the first image feature.
[0105] That is, the regional attention matrix corresponding to each region is reorganized and unpatchified to obtain the attention matrix. The core step of reorganization and unpatchify is to reassemble the segmented patches and restore them to the structure of the original image. This process needs to ensure that the step size of the patch and the image size meet specific conditions to ensure that the patches can completely cover the image and align with each other to ensure lossless restoration of the image.
[0106] After obtaining the attention feature matrix, the attention feature matrix can be input into the MLP in DCFormer, and the MLP can output the initial image features. The initial image features are then input into the regularization layer, and finally the image features are output.
[0107] With respect to step A2 "determining the union and intersection of image features of multiple sample images", when each feature extraction module includes multiple feature extractors, that is, each sample image can extract multiple image features. In a possible implementation, multiple image feature sets can be determined based on multiple image features corresponding to multiple sample images, respectively, wherein each image feature set includes multiple image features of the same scale. That is, the image features of different sample images extracted by each feature extractor are combined into an image feature set. For any image feature set, the union and intersection of multiple image features in the image feature set are determined. That is, the union and intersection of image features are determined for multiple image features of the same scale, and each image feature set can correspond to the determination of the union and intersection of a set of image features.
[0108] With respect to step A3 "inputting the union and intersection of image features into the dual-stream fusion module for feature fusion to obtain fused image features", after obtaining the union and intersection of image features of multiple sample images, the union and intersection of image features are input into the dual-stream fusion module for feature fusion to obtain fused image features. In one possible implementation, the dual-stream fusion module includes a dual-stream adapter (DA) and a multi-scale feature fusion module, and the fused image features are obtained through the dual-stream adapter and the multi-scale feature fusion module.
[0109] In specific implementation, according to the embodiment corresponding to step A2, the image features of different sample images extracted by each feature extractor can form an image feature set, and multiple image feature sets corresponding to multiple feature extractors can be obtained. For each image feature set, the union and intersection of multiple image features in the image feature set can be determined.
[0110] For any image feature set, the union corresponding to the image feature set is input into the dual-stream adapter for dimensional processing to obtain a preprocessed union. For example, the dual-stream adapter includes a convolution kernel, and the union is first increased in dimension and then reduced in dimension by a convolution operation to obtain a preprocessed union. The intersection corresponding to the image feature set is input into the dual-stream adapter for dimensional processing to obtain a preprocessed intersection. For example, the intersection is first increased in dimension and then reduced in dimension by using the convolution kernel in the dual-stream adapter to obtain a preprocessed union. The intersection corresponding to the image feature set is input into the dual-stream adapter for dimensional processing to obtain a preprocessed intersection.
[0111] Based on the dual-stream adapter, the union and the preprocessed intersection are processed to determine the first initial fused image feature. For example, the dual-stream adapter can splice the union of image features and the preprocessed intersection to obtain the first initial fused image feature. Based on the dual-stream adapter, the intersection and the preprocessed union are processed to determine the second initial fused image feature. For example, the dual-stream adapter can splice the intersection of image features and the preprocessed union to obtain the second initial fused image feature. Then the first initial fused image feature and the second initial fused image feature are input into the multi-scale feature fusion module for feature fusion to obtain the fused image feature. Since there are multiple image feature sets, multiple fused image features can be obtained accordingly.
[0112] For ease of understanding, a specific introduction will be given below in conjunction with an application scenario.
[0113] In this application scenario, X={X1,X2,…,X M} represents a set of multimodal sample images, where represents the mth sample image among M multimodal sample images, H, W and C m Respectively represent the height, width and number of channels of the mth sample image. The true label area can be expressed as in Represents the true label area of N categories. m After being input into the feature extraction module, the obtained image features can be expressed as Where I represents the number of feature extractors in the feature extraction module, Represents the sample image X extracted by the i-th feature extractor m The image features of different sample images extracted by each feature extractor are combined into an image feature set Thus, multiple image feature sets are obtained. For any image feature set, the union of multiple image features in the image feature set is determined and intersection
[0114] Optionally, the dual-stream adapter includes two branches, the inputs of the two branches are respectively the union of and intersection The branch operation of first increasing dimension and then decreasing dimension can be defined as The branch operation of first reducing dimension and then increasing dimension is defined as Then take the union As the input branch, output the first initial fused image feature It can be expressed as By intersection As the input branch, output the second initial fusion image feature It can be expressed as Then, the first initial fused image feature and the second initial fused image feature can be input into a multi-scale feature fusion module for feature fusion to obtain a fused image feature.
[0115] In a possible implementation, the multi-scale feature fusion module may include a multi-scale convolution branch and a multi-scale pooling branch, wherein the multi-scale convolution branch includes multiple convolution layers, and the multi-scale pooling branch includes multiple pooling layers, for obtaining fused image features. The following will be specifically described in conjunction with the structure of the multi-scale feature fusion module.
[0116] Since the dual-stream adapter outputs two initial fused image features, the first initial fused image feature and the second initial fused image feature may be firstly spliced to obtain a spliced fused image feature.
[0117] In a possible implementation, concatenating the first initial fused image features and the second initial fused image features may change the number of channels of the image features. Therefore, the concatenated fused image features may be input into a multi-channel 1x1 convolution kernel for convolution, i.e., pointwise convolution (PW), and the number of channels of the concatenated fused image features may be corrected by using the pointwise convolution operation.
[0118] At the same time, the spliced fusion image features can be input into multiple convolution layers respectively to obtain convolution layer fusion image features. Since the multi-scale convolution branch includes multiple convolution layers, the multiple convolution layers can be parallel structures, and the spliced fusion image features can be input into multiple convolution layers respectively, and the output features of the multiple convolution layers and the spliced fusion image features are added to obtain the convolution layer fusion image features. For example, the multi-scale convolution branch includes three parallel convolution layers, including a 3x3 convolution kernel, a 5x5 convolution kernel, and a 7x7 convolution kernel, which are used to perform convolution operations on the spliced fusion image features.
[0119] The spliced fusion image features are respectively input into multiple pooling layers to obtain pooling layer fusion image features. Since the multi-scale pooling branch includes multiple pooling layers, the multiple pooling layers can be parallel structures, and the spliced fusion image features can be respectively input into multiple pooling layers, and the output features of the multiple pooling layers and the spliced fusion image features are added to obtain the pooling layer fusion image features. For example, the multi-scale pooling branch can include parallel average pooling layers AvgPooling, which respectively include 1x1 pooling windows, 2x2 pooling windows, and 3x3 pooling windows, which are used to perform pooling operations on the spliced fusion image features.
[0120] Then, based on the convolutional layer fusion image features and the pooling layer fusion image features, the fusion image features are determined. For example, the convolutional layer fusion image features and the pooling layer fusion image features can be added to obtain the fusion image features.
[0121] In a possible implementation, a 1x1 convolution kernel may be added after each pooling layer, and then the output of the convolution kernel may be upsampled to restore the output of the multi-scale pooling branch to the size before pooling.
[0122] For details, please refer to Figure 5 , which is a schematic diagram of the structure of a dual-stream fusion module provided in an embodiment of the present application.
[0123] The dual-stream fusion module 2012 includes: a dual-stream adapter 501 and a multi-scale feature fusion module 502. The dual-stream adapter 501 includes two branches, and the inputs of the two branches are respectively the union of and intersection Pair Union Perform dimensionality increase first and then dimensionality reduction to obtain the preprocessed set and the intersection Perform dimensionality reduction first and then dimensionality increase to obtain the preprocessed intersection Then based on the union and preprocessed intersection Get the first initial fusion image feature Intersection-based and preprocessed union Get the second initial fusion image feature
[0124] Then the first initial fusion image feature and the second initial fusion image feature Splicing is performed to obtain spliced fusion image features. The multi-scale feature fusion module 502 includes a multi-scale convolution branch and a multi-scale pooling branch. The spliced fusion image features can be input into the convolution kernel of PW Conv1x1 for channel number correction, and then respectively input into three parallel convolution layers in the multi-scale convolution branch. The three convolution layers include a 3x3 convolution kernel, a 5x5 convolution kernel, and a 7x7 convolution kernel. The output features of multiple convolution layers and the spliced fusion image features are added to obtain the convolution layer fusion image features. Then, the convolution kernel of PW Conv1x1 can be used to correct the number of channels for the convolution layer fusion image features.
[0125] At the same time, the spliced fused image features can be input into the convolution kernel of PW Conv1x1 for channel number correction, and then respectively input into the three parallel average pooling layers in the multi-scale pooling branch. The three average pooling layers include a 1x1 pooling window, a 2x2 pooling window, and a 3x3 pooling window. A 1x1 convolution kernel can also be added after each pooling layer, and the output of the convolution kernel is upsampled to restore the output of the multi-scale pooling branch to the size before pooling. Then the output features of multiple pooling layers and the spliced fused image features are added to obtain the pooling layer fused image features. In addition, the pooling kernel of PW Conv1x1 can also be used to correct the number of channels of the pooling layer fused image features.
[0126] Finally, the convolution layer fused image features and the pooling layer fused image features are added to output the fused image features. In a possible implementation, the image features obtained by adding the convolution layer fused image features and the pooling layer fused image features can be input into the linear layer to perform a linear transformation on the input image features. For example, the linear layer can calculate the input image features by matrix multiplication (weight matrix) and adding a bias vector to obtain the fused image features. By adjusting the weights and biases, useful features can be extracted and learned from the input image features, and these features can be mapped to a dimensional space suitable for a specific task.
[0127] With respect to step A4 “inputting the fused image features into the image segmentation function module to obtain the image segmentation result”, the image segmentation function module may perform image segmentation on the fused image features and output the image segmentation result.
[0128] Based on the embodiment corresponding to the above step A3, it can be known that multiple fused image features can be obtained according to the image features of the sample images respectively extracted by multiple feature extractors. Based on this, the embodiment of the present application provides an image segmentation function module, which includes an encoder and a decoder, the encoder includes multiple encoding modules, the decoder includes multiple decoding modules, and except for the last decoding module, the output of each decoding module can be used as the input of the next decoding module, that is, the decoding modules are connected in a front-to-back relationship. The number of multiple encoding modules (decoding modules) is the same as the number of fused image features.
[0129] In a possible implementation, the image segmentation function module can perform image segmentation in the following manner: input the multimodal sample image set into the corresponding first encoding module, and use the output of the first encoding module as the input of the first decoding module. That is, among the multiple encoding modules, there is a first encoding module that takes the multimodal sample image set as input. Then the output of the first encoding module is input into the first decoding module, and the first decoding module can represent the decoding module corresponding to the first encoding module among the multiple decoding modules. For any fused image feature among the multiple fused image features, the fused image feature is input into the second encoding module corresponding thereto, and the output of the second encoding module is used as the input of the corresponding second decoding module, and the second decoding module represents the decoding module corresponding to the second encoding module. Since the number of fused image features is the same as the number of encoding modules, and the input of the first encoding module is a multimodal sample image set, there is a remaining fused image feature among the multiple fused image features that has no corresponding encoding module, and the remaining fused image feature can be input into the corresponding third decoding module. At this time, since each encoding module has a corresponding decoding module, that is, the first encoding module corresponds to the first decoding module, the second encoding module corresponds to the second decoding module, and the number of encoding modules is the same as the number of decoding modules, the third decoding module represents a decoding module in the first decoding module or the second decoding module. Since multiple decoding modules are connected front and back, and the output of the previous decoding module can be used as the input of the next decoding module, the image segmentation result can be output based on the last decoding module. The following will introduce the implementation principle of the image segmentation function module in combination with a specific application scenario.
[0130] See also Figure 6 As shown, Figure 6 A schematic diagram of the structure of an image segmentation function module provided in an embodiment of the present application.
[0131] This application scenario is Figure 3Corresponding to the embodiment shown, the four dual-stream fusion modules 2012 are used to output four fused image features, which are respectively represented as fused image feature F1, fused image feature F2, fused image feature F3, and fused image feature F4. The image segmentation function module 202 includes an encoder and a decoder, wherein the encoder includes four encoding modules, which are respectively encoding module 1, encoding module 2, encoding module 3, and encoding module 4 from left to right. The decoder includes four decoding modules, which are respectively decoding module 1, decoding module 2, decoding module 3, and decoding module 4 from left to right, and the output of the previous decoding module is used as the input of the next decoding module.
[0132] Multiple sample images in the multimodal sample image set are spliced and input into encoding module 1, and the output of encoding module 1 is used as the input of decoding module 4. The fused image feature F1 is input into encoding module 2, and the output of encoding module 2 is used as the input of decoding module 3. The fused image feature F2 is input into encoding module 3, and the output of encoding module 3 is used as the input of decoding module 2. The fused image feature F3 is input into encoding module 4, and the output of encoding module 4 is used as the input of decoding module 1. The fused image feature F4 is input into decoding module 1. Finally, decoding module 4 outputs the image segmentation result Y * .
[0133] After obtaining the image segmentation result, the image segmentation result includes the predicted label area. The following will introduce step S103 "determining the loss function value based on the real label area and the predicted label area".
[0134] Among them, the true label area includes the true label area of one or more categories, and the predicted label area includes the predicted label area of one or more categories. In specific implementation, for the true label area of each category, the Dice loss value between the true label area of the category and the predicted label area of the corresponding category is determined. Then the loss function value is determined based on the Dice loss value of one or more categories. For example, after obtaining the Dice loss value corresponding to the true label area of all categories, the average of all Dice loss values can be calculated as the loss function value.
[0135] Then, the initial unbiased multimodal fusion segmentation network is trained based on the loss function value until the training cutoff condition is met, thereby obtaining a trained unbiased multimodal fusion segmentation network UMF-SegNet. For example, the training cutoff condition may be that the loss function value is less than a preset threshold, or the number of iterative training reaches a preset number.
[0136] In order to test the performance of the unbiased multimodal fusion segmentation network UMF-SegNet, the embodiment of the present application can use the data set used in the Brain Tumor Segmentation Challenge (BraTS2021Challenge), referred to as the BraTs 2021 data set. The data set contains a total of 1470 cases, of which 1251 cases can be used as training sets and 219 cases can be used as test sets. Each case contains four modalities: T1 weighted (T1), T1 weighted contrast agent (T1C), T2 weighted (T2) and fluid attenuated inversion recovery (Flair). The image segmentation effect of the unbiased multimodal fusion segmentation network can be evaluated using four evaluation indicators: Dice coefficient (Dice), sensitivity (SEN), intersection over union (IoU) and positive predictive value (PPV). And the performance of the unbiased multimodal fusion segmentation network UMF-SegNet can be compared with other existing models of image segmentation. For details, see Table 1, which shows the evaluation index results of different models.
[0137] According to Table 1, when Dice, SEN and IoU are used as evaluation indicators, the effect of the unbiased multimodal fusion segmentation network UMF-SegNet is better than other models. When PPV is used as the evaluation indicator, the effect of the multimodal fusion segmentation network UMF-SegNet is slightly lower than some methods, but it also reaches nearly 0.93. And the calculation time of the multimodal fusion segmentation network UMF-SegNet is faster than other models with the same number of parameters.
[0138] Table 1
[0139]
[0140] For details, please refer to Figure 7 As shown, Figure 7 A schematic diagram of an image segmentation result provided in an embodiment of the present application.
[0141] Figure 7 The results of image segmentation on the dataset using the unbiased multimodal fusion segmentation network UMF-SegNet and other models are shown in the figure. Figure 7 It can be seen that the unbiased multimodal fusion segmentation network UMF-SegNet is superior to other models in segmentation details.
[0142] Based on the above method embodiment, the present application embodiment also provides a multimodal image segmentation method. Figure 8 As shown, Figure 8 A flowchart of a multimodal image segmentation method provided in an embodiment of the present application.
[0143] The method may include the following steps:
[0144] S801: Acquire a multimodal image of an object to be processed;
[0145] S802: Input the multimodal image into an unbiased multimodal fusion segmentation network to obtain an image segmentation result of the object to be processed.
[0146] Among them, the unbiased multimodal fusion segmentation network is trained based on the training method of the multimodal image segmentation model described in the above method embodiment.
[0147] Based on the above method embodiment, the present application embodiment also provides a training device for a multimodal image segmentation model. Fig. 9 As shown, Fig. 9 A schematic diagram of a training device for a multimodal image segmentation model provided in an embodiment of the present application.
[0148] The device 900 comprises:
[0149] A first acquisition unit 901 is used to acquire a multimodal sample image set of a target object, wherein the multimodal sample image set includes sample images of multiple modalities and real label regions pre-annotated in the sample images;
[0150] A first image segmentation unit 902 is used to input the multimodal sample image set into an initial unbiased multimodal fusion segmentation network to obtain an image segmentation result of the target object, wherein the image segmentation result includes a predicted label area;
[0151] A loss function determining unit 903, configured to determine a loss function value based on the true label region and the predicted label region;
[0152] The training unit 904 is used to train the initial unbiased multimodal fusion segmentation network based on the loss function value until the training cutoff condition is met, thereby obtaining a trained unbiased multimodal fusion segmentation network; wherein the initial unbiased multimodal fusion segmentation network includes an image fusion function module, and the image fusion function module includes a feature extraction module and a dual-stream fusion module, the feature extraction module is used to extract the image features of the sample image, and the dual-stream fusion module is used to perform feature fusion based on the union and intersection of the multiple image features.
[0153] In a possible implementation, the initial unbiased multimodal fusion segmentation network further includes an image segmentation function module;
[0154] The first image segmentation unit 902 is specifically used to input the multimodal sample image set into the feature extraction module to obtain image features of multiple sample images; determine the union and intersection of the image features of multiple sample images; input the union and the intersection into the dual-stream fusion module for feature fusion to obtain fused image features; input the fused image features into the image segmentation function module to obtain the image segmentation result.
[0155] In a possible implementation, each sample image corresponds to a feature extraction module, the feature extraction module includes a plurality of feature extractors, each feature extractor is used to extract image features of different scales of the sample image;
[0156] The first image segmentation unit 902 is specifically configured to input the sample image into the plurality of feature extractors to obtain the plurality of image features of the sample image, wherein the image features correspond to the feature extractors in a one-to-one manner;
[0157] The first image segmentation unit 902 is specifically used to determine multiple image feature sets based on the image features of the multiple sample images, each image feature set including multiple image features of the same scale; for any image feature set, determine the union and intersection of the multiple image features in the image feature set.
[0158] In a possible implementation, the dual-stream fusion module includes a dual-stream adapter and a multi-scale feature fusion module;
[0159] The first image segmentation unit 902 is specifically used to, for any image feature set, input the union corresponding to the image feature set into the dual-stream adapter for dimensional processing to obtain a preprocessed union; input the intersection corresponding to the image feature set into the dual-stream adapter for dimensional processing to obtain a preprocessed intersection; based on the dual-stream adapter, process the union and the preprocessed intersection to determine a first initial fused image feature; based on the dual-stream adapter, process the intersection and the preprocessed union to determine a second initial fused image feature; input the first initial fused image feature and the second initial fused image feature into the multi-scale feature fusion module for feature fusion to obtain the fused image feature.
[0160] In a possible implementation, the multi-scale feature fusion module includes a multi-scale convolution branch and a multi-scale pooling branch, the multi-scale convolution branch includes a plurality of convolution layers, and the multi-scale pooling branch includes a plurality of pooling layers;
[0161] The first image segmentation unit 902 is specifically used to splice the first initial fused image features and the second initial fused image features to obtain spliced fused image features; input the spliced fused image features to multiple convolutional layers respectively to obtain convolutional layer fused image features; input the spliced fused image features to multiple pooling layers respectively to obtain pooling layer fused image features; determine the fused image features based on the convolutional layer fused image features and the pooling layer fused image features.
[0162] In a possible implementation, the feature extractor includes a shift window Swin Transformer and a dual-stream cross selection transformation model DCFormer, and the DCFormer includes a dual-stream multimodal cross attention DMCSA module and a multi-layer perceptron MLP;
[0163] The first image segmentation unit 902 is specifically used to input the sample image into the SwinTransformer to obtain the first image feature; obtain the second image feature of the sample image adjacent to the sample image; input the first image feature and the second image feature into the DMCSA to obtain the attention feature matrix; input the attention feature matrix into the MLP to obtain the image feature.
[0164] In a possible implementation, the first image segmentation unit 902 is specifically used for the DMCSA to divide the first image feature into multiple non-overlapping areas to obtain a third image feature corresponding to each area, and to divide the second image feature into multiple non-overlapping areas to obtain a fourth image feature corresponding to each area; based on the attention mechanism, the third image feature and the fourth image feature are processed to obtain a first query matrix, a first key matrix, and a first value matrix corresponding to the third image feature, and a second query matrix, a second key matrix, and a second value matrix corresponding to the fourth image feature; based on the first query matrix and the second key matrix, an adjacency matrix is determined; the adjacency matrix is processed using the TopK algorithm to obtain k indexes; the second key matrix is processed based on the k indexes to obtain a key matrix, and the second value matrix is processed based on the k indexes to obtain a value matrix; the first query matrix, the key matrix, and the value matrix are processed using the attention function to obtain a regional attention feature matrix; the regional attention feature matrices corresponding to multiple areas are combined to obtain the attention feature matrix, and the attention matrix has the same scale as the first image feature.
[0165] In a possible implementation, the first image segmentation unit 902 is specifically used to determine a first query mean matrix based on first query matrices corresponding to multiple regions respectively; determine a second key mean matrix based on second key matrices corresponding to multiple regions respectively; and determine the adjacency matrix based on the product of the transpose of the first query mean matrix and the second key mean matrix.
[0166] In a possible implementation, the first image segmentation unit 902 is specifically used to use the attention function to calculate the first query matrix, the key matrix and the value matrix to obtain a weighted feature matrix; perform a convolution operation on the second value matrix to obtain a convolution value matrix; and determine the regional attention feature matrix based on the weighted feature matrix and the convolution value matrix.
[0167] In a possible implementation, the feature extractor further includes: a regularization layer;
[0168] The first image segmentation unit 902 is specifically used to input the attention feature matrix into the MLP to obtain initial image features; and input the initial image features into the regularization layer to obtain the image features.
[0169] In a possible implementation, the image segmentation function module includes an encoder and a decoder, the encoder includes a plurality of encoding modules, the decoder includes a plurality of decoding modules, and the output of each decoding module serves as the input of the next decoding module, and the number of the encoding modules / decoding modules is the same as the number of the fused image features;
[0170] The first image segmentation unit 902 is specifically used to input the multimodal sample image set into the corresponding first encoding module, and use the output of the first encoding module as the input of the first decoding module; input any fused image feature into the corresponding second encoding module, and use the output of the second encoding module as the input of the corresponding second decoding module; input the remaining fused image features into the corresponding third decoding module; and output the image segmentation result based on the last decoding module.
[0171] In a possible implementation, the true label region includes one or more categories of true label regions, and the predicted label region includes one or more categories of predicted label regions;
[0172] The loss function determination unit 903 is specifically used to determine, for each category of the true label area, a Dice loss value between the true label area of the category and the predicted label area of the corresponding category; and determine the loss function value based on the Dice loss values of one or more categories.
[0173] The beneficial effects of the training device for the multimodal image segmentation model provided in the embodiment of the present application can be found in the above method embodiment and will not be repeated here.
[0174] In addition, the present application also provides a multimodal image segmentation device. Fig.10 As shown, Fig.10 A schematic diagram of a multimodal image segmentation device provided in an embodiment of the present application.
[0175] The device 1000 comprises:
[0176] The second acquisition unit 1001 is used to acquire a multimodal image of the object to be processed;
[0177] The second image segmentation unit 1002 is used to input the multimodal image into an unbiased multimodal fusion segmentation network to obtain an image segmentation result of the object to be processed. The unbiased multimodal fusion segmentation network is trained based on the training method of the multimodal image segmentation model described in the above method embodiment.
[0178] Based on the above method embodiment and device embodiment, the present application embodiment further provides an electronic device, which will be described below in conjunction with the accompanying drawings.
[0179] See also Fig.11 , Fig.11 A schematic diagram of an electronic device provided in an embodiment of the present application.
[0180] The device 1100 includes: a memory 1101 and a processor 1102;
[0181] The memory 1101 is used to store relevant program codes;
[0182] The processor 1102 is used to call the program code to execute the multimodal image segmentation model training method or the multimodal image segmentation method described in the above method embodiment.
[0183] In addition, an embodiment of the present application also provides a computer-readable storage medium, which is used to store a computer program, and the computer program is used to execute the training method of the multimodal image segmentation model or the multimodal image segmentation method described in the above method embodiment.
[0184] An embodiment of the present application also provides a computer program product, which includes a computer program / instructions. When the computer program / instructions are executed by a processor, the training method of the multimodal image segmentation model or the multimodal image segmentation method described in the above method embodiment is implemented.
[0185] It should be noted that the computer-readable medium mentioned above in the present application may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0186] The computer program product may be written in any combination of one or more programming languages to write program codes for performing the operations of the embodiments of the present application, including object-oriented programming languages, such as Java, C++, etc., and conventional procedural programming languages, such as "C" language or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as an independent software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0187] It should be noted that the various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same and similar parts between the various embodiments can refer to each other. In particular, for the system or device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiment. The device embodiment described above is only exemplary, in which the units or modules described as separate components may or may not be physically separated, and the components displayed as units or modules may or may not be physical modules, that is, they may be located in one place, or they may be distributed on multiple network units, and some or all of the units or modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative work.
[0188] The flowchart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the method, device and equipment etc. of various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or the flowchart, and the combination of the boxes in the block diagram and / or the flowchart can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0189] It should be understood that in the present application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0190] It should also be noted that, in this application, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the statement "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0191] The above description of the disclosed embodiments enables professionals and technicians in the field to implement or use the present application. Various modifications to these embodiments will be apparent to professionals and technicians in the field, and the general principles defined in this application can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown in the present application, but will conform to the widest range consistent with the principles and novel features disclosed in the present application.
Claims
1. A training method for a multimodal image segmentation model, characterized in that: The method comprises: Acquire a set of multimodal sample images of a target object, wherein the set of multimodal sample images includes sample images of multiple modalities and real label regions pre-annotated in the sample images; Inputting the multimodal sample image set into an initial unbiased multimodal fusion segmentation network to obtain an image segmentation result of the target object, wherein the image segmentation result includes a predicted label area; Determine a loss function value based on the true label area and the predicted label area; Based on the loss function value, the initial unbiased multimodal fusion segmentation network is trained until a training cutoff condition is met to obtain a trained unbiased multimodal fusion segmentation network; Among them, the initial unbiased multimodal fusion segmentation network includes an image fusion function module, and the image fusion function module includes a feature extraction module and a dual-stream fusion module. The feature extraction module is used to extract the image features of the sample image, and the dual-stream fusion module is used to perform feature fusion based on the union and intersection of the multiple image features.
2. The method according to claim 1, characterized in that The initial unbiased multimodal fusion segmentation network also includes an image segmentation function module; The step of inputting the multimodal sample image set into an initial unbiased multimodal fusion segmentation network to obtain an image segmentation result of the target object comprises: Inputting the multimodal sample image set into the feature extraction module to obtain image features of a plurality of the sample images; Determining the union and intersection of image features of a plurality of the sample images; Inputting the union and the intersection into the dual-stream fusion module for feature fusion to obtain fused image features; The fused image features are input into the image segmentation function module to obtain the image segmentation result.
3. The method according to claim 2, characterized in that Each sample image corresponds to a feature extraction module, the feature extraction module includes a plurality of feature extractors, each feature extractor is used to extract image features of different scales of the sample image; The step of inputting the multimodal sample image set into the feature extraction module to obtain image features of a plurality of the sample images includes: Inputting the sample image into the plurality of feature extractors to obtain a plurality of image features of the sample image, wherein the image features correspond to the feature extractors one by one; The determining of the union and intersection of the image features of the plurality of sample images comprises: Based on the image features of the plurality of sample images, determining a plurality of image feature sets, each image feature set including a plurality of image features of the same scale; For any image feature set, a union and an intersection of multiple image features in the image feature set are determined.
4. The method according to claim 3, characterized in that The dual-stream fusion module includes a dual-stream adapter and a multi-scale feature fusion module; The step of inputting the union and the intersection into the dual-stream fusion module for feature fusion to obtain fused image features includes: For any image feature set, inputting the union corresponding to the image feature set into the dual-stream adapter for dimensional processing to obtain a preprocessed union; Inputting the intersection corresponding to the image feature set into the dual-stream adapter for dimensional processing to obtain a preprocessed intersection; Processing the union and the preprocessed intersection based on the dual-stream adapter to determine a first initial fused image feature; Processing the intersection and the preprocessed union based on the dual-stream adapter to determine a second initial fused image feature; The first initial fused image feature and the second initial fused image feature are input into the multi-scale feature fusion module for feature fusion to obtain the fused image feature.
5. The method according to claim 4, characterized in that The multi-scale feature fusion module includes a multi-scale convolution branch and a multi-scale pooling branch, the multi-scale convolution branch includes a plurality of convolution layers, and the multi-scale pooling branch includes a plurality of pooling layers; The step of inputting the first initial fused image feature and the second initial fused image feature into the multi-scale feature fusion module for feature fusion to obtain the fused image feature includes: Splicing the first initial fused image feature and the second initial fused image feature to obtain a spliced fused image feature; Inputting the spliced fused image features into the multiple convolutional layers respectively to obtain convolutional layer fused image features; Inputting the spliced fused image features into the multiple pooling layers respectively to obtain pooling layer fused image features; The fused image feature is determined based on the convolutional layer fused image feature and the pooling layer fused image feature.
6. The method according to claim 3, characterized in that The feature extractor includes a shift window SwinTransformer and a dual-stream cross selection transformation model DCFormer, wherein the DCFormer includes a dual-stream multimodal cross attention DMCSA module and a multi-layer perceptron MLP; The step of inputting the sample image into the plurality of feature extractors to obtain the plurality of image features of the sample image comprises: Inputting the sample image into the Swin Transformer to obtain a first image feature; Acquire a second image feature of a sample image adjacent to the sample image; Inputting the first image feature and the second image feature into the DMCSA to obtain an attention feature matrix; The attention feature matrix is input into the MLP to obtain the image features.
7. The method according to claim 6, characterized in that The step of inputting the first image feature and the second image feature into the DMCSA to obtain an attention feature matrix includes: The DMCSA divides the first image feature into a plurality of non-overlapping regions to obtain a third image feature corresponding to each region, and divides the second image feature into a plurality of non-overlapping regions to obtain a fourth image feature corresponding to each region; The third image feature and the fourth image feature are processed based on an attention mechanism to obtain a first query matrix, a first key matrix, and a first value matrix corresponding to the third image feature, and a second query matrix, a second key matrix, and a second value matrix corresponding to the fourth image feature; determining an adjacency matrix based on the first query matrix and the second key matrix; Using the TopK algorithm to process the adjacency matrix, obtaining k indexes; Processing the second key matrix based on the k indexes to obtain a key matrix, and processing the second value matrix based on the k indexes to obtain a value matrix; Using an attention function to process the first query matrix, the key matrix, and the value matrix to obtain a regional attention feature matrix; The regional attention feature matrices corresponding to multiple regions are combined to obtain the attention feature matrix, and the scale of the attention matrix is the same as that of the first image feature.
8. The method according to claim 7, characterized in that The determining of an adjacency matrix based on the first query matrix and the second key matrix comprises: Determine a first query mean matrix based on first query matrices corresponding to the plurality of regions respectively; Determine a second key mean matrix based on the second key matrices corresponding to the plurality of regions respectively; The adjacency matrix is determined based on a product of the first query mean matrix and a transpose of the second key mean matrix.
9. The method according to claim 4, characterized in that The image segmentation function module includes an encoder and a decoder, the encoder includes a plurality of encoding modules, the decoder includes a plurality of decoding modules, and the output of each decoding module is used as the input of the next decoding module, and the number of the encoding modules / decoding modules is the same as the number of the fused image features; The step of inputting the fused image features into the image segmentation function module to obtain the image segmentation result includes: Inputting the multimodal sample image set into a corresponding first encoding module, and using the output of the first encoding module as the input of a first decoding module; Input any fused image feature into a corresponding second encoding module, and use the output of the second encoding module as the input of a corresponding second decoding module; Input the remaining fused image features into the corresponding third decoding module; The image segmentation result is output based on the last decoding module.
10. A multimodal image segmentation method, characterized in that: The method comprises: Acquire a multimodal image of the object to be processed; The multimodal image is input into an unbiased multimodal fusion segmentation network to obtain an image segmentation result of the object to be processed, wherein the unbiased multimodal fusion segmentation network is trained based on the method described in any one of claims 1 to 9.